Computer vision will dissolve into end-to-end perception-action loops, and 3D and camera poses will become obsolete for training robots. The real bottleneck is paired perception-action data, so video world models help bu
The Physical AI Voices by Herbert Scale Experts.