World models predict what happens next, so robots can be trained, tested and steered in imagination. What is known, who argues what, and the dates that will decide it, with a source for every fact.
World models went from a research idea to a crowded field in under two years. Google DeepMind's Genie 3 generates interactive worlds in real time for a few minutes (Google DeepMind, 5 Aug 2025), and Waymo now builds its driving simulator on it (Waymo, 6 Feb 2026); NVIDIA released Cosmos 3 as an open model that predicts video and robot actions in one system (NVIDIA, 1 Jun 2026); Meta's V-JEPA 2 planned pick-and-place on robots in labs it had never seen after less than 62 hours of robot video (arXiv 2506.09985); 1X made a video world model the policy of its NEO humanoid (1X, 12 Jan 2026); and in China 33 world-model start-ups raised more than 26 billion yuan (36Kr, 22 Jun 2026). Money followed the ideas: Yann LeCun's AMI Labs raised $1.03 billion (SiliconANGLE, 10 Mar 2026). The limits are stated by the makers themselves: interaction that holds for minutes, not hours; generated scenes that look right while contact and geometry are wrong; and imagination that is still too slow to run on the robot. Almost every robot use is still in the lab or at a first pilot.
What is a world model, and which kind does a robot need?
A world model predicts what happens next: given the current state and an action, it predicts the next state, which a robot can use to plan. That is Yann LeCun's definition in his Lemley Lecture (Brown University, 1 Apr 2026).
Fei-Fei Li and World Labs sort world models by what they output: renderers (pixels for people to watch), simulators (geometry, physics and dynamics for programs) and planners (actions toward a goal) (Substack, 3 Jun 2026).
A newer class matters most for robots: the world action model, one network that predicts both future video and the robot's actions; NVIDIA's DreamZero defines it as a model that learns "physical dynamics by predicting future world states and actions" (arXiv 2602.15922).
Predict pixels, or predict in an abstract space?
On the same robot task, Meta's latent V-JEPA 2-AC scored 80% against 0% for a Cosmos video model and planned an action in 16 seconds instead of 4 minutes; on grasping a box it reached only 25% (arXiv 2506.09985; Meta, 11 Jun 2025).
Nearly every robot world model released in 2025 and 2026 predicts video: 1X, DreamZero, DreamDojo, Cosmos, Genie Envisioner and Google DeepMind's Veo evaluator (our note). Their advantage is that a policy can watch the output, so the model doubles as a test bench.
BAAI's Physis-v0.1 predicts the next physical state in latent space, with depth, point clouds and force as inputs (TMTPost, 18 Jun 2026).
Can a world model learn physics, and touch?
Physics-IQ (Google DeepMind) found physical understanding in video models "severely limited, and unrelated to visual realism", with the best model at 24.1 out of 100 at launch (arXiv 2501.09038); a later audit changed 57.6% of its samples and reshuffled the ranking (arXiv 2606.18943).
ByteDance found that video models generalise in distribution but fail outside it, and concluded that "scaling alone is insufficient" for them to uncover physical laws (arXiv 2411.02385); on Meta's IntPhys 2, state-of-the-art models sit at chance while people are near perfect (arXiv 2506.09849).
NVIDIA's own Cosmos paper says "all the WFMs equally struggle with physics adherence" (arXiv 2501.03575).
Can a world model replace real-robot testing?
1X calculates that a world model with 70% accuracy picks the better of two policies 90% of the time when their real success rates differ by 15 points (1X, 16 Jun 2025).
Google DeepMind checked its Veo-based simulator against more than 1,600 real trials of eight policy checkpoints and reports that it predicts the ranking of policies and finds safety failures by red-teaming; predicted absolute success was lower than real (arXiv 2512.10675).
World Labs ran 2,000 simulated and 100 real trials per checkpoint and reports that rankings held, adding "A useful simulation need not match real-world success rates exactly." (World Labs, 28 Jul 2026).
Can generated data replace robot data?
NVIDIA's DreamGen taught a humanoid 22 new behaviours from teleoperation data of one pick-and-place task in one place (arXiv 2505.12705); NVIDIA says GR00T-Dreams produced the data for GR00T N1.5 in 36 hours instead of nearly three months (company figure; NVIDIA, 19 May 2025).
Waymo uses its Genie 3-based model to create events its fleet has never met, such as a tornado or an elephant (Waymo, 6 Feb 2026).
Real robot data still drives the scaling: AgiBot's GE-Act 2.0 went from 17.1% to 44.1% success as robot data grew from 300 to 30,000 hours (authors' figures; arXiv 2609.05588), real progress and still far from shift-ready reliability.
The road to intelligence, or one part of the robot stack?
LeCun left Meta in November 2025 to build world models and said of language models: "They are not a path to human-level intelligence." (Euronews, 20 Nov 2025); AMI Labs names robotics among its targets (AMI Labs).
Demis Hassabis argues for pushing language-model scaling to the maximum while Google DeepMind also runs Genie (Business Insider via AOL, 7 Dec 2025).
The released robot models are converging on hybrids that predict video and actions together: DreamZero, Cosmos 3, GE-Act 2.0, DW0.5 and τ0-WM (our note).