World Models: The AI Breakthrough That Could Dethrone LLMs
Yann LeCun quit Meta, called LLMs a dead end, and raised $1 billion to prove it. Here is why world models are the AI bet every investor needs to understand.
Key Takeaways
- Physics gap: AI models score near chance on physics tests that humans ace at 85-95% accuracy.
- $3.7B bet: AMI Labs, World Labs, and Wayve collectively raised over $3.7 billion for world model research in 2026.
- Robot proof: Meta's V-JEPA 2 hit 65-80% zero-shot manipulation success using just 62 hours of robot training data.
- Hybrid future: Most researchers expect LLMs and world models to converge, not one to replace the other outright.
- LeCun's conviction: He left Meta after a decade, raised $1 billion at a $3.5 billion valuation, to build what he believes LLMs never can.
The Turing Award Winner Who Called LLMs a Dead End
Yann LeCun is not some random contrarian. He won the Turing Award. He co-invented convolutional neural networks. He built Meta AI from the ground up.
And in November 2025, he walked out and said that scaling LLMs toward superintelligence is "complete bullshit." His words.
LeCun's argument is specific: training more on synthetic data, hiring more humans for post-training, tweaking reinforcement learning tricks, "it's just never going to work." Then he raised $1.03 billion at a $3.5 billion valuation for AMI Labs, announced March 2026, to build what he believes actually will work: world models.
What Is a World Model?
A world model is an AI system that builds an internal simulation of how reality works, learning physics, causality, object permanence, and cause-and-effect from video and sensor data rather than text.
Where an LLM predicts the next word in a sequence, a world model predicts the next state of the physical environment. The core question it answers: if I am in state X and take action A, what state comes next?
This matters enormously for anything that needs to interact with reality. A robot arm, an autonomous car, a surgical tool. Statistical knowledge that glass and "break" appear together in text is not the same as actually understanding why it happens.
The Benchmark That Exposes the Gap
Here is the cleanest evidence that LLMs do not understand physics. Meta released a benchmark called IntPhys 2 alongside V-JEPA 2 in June 2025. It tests whether a model can identify physics violations in video: an object teleporting, water flowing upward, a ball passing through a wall.
Humans score 85-95% accuracy. Existing AI models, including the best language models, score at or near chance level.
That is not a slight edge case. That is a fundamental capability gap. LeCun put it plainly: "An LLM doesn't understand that if you push a glass off a table, it will break. It only knows that the words 'glass' and 'break' often appear together in that context."
LLMs can describe a plan for picking up a fragile object. A world model can simulate what happens when the robot grips it too hard.
Where World Models Are Already Winning
Robotics. Meta's V-JEPA 2 (1.2 billion parameters, open-source, June 2025) achieved 65-80% success rates for zero-shot pick-and-place operations with new objects in unseen environments. The model was fine-tuned on just 62 hours of robot demonstration data. That data efficiency does not exist in LLM-only approaches, and it benchmarks 30x faster than NVIDIA Cosmos on certain robotic planning tasks.
Autonomous driving. Wayve, the UK-based autonomous vehicle company backed by Microsoft, NVIDIA, and Uber, reached an $8.6 billion valuation in February 2026 with $1.5 billion in total capital raised. Its GAIA-3 world model (15 billion parameters, December 2025) generates physics-grounded driving scenarios for safety testing across rare and dangerous edge cases. Wayve's AI Driver completed the AI-500 Roadshow in 2025: zero-shot across 500-plus cities in Europe, North America, and Japan, with no city-specific training.
3D environment generation. Google DeepMind's Genie 3, released August 2025, generates navigable 3D environments from text prompts at 24 frames per second and 720p resolution. Fei-Fei Li's World Labs, which raised $1 billion in February 2026 (total funding: $1.23 billion), launched Marble commercially that same month. Marble takes text, photos, or video clips and outputs persistent, downloadable 3D environments with real physics.
LLMs vs. World Models: The Investor's Breakdown
| LLMs | World Models | |
|---|---|---|
| Predicts | Next token in text | Next state of a physical environment |
| Training data | Text corpora | Video, sensors, 3D data |
| Physics understanding | Statistical patterns only | Causal and dynamic modeling |
| Planning | Describes plans in words | Simulates futures before acting |
| Robotics | Not natively capable | Core application |
| Inference compute | 1-8 GPUs | 8-32 GPUs |
| Maturity | Production at scale | Early commercial stage |
How to Evaluate Startups in This Wave
The same discipline that applies to any AI investment wave applies here. AMI Labs co-founder Alexandre LeBrun made the point candidly: "My prediction is that 'world models' will be the next buzzword. In six months, every company will call itself a world model to raise funding."
That is the right framing for due diligence. As covered in how to evaluate AI startups before writing the check, the questions that matter are whether the company owns proprietary data, real research depth, and distribution that BigTech cannot replicate easily. Google DeepMind, Meta, and NVIDIA are all building world models in-house. The startups in this space need a genuine moat.
Unicorn Screener is a data-driven scoring tool that grades startups across founder quality, market size, traction velocity, and competitive position. The public leaderboard tracks the highest-scoring AI and physical AI candidates across every stage. It is a faster way to separate the durable bets from the buzzword plays than reading a hundred pitch decks. The same framework that surfaces strong AI agent startups before they price their next round applies directly to the world model cohort.
The Honest Take: Hybrid, Not Replacement
The strongest researchers are not claiming world models kill LLMs. They are saying world models fill the capability gap LLMs leave open.
Google DeepMind wrote it directly in the Genie 3 announcement: "World models are a key stepping stone on the path to AGI, since they make it possible to train AI agents in an unlimited curriculum of rich simulation environments."
The likely winning architecture combines both. An LLM handles language, reasoning, and instruction. A world model handles physical planning, simulation, and embodied perception. The AV stack already looks like this. The robotics stack is heading there fast.
For investors, the practical question is not which paradigm wins. It is whether the companies raising record seed rounds in 2026 have the research depth to compete with the internal teams at Google, Meta, and NVIDIA, all of which are building these systems right now, at scale, for free.
LeCun left Meta because he believed he could build it faster outside. He raised $1 billion to prove it. The benchmark data backs the premise. The commercial payoff may still be years away.
The glass breaks when it hits the floor. LLMs only know the words appear together. That gap is the entire thesis.
Want to screen startups like a top-tier VC? Score any startup for free with our research-backed evaluation model.