News
Google DeepMind Video Model Leaders: This Is the Missing Link to AGI

Leaders of Google DeepMind's generative video models say video, not just language, is the missing piece needed to build human-level AGI.
Contents
Dumitru Erhan, head of video model work at Google DeepMind, and Shane Gu, who leads Gemini RL and Omni Thinking, said it plainly: language models alone are not enough to build artificial general intelligence. In their view, the missing piece is video models, capable of learning the physics and causality of the world in a way text never can.
What they said
Erhan compared the development trajectory of video models to what previously happened with language models. In his view, video models will follow a similar maturation path: they will steadily reduce hallucinations and, over time, become a reliable world model that machine reasoning can build on.
A video model is a complementary foundation model. It will follow a similar path... it will get much better at reducing hallucinations and become a very reliable world model - Dumitru Erhan, Google DeepMind
Gu went further, calling the video model a foundation model that is absolutely required if the goal is AGI matching human cognitive abilities. In his framing, language alone as a representation of the world is insufficient, because it fails to capture the physical, spatial and temporal continuity in which humans, and machines operating in real environments, actually function.
A world model instead of pure image generation
DeepMind's argument fits into a broader debate that has been running across the industry for months. The company's researchers argue that today's video generators already contain a kind of universal world model that the computer vision field has been searching for for years - instead of merely generating frames, they learn the rules of physics, motion and causality from billions of seconds of footage.
World models carry practical significance beyond filmmaking. They make it possible to train AI agents in a virtually unlimited simulated environment, without the need to collect costly real-world data. That is exactly the direction in which DeepMind is already developing its Genie family of models, which build interactive, real-time generated 3D environments.
Competition and other voices
Not everyone in the industry shares this view. Yann LeCun, Meta AI's chief scientist, has long argued that generative video models predicting pixels are a dead end on the path to true machine intelligence, and promotes an alternative approach based on the JEPA architecture, which learns abstract representations instead of predicting the next frame.
This dispute is not purely academic. The answer to whether video really is the missing link determines where the biggest AI labs will direct the next billions of dollars in infrastructure and talent investment. Google DeepMind, OpenAI with its Sora line, and Chinese competitors are already racing to build increasingly sophisticated video generators, betting that these, rather than yet another large language model, will deliver the breakthrough in machine reasoning about the physical world.
What this means in practice
Erhan gave a concrete example of an application that goes beyond entertainment: translating a product manual from English into Romanian while preserving the original layout, fonts and visual elements. This shows that DeepMind's video and image models aim to reach far beyond generating social media clips - they are meant to understand the layout, context and meaning of visual content.
For companies building products on top of Google's models, this means that future iterations of Gemini Omni and Veo will likely offer increasingly broad developer APIs, combining text, image, audio and video in a single generation pipeline. The Gemini Omni Flash API mentioned by the speakers is a first step in that direction and has already reached some developers.
The debate over video's role on the path to AGI will continue alongside intense market competition. Whether Erhan and Gu's position turns out to be an accurate forecast, or rather a justification for DeepMind's strategic investment in this particular branch of models, will only become clear once we see how quickly video-based world models actually start reducing hallucinations and supporting reasoning in real-world applications.
