Wink Pings

Veo 3's Visual Reasoning: A GPT-3 Moment or Old Wine in a New Bottle?

DeepMind's Veo 3 paper has sparked heated discussions, with some calling it the GPT-3 moment for visual reasoning, while others question its actual capabilities. The article explores the emergent properties of visual reasoning, the Chain-of-Frames concept, and the model's limitations in real-world scenarios.

Shortly after DeepMind released the Veo 3 paper, some were quick to declare it the GPT-3 moment for visual reasoning. The paper claims that through massive video training, Veo 3 demonstrates zero-shot learning capabilities, enabling it to solve tasks not explicitly included in the training data.

![A document titled "Video Models are Zero-Shot Learners and Reasoners 2024-05-25" with text about DeepMind’s Veo 3. A bar graph displays Veo 3’s success rates across perception, modeling, manipulation, and reasoning tasks. The graph uses colored bars for each category.](https://wink.run/image?url=https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FG13sw_FbgAAgIY8%3Fformat%3Djpg%26name%3Dlarge)

However, a closer look reveals several issues. First, this is not the original Veo 3 paper but an analysis report on its capabilities. Second, the so-called "Chain-of-Frames" concept is essentially a rebranding of the language model's "Chain of Thought" applied to video models. The comparison image is telling—the same person's two expressions, paired with old and new terminology, are packaged as a breakthrough.

![This is an image composed of two frames. The upper left shows a man in a red down jacket covering his ears with his hands, looking confused or annoyed. The upper right reads "Chain of Thought." The lower left shows the same man smiling with his eyes closed, giving a thumbs-up with his right hand and a phone gesture with his left, against a yellow wall background. The lower right reads "Chain of Frames."](https://wink.run/image?url=https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FG1rc54IXsAAoU3R%3Fformat%3Djpg%26name%3Dlarge)

Tesla engineers put it bluntly: the real test lies in handling messy real-world videos, not clean lab data. How does Veo 3 perform in continuous 360-degree object rotation tests? The paper doesn't mention it. Meta's UniBench research points out that such visual reasoning capabilities are not stable.

![Black text on a dark background detailing a comparison between Veo 3 and UniBench, discussing visual reasoning and machine learning performance. The text includes numbered points and references to research papers and models.](https://wink.run/image?url=https%3A%2F%2Fpbs.twimg.com%2Fmedia%2FG142QDlWIAAQ-v4%3Fformat%3Dpng%26name%3Dlarge)

The most interesting observation comes from the comments: while everyone is busy discussing "emergent" properties, no one is asking whether these capabilities truly go beyond pattern matching. Veo 3 might pass certain tests, but it's still far from genuine physical understanding and causal reasoning.

Autonomous driving requires stable and reliable visual reasoning, not pretty curves in the lab. DeepMind's fireworks this time are dazzling, but to illuminate the complexities of the real world, sparks alone aren't enough.

发布时间: 2025-09-28 01:37