VLA Is Not Dead—At Least Not Yet
The most important point in the DreamZero / WAM paper is not that “robots can generate video now.” It is that the paper decomposes a robot policy into two problems: first predict what will happen in the future, then infer the action from that future.
Equation 1 makes the structure explicit:
It first generates a future video and then generates actions on top of that video. In other words, the essence of WAM is an explicit forward dynamics model + IDM.
This is different from a conventional VLA that attaches an action head to a VLM and directly maps image / language → action. WAM first generates a future video; an inverse dynamics model then produces actions from that future state. The future video serves as a visual plan, and the IDM translates that plan into robot actions.
That distinction is crucial.
Why WAM Has a Data Advantage
VLA pretraining mainly comes from image-text / VLM data. It learns semantic common sense:
- What is this object?
- What does this sentence mean?
- Where is the target likely to be?
But what robots truly lack is something else: if I take this action, how will the physical world change?
That is where WAM has an advantage.
A video-generation backbone is already pretrained on large-scale video. Video naturally contains time, motion, contact, occlusion, deformation, and causal evolution. Internet video may not have robot-action labels, but it is closely aligned with what a robot ultimately needs to predict: a sequence of states of the physical world.
Robot-data fine-tuning therefore does not have to teach physics from scratch. Instead, it teaches a video model that has already seen extensive world dynamics:
- how this embodiment interacts with the world; and
- which motor commands correspond to a future visual trajectory.
By contrast, VLM pretraining and robot-action prediction are misaligned in a VLA. Image-text data is strong on semantics, but it is not directly aligned with the physical evolution of action sequences.
This also explains why WAM can make better use of diverse, non-repetitive robot data. For a VLA, the same messy robot trajectory may look like noisy state-action pairs. For WAM, every continuous video segment provides dense future-state supervision.
This is not a minor difference. It is a fundamentally different way of using data.
Why an IDM Is Easier to Learn Than a Direct VLA Policy
A direct VLA policy must learn two things at once:
- what should happen in the future; and
- how the robot should move to make it happen.
WAM separates them:
- the video model predicts the future state; and
- the IDM maps that future state to an action.
This turns action learning into something closer to an inverse-dynamics problem, rather than forcing an end-to-end policy fit.
The paper also gives a strong signal: many failures come primarily from video-prediction errors rather than action-extraction errors. In other words, once the future video is correct, the IDM can often recover the action reliably.
WAM may therefore have moved the bottleneck from “can the action head learn?” to “can the world model predict correctly?” That makes its strong performance on unseen tasks and cross-embodiment transfer quite plausible.
Why the Unseen-Task and Cross-Embodiment Results Make Sense
On an unseen task, if a VLA has never observed a similar motion in robot data, it can only guess from the semantic prior of its VLM. But a VLM has not been pretrained on action sequences at scale; it does not know how physical motions such as untying shoelaces, ironing clothes, or shaking hands should unfold.
WAM is different.
It can first use its video-generation backbone to produce a future visual plan. That backbone may already have seen similar motion patterns in large-scale video, after which the IDM converts the visual future into robot actions.
There is nothing mysterious about WAM outperforming VLA on unseen tasks.
The same logic applies to cross-embodiment transfer. If a video-only demonstration from a human or another robot can update WAM’s representation of how the task should visually unfold, it may not need action labels. Action labels are embodiment-specific, while future visual states are closer to embodiment-agnostic.
This is WAM’s most valuable property: it turns large amounts of video without robot-action labels into training signals that can shape a policy.
That is difficult for a conventional VLA to do directly.
But the Cost Is High
WAM’s strength comes from explicit world modeling. Explicitly generating a future video, however, is not free.
The model must truly understand the additional information in how the world evolves, which consumes model capacity.
Table 4 makes this clear:
The 5B DreamZero reaches only 21% task progress, while the 14B model reaches 50%.
A small model struggles to carry explicit video generation and action alignment at the same time. It hallucinates the future, and the action fails with it.
WAM is therefore a large-model system that embeds a world model inside the policy. Its engineering profile differs from many VLAs: a VLA can be cheaper, faster, and easier to deploy, while WAM moves much more capability into the video-generation backbone.
The Awkward Part Is Industrial Deployment
WAM looks excellent in zero-shot and low-shot settings. But if the goal is deployment in a factory, warehouse, restaurant, or home-service scenario, one unavoidable issue remains:
You still have to post-train.
Figure 10 is worth examining closely.
On shirt folding, after 33 hours of post-training data, DreamZero and the pretrained VLA are tied at 92.5%. On fruit packing, after 12 hours, DreamZero reaches 96% and clearly outperforms the VLA.
This creates an awkward tradeoff:
With very little data, WAM is stronger than VLA, but it may still not be reliable enough to deploy. As the amount of data grows, the gap between WAM and a strong VLA may narrow, while WAM retains substantially higher inference cost, latency, and system complexity.
The paper acknowledges that DreamZero needs an entire stack of system optimizations to reduce a 14B autoregressive video-diffusion model from 5.7 seconds to 150 ms, and it uses 2×GB200 GPUs to reach 7 Hz. A single GB200 costs more than RMB 500,000 and requires a direct-liquid-cooling (DLC) system. Meanwhile, current VLAs can run at 20 Hz+ on consumer GPUs.
This is not a small engineering gap. It is a deployment-cost gap.
If an industrial task can already support the collection of dozens of hours of high-quality data, a VLA may remain an exceptionally strong baseline because it is cheap, fast, and stable.
WAM’s advantages are more likely to appear in:
- long-tail tasks;
- generalization to unseen motions and low-repetition demonstrations;
- cross-embodiment transfer; and
- the use of video-only data.
It is not a direct replacement for VLA in every scenario.
My Take
WAM is a highly important direction. Its core contribution is not that “robots can generate videos,” but that it demonstrates:
A video-generation backbone can serve as a physical prior for a robot policy.
This direction is more closely aligned with how the physical world evolves than a conventional VLA, and it can make better use of video at scale. It changes robot learning from “directly fit the action” to “understand the future first, then infer the action.”
That is powerful. But it is too early to declare VLA dead.
VLA’s weakness is physical generalization; WAM’s weaknesses are cost, latency, capacity requirements, and deployment complexity.
A more realistic assessment is:
WAM will raise the ceiling on unseen tasks, low-shot learning, cross-embodiment transfer, and video-data scaling; but in many industrial settings, VLA will continue to thrive because it is cheap, fast, and easy to deploy.
WAM exposes the ceiling of VLA. It does not yet prove that VLA’s engineering value has disappeared.
VLA is not dead. But from now on, robot foundation models cannot tell their whole story through VLM pretraining alone.