OpenGEN-1: Reconstructing GEN-1 from First Principles
This article reconstructs the GEN-1 model from first principles, using evidence from Generalist AI’s releases, demonstrations, and blog posts, together with Pete Florence and Andy Zeng’s writing and interviews.
The central hypothesis is:
GEN-1 is a physical foundation model trained primarily on streams of physical interaction, rather than a VLA adapted from a VLM.
Pete Florence and the Generalist team have emphasized this distinction explicitly: GEN-1 is “not a fine-tuned vision-language model with robot actions bolted on,” and it is “not just a world model” [1]. They describe it as a native foundation model for physical interaction [1].
That claim matters. It suggests that the model is not primarily learning language, object semantics, or passive video prediction. It is learning the structure of physical interaction: perception, action, contact, failure, recovery, speed, and consequence.
From Semantic Common Sense to Physical Common Sense
LLMs and VLMs trained on internet-scale data can absorb enormous amounts of semantic knowledge. They know what an apple is, what color it usually has, where it commonly appears, and what people tend to do with it.
But they lack something deeper: physical common sense [2].
If I push an apple to the left, it may roll left. If a box is in front of it, the box may stop it. If the apple starts slipping from my grasp, I tighten my fingers. These are not merely visual facts. They come from actions and their consequences.
Andy Zeng’s article makes this point directly: if text at scale produces semantic common sense, then physical interaction at scale may produce physical common sense—but only if the data preserves the closed loop [2].
There are two layers:
- Forward physical common sense: If I take this action, what will happen next?
- Inverse / reflexive physical common sense: Given what is happening now, what should I do?
The former resembles a dynamics model or world model. The latter resembles a policy, an inverse-dynamics model, or a reflex. Humans perform both continuously and together. We do not first build a perfect world model, then plan, then act. We perceive and act within the same loop.
A Streaming Physical-Token Architecture
When a robot interacts with the physical world, it receives a stream of sensory input and emits a stream of actions. The natural architecture is therefore not a static image-to-action network, but a streaming sequence model.
One plausible architecture is a block-causal transformer over physical tokens, where each token may represent:
- a context / task token;
- an image token;
- a proprioception token;
- a force / tactile token;
- an action-condition token; or
- a state-query or action-query token.
During inference, the model continuously appends observation chunks and periodically decodes action chunks.
This is highly consistent with Generalist’s own description of GEN-0. The team describes Harmonic Reasoning as training a model to “think and act at the same time” through the interaction of asynchronous streams of sensing and acting tokens [3].
The KV cache of the processed stream acts as short-term memory. Longer-term memory may be managed by the harness through keyframes or compressed state summaries. Generalist has stated that GEN-1 needs “a new form of paged attention” for real-time inference [5]. In transformer serving, PagedAttention is fundamentally a method for managing KV-cache memory [6].
This also matches the shell-game demonstration. The task requires visual memory: with only a wrist camera, an object may leave the field of view, become occluded, or move behind the gripper. A policy that sees only the current frame is insufficient. The model needs short-term working memory over recent observations and states [4].
Action decoding could be implemented in several ways. The simplest is a direct action-chunk head similar to StarVLA-alpha [17]. A stronger option is an OpenPI-style flow action expert, where future actions are represented as continuous noisy action tokens and the model predicts a vector field from noise toward clean actions [7]. Another plausible option is a GR00T-like diffusion-transformer action module, where a fast action model generates smooth motor commands conditioned on visual, language, and proprioceptive context [8].
Training Objective: Action Prediction + Latent Dynamics
A standard VLA objective is an action-prediction loss: the difference between a predicted action and the ground-truth action, or the velocity-prediction objective used in flow matching.
This teaches the inverse problem:
Given the current history, what should I do next?
But action prediction alone may be insufficient for learning physical common sense. The model should also learn the forward problem:
Given the current history and an action, what physical state will appear next?
This is where the world-model idea becomes useful. Full RGB reconstruction may be too expensive and may overemphasize irrelevant pixel detail. The model does not need to reconstruct every background texture; it needs to predict the physical state that matters for action.
A better objective is latent future prediction: align a predicted future latent state with the encoded future state.
This is the technical meaning of “VLA + world model + beyond.” Pete has said that Generalist spent more than a year combining ideas from “VLAs, world models, and more,” because the more capabilities a single model combines, the harder it becomes to classify [1].
For policy inference, previously predicted actions are not strictly necessary if the model has a rich history of images and proprioception; the state trajectory already contains most of their effects. For world-model training, however, clean action-condition tokens are useful because they teach causal dynamics.
If both the encoder and predictor can move freely, latent reconstruction may collapse. The target latent should therefore be anchored:
V-JEPA-style methods use latent-space prediction, a stop-gradient target encoder, masking, and predictor asymmetry to avoid collapse [9]. For a physical foundation model, the action loss is itself an anti-collapse anchor, so the formulation can be simpler.
A compact objective is:
This combines inverse dynamics, latent forward dynamics, proprioception prediction, and contact / force prediction.
Physical-Interaction Data
Internet text and video can provide semantic priors, but they do not contain a complete physical loop. Most internet video lacks precise actions, proprioception, force, grip state, tactile feedback, and recovery trajectories.
Generalist’s claim is that large-scale physical-interaction collection can break the data bottleneck. GEN-1 was trained on more than 500,000 hours of high-fidelity physical-interaction data [5]. More importantly, its foundation model was not trained on robot data. Instead, the data came from low-cost wearable devices used by humans performing millions of activities [5].
This matches Andy’s argument about data. Teleoperation often breaks the sensorimotor loop through latency, limited tactile feedback, and unnatural interfaces. Generalist instead built handheld ergonomic devices that preserve force feedback, allowing operators to begin “reacting” rather than “thinking” [2].
Data diversity should not mean only semantic diversity across objects, scenes, and tasks. It should also include interaction diversity:
- too much force → deformation or impact;
- too little force → slippage; and
- a poor grasp → regrasping.
This is why tactile and force data become essential. Vision tells the robot where the object is. Proprioception tells it where the robot is. Touch and force tell it what physical interaction is actually occurring.
Generalist’s blog states that GEN-1 was trained on more than 500,000 hours of interaction data [5]. If the data was recorded at 10 Hz:
If each timestep is compressed into roughly 64–256 visual tokens, the total is approximately 1.15T–4.61T tokens. This does not yet include proprioception, action, tactile, force, or task tokens. The dataset is therefore already at trillion-token scale.
For comparison with common LLMs and VLMs:
| Model family | Model size | Pretraining tokens | Notes |
|---|---|---|---|
| Llama 2 | 7B / 13B / 70B | 2.0T text tokens | The three reported Llama 2 sizes use the same token count [10]. |
| Llama 3 | 8B / 70B | 15T+ text tokens | Meta reports 15T+ pretraining tokens for both Llama 3 sizes [11]. |
| Qwen2-VL | 2B / 8B / 72B | ~1.4T multimodal tokens | Open-Qwen2VL reports approximately 1.4T multimodal pretraining tokens for Qwen2-VL [12, 13]. |
This scale is sufficient to train a physical foundation model in the 10B-parameter range, especially because the tokens are unusually dense in causal physical information.
It is also consistent with GEN-0’s scaling observation. Generalist reports a phase transition at roughly 7B parameters, with 7B+ models better able to internalize large-scale robot-pretraining data, and notes that GEN-0 had already scaled beyond 10B parameters [3].
The key is not just the token count. One trillion web tokens and one trillion physical-interaction tokens are not equivalent. Web tokens are dense in semantic information. Physical-interaction tokens are dense in contact, action, failure, correction, memory, and consequence.
100 Hz Inference Requires a System
Generalist’s early research preview states that the robot maps pixels and other sensor data to actions at 100 Hz, and that the full hardware-software stack enables reactive, fluid, and precise control [16].
For a 10B model in fp16 / bf16, the weights alone occupy roughly 20 GB. At batch size 1, next-token or action-query decoding is usually constrained by memory bandwidth. Ignoring other overheads, an approximate lower bound is:
| GPU | Memory bandwidth | fp16 10B weight-read lower bound | More realistic optimized range |
|---|---|---|---|
| RTX 5090 | ~1.79 TB/s | ~11 ms | ~11–25 ms |
| RTX PRO 6000 Blackwell | ~1.79 TB/s | ~11 ms | ~11–25 ms |
| L40S | 864 GB/s | ~23 ms | ~25–60 ms |
These are only weight-read lower bounds. Real inference must also pay for KV-cache reads, vision encoding, action-expert steps, kernel overhead, scheduling, robot I/O, and safety / control logic.
A plausible system design is therefore:
Large model:
- runs at a lower policy frequency;
- maintains memory through a KV cache / paged attention;
- decodes action chunks; and
- predicts or tracks a latent physical state.
Low-level harness:
- executes at 100 Hz;
- blends action chunks;
- enforces safety constraints;
- handles impedance, force, and gripper control;
- reacts to tactile and proprioceptive events; and
- smooths and stabilizes behavior between model calls.
This explains why Generalist says that GEN-1 is more accurately described as a system, not merely a model [5].
Conclusion
The most important hidden insight is:
GEN-1 is not merely a better robot policy. It is an attempt to build a pretrained substrate for physical interaction.
This contrasts with other VLAs, which are commonly initialized from pretrained VLMs and co-trained with large amounts of web data to learn semantic common sense. GEN-1 appears to target the missing layer: physical common sense learned from dense streams of real interaction.
Generalist’s demonstrations mainly showcase generalization in physical common sense: contact-rich manipulation, occlusion, recovery, and fluidity. In contrast, many π-family or VLM-based robot models emphasize semantic or task-level generalization: following instructions, handling unseen objects, and transferring across high-level task descriptions.
References
[1] Pete Florence and Generalist AI Team, “Going Beyond World Models & VLAs,” Generalist AI Blog, Apr. 7, 2026.
[2] Andy Zeng, “The Dark Matter of Robotics: Physical Commonsense,” Generalist AI Blog, Jan. 29, 2026.
[3] Generalist AI Team, “GEN-0 / Embodied Foundation Models That Scale with Physical Interaction,” Generalist AI Blog, Nov. 4, 2025.
[4] Generalist AI, “GEN-1 Plays the Shell Game,” LinkedIn post, Apr. 2026.
[5] Generalist AI Team, “GEN-1: Scaling Embodied Foundation Models to Mastery,” Generalist AI Blog, Apr. 2, 2026.
[6] Woosuk Kwon et al., “Efficient Memory Management for Large Language Model Serving with PagedAttention,” arXiv, 2023.
[7] Physical Intelligence / Black et al., “π₀: A Vision-Language-Action Flow Model for General Robot Control,” arXiv, 2024.
[8] NVIDIA et al., “GR00T N1: An Open Foundation Model for Generalist Humanoid Robots,” arXiv, 2025.
[9] Meta AI, “V-JEPA: Revisiting Feature Prediction for Learning Visual Representations from Video,” ICLR / OpenReview, 2024.
[10] Meta, “Llama 2 Model Card,” Hugging Face, 2023.
[11] Meta, “Meta Llama 3 8B Model Card,” Hugging Face, 2024.
[12] Wang et al., “Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources,” arXiv, 2025.
[13] Peng Wang et al., “Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution,” arXiv, 2024.
[14] Shuai Bai et al., “Qwen2.5-VL Technical Report,” arXiv, 2025.
[15] Qwen Team, “Qwen2.5 Technical Report,” arXiv, 2024.
[16] Generalist AI Team, “Research Preview,” Generalist AI Blog, Jun. 17, 2025.
[17] Ye et al., “StarVLA-α: Reducing Complexity in Vision-Language-Action Systems,” arXiv, 2026.