Robot Policies Need the Right Memory, Not an 8K Context Window
RoboTTT, WAM-TTT, RoboSSM, and the boundary between context scaling, meta-learning, and robot memory
RoboTTT extends a robot policy's context to 8K timesteps and reports that task completion rises from 43.9% at 1K to 71.5% at 8K. WAM-TTT lets a World Action Model watch a few human videos at test time, then changes the robot's behavior through fast weights. RoboSSM uses a State-Space Model to extrapolate from 2 demonstrations during training to 32 at test time.[1][2][3]
Put these papers side by side and an attractive conclusion appears: context length will become the next scaling axis for robot foundation models.
My view is more conservative, and more specific. What these papers really establish is not that longer is always better. They show that robot policies are beginning to acquire a state that is computationally manageable and can be updated online. The past can be compressed into an SSM state or written into TTT fast weights, giving the policy a mechanism for belief updates, in-context adaptation, and even extracting a task specification from human video.
Length is only the carrier. The ceiling still depends on three things: whether the context contains new task-relevant information, whether the meta-training distribution covers the mapping from that information to action, and whether the memory representation matches the timescale of the information.
Transformers store, SSMs compress, TTT learns how to write
Start with the mechanisms.
A Transformer takes the most direct route: it stores the key/value representation of every past token and uses attention to look back when needed. The advantage is precise, content-addressable access to history. The disadvantage is equally clear: full-attention training and prefill grow roughly quadratically with context, and even with a KV cache, both memory and per-step reading cost grow with the retained history. At 30 Hz with multiple cameras, a robot stream quickly becomes an expensive video archive.
An SSM does not preserve the entire history. It recursively updates a fixed-size state:
s_t = update(s_{t-1}, x_t)
RoboSSM uses Longhorn to formulate this update as online regression. Each new observation is written into the state while a forgetting factor determines how much of the past survives. Computation over the prompt is linear in length, the rollout state remains fixed in size, and deployment requires no test-time parameter update. The price is lossy compression: once a detail fails to survive in the state, it cannot be retrieved precisely in the way attention can retrieve an old token.
TTT goes one step further. Instead of representing memory as a vector, it represents memory as the parameters W of a small network. Every incoming token triggers a gradient update on W under a self-supervised loss; at read time, the current query calls the fast-weight network. Put differently, an SSM learns an update function, while TTT turns each update into a small learning step.
This helps explain why RoboTTT can outperform a linear recurrent state on long, repetitive visual streams: a nonlinear MLP fast model is a more expressive compressor. But "inference latency does not grow with context length" does not mean the method is free. RoboTTT inserts a TTT layer into all 16 DiT blocks, growing the action head from 538M to 690M parameters, and it performs a gradient-based update at every timestep.[1]
The training distinction also deserves precision. Both TTT papers backpropagate through the inner fast-weight update in an outer loop—gradients through gradients. RoboTTT additionally performs truncated BPTT over long trajectories: gradients stop at segment boundaries while fast weights continue into the next segment. WAM-TTT differentiates through one inner SGD update and then freezes the adapted fast weights during rollout. Strictly speaking, it is not long-sequence TBPTT.[1][2]
If engineering simplicity is the priority, SSM training and deployment are usually cleaner. TTT pays for a richer and more flexible write rule. Both are better suited to streaming control than handing the entire raw history to a Transformer, but both obtain long context by trading full history for fixed-capacity compression. They do not provide unlimited memory for free.
| Method | Where history lives | Test-time adaptation | Main strength | Main cost |
|---|---|---|---|---|
| Transformer | Explicit KV / token history | Attention reads history directly | Precise retrieval and parallel training | Context cost and memory grow with length |
| RoboSSM | Fixed-size SSM state | Updates latent state only; no parameter update | Linear computation and simple deployment | Lossy compression and weaker selective recall |
| RoboTTT | Fast-weight MLP | Gradient update at every step | More expressive write rule | Meta-gradients, TBPTT, stability, and constant overhead |
| WAM-TTT | Video-side fast weights | One update from human video before rollout | Human video directly steers the policy | Depends on a paired meta-training distribution |
The first value of long context is really meta-learning
All three papers can be understood through the same lens: meta-learning.
RoboSSM trains on "several demonstrations + a query rollout from the same task," so it learns to infer the task from the demonstrations. RoboTTT treats human video or failed trajectories as context and correct robot actions as targets, so it learns to update fast weights from context. WAM-TTT makes the structure explicit: the inner loop rewrites memory using human video, and the outer loop checks against the paired robot trajectory whether that memory actually changes the action in the right way.
In-context learning therefore does not mean that a deployed model creates a new skill from nowhere. A more accurate description is that meta-training produces a temporary learner, and the new prompt supplies fresh evidence to that learner.
This is also where robot policies differ most from LLMs. The pretraining distribution of an LLM covers an enormous range of tasks, expressions, and reasoning patterns, so a new prompt often remains within an interpolatable region. Robot data usually spans only a handful of embodiments, scenes, task families, and near-expert trajectories. As the horizon grows, the space of possible observation histories grows almost exponentially. The moment a policy rollout leaves the expert path, it enters histories that training data may never have covered.
BPP offers direct evidence of this failure. Even when the encoder is forced to predict ground-truth history state, validation accuracy can look strong. But under the policy's own rollout distribution, state accuracy falls from 86% to 18%, and success falls from 56% to 19%. The problem is not only that the architecture lacks capacity; expert demonstrations do not cover deployment histories.[4]
Context length can therefore expand the amount of information available to the learner, but it cannot expand the meta-training distribution itself. The RoboSSM authors explicitly acknowledge that broader generalization to new tasks requires a larger and more diverse training corpus. WAM-TTT states the same boundary from another direction: adaptation weakens as the deployment task moves away from the human–robot pairing distribution, and that boundary has not yet been characterized systematically.[2][3]
RoboSSM provides the cleanest evidence among the three for length extrapolation. One model is trained with only 2 demonstrations and continues to improve when the test prompt grows to 32, while the matched Transformer ICRT collapses beyond its training length. But the result currently holds only in LIBERO simulation. The two long-horizon tasks use partial credit: RoboSSM reaches 16.1% / 17.4%, while ICRT gets 0. On ordinary LIBERO-Object, ICRT actually beats RoboSSM (66.3% vs. 57.9%). The evidence supports the narrower claim that SSMs are more robust to prompt-length shift. It does not yet show that SSMs dominate Transformers across the board, or that real-robot in-context learning is solved.[3]
Human video is the most practical interface, but its generalization is not free
Of the ideas in these papers, the most practically valuable one is not 8K. It is turning human video into a test-time interface for a robot policy.
RoboTTT uses a simple construction: concatenate a human video and a robot trajectory recorded under the same configuration. The human segment updates fast weights without an action loss; the robot segment supplies action supervision. At test time, a human video from an unseen Circuit configuration becomes the task specification. RoboTTT completes 6 of 10 trials; GDN completes none.[1]
WAM-TTT turns the same idea into something closer to a reusable system. During meta-training, phase-aligned human–robot pairs align human keys/values with robot queries. At test time, action-free human video updates the fast weights once through video prediction and memory reconstruction; the WAM and the action expert then remain frozen.
The result is persuasive. Across nine real-robot tasks evaluated in new environments, WAM-TTT reaches 46.2% average progress, compared with 32.5% for frozen LDA and only 7.1% for WAM-ICL, which places the same human video directly in the context. "Treat the video as more tokens" and "compile the video into executable memory" are not the same operation.[2]
The most accurate abstraction, however, is not that the robot can imitate any new skill from any human video. The system has learned a compiler from human evidence to a robot-control residual. That compiler still has to be calibrated with paired data.
WAM-TTT uses 2,286 paired human–robot episodes across the same nine task families. The human videos for the New setting are filmed inside the real household environments used later for deployment. RoboTTT's one-shot imitation also stays within the Circuit task family: training covers 20 configurations and testing covers another 60. These papers demonstrate useful task and configuration steering, not open-world skill acquisition.[1][2]
WAM-TTT's data-ratio ablation makes the boundary even clearer. On three tasks, 100 robot + 100 human episodes produce 74.1% progress, while 200 robot + 0 human episodes produce 73.7%. But 10 robot + 190 human episodes reach only 51.4%. Human data can replace some expensive robot collection inside an aligned domain; it cannot replace action grounding.[2]
This remains a strong direction. Human video is cheap, natural, and rich in object choice, operation order, style, and goal configuration. But the first thing that must scale is not the context window. It is the breadth of the paired distribution: more task families, embodiments, camera geometries, contact modes, failure modes, and multiple executable robot strategies for the same human intent.
Is 8K a scaling axis, or just horizon matching?
RoboTTT's most striking result is its context-length curve. As context grows from 128 to 8K timesteps, average completion across three real-robot assembly tasks rises from roughly 0.31 to 0.715; the 1K result is 0.439. It looks like context scaling.[1]
The result is real and important. Elevating it into a general scaling axis is still premature.
First, much of the gain may also come from the training horizon finally matching the rollout horizon. RoboTTT runs at 30 Hz, so 1K timesteps cover only about 33 seconds. The three tasks average roughly one, two, and five minutes. The paper itself notes that below 1K, inference updates fast weights into regions never encountered during training, while the positional embedding is also forced to extrapolate. An 8K sequence spans about 4.5 minutes—long enough to begin covering the tasks chosen specifically for their long horizons.
In other words, moving from 128 to 8K does more than "give the model more memory." It trains recurrent update dynamics for something close to a full episode rather than for a few seconds. For fixed-size SSM state or fast weights, a longer context does not add state parameters. It primarily extends how long the state must survive and remain stable under repeated updates. What scales here looks more like state lifetime / meta-optimization horizon than Transformer-style context capacity.
The training recipe points in the same direction. Every downstream task is post-trained with only 1K context; 8K appears only in sequence pretraining. A matched GDN has the same fixed-size state but does not improve as context grows. Longer sequences are mainly teaching RoboTTT's gradient update to operate continuously for thousands of steps, not merely expanding an addressable memory window.
Second, this is not a strictly compute-matched sweep. Every pretraining run uses 30K optimizer steps, but global batch size is 64 for 4K and below, and 16 for 8K. The 8K run therefore sees roughly twice as many timesteps per optimizer step as the 1K run. At the same time, 4K sees more total timesteps than 8K, so data exposure alone cannot explain the full curve. Still, the experiment is not controlled as a clean scaling law.[1]
Third, the evaluation remains narrow: one YAM bimanual platform, three intentionally long assembly tasks, 10 or 20 rollouts, and a partial-completion rubric for the main result. RoboTTT is the only method to achieve any full success on the five-minute Gear Bot task, but that result is still only 2/10. This is enough to show that long training horizons matter for this regime. It is not enough to show that context length produces broadly predictable gains in the way parameters, data, or compute often do.
The more fundamental question is: how much new information is actually present in the long context?
If 8K frames mostly show the same gripper moving, increasing length adds redundancy and spurious correlations. Across four real-robot tasks in BPP, naive full history averages only 12.8% success—worse than the 14.4% from the current observation alone. Retaining only task-relevant keyframes reaches 53.6%. HALO finds the same non-monotonic pattern: top-8 retrieval reaches 52%, while increasing the read set to top-16 drops performance to 33%. More context is not the same as more information.[4][5]
RoboTTT's curve is therefore better read as a strong existence proof: with a good enough update rule, a robot policy can preserve useful memory for minutes using fixed state and approximately fixed per-step cost. It is not yet evidence that every robot context should be stretched to 8K and expected to keep improving.
What robots actually need is a memory hierarchy
None of this means that low-level policies do not need memory. Quite the opposite: robot control is inherently a POMDP.
Suppose a robot tries twice and fails to pick up a parcel. The current RGB frame may not reveal the cause, but the failure history provides evidence: the package may be slippery, the true friction coefficient may be lower than the prior, or the chosen contact surface may be wrong. The policy should update its belief and, within safety limits, increase normal force, adjust the grasp pose, or select another contact surface. This latent state does not exist in a single observation. Short-term memory is essential.
The same applies to whether a screw is actually tight, whether a drawer has already been searched, whether the previous failure came from grasping or planning, and which stage of an assembly sequence is active. These are exactly the cases where memory resolves state aliasing—and where RoboTTT's gains are most credible.
Information that must survive for minutes or across episodes, however, should not all remain inside low-level in-context state. A more efficient long-horizon architecture assigns different abstractions to different memories:
At the bottom is a second-scale belief state for contact, friction, occlusion, and action outcomes, continuously updated by an SSM, TTT, or RNN. The middle layer is event-level episodic memory: whether a subgoal is complete, where an object was last seen, and under what conditions a failure occurred, represented by keyframes, top-k retrieval, or an explicit memory bank. At the top is planner state: task graph, current subgoal, constraints, and cross-episode affordances, expressed through language or symbolic summaries.
Existing experiments support this decomposition. In Hi-VLA's systematic study, an optimized hierarchy reaches 67.08% on long-horizon tasks versus 25.30% for a flat VLA, and the study recommends re-engaging the high-level planner every 4–8 seconds. More interestingly, giving the planner the complete raw history from the current episode brings almost no benefit, while summaries of affordances learned across episodes do help.[6]
HALO highlights the opposite failure mode: not everything should be compressed into fixed-size state. When the task requires precise recall—"where was the object placed?" or "when was the stove turned on?"—lossy recurrent state may discard the detail. Explicit episodic memory with selective retrieval is a better fit. RoboMME's systematic comparison across 16 memory-dependent tasks likewise finds no universally superior representation: symbolic memory is stronger for counting and event reasoning, while perceptual memory is better for motion and timing.[5][7]
The more plausible direction is therefore not to choose one winner among TTT, SSM, and Transformers, but to compose them:
Use SSM/TTT for low-level belief, sparse attention to retrieve key episodes, and a high-level planner to maintain task state. Only information that can actually change the action deserves to be written into memory.
Conclusion
RoboTTT, WAM-TTT, and RoboSSM are important works. They move long-context robot policies beyond concatenating more frames and toward learning how to write history into state. They also make human video behave like a real prompt: an interface that can change robot behavior at deployment time.
I would not, however, make raw context length an independent scaling objective for robot policies. For any memory design, I would first ask three questions: what latent information is missing from the current observation? How long must that information survive, and how precisely must it be retrieved? Does the training data cover the mapping from that context to the correct action?
Without answers to those questions, 8K is just a longer recording.
With answers, what we need is often not 8K frames, but eight events that matter.
References
[1] RoboTTT: Context Scaling for Robot Policies, 2026.
[2] WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time, 2026.
[3] RoboSSM: Scalable In-context Imitation Learning via State-Space Models, CoRL 2025 / revised 2026.
[4] BPP: Long-Context Robot Imitation Learning by Focusing on Key History Frames, 2026.
[5] Memory Retrieval in Visuomotor Policies for Long-Horizon Robot Control (HALO), 2026.
[6] What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents, 2026.
[7] RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies, 2026.