Back to Blog

Robot Policies Need the Right Memory, Not an 8K Context Window

Robot policy 需要的不是 8K context,而是对的 memory

从 RoboTTT、WAM-TTT、RoboSSM 看 context scaling、meta-learning 与机器人记忆的边界

RoboTTT 把 robot policy 的 context 拉到 8K timesteps,并报告从 1K 到 8K,task completion score 从 43.9% 上升到 71.5%。WAM-TTT 让 World Action Model 在 test time 看几段 human video,就能用 fast weights 改变后续 robot behavior。RoboSSM 则用 State-Space Model,把训练时的 2 个 demonstrations 外推到测试时的 32 个。[1][2][3]

这三篇工作放在一起,很容易得到一个诱人的结论:context length 会成为 robot foundation model 的下一条 scaling axis。

我的判断更保守,也更具体:它们真正证明的,不是「越长越强」,而是 robot policy 终于开始拥有一种可计算、可在线更新的 state。过去可以被压进 SSM state,也可以被写进 TTT fast weights;这让 policy 有机会做 belief update、in-context adaptation,甚至从 human video 里提取 task specification。

但 length 只是载体。决定能力上限的,仍然是三件事:context 里有没有新增的 task-relevant information,meta-training distribution 有没有覆盖这种信息到 action 的映射,以及 memory 的表示方式是否匹配信息的时间尺度。

Transformer 保存,SSM 压缩,TTT 学会如何写入

先把三种机制讲清楚。

Transformer 的做法最直接:保存过去所有 token 的 key/value,需要时再通过 attention 回看。优点是可以精确、content-addressable 地访问历史;缺点也很明显——完整 attention 的训练与 prefill 随 context 近似二次增长,即使有 KV cache,memory 与单步读取成本仍然会随历史变长。对于 30Hz、多相机的 robot stream,这很快就会变成一个昂贵的录像仓库。

SSM 不保存整段历史,而是递归更新一个固定大小的 state:

s_t = update(s_{t-1}, x_t)

RoboSSM 使用 Longhorn,把这个 update 写成一种 online regression:新 observation 一边写入 state,一边由遗忘系数决定保留多少过去。整个 prompt 的计算对长度是线性的,rollout 时的 state 大小固定,也不需要 test-time parameter update。它的代价是压缩有损:一旦某个细节没有被 state 保留下来,之后就无法像 attention 那样精确找回。

TTT 再往前走一步:它不把 memory 做成一个向量,而是做成一个小网络的参数 W。每来一个 token,就用 self-supervised loss 对 W 做一次 gradient update;需要读取时,再用当前 query 去调用这个 fast-weight network。换句话说,SSM 学的是 update function,TTT 则让 update 本身成为一次小型 learning step。

Transformer、SSM 与 TTT 的 memory geometry
Transformer、SSM 与 TTT 的 memory geometry

这解释了为什么 RoboTTT 在长而重复的视觉流上可能比线性 recurrent state 更强:一个 nonlinear MLP fast model 是更有表达力的 compressor。但「inference latency 不随 context length 增长」不等于它没有成本。RoboTTT 给 16 个 DiT block 都加了 TTT layer,action head 从 538M 增到 690M parameters,而且每个 timestep 都要做 gradient-based update。[1]

训练上的差异也值得说准确。两篇 TTT 工作都需要在 outer loop 里对 inner fast-weight update 反传,也就是 gradients through gradients;RoboTTT 还需要沿长轨迹做 truncated BPTT,靠 segment boundary 截断梯度,但把 fast weights 继续传下去。WAM-TTT 则只对一次 inner SGD update 求 meta-gradient,随后在 rollout 中冻结更新后的 fast weights,严格说并不是长序列 TBPTT。[1][2]

所以如果只比较工程性,SSM 的训练与部署通常更干净;TTT 换来的是更强、也更灵活的写入规则。两者都比把全部 raw history 交给 Transformer 更适合 streaming control,但它们是在用固定容量的压缩换取长上下文,而不是免费获得无限记忆。

方法 历史放在哪里 Test time 如何适应 长处 主要代价
Transformer 显式 KV / token history attention 直接读取 精确检索、训练并行 长 context 成本与 memory 增长
RoboSSM 固定大小 SSM state 只更新 latent state,无参数更新 线性计算、部署简单 有损压缩,selective recall 较弱
RoboTTT fast-weight MLP 每步 gradient update 写入规则更有表达力 meta-gradient、TBPTT、稳定性与常数开销
WAM-TTT video-side fast weights rollout 前用 human video 更新一次 human video 可直接 steer policy 依赖 paired meta-training distribution

Long context 的第一种价值,其实是 meta-learning

这几篇论文都可以被统一地看成 meta-learning。

RoboSSM 在训练时看到「若干 demonstration + 同一 task 的 query rollout」,于是学会从 demonstration 推断 task。RoboTTT 把 human video 或失败轨迹当作 context,把正确 robot action 当作 target,于是学会从 context 更新 fast weights。WAM-TTT 更直接:inner loop 用 human video 改写 memory,outer loop 用 paired robot trajectory 检查这段 memory 是否真的能改变 robot action。

因此,所谓 in-context learning 并不是模型在部署时凭空学会了一个新技能。更准确地说,它是在 meta-training 过程中学到了一个 temporary learner;新 prompt 只是给这个 learner 提供新的 evidence。

这也是 LLM 与 robot policy 最大的差别。LLM 的 pretraining distribution 足够宽,大量 task、表达方式和推理 pattern 已经被覆盖,所以一个新 prompt 往往还能落在某个可插值的区域。Robot data 则通常只有少数 embodiment、scene、task family 和 near-expert trajectory。随着 horizon 增长,可能的 observation history 组合近似指数增长;policy rollout 一旦偏离 expert path,就会进入训练数据没有覆盖过的 history。

BPP 对这个问题给出了很直接的证据:即使强行让 encoder 预测 ground-truth history state,validation accuracy 可以很好,但到了 policy 自己的 rollout,state accuracy 从 86% 跌到 18%,success rate 从 56% 跌到 19%。问题不只是 architecture 不够强,而是 expert demonstrations 没有覆盖 deployment histories。[4]

所以 context length 只能扩大「可被利用的信息量」,不能扩大 meta-learning distribution 本身。RoboSSM 的作者也明确承认,更完整的新 task 泛化需要更广、更丰富的 training corpus;WAM-TTT 也把边界写得很清楚:deployment task 离 human–robot pairing distribution 越远,adaptation 越弱,而且这个边界还没有被系统刻画。[2][3]

RoboSSM 是三篇中对 length extrapolation 最干净的一组证据:同一个模型只用 2 个 demonstration 训练,测试 prompt 增加到 32 个时仍能改善,而 matched Transformer ICRT 在超过训练长度后快速崩溃。但这个结论目前只在 LIBERO simulation 上成立。它的两个 long-horizon tasks 使用 partial credit,RoboSSM 是 16.1% / 17.4%,ICRT 是 0;在普通 LIBERO-Object 上,ICRT 反而高于 RoboSSM(66.3% vs. 57.9%)。所以它证明的是 SSM 对 prompt-length shift 更稳,还不是 SSM 全面优于 Transformer,也不是 real-robot in-context learning 已经解决。[3]

Human video 是最实用的接口,但它不是免费的泛化

这三篇里,我认为最有实际价值的不是 8K,而是把 human video 变成 robot policy 的 test-time interface。

RoboTTT 的做法很简洁:把同一 configuration 的 human video 和 robot trajectory 接起来,human 部分只更新 fast weights、不计算 action loss,robot 部分再提供动作监督。测试时,一段 unseen Circuit configuration 的 human video 就是 task specification。它在 10 次测试中完成 6 次,GDN 是 0 次。[1]

WAM-TTT 把这个 idea 做得更像一个可复用系统。它在 meta-training 时用 phase-aligned human–robot pairs,让 human key/value 与 robot query 对齐;test time 只需要 action-free human video,通过 video prediction 与 memory reconstruction 更新一次 fast weights,之后 WAM 与 action expert 都保持冻结。

WAM-TTT:把 human video 写入 fast-weight memory
WAM-TTT:把 human video 写入 fast-weight memory

结果很有说服力:在 9 个 real-robot tasks 的新环境评测中,WAM-TTT 的平均 progress 是 46.2%,frozen LDA 是 32.5%,把同样 human video 直接塞进 context 的 WAM-ICL 只有 7.1%。这说明「把 video 当更多 token」和「把 video 变成可执行 memory」不是一回事。[2]

但这里最准确的抽象,不是 robot 学会了从任何 human video 模仿任何新 skill,而是系统学会了一个 human evidence → robot control residual 的 compiler。这个 compiler 仍然要靠 paired data 标定。

WAM-TTT 使用了 2,286 个 paired human–robot episodes,覆盖相同的 9 个 task families;New setting 的 human videos 还直接拍摄于之后部署的真实 household environment。RoboTTT 的 one-shot imitation 也发生在 Circuit 这一 task family 内:训练覆盖 20 个 configuration,测试另外 60 个。它们展示的是很有用的 task/configuration steering,而不是 open-world skill acquisition。[1][2]

WAM-TTT 的 data-ratio ablation 更能说明问题。在三个 task 上,100 条 robot + 100 条 human episodes 得到 74.1% progress,200 条 robot + 0 human 是 73.7%;但 10 条 robot + 190 条 human 只有 51.4%。Human data 可以在已对齐的 domain 内替代一部分昂贵的 robot collection,却不能替代 action grounding。[2]

这仍然是一个很好的方向。因为 human video 便宜、自然,而且能携带 object choice、操作顺序、style 和 goal configuration。但真正需要 scale 的,首先不是 context window,而是 paired distribution 的 breadth:更多 task family、embodiment、camera geometry、contact mode、failure mode,以及同一个 human intent 对应多种可执行 robot strategy。

8K 是 scaling axis,还是 horizon matching?

RoboTTT 最醒目的结果是下面这条曲线:context 从 128 增加到 8K timesteps,三项 real-robot assembly task 的平均 completion score 从约 0.31 上升到 0.715;1K 是 0.439。曲线看起来很像 context scaling。[1]

RoboTTT context scaling 曲线
RoboTTT context scaling 曲线

这个结果是真的,也很重要。但把它升级成一条通用 scaling axis,目前还太早。

第一,它很大程度上也可能是 training horizon 终于覆盖 rollout horizon。RoboTTT 以 30Hz 运行,1K timesteps 只有约 33 秒,而三个 task 的平均时长分别是 1、2、5 分钟。论文自己也指出,低于 1K 时,inference 会把 fast weights 更新到训练时从未见过的位置,positional embedding 也在外推。8K 约等于 4.5 分钟,恰好开始覆盖这些为长时程而选择的 task。

换句话说,128 → 8K 不只是「给模型更多 memory」,也是「把 recurrent update dynamics 从几秒训练到接近完整 episode」。对于 fixed-size SSM state 或 fast weights,context 变长并没有增加 state 的参数量;它主要增加的是 state 必须稳定存活、连续更新的时间。这里更像是在 scale state lifetime / meta-optimization horizon,而不是 Transformer 意义上的 context capacity。

这一点从训练 recipe 里也能看出来:所有 downstream tasks 最终都只用 1K context post-train,8K 只出现在 sequence pretraining;matched GDN 有同样的 fixed-size state,却没有随 context 变长而改善。更长的 sequence 主要是在教 RoboTTT 的 gradient update 如何连续运行几千次,而不是单纯扩大一个可寻址的 memory window。

第二,这不是严格的 compute-matched sweep。所有 pretraining run 都做 30K optimizer steps,但 4K 及以下的 global batch 是 64,8K 的 global batch 是 16。因此 8K 每个 optimizer step 看到的 timesteps 大约是 1K 的两倍;与此同时,4K 又比 8K 看到更多 timesteps,所以 compute/data exposure 不能单独解释整条曲线,但它确实没有被控制成一个干净的 scaling law。[1]

第三,评测范围仍然很窄:一个 YAM bimanual platform、三个特意选择的长时程 assembly tasks、20 次或 10 次 rollout,主结果使用 partial-completion rubric。RoboTTT 在五分钟 Gear Bot 上是唯一出现 full success 的方法,但也只有 2/10。它足以说明长训练 horizon 对这类任务有用,还不足以说明 context length 会像 parameters、data 或 compute 一样普遍地产生可预测收益。

更根本的问题是:long context 中有多少新增信息?

如果 8K frames 只是同一个夹爪在移动,增加长度主要是在增加冗余和 spurious correlations。BPP 在四个真实机器人任务上,naive full history 的平均 success 只有 12.8%,甚至低于 current observation 的 14.4%;只保留 task-relevant keyframes 后达到 53.6%。HALO 也发现,top-8 retrieval 是 52%,把读取量增加到 top-16 反而跌到 33%。More context 并不等于 more information。[4][5]

因此,RoboTTT 的曲线更适合被解读为一个很强的 existence proof:只要 update rule 足够好,robot policy 可以在固定 state 和近似固定单步成本下,让 memory 存活数分钟。 它还不是「把所有 robot context 拉到 8K 就会持续变强」的证据。

Robot 真正需要的是 memory hierarchy

这并不是说低层 policy 不需要 memory。恰恰相反,robot control 天然是 POMDP。

比如让机器人拿一个包裹。它连续两次都抓不起来,当前 RGB frame 未必能告诉它原因,但失败历史提供了新的 evidence:包装可能很滑,实际 friction coefficient 比先验更低,或者当前接触面与估计不同。policy 应该更新 belief,在安全约束内提高 normal force、调整 grasp pose,或者换一个接触面。这个 latent state 不存在于单帧 observation 里,短期 memory 是必须的。

类似的还有:螺丝是否真的拧紧、抽屉是否已经搜索过、刚才是 grasp failure 还是 planning failure、当前处于 assembly 的哪一个 stage。这些都是 memory 能解除 state aliasing 的地方,也是 RoboTTT 最可信的收益来源。

但分钟级甚至跨 episode 的信息,不应该全部继续塞进 low-level in-context state。长期任务更高效的结构,是让不同 memory 承担不同抽象层级:

Robot memory hierarchy
Robot memory hierarchy

最底层是秒级的 belief state:接触、摩擦、遮挡、动作结果,用 SSM/TTT/RNN 持续更新。中间是 event-level episodic memory:某个 subgoal 是否完成、物体最后出现在哪里、失败发生在什么条件下,用 keyframe、top-k retrieval 或显式 memory bank 保存。最上层是 planner state:task graph、subgoal、constraint、跨 episode affordance,用 language/symbolic summary 表达。

相关实验也支持这个拆分。Hi-VLA 的系统研究里,优化后的 hierarchy 在 long-horizon tasks 上是 67.08%,flat VLA 是 25.30%;它建议 high-level planner 每 4–8 秒重新介入一次。更有意思的是,把当前 episode 的完整 raw history 交给 planner 几乎没有收益,而跨 episode 总结出的 affordance 反而有帮助。[6]

另一面,HALO 也提醒我们不能把一切都压进 fixed-size state。需要精确回忆「物体放在哪里」「什么时候打开炉子」时,有损 recurrent state 可能丢掉细节,显式 episodic memory 加 selective retrieval 更合适。RoboMME 对 16 个 memory-dependent tasks 的系统比较也没有找到一个永远最好的表示:symbolic memory 擅长 counting 与 event reasoning,perceptual memory 更适合 motion 与 timing。[5][7]

所以未来更合理的方向不是在 TTT、SSM 和 Transformer 之间三选一,而是组合它们:

SSM/TTT 维护低层 belief,sparse attention 检索关键 episode,high-level planner 管理 task state;只有真正改变 action 的信息,才值得被写进 memory。

最后

RoboTTT、WAM-TTT 和 RoboSSM 都是重要的工作。它们把 long-context robot policy 从「把更多 frames 拼起来」推进到了「学习如何把历史写成 state」,并且让 human video 真正像 prompt 一样,成为可以在部署时改变 robot behavior 的接口。

但我不会把 raw context length 当成 robot policy 的独立 scaling goal。对于任何 memory 设计,我更愿意先问三个问题:当前 observation 缺失了什么 latent information?这条信息需要活多久、以什么形式被精确读取?训练数据是否覆盖了从这种 context 到正确 action 的映射?

如果这三个问题没有答案,8K 只是更长的录像。
如果它们有答案,很多时候我们需要的不是 8K,而是 8 个对的 event。


References

[1] RoboTTT: Context Scaling for Robot Policies, 2026.

[2] WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time, 2026.

[3] RoboSSM: Scalable In-context Imitation Learning via State-Space Models, CoRL 2025 / revised 2026.

[4] BPP: Long-Context Robot Imitation Learning by Focusing on Key History Frames, 2026.

[5] Memory Retrieval in Visuomotor Policies for Long-Horizon Robot Control (HALO), 2026.

[6] What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents, 2026.

[7] RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies, 2026.

Robot Policies Need the Right Memory, Not an 8K Context Window

RoboTTT, WAM-TTT, RoboSSM, and the boundary between context scaling, meta-learning, and robot memory

RoboTTT extends a robot policy's context to 8K timesteps and reports that task completion rises from 43.9% at 1K to 71.5% at 8K. WAM-TTT lets a World Action Model watch a few human videos at test time, then changes the robot's behavior through fast weights. RoboSSM uses a State-Space Model to extrapolate from 2 demonstrations during training to 32 at test time.[1][2][3]

Put these papers side by side and an attractive conclusion appears: context length will become the next scaling axis for robot foundation models.

My view is more conservative, and more specific. What these papers really establish is not that longer is always better. They show that robot policies are beginning to acquire a state that is computationally manageable and can be updated online. The past can be compressed into an SSM state or written into TTT fast weights, giving the policy a mechanism for belief updates, in-context adaptation, and even extracting a task specification from human video.

Length is only the carrier. The ceiling still depends on three things: whether the context contains new task-relevant information, whether the meta-training distribution covers the mapping from that information to action, and whether the memory representation matches the timescale of the information.

Transformers store, SSMs compress, TTT learns how to write

Start with the mechanisms.

A Transformer takes the most direct route: it stores the key/value representation of every past token and uses attention to look back when needed. The advantage is precise, content-addressable access to history. The disadvantage is equally clear: full-attention training and prefill grow roughly quadratically with context, and even with a KV cache, both memory and per-step reading cost grow with the retained history. At 30 Hz with multiple cameras, a robot stream quickly becomes an expensive video archive.

An SSM does not preserve the entire history. It recursively updates a fixed-size state:

s_t = update(s_{t-1}, x_t)

RoboSSM uses Longhorn to formulate this update as online regression. Each new observation is written into the state while a forgetting factor determines how much of the past survives. Computation over the prompt is linear in length, the rollout state remains fixed in size, and deployment requires no test-time parameter update. The price is lossy compression: once a detail fails to survive in the state, it cannot be retrieved precisely in the way attention can retrieve an old token.

TTT goes one step further. Instead of representing memory as a vector, it represents memory as the parameters W of a small network. Every incoming token triggers a gradient update on W under a self-supervised loss; at read time, the current query calls the fast-weight network. Put differently, an SSM learns an update function, while TTT turns each update into a small learning step.

Three memory geometries: Transformer, SSM, and TTT
Three memory geometries: Transformer, SSM, and TTT

This helps explain why RoboTTT can outperform a linear recurrent state on long, repetitive visual streams: a nonlinear MLP fast model is a more expressive compressor. But "inference latency does not grow with context length" does not mean the method is free. RoboTTT inserts a TTT layer into all 16 DiT blocks, growing the action head from 538M to 690M parameters, and it performs a gradient-based update at every timestep.[1]

The training distinction also deserves precision. Both TTT papers backpropagate through the inner fast-weight update in an outer loop—gradients through gradients. RoboTTT additionally performs truncated BPTT over long trajectories: gradients stop at segment boundaries while fast weights continue into the next segment. WAM-TTT differentiates through one inner SGD update and then freezes the adapted fast weights during rollout. Strictly speaking, it is not long-sequence TBPTT.[1][2]

If engineering simplicity is the priority, SSM training and deployment are usually cleaner. TTT pays for a richer and more flexible write rule. Both are better suited to streaming control than handing the entire raw history to a Transformer, but both obtain long context by trading full history for fixed-capacity compression. They do not provide unlimited memory for free.

Method Where history lives Test-time adaptation Main strength Main cost
Transformer Explicit KV / token history Attention reads history directly Precise retrieval and parallel training Context cost and memory grow with length
RoboSSM Fixed-size SSM state Updates latent state only; no parameter update Linear computation and simple deployment Lossy compression and weaker selective recall
RoboTTT Fast-weight MLP Gradient update at every step More expressive write rule Meta-gradients, TBPTT, stability, and constant overhead
WAM-TTT Video-side fast weights One update from human video before rollout Human video directly steers the policy Depends on a paired meta-training distribution

The first value of long context is really meta-learning

All three papers can be understood through the same lens: meta-learning.

RoboSSM trains on "several demonstrations + a query rollout from the same task," so it learns to infer the task from the demonstrations. RoboTTT treats human video or failed trajectories as context and correct robot actions as targets, so it learns to update fast weights from context. WAM-TTT makes the structure explicit: the inner loop rewrites memory using human video, and the outer loop checks against the paired robot trajectory whether that memory actually changes the action in the right way.

In-context learning therefore does not mean that a deployed model creates a new skill from nowhere. A more accurate description is that meta-training produces a temporary learner, and the new prompt supplies fresh evidence to that learner.

This is also where robot policies differ most from LLMs. The pretraining distribution of an LLM covers an enormous range of tasks, expressions, and reasoning patterns, so a new prompt often remains within an interpolatable region. Robot data usually spans only a handful of embodiments, scenes, task families, and near-expert trajectories. As the horizon grows, the space of possible observation histories grows almost exponentially. The moment a policy rollout leaves the expert path, it enters histories that training data may never have covered.

BPP offers direct evidence of this failure. Even when the encoder is forced to predict ground-truth history state, validation accuracy can look strong. But under the policy's own rollout distribution, state accuracy falls from 86% to 18%, and success falls from 56% to 19%. The problem is not only that the architecture lacks capacity; expert demonstrations do not cover deployment histories.[4]

Context length can therefore expand the amount of information available to the learner, but it cannot expand the meta-training distribution itself. The RoboSSM authors explicitly acknowledge that broader generalization to new tasks requires a larger and more diverse training corpus. WAM-TTT states the same boundary from another direction: adaptation weakens as the deployment task moves away from the human–robot pairing distribution, and that boundary has not yet been characterized systematically.[2][3]

RoboSSM provides the cleanest evidence among the three for length extrapolation. One model is trained with only 2 demonstrations and continues to improve when the test prompt grows to 32, while the matched Transformer ICRT collapses beyond its training length. But the result currently holds only in LIBERO simulation. The two long-horizon tasks use partial credit: RoboSSM reaches 16.1% / 17.4%, while ICRT gets 0. On ordinary LIBERO-Object, ICRT actually beats RoboSSM (66.3% vs. 57.9%). The evidence supports the narrower claim that SSMs are more robust to prompt-length shift. It does not yet show that SSMs dominate Transformers across the board, or that real-robot in-context learning is solved.[3]

Human video is the most practical interface, but its generalization is not free

Of the ideas in these papers, the most practically valuable one is not 8K. It is turning human video into a test-time interface for a robot policy.

RoboTTT uses a simple construction: concatenate a human video and a robot trajectory recorded under the same configuration. The human segment updates fast weights without an action loss; the robot segment supplies action supervision. At test time, a human video from an unseen Circuit configuration becomes the task specification. RoboTTT completes 6 of 10 trials; GDN completes none.[1]

WAM-TTT turns the same idea into something closer to a reusable system. During meta-training, phase-aligned human–robot pairs align human keys/values with robot queries. At test time, action-free human video updates the fast weights once through video prediction and memory reconstruction; the WAM and the action expert then remain frozen.

WAM-TTT writes human video into fast-weight memory
WAM-TTT writes human video into fast-weight memory

The result is persuasive. Across nine real-robot tasks evaluated in new environments, WAM-TTT reaches 46.2% average progress, compared with 32.5% for frozen LDA and only 7.1% for WAM-ICL, which places the same human video directly in the context. "Treat the video as more tokens" and "compile the video into executable memory" are not the same operation.[2]

The most accurate abstraction, however, is not that the robot can imitate any new skill from any human video. The system has learned a compiler from human evidence to a robot-control residual. That compiler still has to be calibrated with paired data.

WAM-TTT uses 2,286 paired human–robot episodes across the same nine task families. The human videos for the New setting are filmed inside the real household environments used later for deployment. RoboTTT's one-shot imitation also stays within the Circuit task family: training covers 20 configurations and testing covers another 60. These papers demonstrate useful task and configuration steering, not open-world skill acquisition.[1][2]

WAM-TTT's data-ratio ablation makes the boundary even clearer. On three tasks, 100 robot + 100 human episodes produce 74.1% progress, while 200 robot + 0 human episodes produce 73.7%. But 10 robot + 190 human episodes reach only 51.4%. Human data can replace some expensive robot collection inside an aligned domain; it cannot replace action grounding.[2]

This remains a strong direction. Human video is cheap, natural, and rich in object choice, operation order, style, and goal configuration. But the first thing that must scale is not the context window. It is the breadth of the paired distribution: more task families, embodiments, camera geometries, contact modes, failure modes, and multiple executable robot strategies for the same human intent.

Is 8K a scaling axis, or just horizon matching?

RoboTTT's most striking result is its context-length curve. As context grows from 128 to 8K timesteps, average completion across three real-robot assembly tasks rises from roughly 0.31 to 0.715; the 1K result is 0.439. It looks like context scaling.[1]

RoboTTT context-scaling curve
RoboTTT context-scaling curve

The result is real and important. Elevating it into a general scaling axis is still premature.

First, much of the gain may also come from the training horizon finally matching the rollout horizon. RoboTTT runs at 30 Hz, so 1K timesteps cover only about 33 seconds. The three tasks average roughly one, two, and five minutes. The paper itself notes that below 1K, inference updates fast weights into regions never encountered during training, while the positional embedding is also forced to extrapolate. An 8K sequence spans about 4.5 minutes—long enough to begin covering the tasks chosen specifically for their long horizons.

In other words, moving from 128 to 8K does more than "give the model more memory." It trains recurrent update dynamics for something close to a full episode rather than for a few seconds. For fixed-size SSM state or fast weights, a longer context does not add state parameters. It primarily extends how long the state must survive and remain stable under repeated updates. What scales here looks more like state lifetime / meta-optimization horizon than Transformer-style context capacity.

The training recipe points in the same direction. Every downstream task is post-trained with only 1K context; 8K appears only in sequence pretraining. A matched GDN has the same fixed-size state but does not improve as context grows. Longer sequences are mainly teaching RoboTTT's gradient update to operate continuously for thousands of steps, not merely expanding an addressable memory window.

Second, this is not a strictly compute-matched sweep. Every pretraining run uses 30K optimizer steps, but global batch size is 64 for 4K and below, and 16 for 8K. The 8K run therefore sees roughly twice as many timesteps per optimizer step as the 1K run. At the same time, 4K sees more total timesteps than 8K, so data exposure alone cannot explain the full curve. Still, the experiment is not controlled as a clean scaling law.[1]

Third, the evaluation remains narrow: one YAM bimanual platform, three intentionally long assembly tasks, 10 or 20 rollouts, and a partial-completion rubric for the main result. RoboTTT is the only method to achieve any full success on the five-minute Gear Bot task, but that result is still only 2/10. This is enough to show that long training horizons matter for this regime. It is not enough to show that context length produces broadly predictable gains in the way parameters, data, or compute often do.

The more fundamental question is: how much new information is actually present in the long context?

If 8K frames mostly show the same gripper moving, increasing length adds redundancy and spurious correlations. Across four real-robot tasks in BPP, naive full history averages only 12.8% success—worse than the 14.4% from the current observation alone. Retaining only task-relevant keyframes reaches 53.6%. HALO finds the same non-monotonic pattern: top-8 retrieval reaches 52%, while increasing the read set to top-16 drops performance to 33%. More context is not the same as more information.[4][5]

RoboTTT's curve is therefore better read as a strong existence proof: with a good enough update rule, a robot policy can preserve useful memory for minutes using fixed state and approximately fixed per-step cost. It is not yet evidence that every robot context should be stretched to 8K and expected to keep improving.

What robots actually need is a memory hierarchy

None of this means that low-level policies do not need memory. Quite the opposite: robot control is inherently a POMDP.

Suppose a robot tries twice and fails to pick up a parcel. The current RGB frame may not reveal the cause, but the failure history provides evidence: the package may be slippery, the true friction coefficient may be lower than the prior, or the chosen contact surface may be wrong. The policy should update its belief and, within safety limits, increase normal force, adjust the grasp pose, or select another contact surface. This latent state does not exist in a single observation. Short-term memory is essential.

The same applies to whether a screw is actually tight, whether a drawer has already been searched, whether the previous failure came from grasping or planning, and which stage of an assembly sequence is active. These are exactly the cases where memory resolves state aliasing—and where RoboTTT's gains are most credible.

Information that must survive for minutes or across episodes, however, should not all remain inside low-level in-context state. A more efficient long-horizon architecture assigns different abstractions to different memories:

A hierarchy for robot memory
A hierarchy for robot memory

At the bottom is a second-scale belief state for contact, friction, occlusion, and action outcomes, continuously updated by an SSM, TTT, or RNN. The middle layer is event-level episodic memory: whether a subgoal is complete, where an object was last seen, and under what conditions a failure occurred, represented by keyframes, top-k retrieval, or an explicit memory bank. At the top is planner state: task graph, current subgoal, constraints, and cross-episode affordances, expressed through language or symbolic summaries.

Existing experiments support this decomposition. In Hi-VLA's systematic study, an optimized hierarchy reaches 67.08% on long-horizon tasks versus 25.30% for a flat VLA, and the study recommends re-engaging the high-level planner every 4–8 seconds. More interestingly, giving the planner the complete raw history from the current episode brings almost no benefit, while summaries of affordances learned across episodes do help.[6]

HALO highlights the opposite failure mode: not everything should be compressed into fixed-size state. When the task requires precise recall—"where was the object placed?" or "when was the stove turned on?"—lossy recurrent state may discard the detail. Explicit episodic memory with selective retrieval is a better fit. RoboMME's systematic comparison across 16 memory-dependent tasks likewise finds no universally superior representation: symbolic memory is stronger for counting and event reasoning, while perceptual memory is better for motion and timing.[5][7]

The more plausible direction is therefore not to choose one winner among TTT, SSM, and Transformers, but to compose them:

Use SSM/TTT for low-level belief, sparse attention to retrieve key episodes, and a high-level planner to maintain task state. Only information that can actually change the action deserves to be written into memory.

Conclusion

RoboTTT, WAM-TTT, and RoboSSM are important works. They move long-context robot policies beyond concatenating more frames and toward learning how to write history into state. They also make human video behave like a real prompt: an interface that can change robot behavior at deployment time.

I would not, however, make raw context length an independent scaling objective for robot policies. For any memory design, I would first ask three questions: what latent information is missing from the current observation? How long must that information survive, and how precisely must it be retrieved? Does the training data cover the mapping from that context to the correct action?

Without answers to those questions, 8K is just a longer recording.
With answers, what we need is often not 8K frames, but eight events that matter.


References

[1] RoboTTT: Context Scaling for Robot Policies, 2026.

[2] WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time, 2026.

[3] RoboSSM: Scalable In-context Imitation Learning via State-Space Models, CoRL 2025 / revised 2026.

[4] BPP: Long-Context Robot Imitation Learning by Focusing on Key History Frames, 2026.

[5] Memory Retrieval in Visuomotor Policies for Long-Horizon Robot Control (HALO), 2026.

[6] What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents, 2026.

[7] RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies, 2026.