Pi 0.6: Supervised Learning in Reinforcement Learning's Clothing
I recently studied the landmark Pi 0.6 paper, in which RL-based Post-training is the key ingredient enabling the leap forward. Why is this form of RL Post-training so powerful? What fundamental assumptions about training large-scale Vision–Language–Action (VLA) models does it overturn?
This article was motivated by exactly that question. Specifically, I wanted to look past the surface-level RL formulation and understand the essence of RECAP, as well as how we should think about its further iterations. To do this, I found it immensely helpful to step back and survey the broader landscape of Reinforcement Learning.
In the process, I discovered something striking: over the past decade, Reinforcement Learning — especially Offline RL — has been steadily converging toward Supervised Learning — sometimes explicitly, sometimes under heavy disguise. RECAP is no exception to this trend; in many ways, it is the culmination of it.
From Classic RL to Supervised Learning: A Gradual Drift
In the early days, Reinforcement Learning and Supervised Learning had a clearly delineated boundary. Classic RL methods — Policy Gradients, Q-learning, Actor–Critic algorithms, and later DQN — were all built around tight interaction loops with the environment. Agents collected data through actions, estimated Value functions via Bellman backups, and improved policies through gradient ascent on expected returns or greedy maximization of Q-values. Supervised Learning, by contrast, was fitting to a fixed dataset, with no notion of interaction, long-term credit assignment, or counterfactual evaluation.
This boundary began to blur when researchers tried to apply RL to settings with limited or no interaction, giving rise to Offline RL. Offline RL immediately exposed a fundamental problem: Bellman backups require evaluating actions never observed in the dataset. This leads to severe extrapolation error and instability, as the Value function assigns arbitrarily high values to out-of-distribution (OOD) actions.
Much of the early Offline RL literature can be understood as patching the classic RL framework without abandoning it. Methods like BCQ and BEAR constrained the action space or regularized the policy to stay near the data distribution. Subsequently, methods like CQL introduced conservative regularization directly into the Value function, keeping the Critic pessimistic about actions outside the dataset's support. These methods still looked like RL — with Value functions, Bellman backups, Critics, Actors — but the goal had shifted: the primary objective was no longer theoretical optimality, but staying within the data manifold.
A second, deeper conceptual shift followed. Rather than optimizing policies directly through RL objectives, many methods began to recast Policy Improvement as supervised regression on the dataset, guided by learned value signals. AWR, AWAC, CRR, and TD3+BC all fall into this category. Even IQL, which retains Bellman structure internally, goes to great lengths to avoid explicitly evaluating unseen actions. Policy learning was looking increasingly like Supervised Learning, with the Value function serving as a weighting or filtering mechanism.
Finally, with methods like Decision Transformer (DT), the Bellman backup disappeared entirely. Reinforcement Learning became Sequence Modeling, trained with a purely supervised objective and conditioned on returns at inference time.
This evolution was not accidental. It reflects a deep empirical lesson: as model and data scale grow, regression is stable while bootstrapping is fragile.
Re-examining Reinforcement Learning: What Really Matters?
Once we see this pattern clearly, a deeper question naturally arises: what core elements of Reinforcement Learning have survived this transformation?
Stripping away historical implementation details, RL's genuine contributions boil down to two things:
- Interaction with the environment: This allows the model to discover its own failure modes.
- Value function: It resolves the multi-modality problem by introducing preference and temporal coherence.
Everything else — Policy Gradients, Q-learning variants, Bellman bootstrapping — is negotiable.
The Value Function as "Preference-Conditioned Supervision"
Let us be precise here: the Value function that modern systems value is not just any Value function — it is one that resolves multi-modality through conditioning.
In many real-world tasks, the action distribution is inherently multi-modal. Consider grasping a cup: there are multiple valid grasp types, trajectories, and styles. If we naively use Supervised Learning to learn the full action distribution, sampling becomes unstable and often incoherent. A simplistic fix is to compress the dataset to expert demonstrations, but this introduces severe bias and fragility — once the policy deviates from the expert manifold, it enters an OOD regime.
The Value function offers a different solution. Instead of compressing the data, we keep all trajectories (successful, partially successful, and failed) and introduce a signal that relates actions to their long-term expected value. In this sense, the Value function acts as a latent preference variable. It allows us to sample "good" behavior from a diverse dataset without destroying that diversity.
This framing immediately connects to modern LLM training. When we prompt an LLM to "write more professionally" or "be more concise," we are conditioning generation on a latent preference variable. RLHF formalizes this through Reward Models; Preference Optimization and Post-training bias generation toward high-reward outputs.
Decision Transformer did the same thing by conditioning on return-to-go. RECAP conditions on "optimality indicators" derived from a learned Value function. The mechanism is identical: Supervised Learning guided by a preference signal.
Interestingly, this convergence is now bidirectional. Robotics increasingly borrows LLM-style conditioning and preference modeling, while LLMs themselves increasingly rely on RL-style Post-training to refine behavior. The two communities are meeting in the middle.
Environment Interaction: Expanding Sample Coverage
The second surviving element of RL is interaction with the environment. No matter how large a fixed dataset, it cannot anticipate every failure mode. A deployed policy will eventually encounter underrepresented states and make systematic errors.
Interaction addresses this by revealing where the model goes wrong. The idea predates modern RL — DAgger already showed that Imitation Learning must be iterative — but RL adds a crucial element: a way to annotate failures with meaning.
When combined with the Value function, interaction becomes extraordinarily powerful. Failed trajectories are no longer useless waste; they become negatively labeled samples. Partial progress still carries learning signal. Interaction generates diverse data; the Value function transforms that data into structured supervision.
The same pattern appears in LLM Post-training: deploy, observe failure cases, collect preference feedback, retrain. Once again: Interaction + Value.
RECAP: Supervised Learning in RL's Clothing
Following this trajectory, we can see clearly: RECAP's essence should not be understood as a return to classic Reinforcement Learning, but rather as a system fundamentally rooted in Supervised Learning, augmented with a carefully curated set of RL ingredients.
Despite its RL-derived formulation, RECAP's operational reality is straightforward:
- The policy is trained via supervised regression on data.
- The Value function is trained via supervised regression on returns.
- There is no Bellman bootstrapping, no Policy Gradient, and no explicit optimization over the action space.
Performance gains come not from solving an RL objective in the traditional sense, but from conditioning a powerful supervised model on a learned preference signal.
In this sense, RECAP is far closer to Supervised Learning than to classic RL. Its success depends primarily on the same factors that govern large-scale supervised systems: data quality, diversity, and coverage; model capacity and inductive biases; and stability of the training objective. These factors dominate performance far more than any subtle mathematical property of the RL formulation.
The remaining RL components are not the engine of optimization but structural tools. The Value function provides a way to impose preference and temporal coherence on multi-modal action distributions. Environmental interaction provides a mechanism to discover failure modes and expand the dataset. Together, they guide Supervised Learning rather than replace it.
What RL Still Contributes — and Directions for Improvement
Viewing RECAP as fundamentally Supervised Learning is not a dismissal of Reinforcement Learning. Rather, it clarifies where RL ideas remain important and where they are most efficiently applied.
One important contribution lies in how the Value function is learned. Even when trained via regression, value estimates are still subject to classic RL trade-offs: variance-bias trade-off, horizon length, and sensitivity to distributional shift. Techniques developed in the RL community — such as variance reduction, multi-step estimation, or uncertainty-aware Critics — remain highly relevant for improving the quality and reliability of the conditioning signal.
Another key area is how the Value signal is used to condition the policy. RECAP currently relies on a binary optimality indicator, which prioritizes robustness and simplicity. While effective in many cases, this inevitably discards information. In scenarios with nuanced trade-offs, long-horizon dependencies, or noisy value estimates, hard thresholds may become brittle or misleading. Here, ideas from Offline RL — such as soft advantage weighting, calibrated preference scores, or uncertainty-aware conditioning — offer clear paths for improvement.
Outlook
This brings us back to the original motivation: understanding what RECAP actually is, why it works, and how to think about improving it. I believe the answer is: RECAP represents a mature convergence point. It extracts the parts of Reinforcement Learning that scale — interaction and Value-based conditioning — and embeds them in a stable, expressive, data-driven Supervised Learning framework.
Understanding RECAP this way allows us to leverage both perspectives simultaneously. When designing models, datasets, and loss functions, we can apply Supervised Learning intuitions; while still drawing on Reinforcement Learning theory to reason about value estimation, preference signals, and long-horizon structure. This combination is not contradictory — it is precisely what makes methods like RECAP so effective.
In this sense, RECAP does not hide Reinforcement Learning behind Supervised Learning. Instead, it shows us: what Reinforcement Learning looks like when distilled down to its truly essential core.
Following my previous post, I read the Pi 0.6 paper carefully to understand RL post-training in depth. Along the way, I reviewed and traced the recent development trajectory of RL, and discovered a fascinating phenomenon: RL has been converging toward supervised learning — and vice versa (a mutual convergence). This is not just the conclusion of this one paper, but reflects the entire development trajectory. In one sentence: the core ideas of RL are being solved through SL methods. I feel Richard Sutton himself would be pleased with where this trend is heading.