Learning Social Navigation from Internet Videos in the Policy State Space

Building a social-navigation simulator normally means building a world: layouts, terrain, 3D assets, and a pedestrian model that decides how people behave. We skip the world. An ordinary monocular walking video is converted directly into the policy's own state space — a metric traversability map plus the pedestrian trajectories actually recorded on that street — and that is enough to train a closed-loop policy.

Overview video

Three minutes, end to end

Video pipeline, replay simulator, Arena transfer, and the Go2 deployment.

Motivation

Two bottlenecks, both made of hand-built content

Social-navigation policies are trained in simulation, because learning them on real robots is costly and potentially unsafe. But the simulators themselves are expensive in two independent ways, and each caps a different kind of diversity:

Scene diversity is capped by construction cost. Layouts, terrain and obstacles are modelled by hand, so a benchmark ships a handful of environments — and a policy only ever sees the streets someone had budget to build.
Behavior diversity is capped by the pedestrian model. People are driven by social force or a learned trajectory predictor, so the policy only meets the encounters that model can produce, not the ones that actually happen on a street.

The obvious fix — reconstruct real scenes from video — replaces the artist with a reconstruction, but keeps the expensive part. Gaussian splats, generated assets and neural renderers all have to render counterfactual observations: the moment the robot steps off the recorded camera path, the simulator must synthesise what it would have seen. Dynamic social scenes make this worse, since people now need appearance models and motion models too.

Simulation in pixel space

hours per scene · renders every step

Simulation in state space

4.5 min per 20 s clip · never renders

Source
LLMdesignervideo

Layouts, terrain and 3D assets are authored, or reconstructed into something renderable.

First-person frame of a city street at t = 4 s. The same walk at t = 8 s.

An ordinary monocular walking video. Nothing else.

World
geometryassetshuman model

A 3D scene plus scripted people, whose behavior is whatever the chosen model produces.

Overhead metric traversability map: green walkable ground, red obstacles, grey unobserved.
static → rigid warp
Overhead view of recorded pedestrian trajectories with velocity arrows.
dynamic → replay at t

A metric traversability map and the trajectories people actually walked.

Observation
A photorealistic rendered street scene from a 3D reconstruction pipeline.

Rendered pixels — re-synthesised for every counterfactual robot pose.

Egocentric traversability map with pedestrian discs, the goal star and the robot's own path.

The robot-frame state itself — a rigid transform away, at any pose.

Policy acts on
pixels
state
traversable non-traversable unobserved people camera path robot
From video to social-navigation simulation. The recorded world is decomposed into a static traversability map, which is rigidly transformed under counterfactual robot motion, and dynamic pedestrian trajectories, which are replayed over time. This lightweight abstraction enables efficient simulation and reinforcement learning without reconstructing or rendering photorealistic scenes.

Key idea

A navigation policy does not need pixels

Local social navigation depends on two things: where the robot can traverse and how nearby pedestrians move. Both are geometric. Neither needs appearance. So instead of reconstructing a scene that can be rendered, we recover only those two quantities and define the simulator's forward dynamics directly on them.

That choice is what makes counterfactual robot motion cheap, because each layer transforms in a way that needs no model:

LayerWhat it holdsUnder counterfactual robot motion
Static Gravity-aligned metric traversability map, accumulated causally A rigid transform into the new robot pose — exact, no rendering
Dynamic Time-indexed pedestrian states (x, y, vₓ, v_y) Replayed at the video's own clock — real speeds, crossings and crowd patterns, no behavior model

World time follows the recording; the robot pose evolves independently from the policy's velocity commands. The two only interact when the recorded state is transformed into the robot's current frame — which is defined for any pose. So the robot can stop, turn, or leave the demonstrator's path entirely and still observe a consistent world state.

The cost of this abstraction. Pedestrians replay rather than react, and the map only covers what the camera saw. Both are measured rather than assumed — and replay turns out to beat both reactive alternatives we tested, including on an independent benchmark whose crowds are reactive.

Method

The pipeline, stage by stage

Every stage below runs on ordinary 4K walking-tour footage with pretrained perception models. One 20 s clip (200 frames at 10 fps) takes 4.5 minutes on a single A100.

a Walking video

Walking video frame at t = 4 s. Walking video frame at t = 6 s. Walking video frame at t = 8 s.

First-person 4K walking tours. The only input.

b Data processing

Person segmentation masks over the street frame. Per-pixel traversability prediction over the street frame. Metric depth map of the street frame.

People (YOLO11 + BoT-SORT), walkability (SAM-TP) and metric depth + camera poses (Depth Anything 3), fused into a gravity-aligned world frame.

c Metric replay step

Robot-frame traversability map with replayed pedestrians, goal and paths.

World time follows the recording; the robot pose integrates from the policy's command with a unicycle model. Map and people are transformed into the robot frame each step.

0.1 m cells±10 m10 Hz ~1,000 steps/s

d Policy

zBt ht1:KgtR at−1
transformer · 4 blocks
203 tokens · d = 192
1.6 M parameters
actor → Beta (v, ω)critic V

180 angular cone tokens + up to 20 pedestrian tokens + goal + previous action. RoPE on the cone tokens only; pedestrian tokens stay permutation-invariant.

Trained with PPO from scratch in the replay simulator — 32 parallel environments, reward on goal progress with collision termination.

e Deployment

Onboard RGB view from the Go2's camera on a brick path. Traversability map and pedestrian tracks built online on the robot.

The same state is rebuilt online from RGB-D and odometry, so the video-trained policy transfers with no modification.

Overview. (a) First-person walking videos are processed by (b) pretrained perception models to recover a gravity-aligned metric traversability map and time-indexed pedestrian states. (c) These quantities define the state-space replay simulator. (d) A lightweight transformer encodes the egocentric map, nearby pedestrians, goal and previous action, and outputs velocity commands. (e) At deployment the same policy state is reconstructed online from onboard perception.

What the simulator actually checks

The map is accumulated causally: the map the robot sees at video time t contains only cells the camera had observed by t, never ground revealed later. On top of that state the simulator enforces the usual robot constraints —

Training is PPO from scratch, 32 parallel environments, at roughly 1,000 environment steps per second on one RTX 5090 — a 100 M-step run takes 26 hours.

Corpus

44 walking tours, 29 cities, no overlap between splits

Three held-out source frames: an evening crowd in Istanbul, a rainy London shopping street, and a snowy street at dusk in Tromsø.
Frames from three held-out source videos: an evening crowd on Istiklal Street (Istanbul), a rainy shopping street (London), and a snow-covered street at dusk (Tromsø). Ordinary first-person walking tours like these are the only input of the pipeline.
Corpus statistics. Train and test videos are disjoint, with no city shared between the splits.
 TrainTest
Source videos (distinct cities)25 (12)19 (17)
20 s episodes extracted8,1473,126
Episodes passing gates6,0752,155
Hours of usable video33.812.0
Pedestrian tracks91,26129,846
Tracks per episode (mean)15.013.8
Pedestrians per frame (mean / max)2.5 / 182.3 / 18
Frames with ≥ 1 pedestrian72%68%
Demonstrator path per episode (m)23.721.7

Result 1 — unseen videos

92.6% on 2,155 episodes from cities it has never seen

Closed-loop in the state-space simulator. The robot starts at the demonstrator's pose at t = 2 s and must reach a point at least 15 m ahead, among the recorded people.

Closed-loop evaluation on the held-out episodes in the state-space simulator.
Success rate (95% Wilson interval)92.6% [91.4, 93.7]
Pedestrian collision3.7%
Static-obstacle collision1.9%
Timeout (18 s)1.8%
SPL0.860
Path efficiency (shortest / traveled)0.928
Time to goal (s)12.2
Mean speed (m/s)1.28
Time in personal zone (< 1.2 m) per episode (s)0.24

For scale: replaying the recorded camera trajectory itself — a real person walking that street — satisfies the same collision criteria in only 72% of episodes, with 21% pedestrian and 7% static collisions. The walker brushes past people far closer than the robot's rules allow.

Result 2 — the central claim

Replaying people beats modelling them

The simulator's main approximation is that pedestrians follow their recorded trajectories instead of reacting to the robot. If that were a serious limitation, a reactive model should train a better policy. It does not.

One held-out episode, five policies

A street in Belgrade, from a video none of these policies has seen. Every clip is the same episode with the same deterministic start; only the training setup differs. Left panel is the policy's egocentric observation, right panel the world map with the path it took.

Trained on replayed people reached the goal 91.4%

Reads the two approaching people early, arcs wide around them, and carries its speed through to the goal — 126 steps.

Trajectron++ peoplecollision90.0%

Trained against a learned predictor. Commits to a line between the two people and contacts one at step 88.

Social-force peoplecollision75.6%

Trained against people who yield to the robot, so it holds course and expects them to move. They don't — step 81.

25% of the videoscollision81.8%

Same method, quarter of the corpus. Starts to turn far too late — fewer videos means fewer crowd configurations seen.

Imitation onlycollision55.3%

Supervised on the demonstrator's own motion, never trained closed-loop. Tracks the walked path straight into the pair.

traversable non-traversable unobserved pedestrians (+ velocity) camera path policy path
One episode is an illustration, not evidence — it is one of the 166 held-out episodes where the replay-trained policy succeeds and all four variants collide. The percentage on each card is that policy's success rate over all 2,155 held-out episodes, which is the claim the clips illustrate. All five are 30 M-step runs evaluated under identical conditions.

The cross-evaluation that settles it

Training on replay could simply be tuning a policy to a quirk of replay. So each policy is scored under all three pedestrian dynamics — the recorded replay, a social-force crowd, and a Trajectron++ crowd. A policy that only works in its own training dynamics would show a strong diagonal. The replay-trained policy is the best or tied-best in every column:

Success rate (%) on the same 2,155 held-out episodes, under three pedestrian dynamics at evaluation time. Rows are what the policy was trained against; columns are what it is tested against. Training on social force costs 16–18 points whenever the crowd is not social force.

Trajectron++ matches replay without improving on it, and costs more to get there: it is an extra model that has to be trained (here from scratch on our own tracks, with no test video touching it) and it drops the simulator from about 1,000 steps/s to 800. Social force is clearly worse. Direct replay is the simplest of the three and the most robust.

More video is what helps

If the recorded crowds are the source of the policy's competence, then adding videos should buy competence — and specifically it should buy pedestrian safety rather than general path quality. It does:

Scaling the corpus from 25% to 100% of the videos lifts success from 81.8% to 91.4%, and the clearest movement is in pedestrian collisions, which fall from 13.5% to 4.4%. Static collisions stay low and flat throughout.
Full ablation table — all 30 M-step variants on the 2,155 held-out episodes
VariantSucc. ↑SPL ↑Ped. coll. ↓Stat. coll. ↓Pers. (s) ↓
Control91.4 [90.1, 92.5]0.8584.41.50.19
IL only55.3 [53.2, 57.4]0.54832.111.20.46
PPO from IL initialization90.0 [88.6, 91.2]0.8595.41.20.45
No spawn jitter90.4 [89.1, 91.6]0.8605.41.60.58
90° FOV map88.9 [87.5, 90.1]0.8464.51.30.09
Proxemics reward85.0 [83.4, 86.5]0.82311.52.10.42
Square-patch tokens89.7 [88.3, 90.9]0.8555.41.90.39
25% training videos81.8 [80.1, 83.3]0.79113.52.40.28
50% training videos88.2 [86.8, 89.5]0.8345.00.90.24
Social-force pedestrians75.6 [73.7, 77.4]0.73821.11.60.44
Trajectron++ pedestrians90.0 [88.7, 91.2]0.8505.71.80.34
Same episodes, pedestrians driven by social-force model
Control93.4 [92.3, 94.4]0.8721.81.50.18
Social-force pedestrians88.7 [87.3, 89.9]0.8597.61.40.59
Trajectron++ pedestrians93.3 [92.1, 94.3]0.8782.31.70.32
Same episodes, pedestrians driven by Trajectron++ model
Control91.9 [90.7, 93.0]0.8664.11.50.20
Social-force pedestrians75.5 [73.6, 77.2]0.73821.81.30.46
Trajectron++ pedestrians91.3 [90.0, 92.4]0.8705.21.60.37

Succ. is success rate; Ped. coll. and Stat. coll. are pedestrian and static-obstacle collision rates; Pers. is time per episode within 1.2 m of a pedestrian. Brackets are 95% Wilson confidence intervals. All variants trained for 30 M steps.

Closed-loop interaction is the thing that matters, not the demonstrations. Imitating the recovered demonstrator motion reaches 55.3%, against 91.4% for PPO from scratch — and initialising PPO from that imitation policy gives no benefit at all (90.0%). The value of the video is not that it contains expert actions; it is that it can be turned into an environment where failures and recoveries that are absent from any demonstration can happen.

Result 3 — independent benchmark

Transfer to Arena with no fine-tuning

Arena / Isaac Sim, four UrbanVerse street scenes, 160 scripted scenarios with reactive social-force pedestrians. A Clearpath Jackal with an OAK-D camera navigates between goals 18–22 m apart. Our policy is dropped in at 10 Hz, unchanged.

Isaac Sim view of a rendered street scene with pedestrians.
What Arena renders.
The traversability map the policy builds online from the OAK-D depth camera.
What our policy uses — rebuilt online from RGB-D. No LiDAR, no prebuilt map, no global plan.
Arena / Isaac Sim results on 160 scenarios. Personal and intimate denote seconds per episode within 1.2 m and 0.45 m of a pedestrian. Every baseline receives more input than we do; all methods are given ground-truth pedestrian positions.
MethodLiDARPlanSucc. ↑ Coll. ↓ (ped/wall)T/O ↓Time ↓SPL ↑ Pers. ↓Intim. ↓
CrowdSurfer✓✓42.5 [35.1, 50.2]50.0 (47.5/2.5)7.546.20.41819.91.9
AttnGraph—✓35.0 [28.0, 42.7]45.6 (11.2/34.4)19.452.00.29910.31.2
HEIGHT✓✓75.0 [67.8, 81.1]25.0 (19.4/5.6)0.046.40.71411.83.5
Ours (SF people)——58.1 [50.4, 65.5]41.9 (39.4/2.5)0.042.60.5598.41.8
Ours (Traj.++ people)——70.0 [62.5, 76.6]25.0 (20.0/5.0)5.051.90.59912.22.3
Ours (replay)——81.2 [74.5, 86.5]15.0 (13.8/1.2)3.846.20.74710.62.1

Against HEIGHT, the strongest baseline, success improves by 6.2 points and collisions fall from 25.0% to 15.0% — both pedestrian contacts (13.8% vs. 19.4%) and wall contacts (1.2% vs. 5.6%). And the gap is not bought by timidity: our policy spends less time in the personal and intimate zones than HEIGHT while succeeding more often. It maintains progress through dense crowds, including the brief close passes that are normal in the recorded videos.

The advantage grows with crowd density

Ours (RGB-D) HEIGHT AttnGraph CrowdSurfer
Arena results by crowd density. Easy = 0–4 pedestrians, medium = 5–9, hard = 10–14. On hard scenarios our policy reaches 74.2% against 53.2% for HEIGHT — a 21-point gap where it matters most. Numbers are the same 160 scenarios as the table above.
Arena is also an out-of-distribution test of the replay assumption. Its crowds react to the robot, so a policy trained on non-reactive replay ought to be at a disadvantage. Instead the replay-trained policy scores 81.2% while the same architecture trained with social-force people reaches 58.1% and with Trajectron++ people 70.0%. Preserving real recorded motion transfers better than approximating it.

Result 4 — physical robot

19 out of 20 trials on a Unitree Go2

No policy fine-tuning, no global map or planner. Traversability is built online from RGB-D and odometry; pedestrians are detected and tracked from LiDAR with a YOLO11 prior.

Unitree Go2 quadruped fitted with an RGB-D camera and a LiDAR unit. The indoor test scene: the robot approaching people in a corridor-like space. The outdoor test scene: the robot on a brick path with a person walking towards it.
(a) Unitree Go2 with RGB-D camera and LiDAR; (b, c) the indoor and outdoor test scenes. Each scene has five scripted encounters, repeated twice per method.
Real-robot results (10 trials per method and scene; means over successes). A trial succeeds if the robot reaches a goal 8–10 m away without a static collision and without coming within 0.3 m of a person.
SceneMethodSuccess ↑Mean time (s)Mean path (m)
IndoorNav2-MPPI5/1028.69.64
Ours (Trajectron++ people)7/1032.110.96
Ours (replay)9/1031.110.14
OutdoorNav2-MPPI3/1018.97.74
Ours (Trajectron++ people)8/1016.18.29
Ours (replay)10/1017.98.59

The failure mode of the baseline is instructive. MPPI treats a detected person as an occupied costmap cell, so it holds its course until someone is already close and then avoids late. Both learned policies receive pedestrian velocities, so they anticipate the encounter and slow or step aside before the crossing point. The replay-trained policy's single failure is indoors, where the 0.5 m/s hardware speed limit — far below the speeds it selects during training — leaves it too little time to clear an approaching person's path.

Limitations

What this representation cannot do