This world needs WebGL and a connection to the Three.js CDN.
Everything measured lives on the 2D evidence page and in the paper.
Protocol explainer · not a live rerunH 15 planK 3 execute↻ replan
Protocol worldkeys 1–2 switch worlds
The planner is imagining. Most of it will not survive.
Watch one receding-horizon cycle: a ghost of fifteen imagined chunks grows out of the agent, the graded endpoint is marked, K chunks become solid, and the rest of the plan is thrown away. Then it happens again.
World
Where the grading mark sitssame ghost, same weights
score ẑH
Executed before replanningmeasured intervals only
Horizon fixed at H = 15plan kept: 3 / 15
Protocol ledgerdefinitional arithmetic
Cycles0
Imagined0
Executed0
Discarded0
Per cycle: 15 chunks imagined, K executed, 15 − K discarded at feedback. Counted from the protocol, not measured from a run.
Measured · paper · PushTN = 200 starts · H = 15 · K = 3
17.0%terminal@H
94.5%prefix@K
86.5%running
Same model, same weights, same starts — only the question changes. The same alignment check moves TwoRoom from 52.0% to 83.5% and Reacher from 20.5% to 100%.
The replanning sweep was measured only at K ∈ {1, 3, 5, 15}. Nothing between those points is interpolated and no curve is fitted: the PushT values are non-monotonic. The full figure is on the 2D page.
Directions from the same statetoggle the layers
Walk it yourselfclick the floor to steer
Try clicking the far side first. The wall is allowed to win once.
When “The detour” and “Walking” are active, focus the 3D world and use the arrow keys to move the target. Pointer steering remains available.
Steering target: none.
Measured · paper · TwoRoom Far Doorheld-out evaluation
Best scalar / value objective26.5%
Waypoint + nearest retrieval92.7%
The 92.7% is the paper's five-fold mean. In the matched failure clip (held-out fold 0, episode 04), the closest approach was 39.70 against the success threshold of distance < 16 — reported metadata from the display export, not a recomputation.
Event logprotocol events only
READYTwo staged worlds. Nothing here replays an episode.
What is exact
Every measured number on this page is quoted from the paper: PushT at H=15, K=3 scores 17.0% under terminal@H, 94.5% under prefix@K, and 86.5% under running cost (N=200 starts). The alignment check moves TwoRoom Far Door from 52.0% to 83.5% and Reacher from 20.5% to 100%. On TwoRoom Far Door, the best scalar/value form reaches 26.5% against 92.7% for waypoint retrieval with a local actuator — a five-fold mean. The sweep covers exactly K ∈ {1, 3, 5, 15}; nothing between those points is interpolated, and no fitted curve is drawn, because the PushT values are non-monotonic. The ledger counts (15 imagined, K executed, 15 − K discarded) are definitional arithmetic on H and K, not outcomes of a run.
What is staged
All motion in this scene is staged. assets/lewmro/matched-trace-receipts.json records numeric_state_trace_available: false — no per-step state traces exist in this repository, so nothing here replays an episode: not the ghost rollouts, not the block, not the drivable agent. The 400-path field is drawn to the described shape of the stored expert dataset — descend to the low doorway, cross the wall, climb to the high goal — and is a schematic of that geometry, not the stored trajectories. The evidence remains the matched-configuration clips and the paper's aggregates on the 2D page. This world explains the mechanism; it does not add a measurement.