Selected workKV / LeWMRO / reader’s tourPaper ↗
ICML 2026 DEMO Workshop · oral

A useful world model can
look broken.

QuestionDoes a low planning score diagnose the learned model, or the interface wrapped around it?

MethodHold the offline-trained model fixed. Change the scoring time, replanning interval, proposal, and actuator.

ResultAligned costs repair standard tasks. Deceptive navigation still needs locally executable waypoints.

BoundaryThe study diagnoses interfaces. It does not propose a universal planner or claim that running costs are new.

Inspect the evaluation chain ↓

A benchmark score belongs to the whole chain.

Select a link to see what the experiments changed and what remained entangled.

Plan H steps. Execute K. Then look again.

Terminal@H judges the end of an imagined sequence. A receding-horizon controller executes only K chunks before new feedback changes the plan.

Task
Measured execution interval K

No interpolation. Every selection is one measured Table 4 cell with N=200 starts.

execute 3, observe, replan
terminal @ H17.0%

Scores the imagined endpoint.

running cost · γ=.9586.5%

Scores progress across the imagined rollout.

At K=3, PushT changes by +69.5 percentage points when only the scoring rule changes.

Same start. Same goal. Different score.

These clips explain one matched PushT episode. They do not establish the aggregate result.

Terminal @ Hillustrative failure
Scores the imagined endpoint at H=15.
Prefix @ Killustrative success
Scores the last executed chunk at K=3.
Running costillustrative success
Weights progress across the full imagined rollout.

The paper aggregates N=200 starts per standard-task cell. This retained display run contains 12 matched starts. The three clips above show one of them and use a reach-once tolerance within the 75-step budget.

The direct route points into a wall.

TwoRoom Far Door forces the first useful move away from the final goal. Choose an interface to see what path it can propose.

Conceptual interface schematic · not an observed proposal or retained trace
startfinal goalconceptual waypointconceptual K-step edge

26.5%single N=200 evaluation

Watch the waypoint redirect the first move.

Scalar valueillustrative failure
No route is supplied.
Nearest waypointillustrative success
The ring is the solver's recorded waypoint.

This pair shows one held-out-fold display episode. The 92.7% result is the five-fold mean over the paper's evaluation, not the outcome of one clip.

Ninety paths shown. Four hundred stored.

The quicklook samples 90 successful expert trajectories for legibility. The stored TwoRoom Far Door dataset contains 400 successful episodes. Each valid route climbs to the far doorway, crosses the wall, then returns toward the high final goal.

TwoRoom Far Door expert trajectories and representative frames
Dataset geometry, not evaluation rollouts.

A diagnostic, not a universal planner.

Read the 8-page paperHidden Failure Modes in Latent World-Model Planning ↗