PlanWAM: Planning-Shaped Future Representations for End-to-End Autonomous Driving

Jinchang Xu1,†, Hongda Yu2,†, Fengwei Dong1, Wenhui Huang3, Xi Wei2,‡, Yongzhi Liu1, Sunan Zhang1, Jirao Wang4, Chen Lv3, Bingbing Li1, Guodong Yin1, Weichao Zhuang1,*

1Southeast University  ·  2NIO  ·  3Nanyang Technological University  ·  4Tongji University

†Equal contribution.  ·  ‡Project leader.  ·  *Corresponding author.

Driving world models usually ask how to predict the future and how to use it. PlanWAM asks which future representation is most useful for planning. Trajectory-planning objectives shape a future latent from privileged future observations; a history-only prior predicts that latent and uses it for trajectory generation and selection.

Zero-shot closed-loop driving on HUGSIM.
01 / Overview

Overview

A future representation that can reconstruct the scene does not necessarily retain the information that changes a trajectory decision. PlanWAM treats future modeling as planning-shaped predictive representation learning: planning defines which future information matters, future prediction decides which history to keep, and the predicted future then improves planning.

A train-only posterior learns the future latent from real future frames under planning losses. A history-only prior learns to predict it and is the only branch used at inference.

93.8 PDMS NAVSIM-v1 navtest
90.9 EPDMS NAVSIM-v2 navtest
38.7 HD-Score HUGSIM, zero-shot
53.7 Route Completion HUGSIM, zero-shot
02 / Motivation

Motivation

Driving world-action models have moved from pixel reconstruction, to latent prediction, to future-conditioned planning. The last step lets a future state z+ enter the planner, but still leaves open how that state should be defined. Existing targets are usually pixels or image latents. PlanWAM defines the target by the planning losses themselves.

Four future-modeling paradigms: pixel reconstruction, latent prediction, future-conditioned planning, and PlanWAM
Figure 1. Four future-modeling paradigms. (a) Pixel reconstruction. (b) Latent prediction, supervised toward DINO, JEPA, or BEV features of the future video. (c) Future-conditioned planning, where the predicted latent is distilled into the planner. (d) PlanWAM: a train-only posterior shapes the future latent with the planning task, and a prior available at inference predicts it.
A

Planning defines the future

The future representation is shaped directly by the trajectory-planning objectives, so it preferentially encodes the future information that affects trajectory decisions.

B

The future decides what history to keep

History is compressed to the dynamic cues needed for future reasoning and planning. What is kept is decided by the future to be predicted, not by retaining the full observation.

C

Foresight has to be deployable

Ground-truth future frames exist only in training. Distillation turns that privileged hindsight into a prior that runs on history alone.

03 / Method

Method

PlanWAM inserts a future latent between history and planning, and lets the planning losses decide what that latent encodes.

PlanWAM prior and posterior branches, two-stage training, and history-only inference
Figure 2. PlanWAM overview. The posterior sees the future and exists only in training; the prior sees history only and is what runs at inference.
01

Temporal Register Pyramid

Compresses multi-frame, multi-view history into a compact representation, giving recent frames more capacity.

02

Planning-shaped future posterior

In training, a posterior reads real future frames. Trajectory generation and scoring shape its latent, with no reconstruction loss.

03

Hindsight-to-Foresight Distillation

A history-only prior learns to predict that latent, turning hindsight available in training into foresight available at deployment.

04

Foresight-conditioned planning

The predicted latent conditions both trajectory generation and trajectory selection.

Stage 1 · posterior L_task = L_plan + L_score
Stage 2 · prior L_train = L_task + λ L_latent
Temporal Register Pyramid compressing history with a Q-Former
Figure 3. Temporal Register Pyramid: recency-aware compression of historical scene registers.
World-model, planner, and scorer transformers
Figure 4. World-model, planner, and scorer transformers. Planner and scorer both attend to history plus the future latent.
04 / Experiments

Experiments

PlanWAM is trained on NAVSIM navtrain (103,288 scenes: 85k train / 18k validation) and evaluated on navtest (12,146 scenes). NAVSIM-v1 uses PDMS. NAVSIM-v2 uses EPDMS on the same split. The same checkpoint is then deployed on HUGSIM, a closed-loop simulator with 436 public scenarios from nuScenes, KITTI-360, Waymo, and PandaSet, without HUGSIM-specific fine-tuning.

4 × 4history frames × cameras
128history tokens
256future latent tokens
K = 64candidates, 8 waypoints
4 splanning horizon

NAVSIM-v2 navtest

The best learned result in each column is shown in bold. PlanWAM reaches 90.9 EPDMS, 1.0 above DriveFuture and 0.6 above the reported Human Agent. It also obtains the best learned NC (99.0) and TTC (98.5).

MethodReferenceNCDACDDCTLCEPTTCLKHCECEPDMS
Human Agent-100.0100.099.8100.087.4100.0100.098.190.190.3
E2E-based planners
TransFuserTPAMI'2396.989.997.899.787.195.492.798.387.276.7
DiffusionDriveCVPR'2598.295.999.499.887.597.396.898.387.784.5
MeanFuserCVPR'2698.397.299.699.887.697.497.398.388.289.5
VLM-based planners
ReCogDriveICLR'2698.395.299.599.887.197.596.698.386.583.6
SGDriveCVPR'2698.694.399.599.986.097.996.198.385.986.2
ExploreVLAECCV'2698.896.299.699.887.198.297.898.386.888.8
World-model-based planners
DriveVLA-W0ICLR'2698.599.198.099.786.498.193.297.958.986.1
DriveLaWCVPR'2698.796.999.699.887.598.397.698.477.488.6
DreamerADECCV'2698.097.299.599.887.897.497.598.372.487.7
Latent-WAMarXiv'2698.197.399.699.887.797.397.698.187.389.3
DriveFuturearXiv'2698.899.199.699.986.698.496.498.374.889.9
PlanWAM (ours)-99.099.299.899.690.898.596.598.170.590.9

NAVSIM-v1 navtest

PlanWAM reaches 93.8 PDMS, 0.7 above DrivoR and 3.1 above the strongest prior world-model planner, DriveFuture. The best learned ego-progress score is 91.0.

MethodNCDACTTCCEPPDMS
Human Agent100.0100.0100.099.987.594.8
E2E-based planners
DiffusionDrive98.296.294.7100.082.288.1
MeanFuser98.697.095.0100.082.889.0
iPad98.698.394.9100.088.091.7
DrivoR98.998.396.2100.089.193.1
VLM-based planners
AutoVLA98.495.698.099.981.989.1
ReCogDrive97.997.394.9100.087.390.8
ExploreVLA98.898.496.599.983.590.4
SGDrive98.697.896.2100.085.891.1
World-model-based planners
WoTE98.596.894.999.981.988.3
Epona97.995.193.899.980.486.2
DreamerAD98.097.294.3100.083.188.7
DriveLaW99.097.196.7100.081.389.1
DriveVLA-W098.799.195.399.383.390.2
DriveFuture98.899.195.4100.084.290.7
PlanWAM (ours)98.998.896.3100.091.093.8

HUGSIM zero-shot closed-loop

E, M, H, and X are Easy, Medium, Hard, and Extreme (80 / 157 / 96 / 103 scenarios). Avg. is the mean over all 436 scenarios, not the unweighted mean of the four difficulties. PlanWAM improves the strongest baselines by 7.8 RC and 9.8 HD-Score. On Hard / Extreme it reaches 51.9 / 37.2 RC and 35.6 / 22.0 HD-Score.

MethodEMHXAvg.
Route Completion
UniAD58.641.240.426.040.6
VAD38.727.025.523.027.9
LTF68.440.736.925.541.4
GTRS-Dense64.250.020.722.338.0
Latent-WAM84.242.530.635.545.9
PlanWAM (ours)87.648.351.937.253.7
HD-Score
UniAD48.729.527.314.328.9
VAD24.39.910.48.212.3
LTF52.824.619.88.124.8
GTRS-Dense55.539.011.714.328.6
Latent-WAM72.524.012.218.128.9
PlanWAM (ours)75.332.935.622.038.7
Qualitative comparison with DrivoR on NAVSIM and zero-shot HUGSIM rollouts
Figure 5. Qualitative planning results. (a) Comparison with DrivoR on NAVSIM. (b) Zero-shot closed-loop rollouts on HUGSIM across driving domains.

Failure analysis

Failed scenes on NAVSIM-v1 fall from 314 to 237. The ones that disappear are collisions, short time-to-collision, and wrong-direction driving. Staying inside the drivable area does not improve. Nearly all of the reduction is a better choice among the candidates: selection errors drop by 87, while proposal collapse increases.

Failure-type comparison and attribution of the reduction from 314 to 237 failures
Figure 6. PDMS = 0 cases, base model (ID0) versus PlanWAM (ID3). (a) NC, TTC, and DDC drop by 62.3%, 52.4%, and 66.7%; DAC rises by 7.3%. (b) Selection error (−87) is the main source of the drop.
05 / Citation

Citation

Please cite this preprint as follows.

@misc{planwam2026,
  title     = {PlanWAM: Planning-Shaped Future Representations for End-to-End Autonomous Driving},
  author    = {Xu, Jinchang and Yu, Hongda and Dong, Fengwei and Huang, Wenhui and Wei, Xi and Liu, Yongzhi and Zhang, Sunan and Wang, Jirao and Lv, Chen and Li, Bingbing and Yin, Guodong and Zhuang, Weichao},
  year      = {2026},
  note      = {Preprint}
}