Learn the predictive representation
Use current–future observations to learn a compact transition representation through inverse dynamics and feature reconstruction.
Privileged future observations · Stage 1 onlyMiniWAM: Learning Compact Future Targets for Efficient World-Action Modeling
Predict what matters for action. MiniWAM learns compact predictive representations, making world-action training more efficient with a 0.25B-parameter policy.
Conventional WAMs predict dense native visual features, which is computationally intensive. MiniWAM instead learns a compact target that is lightweight to predict, preserves useful future information, and emphasizes the dynamics needed for control.
MiniWAM replaces costly native future prediction with compact predictive representations. (a) While conventional WAMs predict high-dimensional native future features, MiniWAM jointly predicts actions and compact PRISM targets learned from the same visual features via inverse dynamics and reconstruction. (b) MiniWAM remains competitive with substantially larger WAMs and VLAs on LIBERO-Plus and RoboTwin 2.0. (c) In controlled comparisons across WAN2.1-VAE and DINOv3 feature spaces, PRISM simultaneously improves the success rate and reduces training time relative to native future-feature prediction.
PRISM — Predictive Representations via Inverse Spatiotemporal Modeling — learns a compact future dynamics representation that improves the performance–efficiency trade-off.
Use current–future observations to learn a compact transition representation through inverse dynamics and feature reconstruction.
Privileged future observations · Stage 1 onlyFreeze the PRISM encoder. Jointly predict compact future targets and robot actions.
Compact prediction targets · Computationally efficientTwo-stage training pipeline of PRISM and MiniWAM. In Stage 1, privileged current–future visual features are compressed by the PRISM encoder into a compact predictive representation Lt, shaped through inverse action prediction and feature reconstruction. The corresponding world tokens remain clean while the action tokens are corrupted for conditional action prediction. In Stage 2, the PRISM encoder is frozen and provides compact world-modeling targets; MiniWAM jointly denoises the world and action tokens using its world and action experts. Solid and hatched tokens denote clean and noised tokens, respectively. Language and proprioceptive conditioning are omitted for clarity. Future observations are not needed at deployment.
The 0.25B MiniWAM policy is evaluated across diverse manipulation benchmarks and is competitive with much larger models.
40 tasks · 2,000 episodes
10,030 perturbed tasks
Clean / Randomized mean · Trained on clean split
Success rate (%) of MiniWAM and substantially larger VLAs and WAMs.
Bar length: reported model size
| Method | Scale | Embodied pretraining |
Backbone init. |
LIBERO | LIBERO-Plus | RoboTwin 2.0 |
|---|---|---|---|---|---|---|
| VLA Vision-language-action models | ||||||
| OpenVLA-OFT | 7B | Yes | VLM | 96.6 | 71.4 | Not reported |
| π0 | 3.3B | Yes | VLM | 93.5 | 70.5 | 31.4 |
| π0-FAST† | 3.3B | Yes | VLM | 85.5 | 61.6 | Not reported |
| UniVLA† | 7B | Yes | VLM | 95.2 | 42.9 | Not reported |
| StarVLA-OFT† | 4B | No | VLM | 96.6 | Not reported | 24.8 |
| Xiaomi-Robotics-0† | 4.7B | Yes | VLM | 98.7 | Not reported | 40.5 |
| π0.5 | 3.3B | Yes | VLM | 96.4 | 85.7 | 58.4 |
| WAM World-action models | ||||||
| WorldVLA† | 7B | No | VLM | 79.1 | 25.0 | Not reported |
| Fast-WAM | 6B | No | VGM | 95.8 | 71.0 | 39.9 |
| DreamWAM | 6B | No | VGM | 96.9 | 75.4 | Not reported |
| AHA-WAM† | 7.2B | No | VGM | Not reported | Not reported | 33.8 |
| MiniWAM Ours | 0.25B¹ | No | None | 97.2 | 73.6 | 46.1 |
† Externally reported. Other baselines are released checkpoints that we evaluate on LIBERO and LIBERO-Plus with the same task configurations and random seeds as MiniWAM. RoboTwin 2.0 reports the official leaderboard mean of clean and randomized evaluation, with training on clean demonstrations only. Scale is the reported model size. VLM: vision-language model; VGM: video generation model; –: not reported.
Controlled comparison on the fixed LIBERO-Plus-2K subset.
MiniWAM is evaluated on the LIBERO, LIBERO-Plus, and RoboTwin 2.0 benchmarks. Below are successful policy rollouts.
What does PRISM keep? Retrieve transitions from different tasks, then compare their actions and observed motion. Similar behavior should remain close even when the scene changes.
Main + wrist views average · 882 queries per view · top-5 neighbors
Loading the recorded nearest neighbors…
Leave-task-out, not leave-scene-out. Every neighbor is from a different task and an episode-disjoint gallery; scenes and objects may still repeat. Each clip shows all 17 retained frames from t to t+16, played slowly at 6 frames/s for inspection. These are recorded LIBERO-90 demonstrations, not policy predictions.
Behavior, measured independently of retrieval. Neighbors are selected by cosine distance in representation space. We then measure their normalized 16-step action RMSE, end-effector displacement-vector error and gripper-sign agreement against the query. Lower errors and higher agreement indicate closer behavior; matching the task label alone does not.
A shared, held-out gallery. The quantitative analysis uses 882 locked queries and an episode-disjoint gallery of 3,618 transitions. Candidates from the query's task are excluded. The archived native readout spatially averages each frame’s features and keeps the temporal sequence; PRISM uses its full-grid posterior mean (3 × 4 × 4 tokens, 64 dimensions). These are the reconstruction-weight-1 analysis encoders. This comparison does not test every possible native-feature readout. Both representations see the observed future. This tests representation geometry, not online policy prediction.
Aggregation. Each card is the equal-weight mean of main- and wrist-view metrics for the selected backbone, computed after retrieval separately in each view. The paper table instead averages the two backbones and reports each view separately. The video examples use the main camera only.
Evidence, with limits. The cards use the complete quantitative cohort; the three clips are selected illustrations. These averages do not imply improvement in every view or example, establish scene invariance, or explain policy success by themselves.
Download the displayed measurements and neighbor identities ↗
We will open-source the code, configs, and checkpoints before November 2026. Stay tuned.
Pinned policy and simulator environments.
A named checkpoint, its statistics, and a fixed protocol.
PRISM representation learning, then world-action training.
@article{chen2026miniwam,
title = {MiniWAM: Learning Compact Future Targets for Efficient World-Action Modeling},
author = {Chen, Jie and Bai, Ruofei and Cai, Yuxin and Zhang, Yifeng and
He, Chengyang and Li, Jun and Yau, Wei-Yun and Sartoretti, Guillaume},
journal = {arXiv preprint arXiv:2610.12194},
year = {2026}
}
1 Parameter counts exclude frozen visual and text encoders.
2 65× refers to the WAN2.1-VAE native-to-PRISM future-token ratio at visual horizon 16 (3,136 vs. 48 tokens per camera).
3 Maximum estimated total training speedup at RoboTwin horizon 32, accounting for Stage 1 under the paper’s equal per-update-time assumption. Excludes offline feature extraction.