A mini world-action model.
A capable robot policy.

MiniWAM: Learning Compact Future Targets for Efficient World-Action Modeling

Jie Chen1,3Ruofei Bai2,3Yuxin Cai2,3Yifeng Zhang1Chengyang He1Jun Li3Wei-Yun Yau3Guillaume Sartoretti1

1 National University of Singapore2 Nanyang Technological University3 A*STAR Institute of Advanced Intelligence and Computing

Predict what matters for action. MiniWAM learns compact predictive representations, making world-action training more efficient with a 0.25B-parameter policy.

arXiv Code & models Coming soon before Nov 2026
Learn the future.
In a smaller space.
0.25BPolicy parameters¹
65×Fewer future tokens²
Up to 8×Estimated training speedup compared with native future prediction³

A smaller prediction target.
A different way to learn.

Conventional WAMs predict dense native visual features, which is computationally intensive. MiniWAM instead learns a compact target that is lightweight to predict, preserves useful future information, and emphasizes the dynamics needed for control.

MiniWAM replaces costly native future prediction with compact predictive representations. (a) While conventional WAMs predict high-dimensional native future features, MiniWAM jointly predicts actions and compact PRISM targets learned from the same visual features via inverse dynamics and reconstruction. (b) MiniWAM remains competitive with substantially larger WAMs and VLAs on LIBERO-Plus and RoboTwin 2.0. (c) In controlled comparisons across WAN2.1-VAE and DINOv3 feature spaces, PRISM simultaneously improves the success rate and reduces training time relative to native future-feature prediction.

Two stages.
One compact representation.

PRISM icon

PRISM — Predictive Representations via Inverse Spatiotemporal Modeling — learns a compact future dynamics representation that improves the performance–efficiency trade-off.

01

Learn the predictive representation

Use current–future observations to learn a compact transition representation through inverse dynamics and feature reconstruction.

Privileged future observations · Stage 1 only
02

Learn the world-action policy

Freeze the PRISM encoder. Jointly predict compact future targets and robot actions.

Compact prediction targets · Computationally efficient

Two-stage training pipeline of PRISM and MiniWAM. In Stage 1, privileged current–future visual features are compressed by the PRISM encoder into a compact predictive representation Lt, shaped through inverse action prediction and feature reconstruction. The corresponding world tokens remain clean while the action tokens are corrupted for conditional action prediction. In Stage 2, the PRISM encoder is frozen and provides compact world-modeling targets; MiniWAM jointly denoises the world and action tokens using its world and action experts. Solid and hatched tokens denote clean and noised tokens, respectively. Language and proprioceptive conditioning are omitted for clarity. Future observations are not needed at deployment.

Compact in size.
Competitive in simulation.

The 0.25B MiniWAM policy is evaluated across diverse manipulation benchmarks and is competitive with much larger models.

LIBERO97.2%

40 tasks · 2,000 episodes

LIBERO-PLUS73.6%

10,030 perturbed tasks

ROBOTWIN 2.046.1%

Clean / Randomized mean · Trained on clean split

Compared with state-of-the-art policies

Success rate (%) of MiniWAM and substantially larger VLAs and WAMs.

Bar length: reported model size

Method Scale Embodied
pretraining
Backbone
init.
LIBERO LIBERO-Plus RoboTwin 2.0
VLA Vision-language-action models
OpenVLA-OFT7B YesVLM96.671.4Not reported
π03.3B YesVLM93.570.531.4
π0-FAST†3.3B YesVLM85.561.6Not reported
UniVLA†7B YesVLM95.242.9Not reported
StarVLA-OFT†4B NoVLM96.6Not reported24.8
Xiaomi-Robotics-0†4.7B YesVLM98.7Not reported40.5
π0.53.3B YesVLM96.485.758.4
WAM World-action models
WorldVLA†7B NoVLM79.125.0Not reported
Fast-WAM6B NoVGM95.871.039.9
DreamWAM6B NoVGM96.975.4Not reported
AHA-WAM†7.2B NoVGMNot reportedNot reported33.8
MiniWAM Ours0.25B¹ NoNone97.273.646.1

† Externally reported. Other baselines are released checkpoints that we evaluate on LIBERO and LIBERO-Plus with the same task configurations and random seeds as MiniWAM. RoboTwin 2.0 reports the official leaderboard mean of clean and randomized evaluation, with training on clean demonstrations only. Scale is the reported model size. VLM: vision-language model; VGM: video generation model; –: not reported.

What should a world-action model predict?

Controlled comparison on the fixed LIBERO-Plus-2K subset.

Without co-training
Native future features
Compact PRISM targets

From a language instruction
to a sequence of actions.

MiniWAM is evaluated on the LIBERO, LIBERO-Plus, and RoboTwin 2.0 benchmarks. Below are successful policy rollouts.

A smaller representation.
Neighbors closer in behavior.

What does PRISM keep? Retrieve transitions from different tasks, then compare their actions and observed motion. Similar behavior should remain close even when the scene changes.

Main + wrist views average · 882 queries per view · top-5 neighbors

WATCH THE RETRIEVED MOTION

Main-view clips. One nearest neighbor per representation.

Loading the recorded nearest neighbors…

Leave-task-out, not leave-scene-out. Every neighbor is from a different task and an episode-disjoint gallery; scenes and objects may still repeat. Each clip shows all 17 retained frames from t to t+16, played slowly at 6 frames/s for inspection. These are recorded LIBERO-90 demonstrations, not policy predictions.

How to read this comparison

Behavior, measured independently of retrieval. Neighbors are selected by cosine distance in representation space. We then measure their normalized 16-step action RMSE, end-effector displacement-vector error and gripper-sign agreement against the query. Lower errors and higher agreement indicate closer behavior; matching the task label alone does not.

A shared, held-out gallery. The quantitative analysis uses 882 locked queries and an episode-disjoint gallery of 3,618 transitions. Candidates from the query's task are excluded. The archived native readout spatially averages each frame’s features and keeps the temporal sequence; PRISM uses its full-grid posterior mean (3 × 4 × 4 tokens, 64 dimensions). These are the reconstruction-weight-1 analysis encoders. This comparison does not test every possible native-feature readout. Both representations see the observed future. This tests representation geometry, not online policy prediction.

Aggregation. Each card is the equal-weight mean of main- and wrist-view metrics for the selected backbone, computed after retrieval separately in each view. The paper table instead averages the two backbones and reports each view separately. The video examples use the main camera only.

Evidence, with limits. The cards use the complete quantitative cohort; the three clips are selected illustrations. These averages do not imply improvement in every view or example, establish scene invariance, or explain policy success by themselves.

Download the displayed measurements and neighbor identities ↗

Start with a checkpoint.
Make the result your own.

We will open-source the code, configs, and checkpoints before November 2026. Stay tuned.

Public code & model links coming soon
THE REPRODUCTION PATH
  1. 01
    Install & verify

    Pinned policy and simulator environments.

  2. 02
    Download & evaluate

    A named checkpoint, its statistics, and a fixed protocol.

  3. 03
    Extract & train

    PRISM representation learning, then world-action training.

Citation

@article{chen2026miniwam,
  title = {MiniWAM: Learning Compact Future Targets for Efficient World-Action Modeling},
  author = {Chen, Jie and Bai, Ruofei and Cai, Yuxin and Zhang, Yifeng and
            He, Chengyang and Li, Jun and Yau, Wei-Yun and Sartoretti, Guillaume},
  journal = {arXiv preprint arXiv:2610.12194},
  year = {2026}
}

1 Parameter counts exclude frozen visual and text encoders.

2 65× refers to the WAN2.1-VAE native-to-PRISM future-token ratio at visual horizon 16 (3,136 vs. 48 tokens per camera).

3 Maximum estimated total training speedup at RoboTwin horizon 32, accounting for Stage 1 under the paper’s equal per-update-time assumption. Excludes offline feature extraction.

Expanded MiniWAM paper figure