DualWAM

Dual-System World Action Models for Asynchronous
Global Planning and Local Refinement

Yixin Zheng1,2,*, Jiangran Lyu2,3,*, Yuntian Deng2,4,*, Kai Liu1,2,
Yizhou Zhou5, Yizhou Wang3, Xiaoguang Zhao1, He Wang2,3,†, Zhizheng Zhang2,†
1Institute of Automation, Chinese Academy of Sciences 2Galbot 3Peking University 4Shanghai Jiao Tong University 5Individual Researcher

*Equal contribution †Co-corresponding authors

DualWAM overview showing role-specialized System 2 global planning and System 1 local refinement
Overview. DualWAM asynchronously coordinates a low-frequency System 2 for long-horizon world-action planning and a high-frequency wrist-conditioned System 1 for local refinement. Their role-specialized observation spaces enable learning from egocentric, robot, and UMI data. Across zero-shot tasks on Galbot G1 and Franka FR3, DualWAM improves success rate and reduces inference latency.

01 / METHOD

Model Architecture

DualWAM pipeline from System 2 high-noise bidirectional generation to System 1 low-noise causal refinement
System 2 denoises a complete long-horizon world-action state through the high-noise interval using bidirectional attention. At threshold τ, temporally aligned wrist-action windows are passed to compact System 1, which completes the low-noise interval using causal attention and fresh wrist observations.

02 / METHOD

Asynchronous Inference Pipeline

System 2 refreshes a global plan in the background while System 1 repeatedly denoises one temporally aligned action chunk.

03 / EVALUATION

Results on Real-World Tasks

Zero-shot Real-World Tasks

Performance overviewSuccess rate · completion time · latency
Zero-shot real-robot success rate, completion time, and critical-path latency results
Zero-shot real-robot performance and critical-path latency. DualWAM displays strong generalization ability in real-world zero-shot tasks, achieving a relatively high success rate with significantly lower latency than all baseline families. Right: Critical-path latency on a logarithmic scale. Cosmos3-Nano-Policy uses the same applicable system optimizations, such as multi-GPU classifier-free guidance and quantization; other baselines use their official inference scripts and default sampling configurations.

Performance Demonstration

More zero-shot tasks

Clean the laptop with a brush.
1×
Elevate the cube to the highest platform.
1×
Hook the cup onto the rack.
1×
Toss the burger in the pan.
1×
Untie the ribbon on the gift.
1×
Remove the hat from the hook.
1×
Close the laptop.
1×
Open the lower drawer of the cooker.
1×
Hit the block with the hammer.
1×
Put the toy pig into the corresponding plate based on its color.
1×

Better zero-shot generalization

Orient the mug with its handle pointing to the pot. π0.5 ✗
1×
Orient the mug with its handle pointing to the pot. DualWAM ✓
1×

Faster response

Extract the straw from the cup. Cosmos-Nano ✗
1×
Extract the straw from the cup. DualWAM ✓
1×

Fine manipulation via wrist-guided refinement

Open the top drawer. Cosmos-Nano ✗
2×
Open the top drawer. DualWAM ✓
2×

Adaptive interaction strategies under unexpected conditions

Leveraging the model’s strong generalization capabilities, our policy can explore multiple interaction approaches to address the sudden problem in exceptional cases.

Extract the straw from the cup. DualWAM ✓
2×
Reveal the object under the green cup. DualWAM ✓
2×

04 / CITATION

BibTeX

@misc{zheng2026dualwamdualsystemworldaction,
  title   = {DualWAM: Dual-System World Action Models for
             Asynchronous Global Planning and Local Refinement},
  author  = {Yixin Zheng and Jiangran Lyu and Yuntian Deng and
             Kai Liu and Yizhou Zhou and Yizhou Wang and
             Xiaoguang Zhao and He Wang and Zhizheng Zhang},
  year    = {2026},
  eprint  = {2609.24868},
  archivePrefix = {arXiv},
  primaryClass = {cs.RO},
  url     = {https://arxiv.org/abs/2609.24868},
}