WAMJET: A Harness for World Action Model Acceleration

1 Max Planck Institute for Intelligent Systems2 The Chinese University of Hong Kong3 Johannes Kepler University Linz
† Corresponding author

Abstract

World Action Models (WAMs) leverage pretrained video foundation models for robot manipulation, but their large backbones and video-action co-prediction are expensive. Although existing acceleration techniques offer many ways to reduce this cost, selecting and composing them requires substantial engineering for each model and hardware platform. To tackle this bottleneck, we present WAMJET, an agentic harness that accelerates WAM inference by equipping coding agents with reusable optimization guidance and measurement and validation tools. WAMJET follows a bottleneck-driven workflow where the agent profiles inference, modifies targeted code, validates effects, and iteratively refines the acceleration stack as bottlenecks shift, while preserving action quality. Experiments span six WAMs, three coding agents, and two GPU architectures. WAMJET achieves up to 9.95× lossless speedup over upstream implementations. Approximation and hardware-aware optimization yield additional latency reductions, with comparable success rates. The results show that WAMJET can produce effective acceleration stacks for WAM deployment.

Why WAMJET

Existing WAM acceleration techniques offer many ways to reduce inference cost, but hand-crafting them for each model and hardware platform requires substantial engineering and may not generalize.

WAMJET addresses this bottleneck with a harness that equips coding agents to develop acceleration strategies tailored to different WAMs and GPU architectures.

Manual optimization coordinates domain experts. A coding agent alone may stop at generic optimizations. WAMJET equips the agent with reusable guidance and tools tailored to WAMs and GPU architectures.

Workflow

WAMJET's bottleneck-driven optimization workflow. The agent reduces startup costs, profiles inference to identify bottlenecks, explores lossless acceleration followed by approximate techniques when permitted, and validates candidates. The agent retains accepted changes, reassesses the remaining bottlenecks, and continues searching within the specified budget.

WAMJET workflow: fast startup, profile bottlenecks, lossless acceleration, permitted approximate acceleration, then validate and retain. Re-profile and repeat within the budget.

Optimizing the sources of latency

Choose a latency source to see where time is spent and how to reduce it.

Lossless
Schematic CPU/GPU pipeline
CPUGPU computeData movementWaiting

Python pseudocode
Before
After

Lossless Acceleration across Different Agents and WAMs

Our first set of experiments evaluates WAMJET's effectiveness for lossless acceleration and its generality across different combinations of coding agents and WAMs on H100.

UpstreamWithout WAMJETWAMJET lossless

Speedup = upstream latency / candidate latency.

WAMJET effectively accelerates inference across different combinations of coding agents and WAMs. All 12 final WAMJET configurations outperform their corresponding no-WAMJET baselines. For LingBot-VA, WAMJET achieves a significant 9.95× speedup with GPT-6-astra, suggesting greater benefits from more advanced coding agents.

More WAMJET optimization rounds lead to better optimization. Even though the first round has already reduced latency quite a lot, a second WAMJET optimization round delivers further speedups across all 12 configurations. Click to see the gains.

Lossless optimization breakdown GPT-6-astra · H100

Total speedup is relative to upstream. Step gain compares with the previous displayed configuration.

Approximate Acceleration across GPU Architectures

Our second set of experiments demonstrates WAMJET's support for architecture-aware optimization.

B200 · Inference latency

UpstreamWAMJET losslessWAMJET approximate
WAMJET supports hardware-aware optimization. The WAMJET-guided agent selects and refines quantization strategies through profiling and empirical search, adapting them to the target model and GPU architecture.

We further evaluate task success on a RoboLab subset and find that WAMJET can preserve action quality and task performance for approximate acceleration, indicating the effectiveness of our iterative validation workflow.

Task success on RoboLab

20 tasks with 10 episodes per task.

DreamZero
Configuration Success rate
Upstream 25.0%
WAMJET lossless 24.0%
WAMJET approximateFP8 26.0%
WAMJET approximateNVFP4 + FP8 25.0%
Cosmos3-Nano-Policy-DROID
Configuration Success rate
Upstream 43.0%
WAMJET lossless 47.0%
WAMJET approximateFP8 45.5%
WAMJET approximateMXFP8 44.5%

Further Optimization of Recent SOTA WAM

Wang et al. recently released OpenWAM, which achieves SOTA performance across several robotic benchmarks. OpenWAM already incorporates several inference optimizations, such as the PyTorch compiler and velocity cache, making it a useful test of whether WAMJET can guide an agent to find further optimization opportunities.

We apply WAMJET-guided GPT-6-astra to optimize OpenWAM on B200, and evaluate on full LIBERO and full LIBERO-plus.

Configuration Latency Speedup LIBERO LIBERO-plus
UpstreamBF16 87.25 ms 1.00× 98.95% 69.27%
WAMJET losslessBF16 32.69 ms 2.67× 99.10% 69.24%
WAMJET approximateFP8 31.70 ms 2.75× 99.25% 69.30%
The WAMJET-guided agent can identify substantial acceleration opportunities even in an already highly optimized inference pipeline.

Citation

BibTeX
@article{chen2026wamjet,
  title={WAMJET: A Harness for World Action Model Acceleration},
  author={Chen, Le and Liu, Lixin and Schneider, Jan and Qiu, Zeju and Guist, Simon and Sch{\"o}lkopf, Bernhard and B{\"u}chler, Dieter},
  journal={arXiv preprint arXiv:2610.03797},
  year={2026}
}

References

  1. S. Ye et al. World action models are zero-shot policies. arXiv:2602.15922, 2026.
  2. T. Yuan, Z. Dong, Y. Liu, and H. Zhao. Fast-WAM: Do world action models need test-time future imagination? arXiv:2603.16666, 2026.
  3. L. Li et al. Causal world modeling for robot control. arXiv:2601.21998, 2026.
  4. M. J. Kim et al. Cosmos Policy: Fine-tuning video models for visuomotor control and planning. arXiv:2601.16163, 2026.
  5. N. Agarwal et al. Cosmos 3: Omnimodal world models for physical AI. arXiv:2606.02800, 2026.
  6. X. Yang et al. RoboLab: A high-fidelity simulation benchmark for analysis of task generalist policies. Proceedings of Robotics: Science and Systems, Sydney, Australia, July 2026.
  7. Y. Wang et al. OpenWAM: An open, modular exploration towards systematic world-action model pretraining. arXiv:2609.07398, 2026.
  8. B. Liu et al. LIBERO: Benchmarking knowledge transfer for lifelong robot learning. Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, 2023.
  9. S. Fei et al. LIBERO-Plus: In-depth Robustness analysis of vision-language-action models. arXiv:2510.13626, 2025.