Abstract
World Action Models (WAMs) leverage pretrained video foundation models for robot manipulation, but their large backbones and video-action co-prediction are expensive. Although existing acceleration techniques offer many ways to reduce this cost, selecting and composing them requires substantial engineering for each model and hardware platform. To tackle this bottleneck, we present WAMJET, an agentic harness that accelerates WAM inference by equipping coding agents with reusable optimization guidance and measurement and validation tools. WAMJET follows a bottleneck-driven workflow where the agent profiles inference, modifies targeted code, validates effects, and iteratively refines the acceleration stack as bottlenecks shift, while preserving action quality. Experiments span six WAMs, three coding agents, and two GPU architectures. WAMJET achieves up to 9.95× lossless speedup over upstream implementations. Approximation and hardware-aware optimization yield additional latency reductions, with comparable success rates. The results show that WAMJET can produce effective acceleration stacks for WAM deployment.
Why WAMJET
Existing WAM acceleration techniques offer many ways to reduce inference cost, but hand-crafting them for each model and hardware platform requires substantial engineering and may not generalize.
WAMJET addresses this bottleneck with a harness that equips coding agents to develop acceleration strategies tailored to different WAMs and GPU architectures.
Workflow
WAMJET's bottleneck-driven optimization workflow. The agent reduces startup costs, profiles inference to identify bottlenecks, explores lossless acceleration followed by approximate techniques when permitted, and validates candidates. The agent retains accepted changes, reassesses the remaining bottlenecks, and continues searching within the specified budget.
Optimizing the sources of latency
Choose a latency source to see where time is spent and how to reduce it.
Lossless Acceleration across Different Agents and WAMs
Our first set of experiments evaluates WAMJET's effectiveness for lossless acceleration and its generality across different combinations of coding agents and WAMs on H100.
Speedup = upstream latency / candidate latency. denotes executing the WAMJET optimization for rounds. Our final reported results use the latency.
WAMJET effectively accelerates inference across different combinations of coding agents and WAMs. All 12 final WAMJET configurations outperform their corresponding no-WAMJET baselines. For LingBot-VA, WAMJET achieves a significant 9.95× speedup with GPT-6-astra, suggesting greater benefits from more advanced coding agents.
More WAMJET optimization rounds lead to better optimization. Even though the first round has already reduced latency quite a lot, a second WAMJET optimization round delivers further speedups across all 12 configurations. Click to see the gains.
Lossless optimization breakdown GPT-6-astra · H100
Total speedup is relative to upstream. Step gain compares with the previous displayed configuration.
Approximate Acceleration across GPU Architectures
Our second set of experiments demonstrates WAMJET's support for architecture-aware optimization.
B200 · Inference latency
We further evaluate task success on a RoboLab subset and find that WAMJET can preserve action quality and task performance for approximate acceleration, indicating the effectiveness of our iterative validation workflow.
Task success on RoboLab
20 tasks with 10 episodes per task.
| Configuration | Success rate |
|---|---|
| Upstream | 25.0% |
| WAMJET lossless | 24.0% |
| WAMJET approximateFP8 | 26.0% |
| WAMJET approximateNVFP4 + FP8 | 25.0% |
| Configuration | Success rate |
|---|---|
| Upstream | 43.0% |
| WAMJET lossless | 47.0% |
| WAMJET approximateFP8 | 45.5% |
| WAMJET approximateMXFP8 | 44.5% |
Further Optimization of Recent SOTA WAM
Wang et al. recently released OpenWAM, which achieves SOTA performance across several robotic benchmarks. OpenWAM already incorporates several inference optimizations, such as the PyTorch compiler and velocity cache, making it a useful test of whether WAMJET can guide an agent to find further optimization opportunities.
We apply WAMJET-guided GPT-6-astra to optimize OpenWAM on B200, and evaluate on full LIBERO and full LIBERO-plus.
| Configuration | Latency | Speedup | LIBERO | LIBERO-plus |
|---|---|---|---|---|
| UpstreamBF16 | 87.25 ms | 1.00× | 98.95% | 69.27% |
| WAMJET losslessBF16 | 32.69 ms | 2.67× | 99.10% | 69.24% |
| WAMJET approximateFP8 | 31.70 ms | 2.75× | 99.25% | 69.30% |
Optimization Progress over Search Time
We plot the optimization progress of GPT-6-astra on DreamZero and OpenWAM using B200. The figures show individual trials and accepted optimization trajectories, illustrating how performance improves through iterative exploration.
DreamZero · B200
Measured trials and accepted optimization trajectory
Citation
@article{chen2026wamjet,
title={WAMJET: A Harness for World Action Model Acceleration},
author={Chen, Le and Liu, Lixin and Schneider, Jan and Qiu, Zeju and Guist, Simon and Sch{\"o}lkopf, Bernhard and B{\"u}chler, Dieter},
journal={arXiv preprint arXiv:2610.03797},
year={2026}
}
References
- S. Ye et al. World action models are zero-shot policies. arXiv:2602.15922, 2026.
- T. Yuan, Z. Dong, Y. Liu, and H. Zhao. Fast-WAM: Do world action models need test-time future imagination? arXiv:2603.16666, 2026.
- L. Li et al. Causal world modeling for robot control. arXiv:2601.21998, 2026.
- M. J. Kim et al. Cosmos Policy: Fine-tuning video models for visuomotor control and planning. arXiv:2601.16163, 2026.
- N. Agarwal et al. Cosmos 3: Omnimodal world models for physical AI. arXiv:2606.02800, 2026.
- X. Yang et al. RoboLab: A high-fidelity simulation benchmark for analysis of task generalist policies. Proceedings of Robotics: Science and Systems, Sydney, Australia, July 2026.
- Y. Wang et al. OpenWAM: An open, modular exploration towards systematic world-action model pretraining. arXiv:2609.07398, 2026.
- B. Liu et al. LIBERO: Benchmarking knowledge transfer for lifelong robot learning. Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, 2023.
- S. Fei et al. LIBERO-Plus: In-depth Robustness analysis of vision-language-action models. arXiv:2510.13626, 2025.