World-action models just got a cheaper imagination path on paper. Authors submitted Sparse-WAM on Sept. 30, 2026 (arXiv:2609.38984), a training-free framework that speeds robot world-action inference by pruning future-frame tokens using action relevance. The stakes are inference cost: dense future frames make joint visual-plus-action denoising expensive even when many pixels never change the control decision.
In plain terms, a world-action model, or WAM, uses a pretrained video model to jointly predict future visual states and actions for robot control. Prior pruning work often kept tokens for visual fidelity alone. Sparse-WAM instead asks which future-frame regions the action tokens attend to, then keeps those action-relevant regions plus cross-frame context.
Authors say action-to-future attention maps overlap a lot between consecutive denoising steps even as future representations keep updating. That observation motivates Action-Guided Token Selection. A naive prune can spend the savings on scoring and packing overhead, so they also introduce Pilot, an engine for lightweight scoring and cross-step reuse of token selections.
On LIBERO with FastWAM-Joint and on RoboLab-120 with Cosmos 3 Edge, authors report inference speedups of about 2.0x and 1.8x over dense eager inference on an NVIDIA RTX 4090, while largely preserving task performance. Those figures are author-reported preprint claims, not peer-reviewed third-party re-runs.
Sparse-WAM is a Sept. 30 Submitted efficiency primary in the world-action lane. It is distinct from InternW0-Delta's Causal Imprint corpus story, from Rolling-WAM's rolling video-action line, and from prior OpenWAM tips. Discover lead InternW0-Delta stays untouched. Method and speedup claims remain author-reported until outside labs reproduce them.