Back to Discover

InternW0-Delta pairs Causal Imprint with a 20K-hour open WAM corpus

InternW0-Delta architecture: Mixture-of-Transformers world-action model with Causal Imprint queries feeding the action expert

InternRobotics posted InternW0-Delta on arXiv on Sept. 25, 2026 (2609.31394), listed on the Mon Sept. 28 cs.RO recent feed. The Mixture-of-Transformers world-action model adds Causal Imprint predictive tokens and a claimed open corpus of more than 20K hours. Authors report LIBERO-Plus 92.8% and RoboTwin Clean2Clean 90.0%.

Open world-action models just got a fresher primary. On Sept. 25, 2026, an InternRobotics team submitted InternW0-Delta (arXiv:2609.31394), and the paper appeared on the Monday Sept. 28 cs.RO recent list with a project page updated the same morning. Authors cast it as a unified World Action Model that folds pretrained visual dynamics, scene semantics, 4D geometric and motion priors, and action generation into one Mixture-of-Transformers stack under frozen-VLM guidance.

In plain terms, Causal Imprint is the method beat. During training, the model learns future-relevant scene changes from future supervision only. At inference it feeds those predictive tokens straight to the action expert, without rolling out future video. A pretrained video expert and an action expert interact under frozen VLM semantic guidance, while a 4D foundation model injects geometric and motion priors through training-only distillation.

For data scale, authors say they curated robot demonstrations, UMI, egocentric human, and Ego2Robot streams into a common state-action representation totaling more than 20K hours of processed training data. They call that, to their knowledge, the largest open-source corpus of its kind, and they pledge to open-source code, weights, infrastructure, the data pipeline, and processed data where licenses permit. Those scale lines are author claims.

On the project-page scoreboard, authors report LIBERO-Plus success of 92.8%, RoboTwin Clean2Clean 90.0%, RoboTwin Clean2Random 71.9%, RoboDojo 23.9, and EBench 49.2%. All of those figures are author-reported from the paper and project page. Independent reproduction outside the authors' protocol is not yet available in the materials reviewed for this article.

Hugging Face Papers indexes the same arXiv record and showed the paper on the Sept. 28 Daily list. This card is a fresher open-WAM primary than Rolling-WAM's rolling video-action denoising line and is distinct from AD-WM's JEPA action-discrimination work and from the prior InternW0 physical world-model paper. Method and corpus claims stay author-reported until outside labs reproduce them.