Back to Discover

RTP revises visual plans under feedback instead of restarting from noise

RTP Figure 1: visual-plan update modes comparing fresh replanning, binary retain/fresh, and learned revision bridge with RoboMME success bars

Beihang and Peking authors posted Revisable Temporal Planning on arXiv on Sept. 28, 2026 (2609.35439), listed on the Tue Sept. 29 cs.RO recent feed. RTP keeps a visual future as a persistent action condition and revises it after execution feedback through a learned revision bridge. Authors report RoboMME 48.6% and RMBench 84.8%.

Closed-loop world-action models just got a fresher revision primary. On Sept. 28, 2026, a Beihang and Peking University team submitted Revisable Temporal Planning, or RTP (arXiv:2609.35439), and the paper appeared on the Tuesday Sept. 29 cs.RO recent list. Authors cast it as a way to keep a predicted visual future as a persistent action condition, then revise that future after execution feedback instead of always regenerating from noise.

In plain terms, the revision bridge is the method beat. During visual generation the model saves intermediate denoising states. After feedback, a learned bridge resumes one of those states and adapts the continuation to current observations, with visual and action supervision tying the revised future to the next control. Time-aware history supplies observed evidence. An adaptive policy then chooses retention, bridge revision, or fresh replanning from new noise before decoding the next action.

Authors instantiate RTP on task-adapted LingBot-VA and evaluate on RoboMME and RMBench. On the paper scoreboard they report task-averaged success of 48.6% on RoboMME and 84.8% on RMBench, the highest among methods compared in those tables. Matched time-aware LingBot-VA fresh replanning sits lower by 8.2 and 4.9 percentage points on the displayed rates. All of those figures are author-reported. Independent reproduction outside the authors' protocol is not yet available in the materials reviewed for this article.

The abs lists a project page, but that URL was still a PLACEHOLDER host returning 404 at dig time. Live demos and project visuals are therefore not claimed here. Lead art uses the paper's public Figure 1 overview from arXiv HTML.

This card is a closed-loop visual-plan revision primary, method-distinct from Rolling-WAM's rolling video-action denoising line, from InternW0-Delta's Causal Imprint Mixture-of-Transformers corpus work, and from AD-WM's JEPA action-discrimination line. RoboFL's federated expert assembly is a separate cs.RO neighbor, not this method. Revision and scoreboard claims stay author-reported until outside labs reproduce them.