An arXiv preprint keeps task memory in a separate agent and trains a vision based action model to follow plain language hints, lifting a hard simulated benchmark from 14.8% to 76.3%.
A new robot architecture called 2AM splits a robot's task memory from the module that actually moves its arm. On a hard simulated benchmark for multi-step manipulation, the design lifts average completion from 14.8% to 76.3% (arXiv 2609.11308).
The paper, posted to arXiv by Yutong Hu of INSAIT and KU Leuven with co-author Renaud Detry, frames the result as a critique of the monolithic vision-language-action default in robot learning. In 2AM, an Agent holds the running history of a task and compiles it into subtask language plus optional 2D grasp, place, and move hints. A separate Action Model, a single RGB-based policy with no memory of past steps, just executes each move. The system uses no depth sensor, no online geometry, and no planner-based object motion.
To make the split work, the team trained the Action Model to tolerate imperfect hints. They augmented demonstrations with structured hint labels and added condition dropout, spatial noise, and temporal jitter. On LIBERO-Mem, the 2AM setup also reported 63.0% relaxed success and 11.8% strict success.
The lever is the interface, not model size. 2AM is still a preprint, the result is on a single simulated suite, and the authors do not claim physical-robot generalization, but the split offers a real alternative to one giant policy trying to remember and act at the same time.