Recova: Agent-guided failure recovery
for autonomous robotic manipulation
Turn a failure that stops data collection into a skill that keeps the next rollout going.
Robots that keep collecting useful experience, with progressively less human intervention.
Real → sim → real → learning.
Select a stage to follow the loopA twin tied to the real workstation.
Reconstruct the scene so task attempts and failure recovery can be explored before returning to hardware.
Output: a reconstructed environment for task and recovery exploration.
The agent can diagnose, write a correction, test it, and try again. Successful programs become reusable recovery procedures and generate training data.
Contact and visual mismatch still need physical validation. New failure cases and human corrections expand what the system can handle next.
Recover the scene.
Keep the experience.
Successful task rollouts
+ human task demonstrations
Verified recoveries
+ human recovery demonstrations
Twin trajectories initialize separate policies. Real DAgger updates happen between collection rounds; policy weights remain fixed within a round.
4 workstations
Monitor rollouts, coordinate recovery, request help when needed, and save experience to the appropriate data pool.
Autonomy improves by expanding recovery coverage. In one four-round tile-drawing study, human takeovers fell from 7 of 8 episodes to 0 of 7. Fully unattended, open-ended data collection remains the goal.
Does the robot invent a new skill online?
In the reported real-robot experiments, autonomous recoveries use the learned recovery policy. The coding agent develops and tests corrective programs in the digital twin; a new real failure can trigger a recovery instruction and a first human demonstration, followed by twin development, training and registration of the new skill.
The demonstrated loop is targeted failure recovery and separate policy improvement. A general agent that autonomously identifies every missing data regime, invents new collection tasks and trains indefinitely is a natural next direction, not an established result here.
The next action depends
on what just happened.
Follow a real recording one turn at a time. Pause on a decision, inspect the program, then see the feedback that changes the next turn.
Replay sources and timing
The adjacent programs and decision summaries explain the control flow; they are not verbatim private reasoning or raw execution logs. Each turn plays its selected footage and holds briefly for reading. Step controls let you inspect it without speeding up the recording.
Build the interface that lets a model learn from embodied interaction.
Build the recovery loop that keeps embodied experience coming.