RECOVA Research talk · Sifei Liu
Real robots · recovery · learning

Recova

Agent-Guided Failure Recovery for Autonomous Robotic Manipulation

Sifei Liu

Recovery keeps the robot working.
Experience makes the policies better.

Real-world mahjong demonstration
01 / 08
The problem

Recovery keeps collection moving.

A capable task policy still stops at an unexpected state.

Continuous collection

What happens
after a failed grasp?

Restore a useful state.
Resume the task.
Keep the experience.

Success alone is not enough to sustain a data collection loop.
01Real workstationImages + robot trajectories
→
02Digital twinGround the scene
→
03Code explorationTry, inspect, revise
↓
06Deploy + collectRecover or ask an operator
←
05Learned policiesTask policy + recovery policy
←
04Successful rolloutsSeparate task / recovery data
↶Physical experience → retrain between rounds → redeploy

Measured as separate collection and DAgger fine-tuning rounds.

Method overview · Code is developed in simulation; learned policies execute on the robot.
02 / 08
Real → sim → programs and data

Develop corrections in a digital twin.

Real camera 2×
MuJoCo replay 2×

Replay → compare views → refine poses and contact

The same recorded robot trajectory, aligned to real camera observations.

Clear the grasp corridor 4×
Schematic recovery program
observe(scene)
move_aside(blocking_ring)
verify(grasp_is_clear)
resume(stacking)
Retain the strategyReusable recovery programs
+
Retain the successful experienceTask trajectories → πθ   ·   Recovery trajectories → ρφ
Matched replay and twin recovery from the original presentation. Program text is illustrative pseudocode.
03 / 08
Real-world execution

Learned policies act; a monitor coordinates.

Camera observations → VLM monitor
Task can proceedTask policy πθLearned motor actions
Recoverable failureRecovery policy ρφInstruction-conditioned correction
Needs assistanceOperator fallbackCorrection / demonstration
↓Verify the recovered state↶ Resume the task
Task experience → task datasetRecovery experience → recovery dataset

Policy weights stay fixed within a collection round.

Between rounds   Human correction → train ρφ → register its instruction
Learned physical recovery 2×

Restore the scene,
then restart the task.

Reported deployment: learned task and recovery policies; VLM monitoring and human fallback.
04 / 08
Real-world DAgger

One operator. Four workstations.

Monitor, recover, hand off when needed — and save the next rollout.

Recovery keeps collection moving; saved experience trains the next policies.

Full recorded panel · five synchronized views · recording at its original 10× playback. Click a station to inspect its recorded decision.
05 / 08
Measured learning loop

Less intervention over collection rounds.

Mahjong draw · task and recovery policies updated between rounds

Task-policy completionHuman takeover
100755025012.5%57.1%75%85.7%87.5%28.6%25%0%Round 1Round 2Round 3Round 48 episodes7 episodes8 episodes7 episodes
Recorded four-round collection study; episode counts shown. Policy updates occur between rounds.
06 / 08
Physical evaluation

Collected experience improves the policy.

Base policy+ real-world DAgger+ recovery
100500Success (%)203525157075858085909085Pencil boxStack ringsDraw tileDiscard tile

Task policy initialized from π0.5 · 20 trials per task and configuration

Four physical tasks, same evaluation protocol. Physical recovery demonstrations: 2×. Simulation benchmarks are separate evaluations.
07 / 08
The mental model

A twin turns failures
into reusable corrections.

Development toolGround a twin

Reconstruct enough of the workstation
to develop and test a correction.

↓
Reusable productsPrograms + learned policies

Keep the strategy and its successful trajectories.

↓
Measured loopDeploy → collect → fine-tune → redeploy

Task execution and recovery improve together.

Keep operating. Keep collecting.
Future direction

Limited Code-as-Policy fallback on the real robot, with validation and operator oversight.

A proposed extension, not a reported result.
Recova · Sifei Liu
08 / 08
Recova · Eight slides