SIFEI LIU / RESEARCH
Robot Brain · Programs, interaction, and learning

Better interaction.
Better experience.
More capable robots.

Two systems that connect model reasoning to physical action—and turn the outcome into the next improvement.

Explore each system: ASENA ↗Recova ↗
01 / ASENACode as Policy + RSI

ASENA: Self-Evolving Agents
for Embodied Navigation

Give a coding model an embodied environment it can act in, inspect, and improve from.

The key idea

The model writes a program. The environment returns physical evidence. The next attempt starts with better code and reusable skills.

The interface is the bridge.

Task → program → embodied interaction
High-level policyCoding model

Read observations
Write and repair programs

task + context + evidence
programs →← feedback
ASENA interface
Workspacecode · notes · reusable skills
Observeimages · geometry · robot state
Actlearned navigation · geometric motion · gestures
Inspecttool replies · errors · recorded artifacts
actions →← evidence
SimulationNavigation & EQATask outcomes and execution feedback
Physical extensionUnitree G1Sensor access, motion checks and operator approval

The interface pattern carries across platforms; the available tools and feedback differ. ASENA-VLN is a learned navigation tool inside this system.

↻

RSI updates the workspace. Execute → inspect → revise code and skills → reuse. The coding model, improver and navigation policy keep fixed weights during this loop.

ASENA / THE TRAINING CONNECTIONSame interaction layer, a different update loop

From parallel task execution
to a possible RL rollout layer.

Workers already execute embodied tasks and return outcomes. For RL, keep that interaction structure and change what consumes the experience.

Fixed coding modelCoding model serves independent worker contextsEach worker has its own task context.
Task allocation · private workspace per worker
01Worker
Private code + skills
↕
Task / environment interaction
02Worker
Private code + skills
↕
Task / environment interaction
…16Worker
Private code + skills
↕
Task / environment interaction
Gather traces, local edits and task outcomes between passes
Fixed improver modelRevise the canonical workspace

Distribute the updated notes and skills to the next pass.

↶Wₖ₊₁ = Improve(Wₖ, traces, feedback)

16 workers in the reported simulation campaigns. This shows the logical worker–environment relationship; the paper does not specify one GPU or one simulator process per worker.

1

Host the trainable model

Replace the API backend with a rollout sampler backed by a hosted VLM; distribute serving across GPUs or machines as needed.

2

Simplify the workspace

Use an episode-scoped execution workspace. Cross-pass skill evolution can be removed; containers are one implementation option.

3

Connect the learner

Export training trajectories and rewards, track policy versions, and refresh sampler weights. Keep episode reset and isolation explicit.

72 → 98%R2R success
65 → 89%RxR success

With ASENA-VLN. Evidence for workspace RSI: 10 passes on recurring 100-task sets, ground-truth simulator feedback and replay allowed. These are not RL training results.

Connecting the environment to RL training

The paper establishes parallel embodied execution and workspace improvement, not a completed RL trainer. A hosted sampler must preserve the multimodal conversation, generated code, tool responses, action boundaries, termination and reward signals. Algorithms that need token log-probabilities must collect them consistently with the sampled policy version.

A multi-turn rollout normally collects experience through the external environment; policy optimization then happens in the learner. It does not backpropagate through arbitrary Python tools, simulator calls or a physical robot. Workspace persistence is a design choice, not a requirement of RL.

GPU placement, simulator allocation, reset behavior, throughput and policy synchronization still require implementation and validation. Private workspaces alone do not establish a security sandbox.

02 / RECOVAKeep collection running

Recova: Agent-guided failure recovery
for autonomous robotic manipulation

Turn a failure that stops data collection into a skill that keeps the next rollout going.

The goal

Robots that keep collecting useful experience, with progressively less human intervention.

Real → sim → real → learning.

Select a stage to follow the loop
→ → →
← New failures and real demonstrations guide the next recovery skill and policy update
01 / Physical grounding

A twin tied to the real workstation.

Reconstruct the scene so task attempts and failure recovery can be explored before returning to hardware.

Output: a reconstructed environment for task and recovery exploration.

Recorded digital-twin demo with the real camera inset.
Why simulation matters

The agent can diagnose, write a correction, test it, and try again. Successful programs become reusable recovery procedures and generate training data.

Why real feedback matters

Contact and visual mismatch still need physical validation. New failure cases and human corrections expand what the system can handle next.

RECOVA / THE DATA ENGINERecovery closes the collection loop

Recover the scene.
Keep the experience.

RUNTask policy
→
MONITORFailure?
→
RESTORERecovery policy
→
VERIFYResume the task
When recovery is missing or unsuccessfulHuman demonstration → save the correction → expand recovery coverage
Task data

Successful task rollouts
+ human task demonstrations

↓
DAgger update → task policy
Recovery data

Verified recoveries
+ human recovery demonstrations

↓
Separate training → recovery policy

Twin trajectories initialize separate policies. Real DAgger updates happen between collection rounds; policy weights remain fixed within a round.

Parallel real-world collection1 operator
4 workstations

Monitor rollouts, coordinate recovery, request help when needed, and save experience to the appropriate data pool.

Mean real-robot task success
Initial task policy23.8%
After DAgger77.5%
+ recovery87.5%
4 tasks · 20 trials per task and configuration
↻

Autonomy improves by expanding recovery coverage. In one four-round tile-drawing study, human takeovers fell from 7 of 8 episodes to 0 of 7. Fully unattended, open-ended data collection remains the goal.

Does the robot invent a new skill online?

In the reported real-robot experiments, autonomous recoveries use the learned recovery policy. The coding agent develops and tests corrective programs in the digital twin; a new real failure can trigger a recovery instruction and a first human demonstration, followed by twin development, training and registration of the new skill.

The demonstrated loop is targeted failure recovery and separate policy improvement. A general agent that autonomously identifies every missing data regime, invents new collection tasks and trains indefinitely is a natural next direction, not an established result here.

03 / WATCH THE LOOPobservation → decision → code → action → feedback

The next action depends
on what just happened.

Follow a real recording one turn at a time. Pause on a decision, inspect the program, then see the feedback that changes the next turn.

Check whether the machine has my snack. Report back silently.
Recorded G1 demonstration
Desk → vending machine → person00:00
Code as PolicyTURN 01 / 07

OBSERVATION

DECISION

PROGRAMReadable pseudocode
Paused · advance at your own pace
State carried forward

Replay sources and timing

Open the complete demo clip ↗

The adjacent programs and decision summaries explain the control flow; they are not verbatim private reasoning or raw execution logs. Each turn plays its selected footage and holds briefly for reading. Step controls let you inspect it without speeding up the recording.

ASENA

Build the interface that lets a model learn from embodied interaction.

Recova

Build the recovery loop that keeps embodied experience coming.