SIFEI LIU / RESEARCH
01 / ASENACode as Policy + RSI

ASENA: Self-Evolving Agents
for Embodied Navigation

Give a coding model an embodied environment it can act in, inspect, and improve from.

The key idea

The model writes a program. The environment returns physical evidence. The next attempt starts with better code and reusable skills.

The interface is the bridge.

Task → program → embodied interaction
High-level policyCoding model

Read observations
Write and repair programs

task + context + evidence
programs →← feedback
ASENA interface
Workspacecode · notes · reusable skills
Observeimages · geometry · robot state
Actlearned navigation · geometric motion · gestures
Inspecttool replies · errors · recorded artifacts
actions →← evidence
SimulationNavigation & EQATask outcomes and execution feedback
Physical extensionUnitree G1Sensor access, motion checks and operator approval

The interface pattern carries across platforms; the available tools and feedback differ. ASENA-VLN is a learned navigation tool inside this system.

↻

RSI updates the workspace. Execute → inspect → revise code and skills → reuse. The coding model, improver and navigation policy keep fixed weights during this loop.

ASENA / THE TRAINING CONNECTIONSame interaction layer, a different update loop

From parallel task execution
to a possible RL rollout layer.

Workers already execute embodied tasks and return outcomes. For RL, keep that interaction structure and change what consumes the experience.

Fixed coding modelCoding model serves independent worker contextsEach worker has its own task context.
Task allocation · private workspace per worker
01Worker
Private code + skills
↕
Task / environment interaction
02Worker
Private code + skills
↕
Task / environment interaction
…16Worker
Private code + skills
↕
Task / environment interaction
Gather traces, local edits and task outcomes between passes
Fixed improver modelRevise the canonical workspace

Distribute the updated notes and skills to the next pass.

↶Wₖ₊₁ = Improve(Wₖ, traces, feedback)

16 workers in the reported simulation campaigns. This shows the logical worker–environment relationship; the paper does not specify one GPU or one simulator process per worker.

1

Host the trainable model

Replace the API backend with a rollout sampler backed by a hosted VLM; distribute serving across GPUs or machines as needed.

2

Simplify the workspace

Use an episode-scoped execution workspace. Cross-pass skill evolution can be removed; containers are one implementation option.

3

Connect the learner

Export training trajectories and rewards, track policy versions, and refresh sampler weights. Keep episode reset and isolation explicit.

72 → 98%R2R success
65 → 89%RxR success

With ASENA-VLN. Evidence for workspace RSI: 10 passes on recurring 100-task sets, ground-truth simulator feedback and replay allowed. These are not RL training results.

Connecting the environment to RL training

The paper establishes parallel embodied execution and workspace improvement, not a completed RL trainer. A hosted sampler must preserve the multimodal conversation, generated code, tool responses, action boundaries, termination and reward signals. Algorithms that need token log-probabilities must collect them consistently with the sampled policy version.

A multi-turn rollout normally collects experience through the external environment; policy optimization then happens in the learner. It does not backpropagate through arbitrary Python tools, simulator calls or a physical robot. Workspace persistence is a design choice, not a requirement of RL.

GPU placement, simulator allocation, reset behavior, throughput and policy synchronization still require implementation and validation. Private workspaces alone do not establish a security sandbox.

03 / WATCH THE LOOPobservation → decision → code → action → feedback

The next action depends
on what just happened.

Follow a real recording one turn at a time. Pause on a decision, inspect the program, then see the feedback that changes the next turn.

Check whether the machine has my snack. Report back silently.
Recorded G1 demonstration
Desk → vending machine → person00:00
Code as PolicyTURN 01 / 07

OBSERVATION

DECISION

PROGRAMReadable pseudocode
Paused · advance at your own pace
State carried forward

Replay sources and timing

Open the complete demo clip ↗

The adjacent programs and decision summaries explain the control flow; they are not verbatim private reasoning or raw execution logs. Each turn plays its selected footage and holds briefly for reading. Step controls let you inspect it without speeding up the recording.

ASENA

Build the interface that lets a model learn from embodied interaction.

Recova

Build the recovery loop that keeps embodied experience coming.