ASENA / Sifei Liu
Code as Policy + recursive self-improvement

ASENA

Self-Evolving Agents
for Embodied Navigation

Sifei Liu

The workspace evolves.
The model weights stay fixed.

G1 · recorded demonstrations
01

The agent composes,
the robot executes

Task + observationsWhat is the goal?
What changed?
RGB · LiDAR · body state
Coding agentChoose and compose
the next computation
Persistent workspaceCode · skills · notes↔ Read, test, revise
Delegate navigation

π Learned VLN

Instruction + RGB history
→ body-frame waypoints

and/or
Construct motion

Programs + geometry

Geometric waypoints
or upper-body joint targets

Recorded robot view
Platform authorityValidate→Operator approval→SONIC / G1

Policy evaluation: simulation. Physical extension: supervised G1 execution.

02

Parallel workers evolve
a shared workspace

Pass k

Canonical workspace Wₖ

Reusable code, skills and experience

Seed private copies
Worker 01Private workspace

Sim task → local edits

Worker 02Private workspace

Sim task → local edits

…Independent contexts

Peer tips during exploration

Worker 16Private workspace

Sim task → local edits

Gather traces, task outcomes and local changes
Between passes
Merge→Test→Prune

Reseed the next pass with Wₖ₊₁

The coding agent, improver and navigation policy all keep fixed weights.

03

The sandbox limits what
generated code can do

Code execution boundary

Worker workspace

Restricted process + every child

Python / NumPy / MuJoCoScoped device + model APIsScratch files + reusable skills
Unix identityNetwork allowlistInherited seccomp

No direct raw-control paths, operator credentials
or writes to the trusted platform.

Physical execution path · outside the agent
Motion proposalWaypoints or joint trajectory
↓
Independent validationBounds · joint limits · self-collision
↓
Operator approvalApprove or return a reason
↓
SONIC → G1Balance / locomotion · independent stop

MuJoCo gesture checks use the robot’s self model, not full environment physics. Deployment remains supervised.

04
Proposed extension

The interaction layer can support
an agent-training loop

Implemented

Workspace RSI

Execution + sim interaction
→ traces and feedback

Update W
Fixed model weights
→

Reuse the
interaction layer

Proposed

Agent training

  1. Trainable sampler
  2. Rollout + reward / termination
  3. Learner + checkpoints
Update θ

A training interface extension; no evaluated RL result is claimed here.

05

ASENA-VLN provides
reusable navigation

Task-level reasoning

Coding agent

Separate learned tool · 4B

ASENA-VLN

Instruction + RGB history
→ body-frame waypoints

R2R

Policy only
ModelSR ↑NE ↓SPL ↑
QwenRobotNav 4B66.94.2260.5
ASENA-VLN68.73.5264.2

RxR

Policy only
ModelSR ↑NE ↓nDTW ↑SPL ↑
QwenRobotNav 4B71.34.1568.661.5
ASENA-VLN70.23.9073.159.7

RxR improves navigation error and path fidelity;
success rate and SPL are slightly lower.

SR, SPL, nDTW: %. NE: meters. ASENA paper, policy-only evaluation.

06

The navigation tool reduces
coding-agent interaction

Agent onlyWith ASENA-VLNTool calls / episode

Claude
Sonnet 5

62.5% fewer calls
103.83
38.91
Success rate50 → 61%+11 points

GPT-6
Astra

51.0% fewer calls
40.80
19.98
Success rate93 → 86%−7 points

Fewer tool calls in both measured cases; different success tradeoffs.

RxR Agentic Split · 100 tasks / configuration · fresh workspaces. Calls are agent–tool interactions, not robot steps or elapsed time.

07

Success grows as the workspace
accumulates experience

R2R + policyR2R agent onlyRxR + policyRxR agent only

16 workers · fixed weights · recurring 100-task sets
Same scenes, instructions and starts; memory and replay allowed.

Measured workspace evolution, not held-out generalization or a formal scaling law.

08

Real runs leave records
we can inspect and reuse

Real MCAP excerpt · camera, 29 joints + body pose

Beyond motion

Agent programs
Tool results
Operator approval
Accepted voice text

Curated records could support

Policy SFT
or coding-agent training

Optional RL belongs in simulation.

Inspection is already available.
Recova uses recovery experience to train policies.

Open the full recording ↗

Additional event channels are recorded across runs; they are not all represented in this camera/body excerpt.

09

ASENA

← → Navigate · O Overview · N Presenter · T Timer · F Full screen

Recorded execution

Robot demonstration