ASENA
Self-Evolving Agents
for Embodied Navigation
Sifei Liu
The workspace evolves.
The model weights stay fixed.
The agent composes,
the robot executes
What changed?RGB · LiDAR · body state
the next computation
π Learned VLN
Instruction + RGB history
→ body-frame waypoints
Programs + geometry
Geometric waypoints
or upper-body joint targets
Policy evaluation: simulation. Physical extension: supervised G1 execution.
02Parallel workers evolve
a shared workspace
Canonical workspace Wₖ
Reusable code, skills and experience
Sim task → local edits
Sim task → local edits
Peer tips during exploration
Sim task → local edits
Reseed the next pass with Wₖ₊₁
The coding agent, improver and navigation policy all keep fixed weights.
03The sandbox limits what
generated code can do
Worker workspace
Restricted process + every child
No direct raw-control paths, operator credentials
or writes to the trusted platform.
MuJoCo gesture checks use the robot’s self model, not full environment physics. Deployment remains supervised.
04The interaction layer can support
an agent-training loop
Workspace RSI
Execution + sim interaction
→ traces and feedback
Reuse the
interaction layer
Agent training
- Trainable sampler
- Rollout + reward / termination
- Learner + checkpoints
A training interface extension; no evaluated RL result is claimed here.
05ASENA-VLN provides
reusable navigation
Coding agent
ASENA-VLN
Instruction + RGB history
→ body-frame waypoints
R2R
Policy only| Model | SR ↑ | NE ↓ | SPL ↑ |
|---|---|---|---|
| QwenRobotNav 4B | 66.9 | 4.22 | 60.5 |
| ASENA-VLN | 68.7 | 3.52 | 64.2 |
RxR
Policy only| Model | SR ↑ | NE ↓ | nDTW ↑ | SPL ↑ |
|---|---|---|---|---|
| QwenRobotNav 4B | 71.3 | 4.15 | 68.6 | 61.5 |
| ASENA-VLN | 70.2 | 3.90 | 73.1 | 59.7 |
RxR improves navigation error and path fidelity;
success rate and SPL are slightly lower.
SR, SPL, nDTW: %. NE: meters. ASENA paper, policy-only evaluation.
06The navigation tool reduces
coding-agent interaction
Claude
Sonnet 5
62.5% fewer callsGPT-6
Astra
51.0% fewer callsFewer tool calls in both measured cases; different success tradeoffs.
RxR Agentic Split · 100 tasks / configuration · fresh workspaces. Calls are agent–tool interactions, not robot steps or elapsed time.
07Success grows as the workspace
accumulates experience
16 workers · fixed weights · recurring 100-task sets
Same scenes, instructions and starts; memory and replay allowed.
Measured workspace evolution, not held-out generalization or a formal scaling law.
08Real runs leave records
we can inspect and reuse
Beyond motion
Agent programs
Tool results
Operator approval
Accepted voice text
Policy SFT
or coding-agent training
Inspection is already available.
Recova uses recovery experience to train policies.
Additional event channels are recorded across runs; they are not all represented in this camera/body excerpt.
09