ASENA presentation script
Sifei Liu · Nine slides · English draft
1,087 spoken words + 50 seconds of visual pauses.
Measured from synthesized narration using macOS Samantha: 8:00 speech plus planned pauses, including a 10-second real-robot replay. This is a pacing guide, not a recording of the presenter. At 120–140 spoken words per minute, the same script and pauses take approximately 8:36–9:54. Questions are additional.| Slide | Topic | Elapsed time |
|---|---|---|
| 01 | ASENA | 0:00–0:33 |
| 02 | The agent composes, the robot executes | 0:33–1:39 |
| 03 | Parallel workers evolve a shared workspace | 1:39–2:56 |
| 04 | The sandbox limits what generated code can do | 2:56–3:58 |
| 05 | The interaction layer can support an agent-training loop | 3:58–4:36 |
| 06 | ASENA-VLN provides reusable navigation | 4:36–5:32 |
| 07 | The navigation tool reduces coding-agent interaction | 5:32–6:32 |
| 08 | Success grows as the workspace accumulates experience | 6:32–7:41 |
| 09 | Real runs leave records we can inspect and reuse | 7:41–8:50 |
ASENA
I’ll present ASENA: Self-Evolving Agents for Embodied Navigation. We built a system that lets a coding agent turn a request into programs, execute them through robot interfaces, and inspect what happened. The persistent workspace is central: useful code, skills and experience survive beyond one attempt. During the workspace evolution I’ll show today, all model weights stay fixed. The improvements come from what the system retains and reuses.
Sources and supporting context
The agent composes, the robot executes
Let’s start with what the system actually does. A task and current observations enter the coding agent. The agent reads its workspace, decides what information it needs, and writes a program that combines the available tools. The task program can change as new evidence arrives.
There are two ways to produce motion here. The program can call our learned navigation policy with a goal and recent camera observations. Or it can compute geometric waypoints and construct other motion proposals, such as an upper-body gesture. These tools can be combined within one task.
The proposal then crosses an execution boundary. Independent checks and operator approval control what reaches the robot. SONIC handles whole-body execution on G1. New observations, outcomes and feedback return to the agent.
That return path matters: the agent can inspect a failed assumption, revise its program, and retain the useful procedure in its workspace.
Sources and supporting context
Parallel workers evolve a shared workspace
We used the same workspace idea to run repeated tasks in simulation. At the beginning of a pass, a canonical workspace seeds sixteen private worker copies. Each worker receives its own task context and executes independently, using code, skills and the available simulator tools.
As workers encounter problems, they inspect feedback and revise their local programs. They can also inspect peer work and exchange tips through a shared bulletin. So parallelism gives us several streams of experience, with some sharing during execution.
The important step is consolidation. Between passes, we gather the traces, task outcomes, notes and local edits. A fixed improver merges useful changes, tests them and prunes the workspace. The next pass starts from that consolidated version. We keep this shared revision as the next reproducible starting point.
What changes is the code and retained experience. The coding model, improver and navigation policy keep their weights fixed. This gives us a concrete way to study whether accumulated experience improves subsequent attempts.
Sources and supporting context
The sandbox limits what generated code can do
The workspace contains generated code, so we also need to define its authority. The restricted process boundary covers the agent and its child processes. It exposes scoped resources and APIs while blocking direct control paths, protected credentials and unauthorized platform writes.
That isolates code execution. Motion checking is a separate layer. A program proposes a movement; external validation checks the applicable bounds, and an operator approves physical execution before SONIC carries it out. These controls remain outside the generated program.
For generated gestures, our MuJoCo checks use the robot’s self model. They check properties such as joint limits and self-collision, without simulating the complete surrounding world. Supervision therefore remains part of deployment. During reflection, motion is disabled while the agent inspects records and revises its skills.
Sources and supporting context
The interaction layer can support an agent-training loop
This structure also gives us a specific connection to agent training. The implemented loop updates the workspace: execute, collect feedback, consolidate code and skills.
A proposed training loop would reuse the interaction layer, but would need a trainable sampler, explicit rollout, reward and termination records, and a learner that publishes updated checkpoints. Its update target would be the coding agent’s weights.
That comparison is a proposed extension. The ASENA results here come from fixed-weight workspace evolution.
Sources and supporting context
ASENA-VLN provides reusable navigation
ASENA-VLN gives this system a reusable navigation tool. It is a four-billion-parameter model, separate from the coding agent. The agent sends an instruction and recent RGB history; the tool returns body-frame waypoints, and the agent checks progress.
Evaluated alone, ASENA-VLN reaches 68.7 percent success on R2R and 70.2 percent on RxR. Compared with the reported monocular Qwen-RobotNav results, navigation error falls on both datasets. R2R also gains success and path efficiency; RxR gains trajectory fidelity, with slightly lower success and efficiency.
This gives us a concrete division of work: the coding agent composes the task, while the specialized model turns a navigation request into a short motion trajectory.
Sources and supporting context
The navigation tool reduces coding-agent interaction
Next, we asked what happens when a coding agent can call that navigation tool. These are fresh-workspace comparisons on the same one hundred RxR tasks.
For Claude Sonnet 5, average tool calls fall from about 104 to 39 per episode, while success rises from 50 to 61 percent. For GPT-6 Astra, calls fall from about 41 to 20, while success decreases from 93 to 86 percent.
So both measured agents use fewer tool interactions: reductions of 62.5 and 51 percent. Only Sonnet gains success in this comparison. A policy call can absorb several navigation decisions, but the combined result depends on the agent and how it uses the tool.
These counts measure agent–tool calls. They should not be read as robot steps or elapsed execution time.
Sources and supporting context
Success grows as the workspace accumulates experience
Now we can look at what changed over repeated passes. Sixteen Claude Sonnet 5 workers ran at high effort, with one hundred tasks per benchmark. Each pass reused the same scenes, instructions and starting states, with all model weights fixed.
The solid curves include the navigation policy; the dashed curves use the agent alone. On R2R, success with the policy rises from 72 to 98 percent, and the agent-only curve reaches 97 percent. On RxR, the corresponding results rise from 65 to 89 percent, and from 48 to 80 percent.
Useful route and scene memory are part of what the workers can retain and replay. A repaired program or tested skill can also become part of the next workspace. Consolidation makes that feedback available to subsequent attempts. Across these recurring tasks, capability improved as the workspace accumulated experience, while the models themselves stayed fixed.
Sources and supporting context
Real runs leave records we can inspect and reuse
Finally, here is the physical evidence the system leaves behind. This MCAP replay aligns a real camera stream, twenty-nine measured joint channels and the robot’s body pose on one clock. We can inspect what the robot saw and how it actually moved.
Beyond these synchronized sensor streams, execution records also contain agent programs, tool results, operator decisions and accepted voice text. That lets us connect a proposed action to permission, execution and its observed outcome.
We already use these records for safety inspection and skill revision. After curation, they could also support policy supervised fine-tuning or coding-agent training. A reinforcement-learning extension would use simulation.
This connects to Recova, where recovery experience is used to train manipulation policies. ASENA supplies inspectable programs, interactions and real robot experience for that broader learning loop.