ASENA: Self-Evolving Agents
for Embodied Navigation
Give a coding model an embodied environment it can act in, inspect, and improve from.
The model writes a program. The environment returns physical evidence. The next attempt starts with better code and reusable skills.
The interface is the bridge.
Task → program → embodied interactionRead observations
Write and repair programs
task + context + evidenceThe interface pattern carries across platforms; the available tools and feedback differ. ASENA-VLN is a learned navigation tool inside this system.
RSI updates the workspace. Execute → inspect → revise code and skills → reuse. The coding model, improver and navigation policy keep fixed weights during this loop.
From parallel task execution
to a possible RL rollout layer.
Workers already execute embodied tasks and return outcomes. For RL, keep that interaction structure and change what consumes the experience.
Distribute the updated notes and skills to the next pass.
Wₖ₊₁ = Improve(Wₖ, traces, feedback)16 workers in the reported simulation campaigns. This shows the logical worker–environment relationship; the paper does not specify one GPU or one simulator process per worker.
Host the trainable model
Replace the API backend with a rollout sampler backed by a hosted VLM; distribute serving across GPUs or machines as needed.
Simplify the workspace
Use an episode-scoped execution workspace. Cross-pass skill evolution can be removed; containers are one implementation option.
Connect the learner
Export training trajectories and rewards, track policy versions, and refresh sampler weights. Keep episode reset and isolation explicit.
With ASENA-VLN. Evidence for workspace RSI: 10 passes on recurring 100-task sets, ground-truth simulator feedback and replay allowed. These are not RL training results.
Connecting the environment to RL training
The paper establishes parallel embodied execution and workspace improvement, not a completed RL trainer. A hosted sampler must preserve the multimodal conversation, generated code, tool responses, action boundaries, termination and reward signals. Algorithms that need token log-probabilities must collect them consistently with the sampled policy version.
A multi-turn rollout normally collects experience through the external environment; policy optimization then happens in the learner. It does not backpropagate through arbitrary Python tools, simulator calls or a physical robot. Workspace persistence is a design choice, not a requirement of RL.
GPU placement, simulator allocation, reset behavior, throughput and policy synchronization still require implementation and validation. Private workspaces alone do not establish a security sandbox.
The next action depends
on what just happened.
Follow a real recording one turn at a time. Pause on a decision, inspect the program, then see the feedback that changes the next turn.
Replay sources and timing
The adjacent programs and decision summaries explain the control flow; they are not verbatim private reasoning or raw execution logs. Each turn plays its selected footage and holds briefly for reading. Step controls let you inspect it without speeding up the recording.
Build the interface that lets a model learn from embodied interaction.
Build the recovery loop that keeps embodied experience coming.