SIFEI LIU / RESEARCH

Spatial VLMs / An interactive guide

Ground the region.
Reason about the world.

Spatial reasoning needs more than recognizing objects. It needs a precise referent, a useful geometric representation, and a way to check the evidence behind an answer.

First, make “this object” unambiguous.

A box or mask identifies exactly which region a question refers to. SpatialRGPT combines region features with language; an optional depth branch adds geometric evidence that RGB alone can leave ambiguous.

Region prompts + optional depth
Schematic of two region-prompted objects, with an optional relative depth map Region 1Region 2Which region is closer to the camera?

Schematic: region prompts and optional relative-depth evidence.

The region is part of the input.

RGB imageShared visual encoder
Region mask / boxSelect the referent
Region pooling → RGB connectorRegion tokens enter the language sequence
Global + region tokens → language modelSpatial relationships and quantitative estimates

Region tokens resolve which objects to compare. The RGB-only model already learns spatial relationships from region-aware supervision.

The training signal matters. An offline pipeline turns RGB images into estimated metric 3D scene graphs, then into region-aware spatial questions and answers.
Ground the question

Boxes and masks distinguish similar objects without relying on a long verbal description.

Learn geometry at scale

Open Spatial Dataset contains 8.7M spatial concepts grounded in 5M regions from 1M images.

Use depth when available

A separate connector lets the model use depth evidence while retaining an RGB-only inference path.

Inside the model: data geometry versus inference inputs

The data pipeline uses detection, segmentation, metric depth, and camera calibration to construct spatial supervision. Those offline 3D scene graphs are not passed into the VLM at inference.

The released inference path accepts RGB, region masks, and optional relative depth. It pools region features and inserts their embeddings at region-token positions. Quantitative answers are learned estimates, not guaranteed geometric measurements.

Original SpatialRGPT paper architecture with RGB and optional depth region features
Original SpatialRGPT architecture. Enlarge ↗

Give regions a place in 3D.

Appearance alone does not say where a region sits in a scene. SR-3D fuses visual features with encoded 3D position, using a common representation design to transfer spatial knowledge from single images to multiple views.

One representation design, two input settings
Schematic of image appearance and normalized 3D positions entering a joint region representationView AA relative-depth position map for one image

Schematic of positional evidence. Encoded positions are normalized; they are not raw metric coordinates.

Start with scalable single-view data.

AppearanceVisual patch features
3D positionRelative-depth-derived positions
Fuse visual + encoded position featuresRetain the spatial evidence at each location
Region / scene tokens → language modelPool selected regions; project scene features

Pretrain with large-scale image data. In the released single-view code, normalized image coordinates and relative depth provide positional features.

The bridge is a representation. Single-view spatial pretraining supplies useful priors for subsequent multi-view training.
Bind appearance to position

A region representation carries both what it looks like and where it lies in the input geometry.

Extend region prompts

Mark an object in selected frames with a box or mask, or supply a 3D box, then reason across views.

Transfer, then specialize

The release provides separate single-view, multi-view scan, and multi-view video checkpoints.

What geometry does the released code actually use?

Single-view preprocessing combines normalized image coordinates with normalized relative depth. For static scenes, the scan path back-projects depth with camera intrinsics, aligns points across views, and normalizes scene coordinates before positional encoding. The paper also uses external geometry models to supply aligned point maps.

These routes share a representation design, not identical geometry or one universal checkpoint. Geometry-aware multi-view inference needs aligned point maps; a plain RGB-video call is not equivalent.

Original SR-3D architecture ↗ · Position construction in code ↗

POSITION FEATURES / TEACHING NOTATIONSource ↗
# Single-view position construction
q = (u_normalized * z_normalized,
     v_normalized * z_normalized,
     z_normalized)

# Fuse encoded position with appearance
features = visual_features + position_encoding(q)
Here u, v, and z lie in [−1, 1]; z comes from relative depth. This is simplified teaching notation, not metric reconstruction.

Bring visual evidence back into reasoning.

GR3D extends region-aware modeling with explicit 2D grounding, implicit region insertion during generation, and monocular 3D grounding. A generated object reference can bring fresh region features into the sequence before reasoning continues.

Ground in 2D → insert region evidence → infer in 3DSee the paper’s examples ↗
Step through locating a chair in 2D, inserting its visual features, and predicting a 3D box chair region Identify the object before predicting its geometry.

Schematic: resolve a region, insert its features, then predict a 3D box.

Resolve the object in image space.

Predict a 2D region for the object being discussed. This gives subsequent inference an object-specific visual reference.

object mention2D regioncontinue

1 / 3 · Detect before lifting

Grounding becomes part of generation. The model can reintroduce visual evidence at the moment it reasons about an entity.
Explicit 2D grounding

Localize an object when the task directly asks for a region.

Implicit 2D grounding

Insert region features into the generation stream as the model refers to an entity.

Monocular 3D grounding

Use object-specific evidence to lift a 2D region into a 3D bounding-box prediction.

Why detect, then lift?

The paper compares predicting a 3D box directly with first grounding the target in 2D. The intermediate region anchors the 3D prediction to the right object and lets the model benefit from abundant 2D grounding supervision.

An auxiliary task predicts 3D point coordinates from region prompts, providing geometric supervision beyond sparse 3D-box labels.

Spatial pretraining freezes the vision encoder; grounded-CoT fine-tuning then updates only the language model. Training uses teacher-forced boxes and features pooled from ground-truth regions; inference uses predicted boxes.

Original GR3D paper diagram of explicit 2D, implicit 2D, and monocular 3D grounding
Original GR3D architecture. Enlarge ↗
SpatialRGPTIdentify the region and learn spatial relationships.
SR-3DBind region appearance to 3D position.
GR3DRefresh grounded evidence during generation.

Turn spatial reasoning into a program.

SpatialClaw adds a different route: keep the VLM fixed, let it write Python, and feed execution results back into the next turn. Like ASENA, code is the action interface; here, the actions inspect and transform spatial evidence.

Write codeChoose an analysis
ExecutePerception + geometry
InspectArrays, plots, errors
↳ Revise the next cell using the observed result

Masks, point maps, and intermediate calculations stay in the same Python session across turns. The agent can combine them into a question-specific computation rather than relying on a prewritten tool for every spatial relationship.

Parallel execution, shared model services
Session 1Session 2Session 3…
VLM inference pool
Main agent + visual queries
Perception service pool
SAM3 + reconstruction

Independent Python state per task. GPU models are shared across tasks; concurrency is configurable.

MULTI-TURN GEOMETRY / TEACHING EXAMPLEActual API ↗
recon = tools.Reconstruct.Reconstruct(InputImages)
show(recon.render_bev())
Feedback available to the next turnInspect the reconstructed scene before deciding how to compare the objects.
Published APIs composed into a small teaching example.
EXTENSION / MULTI-TURN RL
A foundation for multi-turn RL sampling.

Parallel sessions, executable feedback, and answer scoring provide reusable building blocks. A training extension would add complete rollout trajectories, optimizer updates, and synchronized policy weights. The published system is training-free.

EXTENSION / RETAINED SKILLS
Connect to the ASENA direction.

Keeping and evaluating useful programs across tasks could enable an RSI loop. The current SpatialClaw runtime resets task state between questions; cross-task skill evolution would be an added mechanism.