Single-view preprocessing combines normalized image coordinates with normalized relative depth. For static scenes, the scan path back-projects depth with camera intrinsics, aligns points across views, and normalizes scene coordinates before positional encoding. The paper also uses external geometry models to supply aligned point maps.
These routes share a representation design, not identical geometry or one universal checkpoint. Geometry-aware multi-view inference needs aligned point maps; a plain RGB-video call is not equivalent.
Original SR-3D architecture ↗ · Position construction in code ↗
POSITION FEATURES / TEACHING NOTATIONSource ↗
q = (u_normalized * z_normalized,
v_normalized * z_normalized,
z_normalized)
features = visual_features + position_encoding(q)
Here u, v, and z lie in [−1, 1]; z comes from relative depth. This is simplified teaching notation, not metric reconstruction.