Accepted at CoRL 2026

We're delighted to share that GeoAlign has been accepted to CoRL 2026.
Congratulations to the team!

GeoAlign:Beyond Semantics with State-Guided
Spatial Alignment in VLA Models

Yizhi Chen1,2, Zhanxiang Cao3,2, Xinyi Peng1,2, Yixiao Zheng7, Xiaxi Si3,2, Yiheng Li3,2, Liyun Yan3,2, Keqi Zhu4,2, Xueyun Chen5, Shengcheng Fu1,2, Tianyue Zhan3,2, Yufei Jia6, Jinming Yao8, Yan Xie7, Kun Wang7, Cewu Lu3,2, Yue Gao3,2

1Tongji University   2Shanghai Innovation Institute   3Shanghai Jiao Tong University   4Zhejiang University   5Jingdezhen Ceramic University   6Tsinghua University   7HONOR   8University of Science and Technology of China

Tongji University Shanghai Innovation Institute Shanghai Jiao Tong University Zhejiang University Tsinghua University HONOR University of Science and Technology of China
Paper arXiv Video BibTeX Code coming soon
GeoAlign · Project video2 min 40 sec

GeoAlign uses robot state to select action-relevant geometry from RGB observations, connecting semantic understanding with precise manipulation.

Real-world experiments

Accelerated playback
Transparent-container manipulation
Cluttered tabletop

Semantic grounding is not sufficient for
executable manipulation.

Recognizing is not executing

VLA policies can parse the instruction and identify the object, yet fail when success depends on local geometry, clearance, insertion, stable release, or transparent-object boundaries.

Geometry is difficult to observe

Measured depth can be unreliable around transparent, reflective, thin, or annular objects, while explicit 3D inputs add sensing and calibration assumptions.

Robot state guides spatial attention

GeoAlign lets proprioceptive state query an RGB-derived geometry grid, selecting compact spatial evidence for the current action phase.

Abstract

Current Vision-Language-Action (VLA) models often optimize for semantic grounding, whereas executable manipulation requires geometry-aware spatial alignment and dynamic affordance selection. We introduce GeoAlign, a state-guided spatial alignment architecture for VLA policy learning. GeoAlign post-trains an RGB geometry branch with robot-domain RGB-D supervision, yielding RGB-derived Geometry-Enhanced Post-Trained (GEP) features for policy rollout. The robot's proprioceptive state queries the GEP feature grid, producing compact, phase-dependent geometry tokens for action prediction. GeoAlign achieves 99.0% on LIBERO, 85.3% across three SimplerEnv-Fractal tasks, and 78.8% on eight geometry-critical real-world ALOHA tasks, with ablations confirming the value of geometry post-training and proprioceptive-state-guided querying.

Method

RGB-derived geometry, selected by robot state

GeoAlign uses robot-domain RGB-D supervision during offline geometry post-training, but policy rollout uses RGB observations, language, and proprioception rather than measured depth.

GeoAlign method overview: RGB-D supervision post-trains a geometry branch; proprioceptive-state queries extract GEP geometry tokens for an Isaac-GR00T N1.6-3B DiT action head

Geometry-Enhanced Post-Trained feature

Depth supervision adapts the encoder-side representation to robot-workspace geometry while avoiding raw depth inputs at deployment.

State-guided spatial alignment

Proprioceptive queries attend over the dense image-space GEP feature grid, producing compact geometry tokens for the action decoder.

Executable action decoding

The selected geometry context is appended to RGB-language tokens for a flow-matching DiT action head based on the Isaac-GR00T N1.6-3B backbone.

Results

Simulation and real-world gains

GeoAlign is evaluated on LIBERO, SimplerEnv-Fractal, and eight geometry-critical real-world ALOHA tasks.

GeoAlign success rates: 99.0 percent on LIBERO, 85.3 percent on SimplerEnv-Fractal, and 78.8 percent on real-world ALOHA
99.0% LIBERO average success
85.3% SimplerEnv-Fractal average
78.8% Real-world ALOHA average
RGB observation, measured depth, and predicted depth diagnostic for RGB-D supervised GEP features

Depth-Supervised VLA

Learning geometry from depth, acting from RGB

GeoAlign uses depth supervision to improve vision-language-action (VLA) policies. Robot-domain RGB-D pairs post-train a Depth Anything V2 geometry branch, producing RGB-derived GEP features. During policy rollout, the depth prediction head is discarded: the policy takes RGB images, language, and proprioceptive state, without measured depth, point clouds, or predicted depth maps as inputs.

Attention Behavior

State-guided geometry attention focuses on execution regions

During transparent-container manipulation, geometry-query attention highlights the gripper, container boundary, held object, and other local regions that matter for insertion and release.

RGB observations and state-guided geometry attention during manipulation
RGB observations and geometry-query attention during manipulation. Attention maps are qualitative diagnostics.

BibTeX

@misc{chen2026geoalign,
  title         = {GeoAlign: Beyond Semantics with State-Guided Spatial Alignment in VLA Models},
  author        = {Yizhi Chen and Zhanxiang Cao and Xinyi Peng and Yixiao Zheng and Xiaxi Si and Yiheng Li and Liyun Yan and Keqi Zhu and Xueyun Chen and Shengcheng Fu and Tianyue Zhan and Yufei Jia and Jinming Yao and Yan Xie and Kun Wang and Cewu Lu and Yue Gao},
  year          = {2026},
  eprint        = {2606.03240},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  doi           = {10.48550/arXiv.2606.03240},
  url           = {https://arxiv.org/abs/2606.03240},
  note          = {Accepted at CoRL 2026}
}