Recognizing is not executing
VLA policies can parse the instruction and identify the object, yet fail when success depends on local geometry, clearance, insertion, stable release, or transparent-object boundaries.
Accepted at CoRL 2026
We're delighted to share that GeoAlign has been accepted to CoRL 2026.
Congratulations to the team!
1Tongji University 2Shanghai Innovation Institute 3Shanghai Jiao Tong University 4Zhejiang University 5Jingdezhen Ceramic University 6Tsinghua University 7HONOR 8University of Science and Technology of China
GeoAlign uses robot state to select action-relevant geometry from RGB observations, connecting semantic understanding with precise manipulation.
Semantic grounding is not sufficient for
executable manipulation.
VLA policies can parse the instruction and identify the object, yet fail when success depends on local geometry, clearance, insertion, stable release, or transparent-object boundaries.
Measured depth can be unreliable around transparent, reflective, thin, or annular objects, while explicit 3D inputs add sensing and calibration assumptions.
GeoAlign lets proprioceptive state query an RGB-derived geometry grid, selecting compact spatial evidence for the current action phase.
Current Vision-Language-Action (VLA) models often optimize for semantic grounding, whereas executable manipulation requires geometry-aware spatial alignment and dynamic affordance selection. We introduce GeoAlign, a state-guided spatial alignment architecture for VLA policy learning. GeoAlign post-trains an RGB geometry branch with robot-domain RGB-D supervision, yielding RGB-derived Geometry-Enhanced Post-Trained (GEP) features for policy rollout. The robot's proprioceptive state queries the GEP feature grid, producing compact, phase-dependent geometry tokens for action prediction. GeoAlign achieves 99.0% on LIBERO, 85.3% across three SimplerEnv-Fractal tasks, and 78.8% on eight geometry-critical real-world ALOHA tasks, with ablations confirming the value of geometry post-training and proprioceptive-state-guided querying.
Method
GeoAlign uses robot-domain RGB-D supervision during offline geometry post-training, but policy rollout uses RGB observations, language, and proprioception rather than measured depth.
Depth supervision adapts the encoder-side representation to robot-workspace geometry while avoiding raw depth inputs at deployment.
Proprioceptive queries attend over the dense image-space GEP feature grid, producing compact geometry tokens for the action decoder.
The selected geometry context is appended to RGB-language tokens for a flow-matching DiT action head based on the Isaac-GR00T N1.6-3B backbone.
Results
GeoAlign is evaluated on LIBERO, SimplerEnv-Fractal, and eight geometry-critical real-world ALOHA tasks.
Depth-Supervised VLA
GeoAlign uses depth supervision to improve vision-language-action (VLA) policies. Robot-domain RGB-D pairs post-train a Depth Anything V2 geometry branch, producing RGB-derived GEP features. During policy rollout, the depth prediction head is discarded: the policy takes RGB images, language, and proprioceptive state, without measured depth, point clouds, or predicted depth maps as inputs.
Attention Behavior
During transparent-container manipulation, geometry-query attention highlights the gripper, container boundary, held object, and other local regions that matter for insertion and release.
@misc{chen2026geoalign,
title = {GeoAlign: Beyond Semantics with State-Guided Spatial Alignment in VLA Models},
author = {Yizhi Chen and Zhanxiang Cao and Xinyi Peng and Yixiao Zheng and Xiaxi Si and Yiheng Li and Liyun Yan and Keqi Zhu and Xueyun Chen and Shengcheng Fu and Tianyue Zhan and Yufei Jia and Jinming Yao and Yan Xie and Kun Wang and Cewu Lu and Yue Gao},
year = {2026},
eprint = {2606.03240},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
doi = {10.48550/arXiv.2606.03240},
url = {https://arxiv.org/abs/2606.03240},
note = {Accepted at CoRL 2026}
}