ROCKET: Residual-Oriented Multi-Layer Alignment for Spatially-Aware Vision-Language-Action Models

Published in arXiv preprint arXiv:2602.17951, 2026

A multi-layer alignment framework that enhances spatial awareness in Vision-Language-Action models through residual-oriented techniques.