Identifies the action-sufficiency gap in vision-language robot policies and shows it can be bridged by guiding intermediate features to preserve three control-relevant structures.

Topological visualization of GIFT: Guided Intermediate Feature Training via Action-Oriented Structural Supervision for Robotic Manipulation
Brave API

GIFT (Guided Intermediate Feature Training) addresses the action-sufficiency gap—the mismatch between visual richness and control utility in vision-language-action (VLA) and world-action models (WAMs)—by enforcing action-oriented structural supervision on intermediate features. This framework bridges the gap by guiding features to preserve three specific control-relevant structures: geometry (governing motion feasibility), affordance (encoding instruction-relevant entities and end-effector configurations), and goals (grounding instructions in task-relevant regions).

Implemented as an architecture-flexible framework, GIFT applies these structures as training-time constraints through geometry alignment, affordance prediction, and goal-region reconstruction without requiring auxiliary predictions as action inputs. It demonstrates consistent performance gains across diverse policy paradigms, including a semantics-centered VLA, a direct-action WAM, and an inverse-dynamics WAM, particularly excelling in articulated-object interaction and zero-shot transfer under spatial and visual perturbations.

Key outcomes include: Superior Performance: On the RoboCasa benchmark, GIFT variants outperformed counterparts by 8.4 to 12.6 points, with GIFT-WAM-Fast reaching 83.6% success rate. Robust Generalization: In zero-shot transfer to LIBERO-Plus, GIFT models improved success rates by up to 12.6 points over baselines like StarVLA-OFT and Fast-WAM. * Real-World Efficacy: GIFT-WAM-IDM achieved an 87.5% success rate in real-world experiments, significantly outperforming matched baselines in high-precision manipulation tasks.

Generated 29d ago
Open-Weights Reasoning

GIFT addresses a representation-level weakness in vision-language robot manipulation policies: even when a model can understand a scene semantically, its intermediate features may not be action-sufficient for producing reliable motor commands. The paper frames this as an action-sufficiency gap, in which features shaped by language grounding, object recognition, or scene understanding fail to preserve the control-critical structure needed for precise manipulation—such as spatial relations, object state, and action-relevant geometry.

Its proposed solution is Guided Intermediate Feature Training, which adds action-oriented structural supervision to the hidden representations of a vision-language-action policy rather than depending only on end-to-end action prediction. The method guides intermediate features to preserve three control-relevant structures across observations and action contexts, effectively aligning the learned representation with the invariants that downstream control requires. In doing so, GIFT treats control as a first-class representation-learning objective instead of an afterthought attached to a semantically pretrained backbone.

This matters because many modern manipulation policies inherit strong semantic priors from vision-language models yet still struggle with the fine-grained, geometry-sensitive demands of physical interaction. By intervening at the intermediate feature level, GIFT offers a more principled route to robust manipulation than simply scaling data or enlarging action heads. It suggests that robot policy representations should be explicitly constrained to remain control-aware, which can improve generalization, reliability, and sample efficiency in settings where small representational errors lead to large physical failures.

Generated 29d ago
Sources