Introduces the AntiGrounding framework that uses feasibility-filtered short trajectories as both executable plans and rendered prompts for VLM-based natural-language robot manipulation.

Topological visualization of AntiGrounding: Executable Robot Trajectories as Visual Prompts for VLM-Guided Manipulation
Brave API

AntiGrounding is a robotic manipulation framework that lifts executable robot trajectories into the visual representation space of Vision-Language Models (VLMs) to enable decision-making without relying on compressed intermediate representations. By rendering feasibility-filtered short translational trajectories from multiple viewpoints, the system uses structured Visual Question Answering (VQA) to evaluate candidates for safety, task alignment, and physical plausibility. This approach allows the same trajectory to serve as an explicit geometric control object for execution and a visual prompt for multimodal evaluation, achieving 71.25% success in eight real-world tasks, significantly outperforming baselines like π0.5 (50.00%) and PIVOT-style (47.50%). The framework integrates a Real2Sim2Real pipeline with a closed-loop Model Predictive Control (MPC) architecture and includes an offline policy refinement module that leverages historical execution logs to adaptively optimize evaluation prompts.

Generated 14d ago
Open-Weights Reasoning

AntiGrounding targets a core weakness in VLM-guided robot manipulation: vision-language models can reason well about task semantics, but they are often poorly grounded in the robot’s kinematics, workspace constraints, and physically executable action space. The paper introduces a framework in which short candidate robot trajectories are generated, filtered for feasibility, and then used in a dual role: as directly executable low-level plans and as rendered visual prompts for a VLM. By converting candidate motions into visual evidence, the system lets the VLM reason over concrete, physically plausible actions rather than only over natural-language descriptions or static scene observations.

A key insight is that trajectory generation can serve as an interface between high-level semantic understanding and low-level robot control. Feasibility filtering removes invalid or unsafe candidates before they are presented to the VLM, while the VLM can use the rendered trajectory to make task-level judgments grounded in the robot’s actual capabilities. This design reduces the mismatch between what a language model asks for and what the robot can execute, and it makes the decision process more interpretable because the VLM is conditioned on explicit motion evidence rather than abstract text alone.

The work matters because it offers a practical way to improve the reliability of VLM-driven manipulation without requiring the VLM itself to predict precise robot commands. By combining trajectory generation, feasibility checking, and visual prompting, AntiGrounding can complement existing planners and motion generators, turning them into semantically interpretable action candidates. In manipulation settings where pure language grounding is brittle and pure trajectory optimization lacks task-level reasoning, this approach helps align VLM understanding with executable robot behavior.

Generated 14d ago
Sources