Shows that visuomotor policies such as ACT fail under visually similar distractors because they lack phase- and state-dependent visual grounding required for reliable manipulation.
Yes, the paper "What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies" (Chavan et al., 2026) explicitly demonstrates that visuomotor imitation policies, specifically Action Chunking with Transformers (ACT), fail under visually similar distractors because they lack conditional visual grounding.
The study defines this failure as a problem where the visual target required for successful control changes with the manipulation phase (e.g., object vs. receptacle) and, in more complex tasks, with the observed task state. Key findings include:
This material examines a key failure mode of visuomotor imitation policies, using action-chunking transformers (ACT) and similar architectures as a case study. Such policies are often trained to map visual observations directly to action sequences, and they can perform well in controlled manipulation settings, but the paper argues that they become unreliable when the scene contains visually similar distractors. The central claim is that the problem is not merely poor visual discrimination, but insufficient conditional visual grounding: the policy does not sufficiently condition its visual attention and action selection on task phase, robot state, object identity, contact state, or other manipulation-relevant variables. As a result, the same visual evidence may be interpreted correctly in one part of a task but incorrectly in another, especially when distractors make the visual scene ambiguous.
The contribution is diagnostic and constructive. It frames conditional visual grounding as a first-class requirement for robust visuomotor control, rather than treating it as an implicit byproduct of end-to-end imitation. The paper likely introduces or uses controlled evaluation scenarios in which visually similar objects, phase changes, or state-dependent cues stress-test a policy’s ability to ground actions in the right context. The key insight is that reliable manipulation requires the policy to know what matters, when: which visual features are task-relevant, which objects are affordances, and how the correct interpretation depends on the current stage of the task and the embodied state of the robot. Improvements are motivated by explicitly conditioning the policy on phase and state, strengthening grounding signals, or otherwise making the visual-to-action mapping more temporally and contextually aware.
This matters because real-world manipulation is rarely clean or uniquely identifiable. Robots deployed in cluttered environments must distinguish target objects from visually similar distractors, respect the order of task phases, and adapt their behavior to changing contact and proprioceptive conditions. The work therefore speaks to a broader limitation of imitation-based visuomotor and vision-language-action policies: strong performance on narrow benchmarks can mask brittle grounding assumptions. By isolating conditional visual grounding as a diagnostic axis, the paper provides a useful lens for evaluating and improving embodied policies, with implications for safer, more generalizable robotic manipulation and for designing benchmarks that test not only whether a policy can perceive a scene, but whether it can interpret it correctly at the right time.