Presents the One Word, Different Action benchmark that jointly measures robot decision invariance and sensitivity to task-preserving vs. task-changing instructions.

Topological visualization of One Word, Different Action: A Real-Robot Benchmark for Language-Conditioned Embodied Reasoning
Brave API

The One Word, Different Action benchmark is not present in the provided search context. The available results focus on other benchmarks such as VLABench (language-conditioned manipulation with long-horizon reasoning), RoboSemanticBench (diagnosing semantic grounding in VLA models), and RoboBench (evaluating multimodal large language models as embodied brains), but do not mention the specific benchmark described in your query.

Generated 28d ago
Open-Weights Reasoning

The paper introduces One Word, Different Action, a real-robot benchmark for evaluating language-conditioned embodied reasoning. Rather than testing only whether a robot can execute a command, the benchmark probes whether the robot’s decision is grounded in the intended task rather than in superficial linguistic cues. It does this by contrasting instructions that differ by only a single word: some variations are task-preserving, meaning they rephrase the instruction without changing the required behavior, while others are task-changing, meaning the altered word materially changes the intended action.

A key contribution is the benchmark’s focus on two complementary capabilities: decision invariance and language sensitivity. A robust policy should remain stable across task-preserving paraphrases, while still shifting its action appropriately when a minimal lexical change alters the task. This framing helps expose common failure modes in language-conditioned robot learning, such as overfitting to surface wording, ignoring semantically important instruction changes, or relying on spurious cues rather than true task understanding.

This matters because real-world robot instruction following requires more than nominal task success. As vision-language models and vision-language-action policies are increasingly deployed in embodied settings, their behavior must be both robust to irrelevant linguistic variation and responsive to meaningful changes in intent. By providing a controlled, real-robot evaluation of these properties, the benchmark offers a useful diagnostic for developing safer, more reliable, and more semantically grounded language-conditioned robot agents.

Generated 28d ago