Introduces the RideWay benchmark and Efficiency Utility metric that quantify and penalize inefficient yet successful behaviors in stateful ride-hailing agent tool use.

Topological visualization of RideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents
Brave API

RideWay is a Chinese-language benchmark for evaluating tool-using language agents in stateful ridehailing environments, focusing on efficiency rather than just success. It introduces Efficiency Utility, a metric that penalizes successful trajectories for excess user-facing turns and backend tool calls relative to task-specific reference floors.

Key characteristics include: Dual-Axis Penalty: It separately models user-facing turns and backend tool calls, with human-calibrated penalties reflecting that extra dialogue is roughly twice as costly as extra tool use. Success-Gated Scoring: Failed tasks receive zero credit; successful ones are discounted only if they exceed reference-effort thresholds derived from human annotations. Validation: The metric achieves 78.7% accuracy on held-out human preferences, outperforming single-axis metrics like tool-only or turn-only evaluations. Scope: Evaluated on 58 tasks across 24 models, highlighting that successful completion is necessary but insufficient for high-quality agent performance.

Generated 17d ago
Open-Weights Reasoning

RideWay addresses a key limitation in evaluations of tool-using language agents: many benchmarks treat task completion as a binary outcome while ignoring how much effort, latency, or tool usage the agent requires to succeed. The paper introduces a benchmark centered on a stateful ride-hailing environment, where agents must issue tool calls to manage evolving ride requests, user preferences, and order-state changes. Because the environment is stateful, later actions depend on earlier observations, so agents must maintain coherent plans across multiple steps rather than simply producing a final answer.

The central contribution is the Efficiency Utility metric, which measures agent performance by combining task success with the efficiency of the trajectory. It penalizes inefficient but ultimately successful behaviors—such as unnecessary tool calls, redundant lookups, avoidable backtracking, or overly verbose interaction sequences—so that two agents that complete the same task can still be ranked differently based on how cleanly they did so. This provides a more nuanced evaluation signal than success rate alone, especially for settings where tool calls have real costs in terms of time, API usage, or side effects.

The work matters because it pushes agent evaluation toward operationally meaningful performance, not just correctness. RideWay gives researchers a concrete testbed for studying compact, cost-aware tool-use policies in realistic service workflows, while the Efficiency Utility metric offers a reusable idea for other tool-centric agent domains. For production LLM agents, this highlights an important design goal: agents should not only complete tasks, but do so with minimal waste, friction, and resource consumption.

Generated 17d ago
Sources