Introduces DAREBench, a curated set of 233 tasks from 22 benchmarks for deployment-aware and reliable evaluation of models as agents.
DAREBench is a benchmark designed for the deployment-aware and reliable evaluation of models as agents, addressing limitations in existing static benchmarks by focusing on real-world deployment factors like workload profiles, execution environments, and accuracy-cost trade-offs.
It organizes 233 tasks curated from 22 source benchmarks into a 2×3 workload matrix defined by input modality (text vs. multimodal) and execution form (single-step, multi-step, multi-step-plus-tools). The evaluation utilizes a shared OpenClaw execution environment and a unified contract-based protocol with evidence-based score auditing to ensure reliability.
Key features and findings include: Comprehensive Evaluation: Tested 35 models (23 commercial API and 12 locally deployed open-weight) across 7,587 model-task runs. Reliable Scoring: Combines automated, LLM-based, and hybrid scoring, with a meta-judge audit that identified and removed 143 unsupported positive scores caused by hallucinations like deadlock loops or pseudo-tool-call credulity. * Deployment Insights: Results show that no single model dominates all workload groups, and local open-weight models are competitive in specific groups but generally trail frontier commercial models. Model selection should consider workload-specific accuracy-cost trade-offs rather than aggregate scores.
Overview. DAREBench is a curated evaluation suite for assessing models as agents, with an emphasis on deployment-awareness and reliability rather than isolated task accuracy. It consolidates 233 tasks drawn from 22 existing benchmarks into a more compact testbed for comparing how models behave in agentic settings. The motivation is that real-world agent systems are evaluated not only on whether they can solve individual tasks, but on whether they can operate consistently, robustly, and practically when embedded in longer workflows or production-like environments.
Key contributions. The main contribution is a benchmark design that reframes agent evaluation around operational deployability. Instead of treating existing benchmarks as independent leaderboards, DAREBench selects a representative subset of tasks to expose capabilities and failure modes relevant to deployed systems, such as sustained multi-step performance, cross-task consistency, and reliability under realistic evaluation conditions. This makes it easier to compare models as agents without requiring exhaustive runs across all 22 source benchmarks, while still preserving coverage of diverse agentic competencies.
Why it matters. The work matters because current agent evaluations can overstate practical readiness: a model may perform well on selected tasks yet struggle with the consistency, robustness, and end-to-end reliability required for deployment. DAREBench offers a more operationally grounded yardstick for researchers and practitioners who need to decide which models are not only capable, but dependable enough to use as agents. In doing so, it shifts the evaluation focus from narrow peak performance toward the kind of reliability that matters in real-world deployment.