Compendium that catalogs criteria, metrics, and benchmarks used to define and evaluate AI agents across the literature.

Topological visualization of [2609.11018] Defining AI Agents: A Compendium of Criteria, Metrics, and Benchmarks
Brave API

The paper "Defining AI Agents: A Compendium of Criteria, Metrics, and Benchmarks" (arXiv:2609.11018, submitted September 10, 2026) by Mia Lassiter and Brinnae Bent addresses the lack of a standard definition for AI agents by organizing evaluation around five dimensions of agenticness: environmental interaction, learning and adaptation, autonomy, goal-directed behavior, and temporal coherence.

The authors introduce the Agent Compendium (available at agent.duketrustlab.com), a living digital resource that maps specific criteria to existing metrics and benchmarks. Key evaluation aspects include:

  • Environmental Interaction: Measured via consistency (e.g., output semantic similarity, decision pathway stability) and robustness (e.g., AgentDojo, HarmBench, JailbreakBench).
  • Learning and Adaptation: Assessed through agentic capabilities (tool use to common-sense reasoning) and generality levels (from emerging to superhuman performance).
  • Autonomy: Categorized from restricted (AI as a tool) to full autonomy (AI as an agent), with specific action-level descriptors like "human-in-the-loop" or "human-out-of-the-loop."
  • Goal-Directed Behavior: Evaluated by intentionality (goal representation stability), initiative (self-directed agency), and resilience (multi-task recovery rates).
  • Temporal Coherence: Focuses on the active processing window and drift velocity in maintaining context over time.

This framework aims to improve research reproducibility and clarify communication by providing explicit criteria for identifying AI agents and connecting them to concrete evaluation methods.

Generated 23d ago
Open-Weights Reasoning

The material is a structured compendium that maps how “AI agent” is defined and evaluated across the research literature. Rather than treating the term as a single fixed category, it catalogs the criteria commonly used to characterize agenthood—such as goal-directedness, autonomy, perception–action coupling, planning, memory, tool use, adaptivity, and interaction with open-ended environments—and connects those conceptual attributes to concrete metrics and benchmarks. These include task success or completion rates, efficiency, robustness, generalization, cost or latency, safety, and performance on long-horizon, embodied, software, web, or multi-agent tasks.

Its key contribution is to make explicit what is often implicit in agent research: that different systems are being compared under different assumptions about what counts as an agent and what counts as meaningful performance. By organizing definitions, evaluation criteria, and benchmark regimes side by side, the compendium helps expose where the literature overlaps, where terminology is inconsistent, and where evaluation is incomplete. It also highlights important gaps, such as the need for better measures of sustained autonomy, tool integration, human–agent collaboration, resource efficiency, and safety under realistic deployment conditions.

This matters because the field currently lacks a shared vocabulary and a common evaluation substrate for agent systems. For researchers, engineers, and policymakers, the compendium can serve as a reference for selecting appropriate benchmarks, designing more rigorous evaluations, and avoiding comparisons that are misleading because they conflate different notions of agency. In short, it helps turn “AI agent” from a loosely used label into a more analytically tractable and empirically evaluable category.

Generated 23d ago
Sources