Defines an evaluation contract that fixes information boundaries, downstream stacks, and team families to isolate routing effects and compute exact finite-benchmark oracle regret in multi-agent systems.
COVER is an auditable measurement methodology designed to isolate the specific effect of routing in multi-agent systems by fixing a public information boundary, a downstream stack, and a finite legal team family before outcomes are generated. This contract prevents confounding by ensuring that observed performance gaps are attributable to team selection rather than changes in the downstream processing pipeline.
The method computes exact finite-benchmark oracle regret by requiring complete coverage (executing every legal team) for absolute regret, or using the union of distinct selected teams for minimal assumption-free pairwise policy contrasts. This approach distinguishes measurement headroom from achieved routing performance, revealing whether a router successfully converts potential value into actual gains or if synthesis failures are masking routing efficacy.
COVER: Identifiable Evaluation of Coalition Routing addresses a core measurement problem in multi-agent systems: how to tell whether a routing or orchestration layer actually improves performance, rather than merely changing which agents, tools, prompts, or information sources are involved. In typical end-to-end evaluations, gains attributed to “routing” can be confounded by differences in agent capability, downstream model stack, available context, or the set of eligible teams. COVER proposes an evaluation contract that fixes these confounding dimensions: it specifies information boundaries for what each agent may observe, standardizes the downstream stack that executes the routed request, and constrains the team family of admissible coalitions. With those factors held constant, the remaining variation is due to the routing decision itself, making the routing effect more cleanly identifiable.
The paper’s key technical contribution is a way to quantify that effect on a finite benchmark by computing exact oracle regret. The oracle represents the best possible routing behavior within the fixed contract—using knowledge of the benchmark outcomes or ground truth—and regret measures the utility gap between the evaluated router and that optimum. Because the benchmark is finite and the contract fixes the relevant action space, the oracle comparison is not merely an empirical estimate of improvement; it is a precise diagnostic of how much performance the router leaves on the table relative to an idealized coalition-routing policy under the same constraints.
This matters because multi-agent orchestration is often judged by aggregate task accuracy, which can obscure whether improvements come from better routing, stronger underlying agents, richer prompts, or more permissive access to information. COVER provides a controlled, reproducible framework for isolating the value of coalition selection itself. It is therefore useful for comparing routing policies, auditing agent ensembles, and designing benchmarks where performance changes can be attributed to routing decisions rather than to incidental changes in the surrounding multi-agent stack.