Introduces a refreshable anomaly-detection benchmark for LLM multi-agent systems that counters data leakage, pattern expiration, and labeling issues.

Topological visualization of MAADBench: The Refreshable Paradigm for Anomaly Detection in Multi-Agent Systems
Brave API

MAADBench is a refreshable benchmark for anomaly detection in multi-agent systems that addresses data leakage, pattern expiration, and labeling issues through three core mechanisms: sampled-and-coupled generative tasks over a vast task space to prevent leakage, configurable trace generation to handle evolving LLM backbones, and automated deterministic step-level labeling using oracle ground truth to ensure high-fidelity, cost-free annotations.

The benchmark was released with MAADBench-Full, a dataset containing 5,200 step-labeled traces generated across five state-of-the-art LLM backbones. Evaluating 25 anomaly detection methods on this dataset revealed that current approaches struggle with subtle MAS-specific anomalies, rely heavily on supervision, and lack robustness across different LLM backbones.

Key features of MAADBench include: Generative Task Space: Uses approximately $10^{37}$ tasks to make contamination practically infeasible. Reproducible Traces: Allows for continuous refresh of traces under new model configurations. * Deterministic Labeling: Eliminates the need for expensive human or LLM-based annotation by propagating labels from oracle-defined atomic tasks.

Generated 4d ago
Open-Weights Reasoning

MAADBench is presented as a response to a core evaluation problem in LLM-based multi-agent systems: anomaly detection is difficult to benchmark reliably because agent behavior is dynamic, language-mediated, and often non-stationary. Traditional static benchmarks can quickly become unreliable once models, prompts, agent topologies, or attack patterns evolve, and they are vulnerable to data leakage, outdated anomaly signatures, and ambiguous or noisy labels. The paper frames anomaly detection in multi-agent settings as a moving-target problem and proposes a “refreshable” benchmark paradigm in which test scenarios, anomaly instances, and evaluation conditions can be regenerated or updated rather than treated as a fixed dataset.

The key contribution is a benchmark methodology for anomaly detection in multi-agent systems that emphasizes refreshability as a first-class design requirement. Rather than relying on a one-time collection of labeled agent interactions, MAADBench appears to support the continuous or periodic generation of new anomaly patterns while preserving controlled ground truth. This design targets three failure modes: data leakage, where detectors overfit to memorized test cases; pattern expiration, where previously meaningful anomalies become obsolete or easily avoided; and labeling issues, where open-ended agent behavior makes manual annotation costly or inconsistent. By tying evaluation to regenerable scenarios and more reliable ground-truth construction, the benchmark aims to separate detectors that genuinely understand anomalous multi-agent behavior from those that merely fit a stale snapshot.

This matters because LLM multi-agent systems are increasingly used in settings where subtle failures—such as coordination breakdowns, policy violations, unsafe delegation, or malicious agent behavior—can be hard to detect with conventional monitoring. A refreshable benchmark is important for building and evaluating anomaly detectors that remain useful as agent architectures and attack surfaces change. MAADBench therefore provides a more realistic and maintainable evaluation substrate for studying reliability, safety, and monitoring in multi-agent LLM ecosystems, rather than relying on benchmarks that may quickly lose predictive value.

Generated 4d ago
Sources