Shows that multi-agent communication gains cannot be isolated from architecture or reasoning effects under standard evaluations, proposing disentanglement methods.
Beyond Final Accuracy: Auditing Communication in LLM Multi-Agent Systems introduces the Independent–Communicate–Revise (ICR) framework to disentangle communication effects from architectural or reasoning biases. By fixing initial reasoning trajectories and using a no-message revision control, the study isolates the specific impact of message content and receiver policies on answer revision.
The audit reveals that aggregate accuracy conceals distinct behavioral profiles, such as high correction rates paired with low preservation of correct answers. It demonstrates that full reasoning messages increase correction but reduce preservation compared to answer-only messages, and that receiver policies (e.g., structured verification) significantly shift this balance. These findings indicate that communication quality is not an intrinsic property of a channel but a result of the joint interaction between message content and the receiver’s policy.
Beyond Final Accuracy: Auditing Communication in LLM Multi-Agent Systems addresses a core measurement problem in evaluating large language model multi-agent systems: reported improvements from inter-agent communication are often confounded with gains from the surrounding architecture, additional inference steps, role specialization, or stronger reasoning prompts. The paper argues that standard task-accuracy benchmarks do not reveal whether communication itself is causally useful, because the same performance lift can arise from non-communicative changes in the system. By treating final accuracy as an incomplete signal, the work reframes the evaluation question from “does the multi-agent system score higher?” to “what portion of that improvement is attributable to communication, and what portion is due to other design choices?”
A central contribution is a set of disentanglement methods for separating communicative effects from architectural and reasoning effects. Rather than relying on a single end-to-end comparison, the paper proposes controlled evaluations that isolate the role of message exchange—for example, by comparing communication-enabled and communication-disabled configurations while holding the rest of the pipeline fixed, and by using baselines that account for extra computation, prompt structure, or agent roles. The key insight is that communication in LLM multi-agent systems should be audited as a distinct mechanism, with explicit metrics and ablations that expose whether agents are actually exchanging useful information, coordinating work, or merely benefiting from additional model calls and structured scaffolding.
This matters because it changes how the field should interpret and report multi-agent results. If communication gains are not disentangled, researchers may overstate the value of multi-agent interaction, misattribute improvements to prompting or inference depth, and build systems around mechanisms that do not actually contribute to performance. The paper therefore provides a more rigorous diagnostic framework for LLM multi-agent research, encouraging evaluations that distinguish genuine communicative utility from incidental architectural advantages. Its implications extend to benchmark design, system architecture choices, and cost-aware deployment, where understanding whether communication is necessary can guide more efficient and reliable agent-based applications.