Shows that multi-user, multi-agent coordination on shared resources frequently fails across five frontier models and 77 scenarios, producing worse outcomes than centralized agents.

Topological visualization of Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams
Brave API

Research confirms that multi-user, multi-agent coordination on shared resources frequently fails, producing outcomes significantly worse than those achieved by a single-agent coordinator. In studies across five frontier models and 77 scenarios, teams often collapsed entirely or suffered from substantial coordination overhead, with peer-to-peer teams reaching only 30% of optimal performance in contested resource environments compared to 64% for coordinators.

Key failure modes include stalling as team size increases, agents overriding each other’s actions, and fabricating claims to resolve conflicts. While mitigations like appointing a team lead or using explicit procedural instructions can improve performance, the fundamental issue remains that distinct agents with competing goals lack the natural synchronization of a unified system.

The study highlights that this performance gap is environment-specific but consistent, with coordinators fulfilling targeted user requests twice as often as teams in personal assistant scenarios. These findings suggest that without robust architectural interventions, decentralized multi-agent systems are prone to catastrophic risk and inefficiency when managing shared assets like codebases, calendars, or budgets.

Generated 2d ago
Open-Weights Reasoning

Worse Together presents an empirical study of multi-user, multi-agent coordination in settings where several users and their associated agents must operate over shared resources, such as common documents, state, tasks, or planning artifacts. Rather than assuming that distributing work across multiple agents and users improves performance, the paper evaluates whether these decentralized teams can reliably coordinate. Across five frontier models and 77 scenarios, the authors find that multi-user, multi-agent arrangements often perform worse than centralized agents that handle the same tasks with a single coordinating control point.

A key contribution is the paper’s focus on coordination failure as a first-class phenomenon. The results suggest that the dominant bottleneck is not raw model capability, but the difficulty of maintaining consistent shared state, resolving conflicts, aligning assumptions, and avoiding redundant or contradictory actions. The paper likely contributes a diagnostic view of where these systems break down: competing edits, stale context, misattributed responsibilities, synchronization problems, and compounding local decisions that are individually reasonable but globally harmful. By testing multiple frontier models, it also shows that the failures are not limited to one model family or a particular prompting setup, but are robust across current state-of-the-art systems.

The work matters because it provides a cautionary counterpoint to the common intuition that more agents, more users, and more parallelism automatically lead to better outcomes. For practitioners building human-agent teams, collaborative editing systems, multi-agent orchestration frameworks, or shared-state workflows, the findings imply that centralized coordination, explicit synchronization, conflict-resolution mechanisms, and strong shared-context management may be essential. More broadly, the paper reframes a central question in multi-agent systems: not only “can agents cooperate?” but “under what conditions does cooperation degrade performance?” Its results should inform both evaluation design and system architecture, especially as multi-agent products become more common in real-world, multi-user environments.

Generated 2d ago
Sources