Surveys open-source LLM-based MAS practitioners to identify core challenges, root causes, and mitigation strategies.
Based on an empirical study of 944 issues from 21 open-source LLM-based Multi-Agent Systems (MAS), the core challenge is Orchestration & Execution, accounting for 22.0% of reported issues. These failures primarily stem from Coordination & Handoff Issues (62 instances) and Control Flow Errors (40 instances), indicating that runtime coordination is more difficult than individual agent reasoning.
The most frequent root causes are Workflow Problems (29.4%), Tool Integration Problems (11.7%), and Memory Problems (11.1%). Workflow issues involve incorrect state handling and outdated context reuse, while tool integration problems arise from schema mismatches and execution deadlocks. Memory & State Management Issues (14.7%) are also significant, driven by context loss and state inconsistency during agent handoffs.
The predominant solution is to Optimize Workflow, often involving explicit handoff designs, retry limits, and timeouts. For developers, this implies a need for robust workflow design and reliable tool integration. Researchers are encouraged to shift focus from individual agent reasoning to robust orchestration frameworks that ensure reliable inter-agent communication, correct execution ordering, and effective memory management.
This paper examines open-source, LLM-based multi-agent systems (MAS) from a practitioner-oriented perspective, focusing on the issues that emerge when such systems move beyond demos into real engineering contexts. Rather than treating multi-agent performance primarily as a model-capability question, it frames MAS as distributed, stochastic software systems composed of agents, orchestrators, tools, memory, and communication channels. Its central aim is to identify recurring failure modes, trace them to underlying causes, and catalog mitigation strategies used by developers building and operating these systems.
Its main contribution is a diagnostic map that connects practical symptoms—such as coordination breakdowns, inconsistent state, context bloat, hallucinated tool calls, role drift, nondeterministic behavior, and poor observability—to root causes in architecture, prompting, tool design, evaluation, and runtime management. The paper likely emphasizes that many MAS problems are not simply “the model is bad,” but arise from system-level design choices: ambiguous agent responsibilities, unstructured inter-agent communication, weak validation of agent outputs, insufficient logging, and the absence of reproducible evaluation harnesses. Corresponding solutions include explicit orchestration, structured message schemas, centralized state management, idempotent and constrained tool interfaces, guardrails, tracing, regression testing, and human-in-the-loop controls.
This matters because open-source LLM-based MAS frameworks are being adopted quickly, yet production reliability remains uneven and often undocumented. By collecting practitioner experience, the paper offers a more actionable reference for engineers and researchers than isolated benchmark results: it helps teams diagnose failures more systematically, compare design trade-offs, and build evaluation practices that reflect real deployment constraints. In doing so, it contributes to the broader question of how to make LLM-based agent systems debuggable, maintainable, and trustworthy at scale.