Empirically shows role-matched agents from independently trained teams are not interchangeable, measuring performance drops after roster trades on held-out tasks.
Current search results do not contain a study with the title "Testing Interchangeability in LLM Agent Teams Empirically" or specific data on performance drops from roster trades on held-out tasks.
However, related research highlights critical constraints on agent team composition:
The paper investigates whether LLM agents that are assigned similar roles in multi-agent teams are truly interchangeable. It studies independently trained agent teams and performs controlled roster trades, replacing an agent with a role-matched agent from another team, then evaluates the resulting teams on held-out tasks. The central empirical finding is that role-matched agents are not reliably drop-in substitutes: swapping agents produces measurable performance drops, even when the replaced agents appear functionally equivalent from the outside. This suggests that an agent’s usefulness in a team depends not only on its nominal role, but also on team-specific coordination dynamics, learned interaction patterns, and implicit conventions developed during training or prior use.
A key contribution is treating interchangeability as an explicit testable property of multi-agent systems rather than an assumed design principle. The held-out-task evaluation is important because it probes whether swapped teams can generalize beyond the settings in which the original teams were formed or tuned, rather than merely failing on a narrow benchmark. The observed degradation after roster trades implies that multi-agent LLM systems may be less modular and composable than their role descriptions suggest, and that “same role” is a weak proxy for “same functional contribution” in practice.
This matters because LLM agent teams are increasingly being used in settings where agents might be swapped, outsourced, upgraded, or mixed across teams. If agents are not interchangeable, system designers need to account for integration costs, evaluate replacements before deployment, and potentially standardize communication protocols, training regimes, or role specifications. More broadly, the work shifts the evaluation question from “Is each agent competent?” to “Do these agents compose reliably as a team?”, highlighting a practical gap between nominal role design and operational substitutability in multi-agent LLM systems.