Identifies open loops in recursive self-improvement for multi-agent systems and presents a closed-loop method that jointly refines collaboration topology and agent policies.
CollabFlow is a recursive self-improvement (RSI) system that addresses open loops in multi-agent collaboration by jointly refining team topology, communication protocols, and agent policies through a closed-loop training process. It utilizes a trainable Collab-Director to build teams and employs Evidence-Conditioned Communication, where agents refine answers only if the sender provides stronger evidence, thereby preventing error propagation.
The system closes the improvement loop using Collaborative Trajectory Balance (CTB), a flow-based objective that credits entire teams rather than individual construction orders, ensuring diverse and stable team distributions across evolution steps. Evaluated on twelve datasets, CollabFlow outperforms baselines by reducing token usage and maintaining performance across both in-distribution and out-of-distribution tasks, while successfully lifting the performance of other frozen executors.
CollabFlow addresses a gap in recursive self-improvement research for multi-agent systems: many existing approaches improve agent policies while leaving the collaboration structure fixed, or adjust interaction patterns without a stable mechanism for policy learning. The paper frames these as open-loop limitations that can lead to misaligned feedback, brittle coordination, and diminishing returns over repeated self-improvement cycles. It then introduces a closed-loop method in which the system jointly refines both the collaboration topology and the agents’ policies, using task outcomes and interaction signals to update who communicates with whom, how roles are allocated, and how each agent behaves in subsequent iterations.
The central insight is that effective self-improvement in multi-agent settings is not only an internal policy optimization problem; the collaboration architecture itself is a learnable variable. By treating topology as a first-class optimization target, CollabFlow can adapt communication pathways, dependencies, and role structure as tasks evolve, rather than forcing agents to compensate for a static interaction scaffold. This makes the self-improvement process more coherent, because improvements in behavior and improvements in coordination are coupled rather than treated as separate concerns.
This matters because multi-agent systems often fail due to coordination overhead, redundant communication, or poor role assignment, especially when tasks are dynamic or poorly specified. A method that adapts both policies and collaboration structure could improve robustness, sample efficiency, and generalization across tasks, while reducing reliance on manually engineered orchestration. More broadly, CollabFlow offers a more complete model of recursive self-improvement for agent collectives, where the system can iteratively refine not just what agents do, but how they work together.