Identifies three core challenges in self-evolving multi-agent orchestration (post-hoc evolution, credit diffusion, uncalibrated skill admission) and proposes a new method for online, calibrated team adaptation.

Topological visualization of EvoSteer: Online Self-Evolving Graph Orchestration via Reference-Anchored Credit Assignment
Brave API

EvoSteer is a new paradigm for Online Self-Evolving Graph Orchestration that dynamically builds and repairs multi-agent teams during execution, addressing the limitations of post-hoc revision and uncalibrated skill admission. It introduces Anchored Trajectory Balance (AnchorTB), a regression-style loss that assigns credit to actions by balancing subtrajectories against a frozen reference policy, and Validated Skill Admission, which promotes or retires skills only after passing sequential paired tests under a shared statistical budget. Evaluated on twelve benchmarks across reasoning and coding tasks, EvoSteer significantly outperforms workflow search, flow-based, and RL-based orchestration methods by unifying construction, repair, and skill growth into a single learned loop.

Generated 3d ago
Open-Weights Reasoning

EvoSteer addresses a central weakness in self-evolving multi-agent orchestration: existing systems often learn from completed trajectories, spread credit too broadly across agents and skills, and admit new capabilities without reliable evidence that they are useful. The paper frames these issues as three coupled challenges—post-hoc evolution, credit diffusion, and uncalibrated skill admission—and argues that they limit the ability of agent teams to adapt safely and effectively in open-ended settings. Its proposed method treats orchestration as an evolving graph of agents, skills, and control dependencies, and makes adaptation online rather than only after full task completion.

The core technical contribution is reference-anchored credit assignment, which uses reference signals or baseline trajectories to localize and calibrate responsibility for outcomes. By anchoring credit to well-defined references, the system can more precisely attribute success or failure to particular orchestration nodes, routing decisions, or skill invocations, rather than diffusing reward across the whole team. This supports calibrated team adaptation: the system can update its orchestration graph, adjust routing or collaboration structure, and admit new skills only when the evidence for their value is sufficiently strong.

The work matters because it moves self-evolving multi-agent systems closer to practical lifelong adaptation. Instead of relying on brittle post-hoc updates or unprincipled skill accumulation, EvoSteer provides a mechanism for continuous, evidence-based refinement of team structure and capability. For technically oriented readers, its significance lies in combining online orchestration, structured credit assignment, and uncertainty-aware skill admission—a potentially important direction for LLM-based agent teams, workflow orchestration, and other domains where multi-agent systems must improve from experience without destabilizing their own coordination logic.

Generated 3d ago
Sources