arXiv:2609.26080v1 Announce Type: cross Abstract: The model exposed to an application need not be a single checkpoint; it can be a governed fleet. Existing serving systems manage checkpoints and replicas, while multi-agent frameworks compose model calls without defining a stable collective identity, effect authority, or member-level evolution. We present Fusion-MoA, a runtime that exposes indepen
Fusion-MoA Pioneer R1 is a runtime system that exposes a heterogeneous fleet of independently served model cells as a single, OpenAI-compatible model. Unlike existing systems that manage individual checkpoints or compose multi-agent conversations without stable identity, this framework uses a versioned Profile to define the collective's identity, authority, and evolution.
Key operational contracts include: Identity: A versioned Profile binds the public API to a specific set of eligible cells, admission policies, and validation rules. Authority: The system enforces a sole-Executor model for final answers and tool actions, while read-only Analyst cells contribute bounded evidence without independent effect authority. * Evolution: Individual cells can be promoted or rolled back via single-slot qualification, allowing the fleet to evolve without changing the public API or requiring redeployment.
The system was evaluated on an eight-cell deployment across three base-model lineages, demonstrating that the collective could solve 8/10 HMMT problems (compared to the best individual cell solving 6/10) and complete 10/20 Terminal-Bench 2.1 tasks with all tool actions attributable to the single Executor. An executable reference implementation is available at https://github.com/tongjiu123/Fusion-MoA-Pioneer-R1.
The paper argues that the “model” consumed by an application should no longer be treated as a single checkpoint, but as a governed fleet of models. It contrasts two existing abstractions: model-serving systems, which primarily manage checkpoints, replicas, and inference load, and multi-agent frameworks, which compose model calls but often lack a durable notion of the collective as a single accountable entity. The central claim is that production-grade collective intelligence requires a higher-level abstraction that defines a stable collective identity, authority over effects, and mechanisms for member-level evolution.
Its main contribution is Fusion-MoA, presented as a runtime that exposes independent models through a unified, governed interface. The work positions Fusion-MoA as a bridge between serving infrastructure and multi-agent orchestration, with Pioneer R1 serving as the initial or reference realization of that runtime. Rather than merely chaining calls between agents, the system appears to treat the fleet itself as the deployed model, with governance and lifecycle semantics attached to the collective.
This matters because it shifts the design problem from “how do we call multiple models?” to “how do we operate a multi-model entity as a first-class system component?” If fleets become the unit of deployment, then infrastructure must address identity, authority, auditability, upgrade paths, and member-level behavior in the same way that traditional systems address service versioning and policy. The paper is therefore relevant to anyone building governed LLM platforms, multi-agent production systems, or model-orchestration stacks where accountability and evolution are as important as raw capability.