Shows that combining automated harness evolution with lightweight fine-tuning enables smaller models to match frontier performance across seven enterprise agent tasks.
Combining automated harness evolution with lightweight fine-tuning enables smaller models to match frontier performance, but the method of adaptation is critical. Research shows that while evolving a harness significantly boosts smaller model performance, directly imitating a stronger expert’s trajectories under this evolved harness causes performance regression (dropping 4–30 points) because it disrupts the fit between the model’s native planning style and the specialized harness.
To successfully co-evolve, researchers developed an on-policy expert correction pipeline where a meta-level agent identifies failing turns in the weaker model’s own rollouts and has the expert rewrite only those specific turns. This approach preserves the model’s planning distribution while transferring expert knowledge, yielding further gains without the counterproductive regression seen in full-trajectory imitation.
This material examines a joint optimization approach for LLM-based agents in which the model and the agent harness evolve together rather than being treated as separate components. The harness includes the surrounding scaffolding that determines how a model is prompted, how tools are exposed, how state is maintained, and how rollouts are evaluated or corrected. The central idea is that smaller models can be improved not only by fine-tuning on high-quality demonstrations, but also by adapting the harness to surface the model’s failure modes and by using on-policy correction: generating training signal from the model’s own trajectories, then correcting or rewarding them in a way that is directly relevant to the current policy.
The key contribution is a co-evolution loop in which automated harness evolution and lightweight fine-tuning work in tandem. The paper argues that standard imitation or distillation from frontier models can fail because it transfers behavior tied to a particular scaffold, prompting style, or tool interface, leaving the weaker model brittle when the environment or harness changes. By contrast, co-evolving the harness allows the training process to expose task-relevant weaknesses, while on-policy correction updates the model on the distribution it will actually encounter. Across seven enterprise agent tasks, this combination reportedly enables smaller models to match frontier-level performance, suggesting that much of the apparent capability gap can be closed through better model–harness alignment rather than only through scaling the base model.
This matters because enterprise agents must operate reliably in constrained, multi-step, tool-heavy workflows where deployment cost, latency, and adaptability are critical. The work points toward a practical path for building capable agents without relying exclusively on frontier-scale models: continuously refine the orchestration layer and the model on its own corrected behavior. More broadly, it suggests that future agent systems should be designed as coupled model–harness stacks, where evaluation, tooling, and fine-tuning are co-optimized rather than fixed in advance.