Proposes off-policy co-evolution methods that synthesize environments beyond on-policy rollouts to maintain continuous learning signals as frontier models improve.
Recent frameworks propose off-policy environment synthesis to overcome the limitations of static training distributions and on-policy bias in terminal agent learning. Envs-FORGE addresses this by using a Mixed-Integer Linear Program (MILP) to select per-seed synthesis actions (increase, reduce, or diversify complexity) based on verifier-derived pass rates, ensuring generated environments remain near the agent's learning frontier rather than applying fixed prompting recipes.
Similarly, COvolve employs a two-player zero-sum game where an LLM-based Environment Designer generates increasingly challenging executable code to expose policy weaknesses, while a Policy Designer adapts; this adversarial process is stabilized by computing a Mixed-Strategy Nash Equilibrium (MSNE) to prevent catastrophic forgetting and ensure robust generalization.
GenEnv further advances this by establishing a difficulty-aligned co-evolutionary loop between an Agent Policy and an Environment Policy (simulator), using an $\alpha$-Curriculum Reward to dynamically adjust task difficulty. This approach shifts training from static, high-cost real-world interaction to adaptive simulation, yielding significant performance gains across benchmarks like API-Bank and BFCL by maintaining continuous, capability-matched learning signals.
Environment Evolution for Terminal Agents addresses a core scaling problem for agentic systems: as terminal-based agents become stronger, fixed task suites and on-policy rollouts quickly stop providing useful learning signal. The paper frames terminal environments—shell sessions, file systems, processes, command-line tools, and their verifiable state transitions—as the substrate on which agents must keep improving. Its central argument is that once a frontier model saturates a static distribution of tasks, further training mostly reinforces already-learned behaviors rather than exposing the model to new capabilities. By contrast, the work proposes an off-policy co-evolutionary approach in which environments are synthesized, mutated, or otherwise extended beyond the immediate rollouts of the current policy, so that the agent continues to face informative, challenging, and verifiable situations.
The key contribution is a method for maintaining a continuous training and evaluation signal through environment evolution rather than through larger static datasets alone. Instead of relying only on on-policy data generated by the current model, the approach can draw on rollouts from earlier policies, replay useful states, and generate new terminal tasks that are grounded in executable or checkable outcomes. This off-policy perspective is important because it decouples environment generation from the current policy’s limitations: tasks can be derived from past failures, model-generated hypotheses, difficulty-controlled mutations, or verifier-validated state changes. The result is a curriculum that can adapt to improving model capability, preserving a balance between tasks that are too easy and those that are unsolvable or poorly specified.
This matters because it points toward a more scalable path for training and benchmarking terminal agents in the frontier-model regime. Manual task authoring and static benchmarks are unlikely to keep pace with rapid model improvement, and reward-based or preference-based learning is only as useful as the diversity and verifiability of the environments it sees. By co-evolving environments with agents, the work offers a framework for continual capability growth, better generalization across shell and file-system operations, and more robust evaluation of agentic behavior. It also highlights an important safety and methodology concern: as environments become increasingly synthetic, automatic verification, difficulty calibration, and protection against reward hacking become essential to ensure that the evolved tasks remain meaningful rather than merely exploitable.