Presents SafeEvolve, an experience-driven self-evolving framework that jointly optimizes runtime control and intrinsic safety for LLM agents.
The search context does not contain information about a framework named "SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment."
The search results do identify a method called SafeEvolve, but it is distinct from the description provided: SafeEvolve is a method-agnostic governance wrapper introduced in the paper Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents (arXiv:2608.12851). It focuses on skill misevolution, governing the lifecycle of persistent skills by combining critic-localized delete-only repair, lineage-risk retrieval, and safety-aware retirement. It reduces unsafe retrieval and fresh-session harm by 26.7 and 17.3 percentage points, respectively, while preserving benign utility. It operates at write and reuse boundaries without changing the underlying agent or evolution algorithm.
Other related frameworks in the context include: SHE (Safety Harness Evolution): Evolves the agent harness (context, memory, tools) from rollout trajectories to learn safe boundaries. EvoHarness-RL: Learns a self-evolving runtime harness for long-horizon agents via experience consolidation. * On-Policy Self-Evolution via Failure Trajectories: Uses failure trajectories for on-policy safety alignment.
There is no evidence in the provided context that SafeEvolve performs harness-policy co-evolution or jointly optimizes runtime control and intrinsic safety as described in the query.
SafeEvolve introduces an experience-driven, self-evolving framework for improving the safety of LLM agents in interactive, tool-using environments. The paper frames agent safety as a joint problem involving both the model’s intrinsic behavior and the external “harness” that governs execution, tool access, monitoring, and intervention. Rather than treating safety as a static property of the model or a fixed set of guardrails, SafeEvolve uses accumulated agent experience to co-evolve both layers over time.
The key contribution is a coupled optimization loop in which agent trajectories are analyzed for safety-relevant signals and used to update both the runtime control harness and the agent’s safety policy. The harness can be refined with better constraints, monitors, or recovery mechanisms, while the policy can be adjusted to internalize safer decision-making patterns. The central insight is that independently evolving guardrails and model behavior can create misalignment: static controls may miss emerging failure modes, while a model without adaptive runtime support may still act unsafely in novel contexts. Co-evolution allows the system to adapt to new tasks, tools, and risks while preserving useful task performance.
This matters because LLM agents are increasingly deployed in long-horizon settings where handcrafted safety rules and one-off alignment training are insufficient. SafeEvolve positions safety as an ongoing, self-improving control problem rather than a fixed compliance layer. For practitioners, it suggests a practical architecture for continuously auditing, hardening, and adapting deployed agents, reducing unsafe behavior without relying solely on manual rule authoring or repeated full model retraining.