Introduces trajectory-level unlearning to prevent LLM agents from reproducing undesired behaviors in action sequences beyond mere knowledge removal.

Topological visualization of Trajectory Unlearning on LLM-based Agents
Brave API

Trajectory Unlearning addresses the limitation of traditional knowledge-level unlearning, which only suppresses factual outputs, by targeting the specific action sequences (trajectories) an LLM-based agent executes. This approach is critical because agents may still reproduce undesirable or poisoned behaviors through multi-step interactions even if they no longer verbally recall the underlying forbidden knowledge.

To solve this, researchers introduced Group-injected Relative Policy Optimization (GiRPO), which injects forget trajectories into the policy rollout group with penalized rewards and isolated normalization statistics. This method effectively unlearns designated action sequences while preserving the agent’s ability to successfully complete tasks, outperforming existing baselines in both forgetting quality and task utility on benchmarks like ALFWorld and WebShop.

Generated 5d ago
Open-Weights Reasoning

Scope and problem. Trajectory Unlearning on LLM-based Agents extends machine unlearning from static knowledge removal to the behavioral level, where the target of forgetting is an undesired sequence of agent actions rather than a fact, preference, or data point. The paper observes that LLM agents can continue to exhibit harmful or policy-violating workflows—such as repeated tool calls, navigation steps, API interactions, or planning patterns—even after the underlying knowledge has been attenuated. By formalizing trajectories as first-class unlearning targets, it distinguishes between a model no longer being able to state a sensitive fact and a model no longer being able to execute or reproduce a sensitive multi-step behavior in an interactive setting.

Key contributions. The central contribution is a trajectory-level unlearning perspective that treats agent behavior as a learned sequence-generation policy and seeks to suppress the probability of specific undesired action traces while preserving general competence on benign tasks. The material introduces formulations, mitigation techniques, and evaluation probes that test whether an agent still follows a target trajectory under relevant prompts, contexts, and tool configurations. This shifts the focus from recall-style unlearning metrics to behavioral metrics: does the agent still initiate the same first action, recover the same plan, or complete the same end-to-end workflow?

Significance. This matters because agentic failure modes are often compositional and procedural. A single forgotten fact may not prevent a model from chaining innocuous actions into an unsafe outcome, misusing a tool, or automating a prohibited process. By framing unlearning around trajectories, the work offers a more realistic safety and compliance mechanism for LLM agents deployed with memory, tools, and environmental feedback, and it points toward auditing and mitigating not only what agents know, but what they can do.

Generated 5d ago
Sources