Highlights risks of jailbreaks enabling harmful tool use in deployed LLM agents and shows that trajectory-based retrieval in agentic attackers can strengthen attacks but risks reusing misleading experiences due to retrieval bias.
RedEvoAgent is an automatic black-box red-teaming agent that addresses the limitations of trajectory-based retrieval by distilling cross-case attack experiences into a concise, human-readable attack skill. This skill adaptively evolves through tool-effectiveness profiling and a Deciding-Tool Attribution mechanism that isolates the specific tool responsible for successful attacks, mitigating the bias of reusing misleading historical data.
Key advantages over traditional trajectory retrieval include: Reduced Context Overhead: By summarizing strategies rather than retrieving full trajectories, it lowers inference costs and improves interpretability. Validation Ratchet: Skill updates are only accepted if they improve performance on an independent validation set, preventing degradation from noisy or irrelevant past experiences. * Superior Performance: Experiments show RedEvoAgent outperforms fixed attack baselines and other agentic methods (like RedCodeAgent) in attack success rate and tool efficiency across diverse target models and execution harnesses.
RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution examines the security of deployed LLM agents that can execute tools, arguing that the most dangerous failure mode is not merely harmful text generation, but a jailbreak that leads to consequential tool use. The paper frames red-teaming as an iterative, agentic process: an attacker agent generates, executes, and evaluates attack trajectories, then uses accumulated experience to evolve more effective jailbreak strategies. In this setting, the attacker is not just a prompt generator but a decision-making system that learns from prior interactions with the target agent.
A central contribution is the use of trajectory-based retrieval to guide future attacks. By storing past attack trajectories and retrieving contextually similar ones, the system can transfer successful jailbreak patterns and refine its strategy over time. The key insight is that this experience reuse can substantially strengthen attacks, but it also introduces a form of retrieval bias: the agent may retrieve trajectories that look similar to the current situation yet are misleading, overfit, or poorly aligned with the actual target. As a result, the red-teaming agent can become trapped in repetitive or ineffective attack patterns, reusing experiences that appear useful on the surface but degrade performance in practice.
This matters because deployed LLM agents increasingly have access to high-impact tools, making automatic red-teaming both more necessary and more complex. The work highlights a broader tension in agentic learning systems: memory and retrieval can improve adaptability, but they also require careful filtering, validation, and utility-based selection of experiences. For defenders and evaluators, the implication is that safety testing of tool-using agents must account not only for static jailbreak prompts but also for adaptive attackers that learn from prior failures, and that robust experience management will be essential for both attack automation and defense.