Studies white-box pre-deployment auditing of AI agents that combine language models with file-modifying tools, identifying risks from adversarial content and software vulnerabilities.

Topological visualization of AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents
Brave API

AgentXploit is a two-role autonomous auditing system designed for white-box pre-deployment security assessments of AI agents that combine language models with external tools like file modifiers, API callers, or code executors. It separates repository-level attack-path discovery (performed by an Analyzer Agent) from runtime exploitation (performed by an Exploiter Agent), addressing security failures where adversarial content or underlying software vulnerabilities, such as path traversal or command injection, compromise agent behavior.

The framework was evaluated on AgentXploit-Bench, a benchmark containing 72 reproducible vulnerabilities across 12 open-source AI-agent systems. AgentXploit achieved a 59.3% end-to-end success rate, significantly outperforming baseline methods like Codex (38.4%–46.3%), highlighting that repository discovery and runtime exploitation are distinct challenges in end-to-end agent security auditing.

Generated 6d ago
Open-Weights Reasoning

AgentXploit presents an autonomous red-teaming approach for auditing AI agents that combine language models with file-modifying tools. Rather than treating agent security as a narrow prompt-injection problem, the work frames it as a repository-to-runtime risk: the agent’s codebase, the repositories it operates on, the tools it can invoke, and the runtime effects of its actions are all part of the attack surface. Its focus on white-box pre-deployment auditing means the auditor has access to the agent’s implementation, tool interfaces, permissions, and configuration, enabling systematic probing of how the agent might be manipulated before it is deployed.

A key contribution is the treatment of AI-agent risk as an end-to-end interaction between adversarial content and software vulnerabilities. Malicious or misleading content in repository files, documentation, code comments, dependency metadata, or other artifacts can influence the language model’s planning, while weaknesses in the agent’s own code, tooling, or runtime environment can turn that influence into concrete harm—such as unauthorized file modification, command execution, credential exposure, or supply-chain compromise. By automating the generation and validation of these red-team scenarios, the work moves beyond isolated prompt tests toward a more realistic assessment of what an attacker could achieve when the agent is given real repository and execution privileges.

This matters because AI coding and DevOps agents are increasingly deployed with access to sensitive codebases and operational environments. The material highlights that traditional software security checks and simple prompt-injection defenses are likely insufficient for such systems, since the language-model layer can interact with file-modifying tools in ways that are difficult to reason about statically. AgentXploit therefore supports a more rigorous pre-deployment security practice: identifying high-risk agent behaviors, validating whether they can be triggered in practice, and motivating controls such as least-privilege tooling, sandboxing, input provenance checks, runtime monitoring, and verification of file-level side effects.

Generated 6d ago
Sources