Reviews construction, deployment, and assessment of LLM agents automating multi-step security workflows of artifact inspection, hypothesis formation, and tool invocation.
LLM-Based Agents for Software and Systems Security: Approaches, Applications, and Assessment is a systematic literature review of 100 peer-reviewed papers published between January 2023 and March 2026 that examines how large language model agents automate procedural security workflows. The study organizes existing research into three core dimensions: Approach (agent architecture, perception, memory, reasoning, and action space), Application (tasks such as vulnerability detection, penetration testing, and malware analysis), and Assessment (datasets, metrics, and safety measures).
The review highlights that while current agents are capable of multi-step planning and tool invocation, the field lacks systems with bounded authority or fully auditable behavior. Key challenges identified include ensuring evidence-centered architecture, implementing risk-aware autonomy, and developing reproducible benchmarks that evaluate trajectory-level performance rather than just final outcomes.
Key findings from the review include:
This arXiv review examines how large language model agents are being designed to automate multi-step software and systems security tasks. It frames agent pipelines around three core primitives: inspecting artifacts such as source code, binaries, configurations, and logs; forming security hypotheses such as potential vulnerabilities, misconfigurations, or malicious behavior; and invoking external tools such as SAST, DAST, fuzzers, sandboxes, debuggers, or exploit frameworks to test those hypotheses. The paper maps approaches across agent architectures—single-model planners, retrieval-augmented agents, multi-agent hierarchies, and tool-calling loops—and discusses deployment concerns including access control, human oversight, traceability, and integration with CI/CD or SOC workflows.
A key insight is that LLM agents are promising not merely as code classifiers, but as orchestrators that can close the loop between reasoning and execution. Their practical value depends on producing actionable, verifiable findings rather than plausible but unsupported text, so assessment must extend beyond task completion to metrics such as detection precision and recall, false-positive burden, exploitability validation, cost, latency, robustness to adversarial inputs, and safety of tool invocation. The review highlights open problems including hallucinated findings, over-permissioned tool access, prompt injection, non-determinism, benchmark gaps, and the need for reproducible evaluation harnesses.
The material matters because security workflows are long-horizon, tool-heavy, and high-stakes, making agent evaluation especially difficult. By organizing the literature around construction, deployment, and assessment, the paper provides a shared vocabulary for comparing systems and identifying where current agents fall short. For practitioners and researchers, it points toward more disciplined agent design: constrained tool use, evidence-grounded hypotheses, staged autonomy, and metrics that reflect operational security impact rather than superficial benchmark success.