Clusters core research themes around LLM-agent cybersecurity, security, and defense.
A PRISMA-guided systematic review of 59 studies (2022–2026) classifies LLM-based cybersecurity agents across five dimensions: security function, agent architecture, knowledge augmentation, human-in-the-loop posture, and evaluation rigor. The research is heavily skewed toward penetration testing and red teaming (50.8%), while critical operational areas like incident response (8.5%) and compliance verification (1.7%) remain underrepresented.
Architectural patterns are primarily divided into single-agent tool-calling (most prevalent) and multi-agent orchestration, with centralized hierarchical and pipeline/parallel designs showing the strongest performance gains. However, the field is characterized by proof-of-concept systems evaluated in controlled labs, with zero reported production deployments in the reviewed corpus.
Key challenges include hallucination, prompt injection vulnerabilities, and benchmark fragmentation, with current agents described as "able to act but not yet bounded or auditable." Future research priorities emphasize multi-agent orchestration, real-time latency reduction for financial sectors, and robust governance frameworks to address regulatory requirements like DORA and GDPR.
This MDPI systematic review surveys how large language model (LLM) agents are being applied to cybersecurity, organizing the emerging literature into a coherent taxonomy rather than a simple tool inventory. It examines agent designs—single-agent and multi-agent systems, tool-augmented planners, retrieval-augmented pipelines, memory and orchestration patterns—and maps them to cybersecurity tasks such as threat intelligence, log analysis, incident triage, vulnerability discovery, malware analysis, penetration testing, and automated response. A central contribution is distinguishing between LLM agents as defenders, as adversarial or red-team capabilities, and as new attack surfaces, highlighting where current systems remain human-supervised copilots versus where they are positioned as more autonomous operators.
The review also identifies recurring technical and operational challenges: hallucination and ungrounded reasoning, prompt injection and tool misuse, unsafe autonomous actions, privacy leakage through logs and prompts, lack of standardized benchmarks, limited explainability, and difficulty evaluating end-to-end effectiveness in realistic security workflows. It points to gaps in agent safety, guardrail design, human-in-the-loop controls, provenance and auditability, and integration with existing SIEM/SOAR/EDR stacks. By clustering these themes, the paper provides a useful baseline for researchers and practitioners seeking to assess maturity, compare architectures, and prioritize future work.
The material matters because LLM-based agents are rapidly changing the economics and attack surface of security operations: they can compress repetitive analyst work, scale monitoring, and support faster incident response, but they can also amplify errors and create new abuse vectors if deployed without strong controls. For a technically literate audience, the review is valuable as a map of the field, a vocabulary for design trade-offs, and a checklist of open problems that must be solved before LLM agents can be trusted in high-stakes defensive or offensive cybersecurity environments.