Identifies a metering vulnerability in tool-calling LLM agent runtimes where external tool returns can be charged multiple times across conversation turns.
Persistent Billable State is a security vulnerability in tool-calling LLM agents where host runtimes preserve external tool returns across conversation turns, causing providers to meter and bill for the same data repeatedly. This creates a Denial-of-Wallet (DoW) attack vector where a malicious or compromised tool can inject untrusted data that accumulates in the context, leading to recurring victim-billed processing without requiring victim credentials.
Research by Zhang et al. (September 2026) demonstrates that this persistent billable-state boundary allows attackers to convert a single tool admission into recurring spending authority. In evaluations, cumulative input reached 14,293x the first-call input, and retaining raw history increased mean session costs by 21.2–35.9%. The study identifies six attack vectors, including direct recursive baseline and polymorphic mutation, which exploit the lack of host-side safeguards to exhaust session budgets.
Defenses involve governing the pre-reingestion boundary through deterministic history transformation and four host-side invariants that bound prompt mass, context growth, recursive opportunity, and cumulative spend. Effective mitigation strategies include compression of retained state (preserving 10/12 task successes vs. 2/12 under deletion) and progress-authorized policies that limit unchecked context accumulation while maintaining workflow continuity.
This material examines a security and billing flaw in tool-calling LLM agent runtimes, where the system’s persistent conversational or execution state can cause the same external tool result to be treated as a new billable event across multiple turns. It frames the issue as a denial-of-wallet vulnerability: rather than degrading performance or leaking data, an attacker or faulty workflow can repeatedly trigger charges for already-consumed external tool calls, exhausting a user’s credits, API budget, or account balance. The central insight is that agent runtimes often conflate state persistence—keeping tool outputs available for context, retry, or summarization—with billing persistence, thereby creating a metering surface that is not protected by the same guarantees as normal API idempotency.
The paper’s key contribution is to identify and articulate this “persistent billable state” attack surface, showing how ordinary runtime behaviors—such as retaining tool results in history, rehydrating cached outputs, retrying failed steps, or reinterpreting prior tool calls in later prompts—can lead to duplicate or repeated charges. It then outlines defensive patterns, including idempotent billing identifiers, canonical tool-result provenance, charge-once semantics for semantically identical invocations, ledger reconciliation, and tighter lifecycle controls over billable state. The work also highlights the need to separate concerns: an agent may legitimately reuse a tool result for reasoning, but that reuse should not necessarily generate a new metering event unless the external service semantics require it.
This matters because tool-calling agents are increasingly deployed in autonomous, multi-turn, and high-volume settings where they invoke paid external services such as search, data APIs, payment processors, or cloud functions. In that context, metering correctness is not merely an accounting detail but a first-class security property. A denial-of-wallet vulnerability can be used for economic denial of service, budget exhaustion, or abuse by prompt-injected or malicious tool responses, and it can erode trust in agentic systems that operate under delegated spending authority. The material therefore provides a useful lens for developers, platform providers, and auditors to evaluate whether agent runtimes treat billing state with the same rigor as authentication, authorization, and idempotency.