Empirically tests whether LLM decision-component explanations satisfy necessity and sufficiency relative to the component's observable action behavior.
Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence empirically tests whether the factors cited by Large Language Models (LLMs) in their explanations align with the model's observable decision behavior, specifically defining necessity as the change in output when a factor is altered, and sufficiency as the preservation of output when only that factor is retained.
The study evaluated eight models from the Claude, GPT, and Gemini families across two use cases: advisor recommendation and prompt monitoring. Key findings include:
The authors conclude that while LLM explanations contain useful information, they should be treated as testable claims rather than verified accounts of internal reasoning, necessitating external reliability checks for agent oversight.
Summary
This paper evaluates LLM explanations of decision components by asking whether they satisfy two distinct behavioral criteria: necessity and sufficiency. Rather than judging an explanation only by its linguistic plausibility or coherence, the work treats an explanation as a testable claim about a component’s role in producing observable action behavior. Under this framing, a component is necessary if the relevant action would not occur, or would be substantially less likely, without it; it is sufficient if the component, under the relevant conditions, is enough to account for the action. The paper’s central methodological move is to ground these explanatory claims in behavioral evidence, comparing the model’s stated account of a component with what actually changes when that component is present, absent, or altered.
Its key contribution is a more rigorous diagnostic for LLM explanation quality, one that separates two often-conflated standards of causal or functional explanation. An explanation may identify a component that is required for the observed behavior but not enough to predict it, or it may identify a factor that can produce the behavior in isolation but is not actually relied on in the decision at hand. By distinguishing these cases, the paper exposes failure modes that are invisible to purely textual evaluation: explanations can sound precise and internally consistent while failing to track the model’s actual action behavior. This is especially relevant for LLMs used in agentic or decision-making settings, where explanations are often used to audit, debug, or trust system behavior.
The broader significance of the work is that it reframes explanation quality as a measurable alignment between stated components and observable behavior. This matters because LLM explanations are frequently post-hoc rationalizations: they may name plausible causes, goals, or constraints that are correlated with the action but not causally responsible for it. By evaluating explanations against behavioral evidence of necessity and sufficiency, the paper provides a stronger benchmark for claims that LLMs can explain their own decision processes, and it highlights a gap between an LLM’s ability to produce fluent justifications and its ability to provide behaviorally grounded accounts of why it acted.