Empirically tests whether LLM decision-component explanations satisfy necessity and sufficiency relative to the component's observable action behavior.

Topological visualization of Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence
Brave API

Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence empirically tests whether the factors cited by Large Language Models (LLMs) in their explanations align with the model's observable decision behavior, specifically defining necessity as the change in output when a factor is altered, and sufficiency as the preservation of output when only that factor is retained.

The study evaluated eight models from the Claude, GPT, and Gemini families across two use cases: advisor recommendation and prompt monitoring. Key findings include:

  • Partial Agreement: The mean Spearman correlations between cited rankings and measured necessity/sufficiency scores were low, ranging from 0.349 to 0.580 depending on the use case and metric.
  • Unreliable Identification: Cited top-three factors did not reliably identify the three most influential features; an uncited factor scored higher than the lowest-cited factor in 57.6% of advisor responses under necessity and 58.1% under sufficiency.
  • Salience Gaps: In prompt monitoring, frequently cited risk categories (e.g., jailbreak attempts) did not always correspond to the features with the strongest measured influence, indicating a disconnect between citation frequency and actual behavioral impact.
  • Model Variability: Performance varied significantly across models and use cases, with no single model consistently outperforming others across all criteria.

The authors conclude that while LLM explanations contain useful information, they should be treated as testable claims rather than verified accounts of internal reasoning, necessitating external reliability checks for agent oversight.

Generated 27d ago
Open-Weights Reasoning

Summary

This paper evaluates LLM explanations of decision components by asking whether they satisfy two distinct behavioral criteria: necessity and sufficiency. Rather than judging an explanation only by its linguistic plausibility or coherence, the work treats an explanation as a testable claim about a component’s role in producing observable action behavior. Under this framing, a component is necessary if the relevant action would not occur, or would be substantially less likely, without it; it is sufficient if the component, under the relevant conditions, is enough to account for the action. The paper’s central methodological move is to ground these explanatory claims in behavioral evidence, comparing the model’s stated account of a component with what actually changes when that component is present, absent, or altered.

Its key contribution is a more rigorous diagnostic for LLM explanation quality, one that separates two often-conflated standards of causal or functional explanation. An explanation may identify a component that is required for the observed behavior but not enough to predict it, or it may identify a factor that can produce the behavior in isolation but is not actually relied on in the decision at hand. By distinguishing these cases, the paper exposes failure modes that are invisible to purely textual evaluation: explanations can sound precise and internally consistent while failing to track the model’s actual action behavior. This is especially relevant for LLMs used in agentic or decision-making settings, where explanations are often used to audit, debug, or trust system behavior.

The broader significance of the work is that it reframes explanation quality as a measurable alignment between stated components and observable behavior. This matters because LLM explanations are frequently post-hoc rationalizations: they may name plausible causes, goals, or constraints that are correlated with the action but not causally responsible for it. By evaluating explanations against behavioral evidence of necessity and sufficiency, the paper provides a stronger benchmark for claims that LLMs can explain their own decision processes, and it highlights a gap between an LLM’s ability to produce fluent justifications and its ability to provide behaviorally grounded accounts of why it acted.

Generated 27d ago
Sources