Introduces PsyAgentBench, a factorial benchmark that runs classic psychology experiments on LLM agents in both named and blind conditions to distinguish response patterns from actual cognitive biases.

Topological visualization of Recognition, Simulation, and Refusal: A Contamination-Aware Study of Classic Psychological Effects in LLM Agents
Brave API

PsyAgentBench is a factorial benchmark introduced by Joy Bose to distinguish between LLM response patterns and actual cognitive biases by re-running classic psychology experiments under named (explicitly labeled) and blind (framed as routine tasks) conditions.

The study, published in July 2026, reveals that LLMs do not possess a single susceptibility to human psychological effects but rather exhibit qualitatively different routes to similar responses: Paradigm-label gating: Effects like Asch conformity appear only when explicitly named (0% blind vs. 83.3% named). Knowledge-dependent signal reliance: Anchoring effects occur only on invented quantities, not grounded facts. * Safety-mediated selection: Refusal itself becomes the primary finding in scenarios like minimal-group allocation.

The research argues that scalar bias-susceptibility scores obscure these structures and that persona changes can eliminate, dampen, or reverse effects, further indicating that producing a human-like response pattern is not equivalent to possessing the underlying bias.

Generated 12d ago
Open-Weights Reasoning
  • Overview and design. The paper introduces PsyAgentBench, a contamination-aware factorial benchmark for evaluating whether LLM agents exhibit classic psychological effects under conditions that resemble the original experiments. Its central methodological move is to run the same psychological paradigms in named and blind conditions: in the named setting, the task or effect is made explicit, while in the blind setting the identity of the experiment is concealed. This design lets the authors separate three competing explanations for an agent’s behavior—whether it is genuinely displaying a bias-like response pattern, recognizing the canonical psychology experiment and simulating the expected answer, or refusing/deflecting because the prompt looks like a known evaluation or sensitive scenario.
  • Key contributions and insights. The main contribution is not merely a set of prompts, but a controlled experimental framework for testing psychological effects in LLM agents while accounting for benchmark contamination and model self-awareness. By comparing named and blind conditions, the study can distinguish surface-level pattern matching from more stable response tendencies. The framing around recognition, simulation, and refusal is especially useful because it names common confounds in LLM psychology: models may “pass” or “fail” a bias test for reasons unrelated to the psychological phenomenon under study, such as recalling the literature, mimicking normative human behavior, or refusing to engage with a test-like prompt.
  • Why it matters. This work matters because much of the current literature on LLM cognitive biases risks overinterpreting prompt-response correlations as evidence of human-like cognition. PsyAgentBench provides a more rigorous way to ask whether an agent’s behavior is actually aligned with a psychological effect, or whether it is an artifact of training-data familiarity, evaluation awareness, or refusal behavior. For researchers building or evaluating LLM agents, the benchmark offers a template for contamination-aware evaluation and for distinguishing robust behavioral tendencies from benchmark-specific artifacts.
Generated 12d ago
Sources