Presents CompMat-Bench, a benchmark of 94 tasks from recent computational materials papers that evaluates agents without repeating expensive simulations.
CompMat-Bench is a benchmark designed to evaluate AI agents on 94 tasks derived from 15 recently published computational materials science studies. To avoid the high cost of repeating simulations, the benchmark precomputes expensive simulations and evaluates agents solely on their ability to prepare input files and analyze the resulting outputs.
The benchmark employs fixed rules for grading rather than an LLM judge, ensuring consistent and objective assessment. Agents are tested under four conditions that vary by task span (single tasks vs. multi-step workflows) and methodological guidance (full vs. reduced).
Key findings indicate that while top-performing agents achieve pass rates between 66.0% and 91.5% on single tasks with full guidance, performance declines with longer workflows and reduced guidance. Failure analysis reveals that most errors are scientific in nature rather than issues with software usage.
CompMat-Bench is a benchmark for evaluating AI agents on realistic tasks in computational materials science. Rather than requiring agents to run new, computationally expensive simulations—such as large-scale density functional theory, molecular dynamics, or high-throughput screening—it curates 94 tasks drawn from recent computational materials papers and frames them so that agents can be assessed using the information, data, and computational context already available in the published work. The benchmark is therefore oriented toward practical scientific competence: can an agent understand a materials-science problem, identify the relevant inputs and assumptions, construct or adapt a computational analysis, and draw conclusions that are consistent with the published result?
A key contribution is the move from generic coding or reasoning benchmarks toward a domain-specific, task-grounded evaluation for AI agents in materials research. Computational materials science is especially challenging for agents because it spans literature interpretation, numerical modeling, unit and property reasoning, approximation awareness, and often heterogeneous toolchains. By anchoring tasks to real papers and avoiding the need to rerun costly simulations, CompMat-Bench makes it possible to evaluate agents at scale while preserving the scientific complexity of the original problems. The result is a more tractable and reproducible way to measure whether agents can perform useful work in a field where simulation cost, methodological nuance, and domain knowledge all matter.
This matters because AI agents are increasingly positioned as assistants for scientific discovery, but their usefulness in materials science cannot be inferred from general-purpose benchmarks alone. CompMat-Bench provides a concrete testbed for identifying what agents can and cannot do in a high-value, compute-intensive domain: whether they can navigate published methods, handle quantitative uncertainty, reason about model limitations, and produce scientifically defensible analyses. As a benchmark, it helps researchers and developers calibrate expectations, prioritize agent capabilities, and track progress toward systems that can meaningfully contribute to computational materials workflows.