Explores whether AI agents can discover physical laws like Newton’s gravitation via iterative data analysis, mathematical pattern extraction, and theory refinement against observations.

Topological visualization of EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
Brave API

EurekaBench is a benchmark designed to evaluate AI agents' ability to conduct long-horizon experiments and discover underlying mechanisms that yield scientific insights, rather than merely optimizing for predictive accuracy.

  • Scope: It contains 26 tasks across six domains (neuroscience, computer science, chemistry, astrophysics, geophysics, and plasma physics), totaling 306 scientific insight questions.
  • Evaluation: Agents are tested on three axes: adherence to scientific constraints, predictive accuracy, and the derivation of scientific insights (SI).
  • Key Finding: While top AI agents (e.g., GPT 6 Astra, Claude Fable 5.1) can approach human-level predictive accuracy (~47-48%), they significantly underperform in scientific insight scores (~29-42%) compared to human scientists (69.7%), often treating discovery as optimization rather than interpretive understanding.
  • Methodology: The benchmark uses an agentic judge to evaluate whether discovered mechanisms satisfy domain-specific assumptions and can answer specific insight questions derived from expert collaboration.
Generated 2d ago
Open-Weights Reasoning

EurekaBench is a benchmark for evaluating whether AI agents can perform a more agentic form of scientific discovery than standard data fitting or retrieval-based reasoning. The material frames discovery as an iterative loop in which an agent must analyze observations, extract mathematical or structural patterns, propose candidate theories, and refine those theories in light of new or residual evidence. A representative task class involves recovering physical laws from empirical data, such as the functional form of Newton’s law of gravitation, where success requires not merely predicting values but identifying a compact, interpretable, and physically meaningful relationship.

A key contribution of the work is its focus on the full discovery pipeline rather than isolated capabilities such as regression, symbolic regression, or prompt-based equation guessing. The benchmark is positioned to measure whether agents can sustain hypothesis-driven search, distinguish genuine regularities from spurious correlations, and improve theories over multiple rounds of evidence. This is important because many existing evaluations of “scientific reasoning” in large language models test narrow inference or pattern recognition, while leaving open whether systems can operate as autonomous or semi-autonomous scientific agents.

The work matters because it provides a more demanding yardstick for agentic AI in scientific domains. If agents can reliably discover or rediscover foundational laws from data, it suggests they may be useful for accelerating exploration in under-theorized areas, generating candidate models, and narrowing the search space for human researchers. Conversely, the benchmark also helps clarify the current limits of LLM-based agents, particularly in long-horizon reasoning, self-critique, and the integration of mathematical structure with empirical constraints.

Generated 2d ago
Sources