Explores whether AI agents can discover physical laws like Newton’s gravitation via iterative data analysis, mathematical pattern extraction, and theory refinement against observations.
EurekaBench is a benchmark designed to evaluate AI agents' ability to conduct long-horizon experiments and discover underlying mechanisms that yield scientific insights, rather than merely optimizing for predictive accuracy.
EurekaBench is a benchmark for evaluating whether AI agents can perform a more agentic form of scientific discovery than standard data fitting or retrieval-based reasoning. The material frames discovery as an iterative loop in which an agent must analyze observations, extract mathematical or structural patterns, propose candidate theories, and refine those theories in light of new or residual evidence. A representative task class involves recovering physical laws from empirical data, such as the functional form of Newton’s law of gravitation, where success requires not merely predicting values but identifying a compact, interpretable, and physically meaningful relationship.
A key contribution of the work is its focus on the full discovery pipeline rather than isolated capabilities such as regression, symbolic regression, or prompt-based equation guessing. The benchmark is positioned to measure whether agents can sustain hypothesis-driven search, distinguish genuine regularities from spurious correlations, and improve theories over multiple rounds of evidence. This is important because many existing evaluations of “scientific reasoning” in large language models test narrow inference or pattern recognition, while leaving open whether systems can operate as autonomous or semi-autonomous scientific agents.
The work matters because it provides a more demanding yardstick for agentic AI in scientific domains. If agents can reliably discover or rediscover foundational laws from data, it suggests they may be useful for accelerating exploration in under-theorized areas, generating candidate models, and narrowing the search space for human researchers. Conversely, the benchmark also helps clarify the current limits of LLM-based agents, particularly in long-horizon reasoning, self-critique, and the integration of mathematical structure with empirical constraints.