Audits 22 frontier LLMs across 12 molecular regression benchmarks and finds widespread but dataset-specific verbatim retrieval of published values rather than genuine property prediction.
Molecular Déjà Vu reveals that 22 frontier LLMs frequently exhibit verbatim retrieval of published molecular property values rather than genuine prediction, a phenomenon that is widespread but benchmark-specific.
This paper presents an empirical audit of frontier large language models on molecular regression tasks, where the goal is to predict numeric molecular properties from chemical structures or descriptions. Across 22 frontier LLMs and 12 molecular regression benchmarks, the authors examine not only aggregate predictive accuracy, but also whether model outputs reproduce published values at the digit level. The central question is whether strong benchmark performance reflects genuine learned structure–property reasoning or, instead, verbatim retrieval of values already present in the model’s training data or accessible public literature.
The key finding is that digit-level retrieval is widespread but highly dataset-specific. On some benchmarks, models appear to recover exact or near-exact published values, suggesting that apparent predictive skill may be substantially inflated by benchmark contamination. On other datasets, this effect is weaker, implying that model rankings and conclusions about LLM capability can depend strongly on the provenance and public visibility of the evaluation set. The paper therefore highlights an important confound in molecular AI evaluation: standard metrics such as RMSE, MAE, or correlation can reward memorized answers in the same way they reward true generalization, making it difficult to distinguish retrieval from prediction without explicit leakage diagnostics.
This matters because molecular regression is often framed as a test of whether language models can function as scientific instruments for property prediction, drug discovery, or materials design. If performance is driven by retrieval of published values, then models may fail on novel molecules, unpublished measurements, or private experimental data—precisely the settings where scientific utility is most needed. The work underscores the need for decontaminated, temporally held-out, or privately sourced benchmarks, as well as evaluation protocols that report retrieval diagnostics alongside predictive metrics. More broadly, it cautions against overinterpreting strong LLM scores on public chemistry benchmarks as evidence of robust molecular understanding.