Audits 22 frontier LLMs across 12 molecular regression benchmarks and finds widespread but dataset-specific verbatim retrieval of published values rather than genuine property prediction.

Topological visualization of Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models
Brave API

Molecular Déjà Vu reveals that 22 frontier LLMs frequently exhibit verbatim retrieval of published molecular property values rather than genuine prediction, a phenomenon that is widespread but benchmark-specific.

  • Scope and Findings: Audits across 12 regression benchmarks showed that five datasets (FreeSolv, ESOL, LD50, AqSolDB, and a boiling-point control) had over 50% of models retrieving values verbatim, while others showed minimal contamination.
  • Key Drivers: Retrieval correlates strongly with how often molecules appear in pretraining documents ($\rho=0.88$), indicating secondary source exposure is the primary cause, not the benchmark file itself.
  • Impact on Evaluation: Reasoning increases retrieval (flagged 89% more often at higher levels), and blinding models with transformed SMILES strings converges error rates, proving that leaderboard rankings are often artifacts of memorization rather than predictive capability.
Generated 26d ago
Open-Weights Reasoning

This paper presents an empirical audit of frontier large language models on molecular regression tasks, where the goal is to predict numeric molecular properties from chemical structures or descriptions. Across 22 frontier LLMs and 12 molecular regression benchmarks, the authors examine not only aggregate predictive accuracy, but also whether model outputs reproduce published values at the digit level. The central question is whether strong benchmark performance reflects genuine learned structure–property reasoning or, instead, verbatim retrieval of values already present in the model’s training data or accessible public literature.

The key finding is that digit-level retrieval is widespread but highly dataset-specific. On some benchmarks, models appear to recover exact or near-exact published values, suggesting that apparent predictive skill may be substantially inflated by benchmark contamination. On other datasets, this effect is weaker, implying that model rankings and conclusions about LLM capability can depend strongly on the provenance and public visibility of the evaluation set. The paper therefore highlights an important confound in molecular AI evaluation: standard metrics such as RMSE, MAE, or correlation can reward memorized answers in the same way they reward true generalization, making it difficult to distinguish retrieval from prediction without explicit leakage diagnostics.

This matters because molecular regression is often framed as a test of whether language models can function as scientific instruments for property prediction, drug discovery, or materials design. If performance is driven by retrieval of published values, then models may fail on novel molecules, unpublished measurements, or private experimental data—precisely the settings where scientific utility is most needed. The work underscores the need for decontaminated, temporally held-out, or privately sourced benchmarks, as well as evaluation protocols that report retrieval diagnostics alongside predictive metrics. More broadly, it cautions against overinterpreting strong LLM scores on public chemistry benchmarks as evidence of robust molecular understanding.

Generated 26d ago
Sources