Introduces the RATIO benchmark defining literature relevance via three ideation moves (Address, Generalize, Realize) for retrieval that supports scientific inspiration at varying abstraction levels.

Topological visualization of RATIO: A Benchmark for Retrieval Across Typed Ideation Operations in Scientific Literature
Brave API

RATIO (Retrieval Across Typed Ideation Operations) is a large-scale benchmark introduced by Maayan Sharon and Tom Hope in August 2026 that defines scientific literature relevance through three specific ideation moves: Address, Broaden, and Specify.

  • Address: Retrieves potential approaches or insights that directly respond to a stated problem in the query.
  • Broaden: Retrieves more general formulations or wider scope directions, effectively "zooming out" from the specific problem.
  • Specify: Retrieves concrete instantiations or detailed realizations, effectively "zooming in" from the general concept.

The benchmark is constructed from millions of full-text scientific papers in Computer Science literature using discourse-marker distant supervision, validated through LLM and human vetting. Experiments show that relation-specific fine-tuning substantially improves retrieval performance, though a significant gap remains compared to human-level inspiration retrieval capabilities.

Generated Aug 30, 2026
Open-Weights Reasoning

RATIO: A Benchmark for Retrieval Across Typed Ideation Operations in Scientific Literature introduces a benchmark for scientific literature retrieval that reframes relevance in terms of ideation operations rather than topical or lexical similarity alone. Instead of asking only which papers are most semantically similar to a query, RATIO asks what kind of intellectual relationship a paper has to the target idea or problem. This makes the benchmark especially relevant for discovery settings where a researcher may need not only direct prior work, but also broader abstractions or concrete instantiations that can inspire new directions.

The central contribution is a typology of three retrieval-relevant moves: Address, Generalize, and Realize. Address captures literature that directly engages a specific problem, method, phenomenon, or question. Generalize captures literature that lifts the target idea into a broader class, principle, or theoretical frame. Realize captures literature that grounds the idea in a concrete setting, such as an implementation, experiment, dataset, application, or example. By defining relevance through these typed operations, RATIO shifts the evaluation target from “nearest neighbor” retrieval to more functionally meaningful retrieval: systems must distinguish papers that solve the same problem, papers that abstract over it, and papers that instantiate it in a new context.

This matters because scientific inspiration often depends on moving across levels of abstraction. A useful literature search should not merely return the most similar work, but should also surface work that can reframe, extend, or concretize an idea. RATIO provides a structured way to evaluate whether retrieval systems can support that process, making it valuable for benchmarking literature search, research assistants, and AI-for-science tools. More broadly, it offers a principled way to measure retrieval systems on their ability to support creative and cross-domain scientific reasoning rather than narrow topical matching.

Generated Aug 30, 2026
Sources