Presents Paper2Agent, an open-source framework that converts papers plus codebases into tested, interactive AI agents runnable on new datasets.
Paper2Agent is an open-source, multi-agent AI framework developed by Stanford researchers that automatically transforms static research papers and their associated codebases into interactive, runnable AI agents. Published in Nature on September 16, 2026, the system converts papers into Model Context Protocol (MCP) servers, allowing users to query methods, reproduce results, and analyze new datasets via natural language without manual coding.
The framework uses a six-step pipeline involving specialized agents to extract code, configure environments, wrap methods as executable tools, and validate outputs against reference results. In testing across 100 computational biology papers, it successfully converted 74 into working agents with 593 validated tools, achieving 91.2% accuracy on benchmark queries. Notable case studies include an AlphaGenome agent that generated 22 tools in 45 minutes for roughly $14 and achieved 100% accuracy on novel queries, and collaborative agents that identified causal genes for diseases like psoriasis and ADHD.
The IEEE Spectrum piece describes Paper2Agent, an open-source framework designed to bridge the gap between published AI research and practical use. Instead of treating a paper as a static document and its code as a disposable artifact, the framework attempts to convert both into a working, testable AI agent. In doing so, it targets a common reproducibility problem: many research methods are described in prose, implemented in unevenly structured code, and validated only on the authors’ original datasets, making them difficult to adopt or adapt.
A key contribution is the framing of a research paper and its codebase as inputs to an executable software agent. The system appears to automate or assist in the process of extracting the method’s intent from the paper, aligning it with the code’s actual behavior, and producing an interface that can be run against new data. The emphasis on “tested” agents is especially important: it suggests that the goal is not merely code generation, but the creation of a validated tool that can be trusted for downstream experimentation. This shifts the unit of reuse from a single script or model checkpoint to a more self-contained, interactive capability.
The broader significance is that Paper2Agent could make research artifacts more actionable and composable. For technically literate users, this means less time spent reverse-engineering papers, debugging incomplete repositories, and manually adapting code to new datasets, and more time spent evaluating methods, extending them, or applying them to novel problems. More generally, it points toward a research ecosystem in which published methods are not only readable but runnable—turning the literature into a living library of AI tools rather than a collection of static reports.