Defines recursive self-improvement (RSI) for LLMs and outlines a five-stage autonomy roadmap, using the Headroom-Closed Index to diagnose current model limitations.
The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement (arXiv:2609.11873, submitted September 10, 2026) defines Recursive Self-Improvement (RSI) as an autonomous, closed-loop process where an AI system identifies its own limitations, develops and validates improvements, and uses those capabilities to improve the improvement process itself.
The paper introduces the Headroom-Closed Index (HCI) to reveal the limitations of existing Large Language Models (LLMs) and proposes a five-stage autonomy roadmap for achieving genuine RSI:
1. Improvement-Execution Autonomy (L1): The AI executes prescribed improvements. 2. Improvement-Strategy Autonomy (L2): The AI selects the strategies for improvement. 3. Experience-Acquisition Autonomy (L3): The AI determines what learning experience to acquire next. 4. Environment-Adaptation Autonomy (L4): The AI adapts through deployment and environmental feedback. 5. Recursive Meta-Improvement (L5): The AI improves the mechanisms that govern subsequent improvements.
This framework distinguishes RSI from simpler automation or fine-tuning by emphasizing the transfer of responsibility for improvement decisions from external designers to the AI system itself.
The paper examines what it would mean for large language models to achieve genuine recursive self-improvement (RSI), as distinct from ordinary fine-tuning, human-supervised optimization, or iterative prompt engineering. It frames RSI as a closed-loop process in which a model’s outputs help produce a successor model with improved capabilities, and it uses that framing to assess where current LLMs stand relative to the capabilities required for such a loop to become self-sustaining.
Its core contributions are conceptual and diagnostic. The material provides a definition of RSI tailored to LLM-based systems, distinguishes different levels of autonomy, and outlines a five-stage roadmap progressing from human-directed model improvement toward increasingly self-directed optimization. It also uses the Headroom-Closed Index to diagnose present limitations by measuring how much capability “headroom” remains before an LLM could close the self-improvement loop. This turns a speculative capability question into a more structured assessment of missing model, tooling, evaluation, and governance capabilities.
The work matters because it reframes RSI from an abstract safety concern into a testable capability framework. For researchers, the roadmap and index offer a shared vocabulary for identifying what current systems can and cannot do. For safety, policy, and deployment planning, it provides a way to reason about thresholds at which self-improvement may become harder to supervise, without relying on vague claims about “superintelligence.” In short, it supplies a structured account of what genuine RSI would require and how far existing LLMs are from that condition.