Proposes Broad and Deep Recursive Self-exploration methods enabling RSIAgent to recursively improve open-source models past closed-source frontiers on OSWorld-v2 and Agent’s Last Exam.

Topological visualization of RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments
Brave API

RSIAgent is a training-free multi-agent framework that enables recursive self-improvement in new digital environments without updating model parameters. It coordinates curriculum, actor, and verifier agents to autonomously explore, validate outcomes, and construct reusable causal memory.

The framework employs a broad-then-deep exploration strategy: Broad Recursive Self-exploration (BRS) rapidly builds diverse environment coverage through parallel task execution. Deep Recursive Self-exploration (DRS) targets hard cases, hidden constraints, and boundary conditions using focused, target-driven practice.

This approach allows open-source models Kimi-K3 and GLM-5.3 to outperform frontier closed-source models like GPT-6 Astra and Claude Opus 5 on OSWorld-v2 and Agent’s Last Exam. Specifically, RSIAgent achieved a partial score of 78.98% on OSWorld 2.0 and 84.82% on Agents’ Last Exam, surpassing GPT-6 Astra’s scores of 72.60% and 82.26%, respectively.

Generated 17d ago
Open-Weights Reasoning

Overview. RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments presents an agent framework for recursively improving open-source models through autonomous exploration in new or under-specified environments. The core idea is to treat agent improvement as an iterative loop: the agent explores, collects experience, extracts reusable policies or skills, and updates itself so that the next round of exploration starts from a stronger baseline. Rather than relying primarily on static datasets or manually curated demonstrations, RSIAgent emphasizes self-generated experience as the engine of improvement, targeting hard, long-horizon agent benchmarks such as OSWorld-v2 and Agent’s Last Exam.

Key contributions. The paper introduces Broad and Deep Recursive Self-exploration as complementary mechanisms for scaling this process. “Broad” exploration is aimed at expanding coverage across tasks, states, tools, and failure modes, while “deep” exploration focuses on iteratively refining promising trajectories, diagnosing errors, and distilling them into more reliable strategies. Together, these methods allow RSIAgent to convert raw interaction data into progressively better behavior without assuming a fixed environment distribution. The reported result is especially notable because it suggests that open-source models can be recursively improved to the point of surpassing closed-source frontier baselines on these challenging agent benchmarks.

Why it matters. This work matters because it points toward a more scalable path for agent development: instead of depending on ever-larger human-labeled corpora or proprietary model access, an agent can bootstrap capability through structured self-exploration and self-improvement. That is particularly relevant for settings where environments are novel, tasks are compositional, and deployment conditions differ from training time. At the same time, the results also sharpen an important research and safety question: if agents can autonomously improve in complex environments, we need equally strong methods for evaluating, constraining, and auditing the quality and reliability of the improvements they generate.

Generated 17d ago
Sources