Introduces ExecCritic, a test-verify-revise scaffold plus role-specific RL that prevents coding agents from generating mutually reinforcing but incorrect patches and tests.

Topological visualization of ExecCritic: Learn to Test, Test to Improve for Coding Agents
Brave API

ExecCritic is a framework that separates test construction from source-code repair to prevent coding agents from generating mutually reinforcing but incorrect patches. It employs a test–verify–revise scaffold where a Test agent generates repository-native tests, a harness freezes and executes them, and a Repair agent revises source code based on execution feedback without altering the tests.

The framework utilizes role-specific reinforcement learning: Learn to Test: The Test agent is trained to produce behaviorally valid tests that distinguish correct patches from incorrect ones. Test to Improve: The Repair agent learns both direct task resolution and feedback-guided revision using the fixed test feedback.

On SWE-bench Verified, composing post-trained Qwen agents achieved a 72.6% resolved rate, an 11.4-point gain over the no-test baseline, demonstrating that independent, fixed tests significantly improve repair quality compared to coupled agent trajectories.

Generated 25d ago
Open-Weights Reasoning

ExecCritic addresses a core reliability problem in autonomous coding agents: when an agent writes both a code patch and the tests used to validate it, the two artifacts can become jointly wrong in a self-consistent way. A patch that implements a misunderstanding of the specification may be paired with tests that encode the same misunderstanding, so the tests pass even though the solution is incorrect. The paper proposes a test–verify–revise scaffold that makes executable testing an explicit part of the agent’s improvement loop: the agent generates candidate tests, runs them against the current patch, and uses the resulting pass/fail evidence to revise either the patch, the tests, or both. This shifts validation from “the model believes the code is correct” to “the code survives an executable check.”

The key methodological contribution is to pair this scaffold with role-specific reinforcement learning rather than training a single monolithic coding model. The system learns distinct behaviors for different roles—such as generating discriminative, specification-faithful tests and producing patches that satisfy valid tests—using reward signals that are aligned with executable verification. This design is intended to prevent the collapse of tests into trivial or overly permissive checks that merely reinforce the patch’s current behavior. In effect, the critic learns to test in a way that can expose failures, while the solver learns to improve its patch in response to concrete execution evidence.

This matters because coding agents are increasingly expected to propose and validate software changes autonomously, and self-consistency alone is a weak guarantee of correctness. By separating generation from verification and grounding both in execution, ExecCritic offers a practical mechanism for reducing false-positive validation and improving patch quality. More broadly, the work highlights a useful design principle for agentic software engineering: robust improvement requires not only generating code, but also learning to test that code in ways that are executable, discriminative, and resistant to self-reinforcing errors.

Generated 25d ago
Sources