Presents SWE-Gate, a repository-level benchmark that evaluates coding agents on both functional test passage and review-derived acceptance constraints.

Topological visualization of SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents
Brave API

Based on the provided search context, there is no benchmark named "SWE-Gate." The search results do not contain any information regarding a benchmark by this name that evaluates agents on "review-derived acceptance constraints."

The context does, however, discuss several related benchmarks that address the limitation of relying solely on functional tests:

  • SWE-Bench evaluates agents on real GitHub issues using fail-to-pass and pass-to-pass tests to verify bug resolution and regression prevention. However, researchers note that manually written tests can be insufficient, allowing patches to pass without resolving the underlying issue.
  • SWE Atlas extends evaluation beyond issue resolution to include Codebase Q&A, Test Writing, and Refactoring. It uses expert-authored rubrics to assess engineering rigor, including code placement, anti-patterns, and maintainability, rather than just test passage.
  • SWE-EVO focuses on long-horizon software evolution, requiring agents to interpret high-level specifications and plan multi-step modifications across versions, capturing broader regression risks.
  • SWE-CI shifts the paradigm from static functional correctness to dynamic long-term maintainability by evaluating agents within a Continuous Integration loop over extensive commit histories.

While the concept of evaluating agents on acceptance criteria beyond functional tests is present in these benchmarks, the specific entity "SWE-Gate" is not supported by the provided data.

Generated Sep 4, 2026
Open-Weights Reasoning

SWE-Gate is a repository-level benchmark for evaluating software engineering agents beyond the narrow criterion of passing existing functional tests. The core premise is that test passage is a necessary but incomplete signal of engineering quality: an agent may produce a patch that satisfies unit or integration tests while still violating acceptance constraints that would emerge in a real code-review process, such as API compatibility, edge-case handling, architectural invariants, or maintainability requirements. The benchmark addresses this gap by deriving additional acceptance constraints from review signals and evaluating patches against both test outcomes and those review-derived gates.

The key contribution is a more faithful evaluation protocol that treats “done” as a multi-criteria engineering decision rather than a purely test-driven one. By framing acceptance as a set of gates, SWE-Gate makes it possible to distinguish patches that are merely test-passing from those that are plausibly merge-ready. Its central insight is that current coding agents can exhibit a significant divergence between functional correctness and review-level correctness, suggesting that test-only benchmarks may overstate agent capability and understate the risk of subtle, review-relevant failures.

This matters because it shifts the evaluation focus for autonomous software engineering systems toward the standards actually used in professional development workflows. If agents are to be trusted with repository-level changes, benchmarks need to measure not only whether code runs as tested, but whether it satisfies the broader constraints that developers, reviewers, and downstream systems expect. SWE-Gate therefore provides a useful step toward more realistic, review-aware evaluation of coding agents.

Generated Sep 4, 2026
Sources