Presents SWE-Gate, a repository-level benchmark that evaluates coding agents on both functional test passage and review-derived acceptance constraints.
Based on the provided search context, there is no benchmark named "SWE-Gate." The search results do not contain any information regarding a benchmark by this name that evaluates agents on "review-derived acceptance constraints."
The context does, however, discuss several related benchmarks that address the limitation of relying solely on functional tests:
While the concept of evaluating agents on acceptance criteria beyond functional tests is present in these benchmarks, the specific entity "SWE-Gate" is not supported by the provided data.
SWE-Gate is a repository-level benchmark for evaluating software engineering agents beyond the narrow criterion of passing existing functional tests. The core premise is that test passage is a necessary but incomplete signal of engineering quality: an agent may produce a patch that satisfies unit or integration tests while still violating acceptance constraints that would emerge in a real code-review process, such as API compatibility, edge-case handling, architectural invariants, or maintainability requirements. The benchmark addresses this gap by deriving additional acceptance constraints from review signals and evaluating patches against both test outcomes and those review-derived gates.
The key contribution is a more faithful evaluation protocol that treats “done” as a multi-criteria engineering decision rather than a purely test-driven one. By framing acceptance as a set of gates, SWE-Gate makes it possible to distinguish patches that are merely test-passing from those that are plausibly merge-ready. Its central insight is that current coding agents can exhibit a significant divergence between functional correctness and review-level correctness, suggesting that test-only benchmarks may overstate agent capability and understate the risk of subtle, review-relevant failures.
This matters because it shifts the evaluation focus for autonomous software engineering systems toward the standards actually used in professional development workflows. If agents are to be trusted with repository-level changes, benchmarks need to measure not only whether code runs as tested, but whether it satisfies the broader constraints that developers, reviewers, and downstream systems expect. SWE-Gate therefore provides a useful step toward more realistic, review-aware evaluation of coding agents.