Presents the E2E-SWE benchmark for repository-scale coding agents, requiring system-level reasoning with precisely specified, implementation-independent tasks.
E2E-SWE is a benchmark introduced in September 2026 that evaluates whether coding agents can build complete, functional software repositories from scratch using only natural-language specifications and an empty workspace. The benchmark consists of 186 tasks spanning 11 programming languages, where agents must produce installable projects that pass hidden test suites without network access or prior codebase context.
Evaluation of 13 frontier models reveals significant performance variation, with pass@1 scores ranging from 11.7% to 67.7%. Claude Opus 5 achieved the highest performance at 67.7%, followed by Claude Fable 5 (65.3%) and GPT-5.6 Sol (61.2%). The benchmark highlights that whole-repository generation requires long-horizon, front-loaded planning and system-level reasoning, leaving considerable headroom for future model improvements.
E2E-SWE is a benchmark for evaluating large language models as repository-scale coding agents, with a focus on the ability to build complete, working codebases from scratch rather than editing or patching existing projects. The benchmark targets end-to-end software engineering tasks that require system-level reasoning: an agent must interpret high-level specifications, decompose them into coherent components, make architectural and interface decisions, and produce integrated code that can be built and executed. By emphasizing precisely specified, implementation-independent tasks, E2E-SWE aims to reduce reliance on memorized patterns, template copying, or narrow function-level completion, and instead measure whether models can design software that satisfies functional and structural requirements at the level of a real project.
Its key contribution is a shift in evaluation from local code generation to repository-level construction. Instead of asking a model to fill in a function or fix a localized bug, E2E-SWE tests whether it can reason across files, manage dependencies, define APIs, maintain consistency, and deliver a codebase that works as a system. This is important because many existing coding benchmarks may overstate an agent’s practical software-engineering ability by focusing on tasks that are easier to solve with pattern matching or limited context. E2E-SWE therefore provides a more demanding test of long-horizon planning, abstraction, integration, and verification—capabilities that are central to realistic agentic coding workflows.
The benchmark matters because it exposes the gap between code completion and autonomous software construction. As LLM agents are increasingly expected to generate multi-file applications, services, or tools, the ability to produce a coherent, buildable system is a critical capability. E2E-SWE gives researchers and practitioners a structured way to measure that capability, encouraging progress in planning, self-verification, debugging, and system design. In short, it frames coding-agent evaluation around a more end-to-end question: can the model not only write plausible code, but assemble it into a functioning software system?