Presents EngiWorld, a benchmark of 1,301 expert tasks across six engineering domains that evaluates agents on complete geometry- and physics-aware design loops.
EngiWorld is the first benchmark structured around the complete engineering design loop, featuring 1,301 expert-curated tasks across six domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms. It utilizes an artifact-centric evaluation methodology with a unified domain-verifier suite to programmatically check geometric validity, physical feasibility, and rule compliance.
Evaluation of seven frontier models reveals a substantial capability gap: the strongest model, Claude Opus 5, achieved an EngiScore of only 44.3, and just 3.6% of multi-software attempts succeeded. The benchmark covers both GUI and CLI interfaces and includes six task types ranging from software-selection to open-ended tasks, highlighting that reliable end-to-end automation of professional engineering workflows remains out of reach for current AI agents.
EngiWorld is a benchmark for evaluating frontier agents in professional engineering settings, where success requires more than answering domain questions or writing isolated code. It contains 1,301 expert-level tasks across six engineering domains and is structured around complete geometry- and physics-aware design loops: agents must interpret engineering requirements, construct or modify geometric designs, apply relevant physical constraints, and iterate based on analysis or simulation feedback. This makes the benchmark significantly closer to real engineering practice than static QA-style evaluations, because a solution is not just semantically plausible but must remain coherent across geometry, physics, manufacturability, and task objectives.
The key contribution is a rigorous, closed-loop evaluation framework for agentic engineering capability. By forcing agents to operate across the full design cycle, EngiWorld exposes the difference between local competence—such as manipulating CAD parameters or running a single analysis—and the long-horizon planning, error recovery, and constraint tracking needed for professional work. The benchmark therefore provides a more realistic yardstick for what current frontier agents can actually deliver in high-stakes technical environments, where small design errors can cascade into safety, cost, or feasibility failures.
It matters because engineering is one of the domains where autonomous agents could create large practical value, but also where unreliable behavior is especially costly. EngiWorld gives researchers and builders a common testbed for measuring progress toward trustworthy engineering agents, and it highlights the remaining gap between useful assistance on bounded subtasks and the level of reliability required for end-to-end professional design workflows.