the entrance to AI for hardware

A Benchmark for Programmatic CAD

Factory part drawings, vendor machines with their assembly drawings, and assembled circuit boards in; parametric CAD out. About 100 public and 100 private cases across six tasks, scored on the geometry and the wiring the model builds.

6 tasks · ~100 public · ~100 private

Four orthographic views in, a CadQuery program out — graded on executed geometry, not appearances. 17,900 verified parts across 106 industrial families.

4 matched tasks
BenchCAD T1 · drawing → part
impeller, a BenchCAD 1.0 part family, turning coil spring, a BenchCAD 1.0 part family, turning double simplex sprocket, a BenchCAD 1.0 part family, turning spline hub, a BenchCAD 1.0 part family, turning handwheel, a BenchCAD 1.0 part family, turning
partcamera mounting bracket
taskT1
parts1

One case per BenchCAD 2.0 task: five reference models, and the wiring a model recovered from a board, drawn as a schematic.Five of the 106 BenchCAD 1.0 part families.

developed by: chosen by:
progress · Vision2Code

Where the Frontier Is

Every frontier model on Vision2Code in 2026, in two views — against release date, and against what a task costs.

OpenAI Anthropic Google Kimi Grok no tools with tools (agentic) frontier, no tools frontier, with tools 0.900.800.700.600.500.400.300.20 IoU-score ↑ Jan 2026FebMarAprMayJunJulAugSep release date → Claude Sonnet 4.6 (max) · no tools · IoU-score 0.2220 (re-graded) GPT-5.3 (max) · no tools · IoU-score 0.1793 (re-graded) · released Feb 2026 Gemini 3.1 Pro (thinking) · no tools · IoU-score 0.2890 (re-graded) Claude Opus 4.7 (max) · no tools · IoU-score 0.2692 (re-graded) GPT-5.5 (max) · no tools · vendor-reported 0.444 — OpenAI GPT-5.6 launch table, not re-graded; harness undisclosed Claude Opus 4.8 (max) · no tools · self-reported voxel IoU 0.273 (full set) — Anthropic Fable 5 / Mythos 5 system card, not re-graded Claude Mythos Preview (max) · no tools · self-reported voxel IoU 0.355 (full set) — Anthropic system card, not re-graded Claude Mythos 5 (max) · no tools · self-reported voxel IoU 0.384 (full set) — Anthropic system card, not re-graded Claude Sonnet 5 (max) · no tools · self-reported voxel IoU 0.322 (1,000-file subset) — Anthropic Opus 5.5 system card fig. 8.13.2.A, not re-graded Grok 4.5 (high) · no tools · IoU-score 0.3194 (re-graded on our scorer) · released Jul 8 2026 GPT-5.6 Terra (max) · no tools · vendor-reported 0.623 — OpenAI launch table, not re-graded; harness undisclosed GPT-5.6 Sol (max) · no tools · vendor-reported 0.706 — OpenAI launch table, not re-graded; harness undisclosed GPT-5.6 Luna (max) · no tools · vendor-reported 0.631 — OpenAI launch table, not re-graded; harness undisclosed Kimi K3 (max) · no tools · IoU-score 0.3670 (re-graded, we ran it) Claude Opus 5 (max) · no tools · self-reported voxel IoU 0.497 (1,000-file subset) — Anthropic Opus 5.5 system card fig. 8.13.2.A, not re-graded Grok 4.6 (xhigh) · no tools · IoU-score 0.3638 (re-graded on our scorer; run arranged by SpaceXAI) Claude Fable 5.1 (max) · no tools · self-reported voxel IoU 0.606 (1,000-file subset) — Anthropic Opus 5.5 system card fig. 8.13.2.A, not re-graded Claude Opus 5.5 (max) · no tools · self-reported voxel IoU 0.730 (1,000-file subset) — Anthropic Opus 5.5 system card fig. 8.13.2.A, not re-graded Claude Sonnet 5.5 (max) · no tools · self-reported voxel IoU 0.747 (1,000-file subset) — Anthropic Sonnet 5.5 system card fig. 8.13.2.A, not re-graded GPT-5.5 (max) · with tools · vendor-reported 0.558 — OpenAI GPT-5.6 launch table, Python tool; split and attempt budget unstated Claude Opus 4.8 (max) · with tools · self-reported voxel IoU 0.518 (1,000-file subset, Python tools) — Anthropic system card fig. 8.16.4.B Claude Mythos Preview (max) · with tools · self-reported voxel IoU 0.610 (1,000-file subset, Python tools) — Anthropic system card fig. 8.16.4.B Claude Mythos 5 (max) · with tools · self-reported voxel IoU 0.650 (1,000-file subset, Python tools) — Anthropic system card fig. 8.16.4.B Claude Sonnet 5 (max) · with tools · self-reported voxel IoU 0.519 (1,000-file subset, Python tools) — Anthropic Opus 5.5 system card fig. 8.13.2.A Grok 4.5 (high) · with tools · 0.7771 — our own agentic run on our scorer, Python sandbox GPT-5.6 Terra (max) · with tools · vendor-reported 0.782 — OpenAI launch table, Python tool; split and attempt budget unstated GPT-5.6 Sol (max) · with tools · vendor-reported 0.834 — OpenAI launch table, Python tool; split and attempt budget unstated GPT-5.6 Luna (max) · with tools · vendor-reported 0.739 — OpenAI launch table, Python tool; split and attempt budget unstated Claude Opus 5 (max) · with tools · self-reported voxel IoU 0.899 (1,000-file subset, Python tools) — Anthropic Opus 5.5 system card fig. 8.13.2.A Grok 4.6 (xhigh) · with tools · 0.8055 — our own agentic run on our scorer, Python sandbox Claude Fable 5.1 (max) · with tools · self-reported voxel IoU 0.926 (1,000-file subset, Python tools) — Anthropic Opus 5.5 system card fig. 8.13.2.A GPT-6 Astra · with tools · 0.959 (1,000-file subset) — OpenAI GPT-6 Astra launch coverage; no no-tools figure published Claude Opus 5.5 (max) · with tools · self-reported voxel IoU 0.962 (1,000-file subset, Python tools) — Anthropic Opus 5.5 system card fig. 8.13.2.A Claude Sonnet 5.5 (max) · with tools · self-reported voxel IoU 0.963 (1,000-file subset, Python tools) — Anthropic Sonnet 5.5 system card fig. 8.13.2.A Sonnet 4.6 GPT-5.3 Gemini 3.1 Pro Opus 4.7 GPT-5.5 Opus 4.8 Mythos Preview Mythos 5 Sonnet 5 Grok 4.5 Terra GPT-5.6 Sol Luna Kimi K3 Opus 5 Grok 4.6 GPT-6 Astra Fable 5.1 Opus 5.5 Sonnet 5.5

Setup. Score is IoU-score (voxel IoU × exec%) for rows we ran; vendor-reported rows are voxel IoU on a random 1,000-file subset, so they are not on the same split. Hollow on the dashed line is the plain run, filled on the solid line is the agentic run — a Python sandbox with the task images and CAD libraries, where the model renders, measures and iterates before submitting. Frontier lines connect the running record high in each setting. GPT-6 Astra has an agentic point only; no no-tools figure is published. Points sharing a release date are nudged apart. Hover any point for its effort setting, split and source.

industry-grounded

Built on Real Engineering

BenchCAD cases are not synthetic shapes. Version 1.0 anchors its part families to industrial standards, and version 2.0 takes its cases from real engineering work, so a correct answer is a part, an assembly or a wiring that an engineer would accept, not a similar-looking shape.

BenchCAD 1.0 · standard parts

About half the part families, 52 of 106, are anchored to real specification tables from 47 ISO, DIN, EN, ASME and IEC codes. The rest follow common engineering practice.

ISO DIN EN ASME IEC
52 / 106 standard-anchored families 47 specification codes
BenchCAD 2.0-preview · real engineering work

Cases come from factory part drawings, vendor machines with their assembly drawings, and assembled circuit boards. Models build parts and assemblies and recover a board's wiring, and are scored on the geometry and connections they produce.

part drawings assemblies circuit boards
~100 public · ~100 private cases 6 tasks
benchmark
BenchCAD
verified parts
17,900
families
106
standards
47
CadQuery ops
>40
grading
deterministic