the entrance to AI for hardware

A Benchmark for Programmatic CAD

Four orthographic views in, a CadQuery program out — graded on executed geometry, not appearances. 17,900 verified parts across 106 industrial families.

4 matched tasks
BenchCAD · part #074 · ISO 606
double_simplex_sprocket — a dimensioned BenchCAD part (ISO 606 duplex roller-chain sprocket)
partsimplex_sprocket
teethz 24×2
rev0.1

Schematic — annotations are illustrative, not the part's actual dimensions.

developed by: chosen by:
progress · Vision2Code

Where the Frontier Is

Every frontier model on Vision2Code in 2026, in two views — against release date, and against what a task costs.

OpenAI Anthropic Google Kimi Grok no tools with tools (agentic) frontier, no tools frontier, with tools 0.900.800.700.600.500.400.300.20 IoU-score ↑ Jan 2026FebMarAprMayJunJulAugSep release date → Claude Sonnet 4.6 (max) · no tools · IoU-score 0.2220 (re-graded) GPT-5.3 (max) · no tools · IoU-score 0.1793 (re-graded) · released Feb 2026 Gemini 3.1 Pro (thinking) · no tools · IoU-score 0.2890 (re-graded) Claude Opus 4.7 (max) · no tools · IoU-score 0.2692 (re-graded) GPT-5.5 (max) · no tools · vendor-reported 0.444 — OpenAI GPT-5.6 launch table, not re-graded; harness undisclosed Claude Opus 4.8 (max) · no tools · self-reported voxel IoU 0.273 (full set) — Anthropic Fable 5 / Mythos 5 system card, not re-graded Claude Mythos Preview (max) · no tools · self-reported voxel IoU 0.355 (full set) — Anthropic system card, not re-graded Claude Mythos 5 (max) · no tools · self-reported voxel IoU 0.384 (full set) — Anthropic system card, not re-graded Claude Sonnet 5 (max) · no tools · self-reported voxel IoU 0.266 (1,000-file subset) — Anthropic Sonnet 5 system card, not re-graded Grok 4.5 (high) · no tools · IoU-score 0.3194 (re-graded on our scorer) · released Jul 8 2026 GPT-5.6 Terra (max) · no tools · vendor-reported 0.623 — OpenAI launch table, not re-graded; harness undisclosed GPT-5.6 Sol (max) · no tools · vendor-reported 0.706 — OpenAI launch table, not re-graded; harness undisclosed GPT-5.6 Luna (max) · no tools · vendor-reported 0.631 — OpenAI launch table, not re-graded; harness undisclosed Kimi K3 (max) · no tools · IoU-score 0.3670 (re-graded, we ran it) Claude Opus 5 (max) · no tools · self-reported voxel IoU 0.366 (1,000-file subset) — Anthropic Opus 5 system card, not re-graded Grok 4.6 (xhigh) · no tools · IoU-score 0.3638 (re-graded on our scorer; run arranged by SpaceXAI) Claude Fable 5.1 (max) · no tools · self-reported voxel IoU 0.437 (1,000-file subset) — Anthropic Fable 5.1 / Mythos 5.1 system card fig. 8.14.2.A, not re-graded GPT-5.5 (max) · with tools · vendor-reported 0.558 — OpenAI GPT-5.6 launch table, Python tool; split and attempt budget unstated Claude Opus 4.8 (max) · with tools · self-reported voxel IoU 0.518 (1,000-file subset, Python tools) — Anthropic system card fig. 8.16.4.B Claude Mythos Preview (max) · with tools · self-reported voxel IoU 0.610 (1,000-file subset, Python tools) — Anthropic system card fig. 8.16.4.B Claude Mythos 5 (max) · with tools · self-reported voxel IoU 0.650 (1,000-file subset, Python tools) — Anthropic system card fig. 8.16.4.B Claude Sonnet 5 (max) · with tools · self-reported voxel IoU 0.373 (1,000-file subset, Python tools) — Anthropic Sonnet 5 system card Grok 4.5 (high) · with tools · 0.7771 — our own agentic run on our scorer, Python sandbox GPT-5.6 Terra (max) · with tools · vendor-reported 0.782 — OpenAI launch table, Python tool; split and attempt budget unstated GPT-5.6 Sol (max) · with tools · vendor-reported 0.834 — OpenAI launch table, Python tool; split and attempt budget unstated GPT-5.6 Luna (max) · with tools · vendor-reported 0.739 — OpenAI launch table, Python tool; split and attempt budget unstated Claude Opus 5 (max) · with tools · self-reported voxel IoU 0.821 (1,000-file subset, Python tools) — Anthropic Opus 5 system card fig. 8.12.2.A Grok 4.6 (xhigh) · with tools · 0.8055 — our own agentic run on our scorer, Python sandbox Claude Fable 5.1 (max) · with tools · self-reported voxel IoU 0.843 (1,000-file subset, Python tools) — Anthropic Fable 5.1 / Mythos 5.1 system card fig. 8.14.2.A GPT-6 Astra · with tools · 0.959 (1,000-file subset) — OpenAI GPT-6 Astra launch coverage; no no-tools figure published Sonnet 4.6 GPT-5.3 Gemini 3.1 Pro Opus 4.7 GPT-5.5 Opus 4.8 Mythos Preview Mythos 5 Sonnet 5 Grok 4.5 Terra GPT-5.6 Sol Luna Kimi K3 Opus 5 Grok 4.6 GPT-6 Astra Fable 5.1

Setup. Score is IoU-score (voxel IoU × exec%) for rows we ran; vendor-reported rows are voxel IoU on a random 1,000-file subset, so they are not on the same split. Hollow on the dashed line is the plain run, filled on the solid line is the agentic run — a Python sandbox with the task images and CAD libraries, where the model renders, measures and iterates before submitting. Frontier lines connect the running record high in each setting. GPT-6 Astra has an agentic point only; no no-tools figure is published. Points sharing a release date are nudged apart. Hover any point for its effort setting, split and source.

industry-grounded

Built on Real Engineering Standards

BenchCAD parts aren't synthetic primitives. About half the families — 52 of 106 — are anchored to real specification tables drawn from 47 ISO · DIN · EN · ASME · IEC codes, so a correct reconstruction is a spec-faithful, manufacturable part rather than a merely similar-looking shape. The rest follow common engineering practice.

ISO DIN EN ASME IEC 52 / 106 standard-anchored families 47 specification codes
examples

What Models Are Asked to Build

A sample of the 106 industrial part families — gears, springs, fittings, fasteners, and more. A model sees only multi-view renders like these and must recover the CadQuery program that rebuilds each one.

benchmark
BenchCAD
verified parts
17,900
families
106
standards
47
CadQuery ops
>40
grading
deterministic