the entrance to AI for hardware

BenchCAD: A Benchmark for Programmatic CAD

Programmatic CAD is the entrance to AI for hardware — the parametric code that turns a model's output into a manufacturable part. Reconstructing one is a hard multimodal problem: cross-modal alignment, 3D spatial reasoning, and recovering the exact numbers. BenchCAD measures it on 17,900 execution-verified CadQuery parts across 106 industrial families.

4 matched tasks
BenchCAD · part #074 · ISO 606
double_simplex_sprocket — a dimensioned BenchCAD part (ISO 606 duplex roller-chain sprocket)
partsimplex_sprocket
teethz 24×2
rev0.1

Schematic — annotations are illustrative, not the part's actual dimensions.

developed by: chosen by: Anthropic OpenAI
leaderboard
A frontier test of vision, language & code

Image→CadQuery, design QA, and editing — probing image–text–code alignment and 3D spatial reasoning across 50+ model × setting runs.

why it matters
The on-ramp to AI for hardware

Parametric code is how AI designs and edits real, manufacturable parts — where software crosses into the physical world.

progress · Vision2Code

The Frontier Over Time

Zoomed to 2026 — every frontier model since January. OpenAI and Anthropic now report BenchCAD in their own system cards and launch tables, each pushing the number higher — GPT-5.6 Sol self-reports 0.706. When every major lab is racing the same benchmark and publishing the gains itself, AI for hardware has become a strategic front — and BenchCAD is the yardstick they measure against. (Circles are re-graded on our scorer; diamonds are vendors' self-reported voxel IoU — a different metric, not re-graded.)

OpenAI Anthropic Google Kimi frontier (reported) 0.700.600.500.400.300.20 IoU-score ↑ Jan 2026FebMarAprMayJunJul release date → re-graded (thinking) vendor-reported · not re-graded frontier ≈ +0.08 IoU / mo Claude Sonnet 4.6 (thinking) · IoU-score 0.2220 (re-graded) Gemini 3.1 Pro (thinking) · IoU-score 0.2890 (re-graded) Claude Opus 4.7 (thinking) · IoU-score 0.2692 (re-graded) GPT-5.3 (thinking) · IoU-score 0.1793 (re-graded) · released Feb 2026 Kimi K3 (thinking) · IoU-score 0.3670 (re-graded, we ran it) GPT-5.5 · vendor-reported 0.444 — OpenAI GPT-5.6 launch table, not re-graded; harness undisclosed Claude Opus 4.8 · self-reported voxel IoU 0.273 — Anthropic system card (full set), not re-graded Claude Sonnet 5 · self-reported voxel IoU 0.266 — Anthropic system card (full set), not re-graded Claude Mythos Preview · self-reported voxel IoU 0.355 — Anthropic system card (full set), not re-graded Claude Mythos 5 · self-reported voxel IoU 0.384 — Anthropic system card (full set), not re-graded GPT-5.6 Terra · vendor-reported 0.623 — OpenAI launch table, not re-graded; harness undisclosed GPT-5.6 Luna · vendor-reported 0.631 — OpenAI launch table, not re-graded; harness undisclosed GPT-5.6 Sol · vendor-reported 0.706 — OpenAI launch table, not re-graded; harness undisclosed Claude Opus 5 · self-reported voxel IoU 0.366 (1,000-file subset, no tools) — Anthropic system card, not re-graded. With tools: 0.821. Sonnet 4.6 Gemini 3.1 Pro Opus 4.7 GPT-5.3 GPT-5.5 Opus 4.8 Sonnet 5 Mythos Preview Mythos 5 GPT-5.6 Sol Terra · Luna Opus 5 Kimi K3

IoU-score (IoU × exec%) vs. release date, 2026. Circles are re-graded by our scorer (thinking setting); diamonds are vendor self-reported voxel IoU — Anthropic system cards, and OpenAI's GPT-5.6 launch table.

industry-grounded

Built on Real Engineering Standards

BenchCAD parts aren't synthetic primitives. About half the families — 52 of 106 — are anchored to real specification tables drawn from 47 ISO · DIN · EN · ASME · IEC codes, so a correct reconstruction is a spec-faithful, manufacturable part rather than a merely similar-looking shape. The rest follow common engineering practice.

ISO DIN EN ASME IEC 52 / 106 standard-anchored families 47 specification codes
examples

What Models Are Asked to Build

A sample of the 106 industrial part families — gears, springs, fittings, fasteners, and more. A model sees only multi-view renders like these and must recover the CadQuery program that rebuilds each one.

capability-decomposed

Four Matched Tasks

Matched because all four are built on the same parts — Vision QA and Code QA even ask identical questions, once from an image and once from the code. Holding the part fixed pins each failure to a single ability — visual recognition, parametric abstraction, or code synthesis — instead of one lumped score.

benchmark
BenchCAD
verified parts
17,900
families
106
standards
47
CadQuery ops
>40
grading
deterministic