Programmatic CAD is the entrance to AI for hardware — the parametric code that turns a model's output into a manufacturable part. Reconstructing one is a hard multimodal problem: cross-modal alignment, 3D spatial reasoning, and recovering the exact numbers. BenchCAD measures it on 17,900 execution-verified CadQuery parts across 106 industrial families.

Schematic — annotations are illustrative, not the part's actual dimensions.
Image→CadQuery, design QA, and editing — probing image–text–code alignment and 3D spatial reasoning across 50+ model × setting runs.
Parametric code is how AI designs and edits real, manufacturable parts — where software crosses into the physical world.
Zoomed to 2026 — every frontier model since January. OpenAI and Anthropic now report BenchCAD in their own system cards and launch tables, each pushing the number higher — GPT-5.6 Sol self-reports 0.706. When every major lab is racing the same benchmark and publishing the gains itself, AI for hardware has become a strategic front — and BenchCAD is the yardstick they measure against. (Circles are re-graded on our scorer; diamonds are vendors' self-reported voxel IoU — a different metric, not re-graded.)
IoU-score (IoU × exec%) vs. release date, 2026. Circles are re-graded by our scorer (thinking setting); diamonds are vendor self-reported voxel IoU — Anthropic system cards, and OpenAI's GPT-5.6 launch table.
BenchCAD parts aren't synthetic primitives. About half the families — 52 of 106 — are anchored to real specification tables drawn from 47 ISO · DIN · EN · ASME · IEC codes, so a correct reconstruction is a spec-faithful, manufacturable part rather than a merely similar-looking shape. The rest follow common engineering practice.
A sample of the 106 industrial part families — gears, springs, fittings, fasteners, and more. A model sees only multi-view renders like these and must recover the CadQuery program that rebuilds each one.














Matched because all four are built on the same parts — Vision QA and Code QA even ask identical questions, once from an image and once from the code. Holding the part fixed pins each failure to a single ability — visual recognition, parametric abstraction, or code synthesis — instead of one lumped score.
Four orthographic views → a CadQuery program, re-executed and scored by a single IoU-score (voxel IoU × exec%).
Numeric geometric reasoning from rendered views, broken out across the L1–L4 capability hierarchy.
The same questions conditioned on CadQuery source — the matched-pair gap isolates seeing from reading.
Instruction-guided program editing, scored by headroom-normalised improvement over the original→target gap.