the entrance to AI for hardware

A Benchmark for Programmatic CAD

A benchmark for code-based 3D modeling on professional tasks.

BenchCAD T1 · drawing → part
impeller, a BenchCAD 1.0 part family, turning coil spring, a BenchCAD 1.0 part family, turning double simplex sprocket, a BenchCAD 1.0 part family, turning spline hub, a BenchCAD 1.0 part family, turning handwheel, a BenchCAD 1.0 part family, turning
taskT1
parts1
chosen by:
Progress · Vision2Code

Where the frontier is

Vision2Code · four rendered views in, a CadQuery program out
OpenAI Anthropic Google Kimi Grok no tools with tools best to date, no tools best to date, with tools 0.900.800.700.600.500.400.300.20 score Jan 2026FebMarAprMayJunJulAugSepOct release date Claude Sonnet 4.6 (max) · no tools · IoU-score 0.2220 (re-graded) GPT-5.3 (max) · no tools · IoU-score 0.1793 (re-graded) · released Feb 2026 Gemini 3.1 Pro (thinking) · no tools · IoU-score 0.2890 (re-graded) Claude Opus 4.7 (max) · no tools · IoU-score 0.2692 (re-graded) GPT-5.5 (max) · no tools · vendor-reported 0.444 — OpenAI GPT-5.6 launch table, not re-graded; harness undisclosed Claude Opus 4.8 (max) · no tools · self-reported voxel IoU 0.273 (full set) — Anthropic Fable 5 / Mythos 5 system card, not re-graded Claude Mythos Preview (max) · no tools · self-reported voxel IoU 0.355 (full set) — Anthropic system card, not re-graded Claude Mythos 5 (max) · no tools · self-reported voxel IoU 0.384 (full set) — Anthropic system card, not re-graded Claude Sonnet 5 (max) · no tools · self-reported voxel IoU 0.322 (1,000-file subset) — Anthropic Opus 5.5 system card fig. 8.13.2.A, not re-graded Grok 4.5 (high) · no tools · IoU-score 0.3194 (re-graded on our scorer) · released Jul 8 2026 GPT-5.6 Terra (max) · no tools · vendor-reported 0.623 — OpenAI launch table, not re-graded; harness undisclosed GPT-5.6 Sol (max) · no tools · vendor-reported 0.706 — OpenAI launch table, not re-graded; harness undisclosed GPT-5.6 Luna (max) · no tools · vendor-reported 0.631 — OpenAI launch table, not re-graded; harness undisclosed Kimi K3 (max) · no tools · IoU-score 0.3670 (re-graded, we ran it) Claude Opus 5 (max) · no tools · self-reported voxel IoU 0.497 (1,000-file subset) — Anthropic Opus 5.5 system card fig. 8.13.2.A, not re-graded Grok 4.6 (xhigh) · no tools · IoU-score 0.3638 (re-graded on our scorer) Claude Fable 5.1 (max) · no tools · self-reported voxel IoU 0.606 (1,000-file subset) — Anthropic Opus 5.5 system card fig. 8.13.2.A, not re-graded Claude Opus 5.5 (max) · no tools · self-reported voxel IoU 0.730 (1,000-file subset) — Anthropic Opus 5.5 system card fig. 8.13.2.A, not re-graded Claude Sonnet 5.5 (max) · no tools · self-reported voxel IoU 0.747 (1,000-file subset) — Anthropic Sonnet 5.5 system card fig. 8.13.2.A, not re-graded Claude Haiku 5.5 (max) · no tools · self-reported voxel IoU 0.670 (1,000-file subset) — Anthropic Haiku 5.5 system card fig. 8.9.2.A, not re-graded GPT-5.5 (max) · with tools · vendor-reported 0.558 — OpenAI GPT-5.6 launch table, Python tool; split and attempt budget unstated Claude Opus 4.8 (max) · with tools · self-reported voxel IoU 0.518 (1,000-file subset, Python tools) — Anthropic system card fig. 8.16.4.B Claude Mythos Preview (max) · with tools · self-reported voxel IoU 0.610 (1,000-file subset, Python tools) — Anthropic system card fig. 8.16.4.B Claude Mythos 5 (max) · with tools · self-reported voxel IoU 0.650 (1,000-file subset, Python tools) — Anthropic system card fig. 8.16.4.B Claude Sonnet 5 (max) · with tools · self-reported voxel IoU 0.519 (1,000-file subset, Python tools) — Anthropic Opus 5.5 system card fig. 8.13.2.A Grok 4.5 (high) · with tools · 0.7771 — our own agentic run on our scorer, Python sandbox GPT-5.6 Terra (max) · with tools · vendor-reported 0.782 — OpenAI launch table, Python tool; split and attempt budget unstated GPT-5.6 Sol (max) · with tools · vendor-reported 0.834 — OpenAI launch table, Python tool; split and attempt budget unstated GPT-5.6 Luna (max) · with tools · vendor-reported 0.739 — OpenAI launch table, Python tool; split and attempt budget unstated Claude Opus 5 (max) · with tools · self-reported voxel IoU 0.899 (1,000-file subset, Python tools) — Anthropic Opus 5.5 system card fig. 8.13.2.A Grok 4.6 (xhigh) · with tools · 0.8055 — our own agentic run on our scorer, Python sandbox Claude Fable 5.1 (max) · with tools · self-reported voxel IoU 0.926 (1,000-file subset, Python tools) — Anthropic Opus 5.5 system card fig. 8.13.2.A GPT-6 Astra · with tools · 0.959 (1,000-file subset) — OpenAI GPT-6 Astra launch coverage; no no-tools figure published Claude Opus 5.5 (max) · with tools · self-reported voxel IoU 0.962 (1,000-file subset, Python tools) — Anthropic Opus 5.5 system card fig. 8.13.2.A Claude Sonnet 5.5 (max) · with tools · self-reported voxel IoU 0.963 (1,000-file subset, Python tools) — Anthropic Sonnet 5.5 system card fig. 8.13.2.A Claude Haiku 5.5 (max) · with tools · self-reported voxel IoU 0.870 (1,000-file subset, Python tools) — Anthropic Haiku 5.5 system card fig. 8.9.2.A, not re-graded Sonnet 4.6 GPT-5.3 Gemini 3.1 Pro Opus 4.7 GPT-5.5 Opus 4.8 Mythos Preview Mythos 5 Sonnet 5 Grok 4.5 Terra GPT-5.6 Sol Luna Kimi K3 Opus 5 Grok 4.6 GPT-6 Astra Fable 5.1 Opus 5.5 Sonnet 5.5 Haiku 5.5
Leaderboard

The top of the board

The ten best models on Vision2Code. Scores run from 0 to 1, and a program that fails to run scores 0.


All four tasks
    * Reported by frontier labs
    Latest

    News and writing

    Citation

    @misc{benchcad2026,
      title        = {BenchCAD: A Comprehensive, Industry-Standard Benchmark for Programmatic CAD},
      author       = {Zhang, Haozhe and Liu, Kaichen and Chen, Miaomiao and Li, Lei
                      and Yang, Shaojie and Peng, Cheng and Chen, Hanjie},
      year         = {2026},
      eprint       = {2605.10865},
      archivePrefix= {arXiv},
      primaryClass = {cs.CV},
      url          = {https://arxiv.org/abs/2605.10865}
    }