updates

News

2026-07-24

Anthropic evaluated its Opus 5 model on BenchCAD

Anthropic's Opus 5 system card reports it on BenchCAD's Vision2Code subset — 0.366 without tools, 0.821 with — and it upstreams two fixes to our reference (a camera-position typo and raw-shape grading).

BenchCAD Vision2Code (1,000-file subset) voxel IoU, no-tools vs tools: GPT-5.6 Sol 0.706 / 0.834, Claude Opus 4.8 0.277 / 0.521, Claude Mythos 5 0.378 / 0.678, Claude Sonnet 5 0.267 / 0.376, Claude Opus 5 0.366 / 0.821.

The story isn't the single number — it's the slope. On BenchCAD, Claude's scores now scale steeply with test-time compute, especially once the model has tools to crop and verify its own renders. More effort buys markedly more accuracy.

BenchCAD Vision2Code voxel IoU vs mean API cost per attempt on a log scale (test-time compute), with and without tools, for Claude Opus 4.8, Mythos 5, Sonnet 5 and Opus 5. Opus 5 with tools climbs from about 0.27 to 0.82 as test-time compute increases.
read the system card
2026-07-09

OpenAI evaluated its GPT-5.6 family on BenchCAD

OpenAI lists BenchCAD in its GPT-5.6 launch table, next to OSWorld and BrowseComp — the second frontier lab after Anthropic to report it in a flagship launch, and another sign AI for hardware is now a strategic front.

GPT · run by OpenAI Claude · Anthropic's own numbers, republished no tools Python tool GPT-5.6 Sol GPT-5.6 Luna GPT-5.6 Terra GPT-5.5 Mythos 5 Mythos Preview Opus 4.8 GPT-5.6 Sol · 70.6 — OpenAI launch table, no tools. Split, metric and attempt budget unstated. GPT-5.6 Luna · 63.1 — OpenAI launch table, no tools. Split, metric and attempt budget unstated. GPT-5.6 Terra · 62.3 — OpenAI launch table, no tools. Split, metric and attempt budget unstated. GPT-5.5 · 44.4 — OpenAI launch table, no tools (self-reported). Claude Mythos 5 · 38.4 = Anthropic's 0.384 voxel IoU on the FULL set (system card fig. 8.16.4.A) Claude Mythos Preview · 35.5 = Anthropic's 0.355 voxel IoU on the FULL set (system card fig. 8.16.4.A) Claude Opus 4.8 · 27.3 = Anthropic's 0.273 voxel IoU on the FULL set (system card fig. 8.16.4.A) GPT-5.6 Sol · 83.4 — OpenAI launch table, Python tool. Split, metric and attempt budget unstated. GPT-5.6 Luna · 73.9 — OpenAI launch table, Python tool. Split, metric and attempt budget unstated. GPT-5.6 Terra · 78.2 — OpenAI launch table, Python tool. Split, metric and attempt budget unstated. GPT-5.5 · 55.8 — OpenAI launch table, Python tool. Split, metric and attempt budget unstated. Claude Mythos 5 · 65 = Anthropic's 0.650 voxel IoU on the 1,000-file SUBSET (system card fig. 8.16.4.B) — a different split from the 38.4 beside it Claude Mythos Preview · 61 = Anthropic's 0.610 voxel IoU on the 1,000-file SUBSET (system card fig. 8.16.4.B) — a different split from the 35.5 beside it Claude Opus 4.8 · 51.8 = Anthropic's 0.518 voxel IoU on the 1,000-file SUBSET (system card fig. 8.16.4.B) — a different split from the 27.3 beside it 70.663.162.344.438.435.527.3 83.473.978.255.865.061.051.8

Scores exactly as printed in OpenAI's launch table (computer-use section, 2026-07-09), self-reported voxel IoU. The Claude bars are Anthropic's own published figures, republished by OpenAI. Hover any bar for its source and split.

read the launch post the re-graded leaderboard the 2.0 open call
2026-07-03

Open call: BenchCAD 2.0 — hard cases and a Python environment

We're designing the next major version of BenchCAD in the open, and the scope is set: hard cases — real-work, multi-feature parts selected because today's frontier fails them — and a sandboxed Python environment where models render, measure and iterate before submitting, making the “with tools” setting Anthropic already reports a first-class, comparable leaderboard division. Propose a task or family from your domain and help shape the release.

Also in design as its own track: BenchCAD-Assembly — from single parts to multi-part assemblies, scored on assembled geometry. Its call is coming.

read the 2.0 call
2026-06-30

Anthropic evaluated its Sonnet 5 model on BenchCAD

Anthropic's Sonnet 5 system card reports it on BenchCAD's Vision2Code subset — 0.266 without tools, 0.373 with — another flagship model card measuring on the benchmark.

BenchCAD Vision2Code (1,000-file subset) voxel IoU, no-tools vs Python-tools: Claude Sonnet 4.6 0.267 / 0.327, Claude Mythos Preview 0.359 / 0.611, Claude Opus 4.8 0.280 / 0.535, Claude Sonnet 5 0.266 / 0.373.
read the system card
2026-06-10

Anthropic evaluated its Fable 5 / Mythos 5 model on BenchCAD

Anthropic's Claude Fable 5 / Mythos 5 system card evaluates on BenchCAD Vision2Code: their strongest model reaches ~0.38 voxel IoU on the full set, and ~0.65 with Python tools on a 1,000-file subset. The first frontier lab to fold execution-grounded CAD into how it measures its models — an early marker of what has since become a strategic front.

Anthropic system card figure — BenchCAD Vision2Code (full): Claude Mythos Preview 0.355, Claude Opus 4.8 0.273, Claude Mythos 5 0.384 voxel IoU Anthropic system card figure — BenchCAD Vision2Code (1,000-file subset): no-tools vs Python-tools voxel IoU for Claude Mythos Preview, Claude Opus 4.8, Claude Mythos 5
read the system card