updates

News

2026-09-22

Claude Opus 5.5 reports 0.962 on BenchCAD, the new top score

Anthropic's Opus 5.5 system card reports 0.730 without tools and 0.962 with tools on the 1,000-file Vision2Code subset, the highest score published in either setting.

BenchCAD Vision2Code (1,000-file subset) voxel IoU, without tools vs with tools: GPT-5.6 Sol 0.706 / 0.834, GPT-6 Astra 0.959 with tools, Claude Sonnet 5 0.322 / 0.519, Claude Opus 5 0.497 / 0.899, Claude Fable 5.1 0.606 / 0.926, Claude Opus 5.5 0.730 / 0.962.

The same card updates Anthropic's earlier models: Fable 5.1 now 0.606 / 0.926, Opus 5 0.497 / 0.899, Sonnet 5 0.322 / 0.519.

BenchCAD Vision2Code with tools, voxel IoU vs cost per task on a log scale, for Claude Sonnet 5, Opus 5, Fable 5.1 and Opus 5.5. Opus 5.5 reaches 0.962 at about $6 per task and sits above the other curves at every cost.
read the system card
2026-09-03

GPT-6 Astra reports 0.959 on BenchCAD, the highest agentic score published

OpenAI's GPT-6 Astra launch lists BenchCAD at 0.959 with a Python tool on the 1,000-file Vision2Code subset — the second lead change in three days, after Claude Fable 5.1's 0.843 on September 1.

OpenAI's BenchCAD (python tool) chart from the GPT-6 Astra launch post: mean voxel IoU against API cost per task for GPT-6 Astra, GPT-5.6 Sol and Claude Fable 5.1. Astra reaches 95.9% at under $2, Sol 83.3% at about $3, Fable 5.1 84.3% at about $12.

Two caveats we are keeping visible. OpenAI has published no no-tools figure, so Astra carries an agentic score only and takes no rank in the plain column. And OpenAI notes its own Claude comparison runs “used modified evaluation settings” — the Claude numbers on our leaderboard come from Anthropic's system cards, not from OpenAI's table.

read the launch post the launch post on X
2026-09-01

Claude Fable 5.1 takes the BenchCAD agentic top spot

Anthropic's Fable 5.1 system card reports it on BenchCAD's Vision2Code subset — 0.437 without tools and 0.843 with, the highest with-tools score any lab has published, narrowly past GPT-5.6 Sol's 0.834.

BenchCAD Vision2Code (1,000-file subset) voxel IoU, without tools vs with tools: GPT-5.6 Sol 0.706 / 0.834, Claude Fable 5 0.376 / 0.675, Claude Sonnet 5 0.267 / 0.376, Claude Opus 5 0.366 / 0.821, Claude Fable 5.1 0.437 / 0.843.

Without tools the order is unchanged — Sol still leads 0.706 to 0.437. The whole gain sits in the agentic setting, and it keeps scaling with test-time compute: with tools, Fable 5.1 nearly doubles its own no-tools score.

BenchCAD Vision2Code voxel IoU vs cost per task on a log scale, with tools (solid) and without (dashed), for Claude Fable 5, Sonnet 5, Opus 5 and Fable 5.1. Fable 5.1 with tools climbs to about 0.84 as test-time compute increases, while its no-tools curve flattens near 0.44.
read the system card
2026-07-24

Anthropic evaluated its Opus 5 model on BenchCAD

Anthropic's Opus 5 system card reports it on BenchCAD's Vision2Code subset — 0.366 without tools, 0.821 with — and it upstreams two fixes to our reference (a camera-position typo and raw-shape grading).

BenchCAD Vision2Code (1,000-file subset) voxel IoU, no-tools vs tools: GPT-5.6 Sol 0.706 / 0.834, Claude Opus 4.8 0.277 / 0.521, Claude Mythos 5 0.378 / 0.678, Claude Sonnet 5 0.267 / 0.376, Claude Opus 5 0.366 / 0.821.

The story isn't the single number — it's the slope. On BenchCAD, Claude's scores now scale steeply with test-time compute, especially once the model has tools to crop and verify its own renders. More effort buys markedly more accuracy.

BenchCAD Vision2Code voxel IoU vs mean API cost per attempt on a log scale (test-time compute), with and without tools, for Claude Opus 4.8, Mythos 5, Sonnet 5 and Opus 5. Opus 5 with tools climbs from about 0.27 to 0.82 as test-time compute increases.
read the system card
2026-07-09

OpenAI evaluated its GPT-5.6 family on BenchCAD

OpenAI lists BenchCAD in its GPT-5.6 launch table, next to OSWorld and BrowseComp — the second frontier lab after Anthropic to report it in a flagship launch, and another sign AI for hardware is now a strategic front.

GPT · run by OpenAI Claude · Anthropic's own numbers, republished no tools Python tool GPT-5.6 Sol GPT-5.6 Luna GPT-5.6 Terra GPT-5.5 Mythos 5 Mythos Preview Opus 4.8 GPT-5.6 Sol · 70.6 — OpenAI launch table, no tools. Split, metric and attempt budget unstated. GPT-5.6 Luna · 63.1 — OpenAI launch table, no tools. Split, metric and attempt budget unstated. GPT-5.6 Terra · 62.3 — OpenAI launch table, no tools. Split, metric and attempt budget unstated. GPT-5.5 · 44.4 — OpenAI launch table, no tools (self-reported). Claude Mythos 5 · 38.4 = Anthropic's 0.384 voxel IoU on the FULL set (system card fig. 8.16.4.A) Claude Mythos Preview · 35.5 = Anthropic's 0.355 voxel IoU on the FULL set (system card fig. 8.16.4.A) Claude Opus 4.8 · 27.3 = Anthropic's 0.273 voxel IoU on the FULL set (system card fig. 8.16.4.A) GPT-5.6 Sol · 83.4 — OpenAI launch table, Python tool. Split, metric and attempt budget unstated. GPT-5.6 Luna · 73.9 — OpenAI launch table, Python tool. Split, metric and attempt budget unstated. GPT-5.6 Terra · 78.2 — OpenAI launch table, Python tool. Split, metric and attempt budget unstated. GPT-5.5 · 55.8 — OpenAI launch table, Python tool. Split, metric and attempt budget unstated. Claude Mythos 5 · 65 = Anthropic's 0.650 voxel IoU on the 1,000-file SUBSET (system card fig. 8.16.4.B) — a different split from the 38.4 beside it Claude Mythos Preview · 61 = Anthropic's 0.610 voxel IoU on the 1,000-file SUBSET (system card fig. 8.16.4.B) — a different split from the 35.5 beside it Claude Opus 4.8 · 51.8 = Anthropic's 0.518 voxel IoU on the 1,000-file SUBSET (system card fig. 8.16.4.B) — a different split from the 27.3 beside it 70.663.162.344.438.435.527.3 83.473.978.255.865.061.051.8

Scores exactly as printed in OpenAI's launch table (computer-use section, 2026-07-09), self-reported voxel IoU. The Claude bars are Anthropic's own published figures, republished by OpenAI. Hover any bar for its source and split.

read the launch post the re-graded leaderboard the 2.0 open call
2026-07-03

Open call: BenchCAD 2.0 — hard cases and a Python environment

We're designing the next major version of BenchCAD in the open, and the scope is set: hard cases — real-work, multi-feature parts selected because today's frontier fails them — and a sandboxed Python environment where models render, measure and iterate before submitting, making the “with tools” setting Anthropic already reports a first-class, comparable leaderboard division. Propose a task or family from your domain and help shape the release.

Also in design as its own track: BenchCAD-Assembly — from single parts to multi-part assemblies, scored on assembled geometry. Its call is coming.

read the 2.0 call
2026-06-30

Anthropic evaluated its Sonnet 5 model on BenchCAD

Anthropic's Sonnet 5 system card reports it on BenchCAD's Vision2Code subset — 0.266 without tools, 0.373 with — another flagship model card measuring on the benchmark.

BenchCAD Vision2Code (1,000-file subset) voxel IoU, no-tools vs Python-tools: Claude Sonnet 4.6 0.267 / 0.327, Claude Mythos Preview 0.359 / 0.611, Claude Opus 4.8 0.280 / 0.535, Claude Sonnet 5 0.266 / 0.373.
read the system card
2026-06-10

Anthropic evaluated its Fable 5 / Mythos 5 model on BenchCAD

Anthropic's Claude Fable 5 / Mythos 5 system card evaluates on BenchCAD Vision2Code: their strongest model reaches ~0.38 voxel IoU on the full set, and ~0.65 with Python tools on a 1,000-file subset. The first frontier lab to fold execution-grounded CAD into how it measures its models — an early marker of what has since become a strategic front.

Anthropic system card figure — BenchCAD Vision2Code (full): Claude Mythos Preview 0.355, Claude Opus 4.8 0.273, Claude Mythos 5 0.384 voxel IoU Anthropic system card figure — BenchCAD Vision2Code (1,000-file subset): no-tools vs Python-tools voxel IoU for Claude Mythos Preview, Claude Opus 4.8, Claude Mythos 5
read the system card