projects · in progress · open call

BenchCAD 2.0

The next major version, designed in the open — the way the best community benchmarks are built. The scope is set.

2.0 · hard cases
Real-work complexity

Multi-feature parts drawn from day-to-day mechanical work — hole patterns, threads, fillets, mating features — selected because today's frontier fails them. Grading unchanged: one IoU-score.

2.0 · python environment
Agentic setting, first-class

A sandboxed Python environment with the task images and CAD / vision libraries — models render, measure and iterate before submitting. Formalizes the “with tools” setting Anthropic already reports into a comparable leaderboard division.

Contribute — and get credited

2.0 is built from complex, parametrically-generatable cases — a real part, hard enough that today's frontier fails it, and driven by a spec table rather than a one-off shape. Propose them from your day-to-day domain:

How it works

The 2.0 dev set is open — 150 public part families with source, specs and validation tools. Start from a part and its drawing.

propose

Open a case proposal

The part, a datasheet or drawing with its dimension table, and why it's hard — threads, hole patterns, fillets, mating features. Open an issue on the 2.0 repo, or claim it first in #contribute-parts on Discord. Credit counts from the issue timestamp — propose before you build.

build

Write the generator

Add a designs/<family>/ with its part.py and spec.py — a parametric CadQuery generator against the family spec. CONTRIBUTING.md has the layout and the local checks to run.

land

Reviewed, merged, credited

A maintainer reviews the generator against the reference drawing — equations, physical plausibility, and whether the views reconstruct (REVIEWING.md). Accepted cases accumulate toward paper recognition.

A separate 300-family held-out set stays private, so scores measure generalization.

the 2.0 repo propose a case discuss on Discord