Outside the SI Score

Spatial & 3D building

Can a frontier model build a physical-looking thing — a CAD part from four drawings, a voxel galleon from one sentence? These two benchmarks measure exactly that, each with its own method. We show their own order and numbers, credited and linked; neither feeds the SI Score.

Why outside the SI Score? The SI Score ranks models on published text benchmarks. Spatial building has no widely adopted, openly licensed numeric feed yet: BenchCAD's scores here are largely lab self-reports (marked per row), and MineBench's order is human preference, not accuracy. Featuring them separately keeps the composite clean while still showing who builds well.

BenchCAD: programmatic CAD

benchcad.com

Vision2Code: four orthographic views in, a CadQuery program out. The metric is voxel IoU — the executed program's solid compared voxel-by-voxel against the ground-truth part; programs that fail to execute score 0 (IoU-score = voxel IoU × exec rate). In the with tools setting the model gets a Python sandbox to render, measure and iterate before submitting. The two settings are different evaluations and are never merged.

With tools agentic Python sandbox

#ModelReleasedTestedVoxel IoURun by
1 Claude Sonnet 5.5 Anthropic · max effort 2026-09 2026-09 96.3% Lab-reported
2 Claude Opus 5.5 Anthropic · max effort 2026-09 2026-09 96.2% Lab-reported
3 GPT-6 Astra OpenAI 2026-09 2026-09 95.9% Lab-reported
4 Claude Fable 5.1 Anthropic · max effort 2026-09 2026-09 92.6%* Lab-reported
5 Claude Opus 5 Anthropic · max effort 2026-07 2026-09 89.9% Lab-reported
6 Claude Haiku 5.5 Anthropic · max effort 2026-10 2026-10 87% Lab-reported
7 GPT-5.6 Sol OpenAI · max effort 2026-07 2026-07 83.4% Lab-reported
8 Grok 4.6 SpaceXAI · xhigh effort 2026-08 2026-08 80.55% BenchCAD-run
9 GPT-5.6 Terra OpenAI · max effort 2026-07 2026-07 78.2% Lab-reported
10 Grok 4.5 SpaceXAI · high effort 2026-07 2026-08 77.71% BenchCAD-run
11 GPT-5.6 Luna OpenAI · max effort 2026-07 2026-07 73.9% Lab-reported
12 Claude Mythos 5 Anthropic · max effort 2026-06 2026-06 65% Lab-reported
13 Claude Mythos Preview Anthropic · max effort 2026-06 2026-06 61% Lab-reported
14 GPT-5.5 OpenAI · max effort 2026-04 2026-04 55.8% Lab-reported
15 Claude Sonnet 5 Anthropic · max effort 2026-06 2026-09 51.9% Lab-reported
16 Claude Opus 4.8 Anthropic · max effort 2026-05 2026-06 51.8% Lab-reported
17 Claude Haiku 4.5 Anthropic · thinking effort 2025-10 2026-10 21% Lab-reported

* BenchCAD lists Claude Fable 5.1 at 92.6% with tools (Anthropic's Opus 5.5 system card, fig. 8.13.2.A). OpenAI's GPT-6 Astra launch post quotes 84.3% for the same model — we show the benchmark's number and keep both contexts visible. Lab-reported rows are vendor self-reported voxel IoU, not re-graded by BenchCAD; Anthropic's figures use a random 1,000-file subset, so their split is not the same as the no-tools column beside them.

No tools single shot, IoU-score × execution rate

#ModelReleasedTestedIoU-scoreRun by
1 Claude Sonnet 5.5 Anthropic · max effort 2026-09 2026-09 74.7% Lab-reported
2 Claude Opus 5.5 Anthropic · max effort 2026-09 2026-09 73% Lab-reported
3 GPT-5.6 Sol OpenAI · max effort 2026-07 2026-07 70.6% Lab-reported
4 Claude Haiku 5.5 Anthropic · max effort 2026-10 2026-10 67% Lab-reported
5 GPT-5.6 Luna OpenAI · max effort 2026-07 2026-07 63.1% Lab-reported
6 GPT-5.6 Terra OpenAI · max effort 2026-07 2026-07 62.3% Lab-reported
7 Claude Fable 5.1 Anthropic · max effort 2026-09 2026-09 60.6% Lab-reported
8 Claude Opus 5 Anthropic · max effort 2026-07 2026-09 49.7% Lab-reported
9 GPT-5.5 OpenAI · max effort 2026-04 2026-04 44.4% Lab-reported
10 Claude Mythos 5 Anthropic · max effort 2026-06 2026-06 38.4% Lab-reported
11 Kimi K3 Moonshot · max effort · Open weights 2026-07 2026-08 36.7% BenchCAD-run
12 Grok 4.6 SpaceXAI · xhigh effort 2026-08 2026-08 36.38% BenchCAD-run
13 Claude Mythos Preview Anthropic · max effort 2026-06 2026-06 35.5% Lab-reported
14 Grok 4.6 SpaceXAI · high effort 2026-08 2026-08 35.1% BenchCAD-run
15 Claude Sonnet 5 Anthropic · max effort 2026-06 2026-09 32.2% Lab-reported
16 Grok 4.5 SpaceXAI · high effort 2026-07 2026-08 31.94% BenchCAD-run
17 Gemini 3.1 Pro Google · thinking effort 2026-02 2026-06 28.9% BenchCAD-run
18 Gemini 3.1 Pro Google 2026-02 2026-06 27.79% BenchCAD-run
19 Claude Opus 4.8 Anthropic · max effort 2026-05 2026-06 27.3% Lab-reported
20 Claude Opus 4.7 Anthropic · max effort 2026-03 2026-06 26.92% BenchCAD-run
21 Claude Opus 4.7 Anthropic 2026-03 2026-06 26.17% BenchCAD-run
22 Claude Sonnet 4.6 Anthropic · max effort 2026-01 2026-06 22.2% BenchCAD-run
23 Claude Sonnet 4.6 Anthropic 2026-01 2026-06 19.2% BenchCAD-run
24 GPT-5.3 OpenAI 2026-02 2026-06 18.73% BenchCAD-run
25 GPT-4o OpenAI 2024-05 2026-06 18.23% BenchCAD-run
26 GPT-5.3 OpenAI · max effort 2026-02 2026-06 17.93% BenchCAD-run
27 Claude Haiku 4.5 Anthropic · thinking effort 2025-10 2026-10 15.5% Lab-reported
28 OpenAI o3 OpenAI 2025-04 2026-06 12.18% BenchCAD-run
29 GPT-4o OpenAI · blank img effort · Control — 2026-06 6.98% BenchCAD-run
30 Moonshot v1-128k Moonshot · Open weights 2025-02 2026-06 1.6% BenchCAD-run
31 Moonshot v1-8k Moonshot · Open weights 2025-02 2026-06 1.27% BenchCAD-run
32 Qwen3-VL-2B Qwen · Open weights 2025-09 2026-06 0.05% BenchCAD-run

GPT-6 Astra has no published no-tools score, so it appears only in the with-tools table. Data: BenchCAD (CC BY 4.0), snapshot Oct 9, 2026. Machine-readable source.

MineBench: voxel builds by vote

minebench.ai

Models build Minecraft-style voxel structures from a text prompt. There is no objectively correct answer: visitors compare two anonymous builds and vote, and a global Bradley–Terry model turns those votes into a rating on a 400-point Elo scale with a 95% confidence interval. Snapshot of the public leaderboard, Oct 9, 2026.

#ModelRatingRecordVotesStatus
1 GPT 6 Astra Pro openai 2,279 ±17.7 W 3607 L 567 D 261 4,444 Stable
2 GPT 6.1 Sol Pro openai 2,212 ±21.8 W 1280 L 480 D 164 1,935 Established
3 Claude Opus 5.5 anthropic 2,187 ±19.7 W 1632 L 668 D 167 2,478 Stable
4 GPT 6 Sol Pro openai 2,094 ±18.8 W 1416 L 678 D 139 2,248 Stable
5 Claude Opus 5 anthropic 2,052 ±11.9 W 4172 L 2713 D 604 7,540 Stable
6 Claude Sonnet 5.5 anthropic 2,039 ±19 W 1423 L 612 D 109 2,161 Stable
7 GPT 5.6 Sol Pro openai 2,038 ±11.4 W 4368 L 2862 D 700 7,968 Stable
8 Claude Fable 5.1 anthropic 1,968 ±12.6 W 2639 L 1710 D 289 4,667 Stable
9 GPT 5.5 Pro openai 1,960 ±9 W 7240 L 4018 D 1034 12,364 Stable
10 Claude Fable 5 anthropic 1,906 ±9.1 W 7120 L 3533 D 658 11,370 Stable
11 GPT 5.5 openai 1,883 ±8.3 W 6738 L 5280 D 1008 13,125 Stable
12 DeepSeek V4.1 Flash deepseek 1,856 ±17.2 W 1905 L 818 D 122 2,894 Stable
13 Grok 4.6 xai 1,852 ±14 W 1630 L 1384 D 165 3,230 Stable
14 Gemini 3.8 Flash gemini 1,845 ±16.7 W 1983 L 837 D 128 2,988 Stable
15 GPT 5.4 Pro openai 1,829 ±7.6 W 9698 L 4714 D 887 15,437 Stable

Top 15 of 72 ranked models in our snapshot. The live board moves with every vote.

Real builds the top five models' three strongest prompts — drag to orbit, pinch or scroll to zoom

An underwater shipwreck: a wooden galleon on its side on the ocean floor, holes in the hull, coral and seaweed growing on it, treasure chests spilling gold, and fish swimming around

GPT 6 Astra Pro

Captured Oct 9, 2026 · prompt has 258 votes · this build on MineBench

A floating island ecosystem: a chunk of earth suspended in air with waterfalls pouring off multiple edges, a small forest on top, exposed roots and rocks hanging underneath, and smaller floating rocks nearby connected by ancient chain bridges

GPT 6 Astra Pro

Captured Oct 9, 2026 · prompt has 256 votes · this build on MineBench

A treehouse village: three large treehouses in adjacent trees connected by rope bridges, each house with different architecture (one rustic, one elvish with curved lines, one modern with clean angles), rope ladders down, and lanterns hanging from branches

GPT 6 Astra Pro

Captured Oct 9, 2026 · prompt has 255 votes · this build on MineBench

A treehouse village: three large treehouses in adjacent trees connected by rope bridges, each house with different architecture (one rustic, one elvish with curved lines, one modern with clean angles), rope ladders down, and lanterns hanging from branches

GPT 6.1 Sol Pro

Captured Oct 9, 2026 · prompt has 80 votes · this build on MineBench

An underwater shipwreck: a wooden galleon on its side on the ocean floor, holes in the hull, coral and seaweed growing on it, treasure chests spilling gold, and fish swimming around

GPT 6.1 Sol Pro

Captured Oct 9, 2026 · prompt has 79 votes · this build on MineBench

A floating island ecosystem: a chunk of earth suspended in air with waterfalls pouring off multiple edges, a small forest on top, exposed roots and rocks hanging underneath, and smaller floating rocks nearby connected by ancient chain bridges

GPT 6.1 Sol Pro

Captured Oct 9, 2026 · prompt has 72 votes · this build on MineBench

A classic arcade cabinet with a joystick and three buttons on the control panel, a screen showing simple graphics, coin slot on the front, and artwork on the sides

Claude Opus 5.5

Captured Oct 9, 2026 · prompt has 115 votes · this build on MineBench

A massive world tree: an enormous trunk with roots visible above ground forming archways, multiple levels of thick branches like platforms, glowing fruit hanging from smaller branches, and vines draping down

Claude Opus 5.5

Captured Oct 9, 2026 · prompt has 115 votes · this build on MineBench

A steampunk airship with a wooden hull, large brass propellers on each side, a balloon made of patchwork fabric above the deck, hanging ropes and ladders, and a glass-enclosed bridge at the front

Claude Opus 5.5

Captured Oct 9, 2026 · prompt has 114 votes · this build on MineBench

A phoenix rising from flames: wings fully spread upward, tail feathers flowing down like fire, head raised to the sky, made of red, orange, and gold blocks with glowstone accents

GPT 6 Sol Pro

Captured Oct 9, 2026 · prompt has 111 votes · this build on MineBench

An underwater shipwreck: a wooden galleon on its side on the ocean floor, holes in the hull, coral and seaweed growing on it, treasure chests spilling gold, and fish swimming around

GPT 6 Sol Pro

Captured Oct 9, 2026 · prompt has 108 votes · this build on MineBench

The builds are the byte-exact artifacts MineBench serves its own visitors (checks SHA-verified), shown with flat colors — MineBench's Faithful texture pack is separately licensed and not redistributed. Block colors are our approximation. Builds carry no stated license on the host; the MineBench repo is MIT. Credit: MineBench by Ammaar Alam.

Sources and licenses

  • BenchCAD — programmatic CAD benchmark (Vision2Code, Vision QA, Code QA, Code Edit). Scores from the public leaderboard JSON; benchmark data is CC BY 4.0, code MIT. Rows marked “Lab-reported” are vendor self-reports, not re-graded; others are BenchCAD's own runs. Leaderboard · Repo · Dataset · snapshot Oct 9, 2026
  • MineBench by Ammaar Alam — voxel building arena, blind human votes scored with a global Bradley–Terry model. Leaderboard order and the 15 builds shown are a one-time manual snapshot from the public pages, credited per build; hosted outputs carry no stated license, the repo is MIT. Site · Leaderboard · Repo (MIT) · snapshot Oct 9, 2026

Both snapshots refresh manually, on occasion — they are not part of the daily pipeline. Model names link to our catalog pages only where the benchmark's name matches a cataloged model.

What changed

Releases · RSS feed

Browser alerts

What to be alerted about
RSS feed

Release emails