Which models can actually do CAD?
14 models, 3 tasks, 84 attempts. Each attempt is graded on whether it built, on measured geometry, and on a checklist an AI judge answers from the renders. Click any number to open the attempt behind it.
All models
Rank by
| # | Model | Cost and time | Overall | Strict | Gate pass |
|---|---|---|---|---|---|
| 1 | GPT-5.6 LunaOpenAI | 100 | — | 100% | |
| 2 | Gemini 3.8 FlashGoogle | 100 | — | 100% | |
| 3 | GPT-5.6 TerraOpenAI | 100 | — | 100% | |
| 4 | GPT-5.6 SolOpenAI | 100 | — | 100% | |
| 5 | Sonnet 5Anthropic | 100 | — | 100% | |
| 6 | Grok 4.6xAI | 100 | — | 100% | |
| 7 | Opus 5Anthropic | 100 | — | 100% | |
| 8 | Qwen 3.8 MaxAlibaba | 100 | — | 100% | |
| 9 | Fable 5.1Anthropic | 100 | — | 100% | |
| 10 | Opus 4.8Anthropic | 100 | — | 100% | |
| 11 | Gemini 3.1 ProGoogle | 99 | — | 100% | |
| 12 | Fable 5Anthropic | 99 | — | 100% | |
| 13 | GLM 5.3 Flash (Z.ai) | 95 | — | 100% | |
| 14 | Kimi K3 (Moonshot) | 0 | — | 0% |
Overall gives partial credit at every layer, so the top of the board bunches. Strict counts only the attempts that got everything right. Gate pass is the share whose part was watertight, hole-correct and overlap-free. Hover any bar for its number. Ties break by cost, then time.
Browse the tasks
How the score is made
Built 30%, measured geometry 50%, AI judge 20%. A failed gate zeroes geometry for that attempt. The AI judges are Opus 4.8. Read the methodology.
This is an older run. See the current results.
Run screen-14-fixed2026-09-04spend $24.32, 3 unpricedOther runs: 2026-09-10T19-44-23-baseline-2026-09, 2026-09-04T20-02-10-stage2-hard-cases