partforge

Which models can actually do CAD?

5 models, 54 tasks, 675 attempts. Each attempt is graded on whether it built, on measured geometry, and on a checklist an AI judge answers from the renders. Click any number to open the attempt behind it.

Quality
against
0%25%50%75%100%$0.00$0.30$0.60$0.90$1.200s50s1m 40s2m 30s3m 20s4m 10sFable 5.1GPT-5.6 SolGPT-6 AstraGemini 3.8 FlashOpus 5
Gemini 3.8 Flash sits alone in the top-left corner: first on human rank and the cheapest to run. 100% is first on every task, 0% is last, 50% is tied with the field.

Where the score and the reviewer disagree

Overall gives partial credit at every layer. Once every model can do most of a task the scores bunch: this run's five sit inside 2.4 points, and what little is left does not match what the reviewer saw. Strict reads the same evidence as a pass or a fail, and it does separate them.

Overall scoreHuman rankingFable 5.1: 1 by overall score, 3 by the reviewer's ranking, 2 places lowerGPT-5.6 Sol: 2 by overall score, 5 by the reviewer's ranking, 3 places lowerGPT-6 Astra: 3 by overall score, 2 by the reviewer's ranking, 1 place higherGemini 3.8 Flash: 4 by overall score, 1 by the reviewer's ranking, 3 places higherOpus 5: 5 by overall score, 4 by the reviewer's ranking, 1 place higher1Fable 5.12GPT-5.6 Sol3GPT-6 Astra4Gemini 3.8 Flash5Opus 5Gemini 3.8 Flash1GPT-6 Astra2Fable 5.13Opus 54GPT-5.6 Sol5

All models

Rank by
#ModelQualityCost and timeOverallStrictHuman rankGate pass
1Fable 5.1Anthropic
9568%49%98%
2GPT-5.6 SolOpenAI
9368%43%95%
3GPT-6 AstraOpenAI
9369%50%94%
4Gemini 3.8 FlashGoogle
9371%61%96%
5Opus 5Anthropic
9265%47%95%

Overall gives partial credit at every layer, so the top of the board bunches. Strict counts only the attempts that got everything right. Gate pass is the share whose part was watertight, hole-correct and overlap-free. Hover any bar for its number. Ties break by cost, then time. In Quality, the bar marked with a person is one reviewer's ranking of the cases they ordered by hand. First place is 100%, last is 0%, ties share, and the tick marks 50%. That reviewer found every first attempt acceptable. Neither affects Overall or Strict.

Browse the tasks

How the score is made

Built 30%, measured geometry 50%, AI judge 20%. A failed gate zeroes geometry for that attempt. The AI judges are Opus 4.8 and GPT-5.6 Sol. Read the methodology.

Run baseline-2026-092026-09-10commit 41aec7daspend $339.70 + $48.02 AI judging, 7 unpricedOther runs: 2026-09-04T20-02-10-stage2-hard-cases, 2026-09-04T19-48-31-screen-14-fixed