OpenDesignArena
Find the right AI model for design.
Compare quality, speed, and cost with real evaluations.
Best value
Strong design performance at the lowest cost.
In this evaluation, DeepSeek V4.1 Flash achieved 98% of top-ranked GPT-6 Astra’s average score, at just 1% of its average cost.
- Cost per artifact
- $0.023vs $1.61
- Evaluation average / 100
- 81.2vs 82.7
Highest score
Fastest delivery
Compare models
Pick two or three models and see all five dimensions on one chart. Different shapes mean different strengths.
Five-dimension profile
Farther out is better on all five axes: stronger results, faster completion, and lower cost.
Mean on each axis across all 13 models
Choose by preference
Balance evaluated quality, cost, and speed to find the model that suits your needs.
- Quality
- 50%
- Cost
- 30%
- Speed
- 20%
- 1 DeepSeek V4.1 Flash 78.8 Average score: 81.2/100; Average cost: $0.023; Average time: 5.3 min
- 2 DeepSeek V4 Flash 68.6 Average score: 70.6/100; Average cost: $0.031; Average time: 11.1 min
- 3 GPT-5.6 Sol 65.7 Average score: 77.6/100; Average cost: $0.544; Average time: 3.2 min
- 4 Gemini 3.8 Flash 64.7 Average score: 68.9/100; Average cost: $0.210; Average time: 3.8 min
- 5 DeepSeek V4 Pro 64.6 Average score: 72.9/100; Average cost: $0.061; Average time: 17.7 min
- 6 Muse Spark 1.3 63.2 Average score: 66.6/100; Average cost: $0.267; Average time: 3.3 min
- 7 GLM-5.3 Flash 61.4 Average score: 69.8/100; Average cost: $0.064; Average time: 23.5 min
- 8 GPT-6 Astra 57.5 Average score: 82.7/100; Average cost: $1.61; Average time: 11.1 min
- 9 Grok 4.6 56.5 Average score: 72.4/100; Average cost: $0.599; Average time: 11.4 min
- 10 Hunyuan H4 Preview 54.6 Average score: 74.6/100; Average cost: $0.418; Average time: 29.2 min
- 11 Qwen 3.8-Max 52.3 Average score: 72.0/100; Average cost: $0.595; Average time: 26.4 min
- 12 Claude Fable 5.1 52.1 Average score: 80.3/100; Average cost: $3.66; Average time: 12.8 min
- 13 Kimi K3 52.1 Average score: 65.7/100; Average cost: $0.614; Average time: 14.0 min
Choose by use case
Pick a design scenario; each of the five dimensions below recommends one model.
GPT-6 Astra
- Avg. score Average scoreHigher is better The mean score across all evaluated artifacts, out of 100.
- 82.7/100
- Avg. requirement Requirement fulfillmentHigher is better Checks whether required pages, content, states, and interactions are present, out of 30.
- 26.5/30
- Avg. design quality Design qualityHigher is better Rates layout, hierarchy, style, color, and image fit, out of 70.
- 56.2/70
- Avg. delivery rate Delivery rateHigher is better The share of evaluated artifacts scoring 80 or above.
- 60.0%
- Mean time Average completion timeLower is better Total execution time across scored artifacts divided by their count.
- 11.1 min
- Mean cost / artifact Cost per artifactLower is better Estimated from official list prices and measured token usage.
- $1.61/run
Claude Fable 5.1
- Avg. design quality
- 53.2/70
- Avg. score
- 80.3/100
- Avg. delivery rate
- 56.7%
- Avg. requirement
- 27.1/30
- Mean time
- 12.8 min
- Mean cost / artifact
- $3.66/run
DeepSeek V4.1 Flash
- Avg. requirement
- 28.4/30
- Avg. score
- 81.2/100
- Avg. delivery rate
- 57.7%
- Avg. design quality
- 52.8/70
- Mean time
- 5.3 min
- Mean cost / artifact
- $0.023/run
GPT-5.6 Sol
- Mean time
- 3.2 min
- Avg. score
- 77.6/100
- Avg. delivery rate
- 53.3%
- Avg. requirement
- 25.8/30
- Avg. design quality
- 51.8/70
- Mean cost / artifact
- $0.544/run
Hunyuan H4 Preview
- Mean cost / artifact
- $0.418/run
- Avg. score
- 74.6/100
- Avg. delivery rate
- 40.0%
- Avg. requirement
- 26.8/30
- Avg. design quality
- 47.8/70
- Mean time
- 29.2 min
Quality and cost
Compare model performance and the cost of achieving it.
Quality ranking
Model quality vs cost
11 models in the pink area cost more and score lower than DeepSeek V4.1 Flash.
Scenario performance
Viewing Web app (Multi-page structure and flows); GPT-5.6 Sol leads at 83.1, with a 66.7% delivery rate and $0.537 per run.
Viewing Mobile app (Touch experience and responsive layout); Hunyuan H4 Preview leads at 89.5, with a 83.3% delivery rate and $0.404 per run.
Viewing Desktop app (Desktop workflows and complex interactions); DeepSeek V4 Flash leads at 90.3, with a 83.3% delivery rate and $0.027 per run.
Viewing Dashboard (Information density and hierarchy); GPT-6 Astra leads at 86.3, with a 66.7% delivery rate and $1.65 per run.
Viewing Landing page (Brand expression and conversion); Qwen 3.8-Max leads at 82.8, with a 66.7% delivery rate and $0.609 per run.
Web app ranking
Web app · Model quality vs cost
10 models in the pink area cost more and score lower than DeepSeek V4.1 Flash.
Mobile app ranking
Mobile app · Model quality vs cost
8 models in the pink area cost more and score lower than DeepSeek V4.1 Flash.
Desktop app ranking
Desktop app · Model quality vs cost
9 models in the pink area cost more and score lower than DeepSeek V4.1 Flash.
Dashboard ranking
Dashboard · Model quality vs cost
3 models in the pink area cost more and score lower than DeepSeek V4.1 Flash.
Landing page ranking
Landing page · Model quality vs cost
9 models in the pink area cost more and score lower than DeepSeek V4.1 Flash.
Model efficiency
Compare completion time, cache hit rate, and token consumption. All figures are measured means.
Cache 81% · Input 318.7K · Output 15.5K
Cache 89% · Input 1462.8K · Output 29.2K
Cache 94% · Input 715.7K · Output 63.6K
Cache 80.7% · Input 236.6K · Output 14.5K
Cache 90.9% · Input 3085.9K · Output 46.0K
Cache 91.3% · Input 1422.1K · Output 24.7K
Cache 58.7% · Input 517.6K · Output 10.1K
Cache 79% · Input 836.8K · Output 28.6K
Cache 91.9% · Input 2809.3K · Output 32.7K
Cache 71.3% · Input 943.7K · Output 28.4K
Cache 70.9% · Input 392.3K · Output 27.9K
Cache 31.5% · Input 235.9K · Output 13.4K
Cache 71.7% · Input 427.1K · Output 16.1K
Evaluation method
Score requirement fulfillment and design quality, with efficiency and cost reported separately.
Tasks and samples
Models are evaluated on shared design tasks across five scenarios: web apps, mobile apps, desktop clients, dashboards and admin panels, and websites and landing pages. These results describe performance on practical prototype-generation tasks, not general model capability.
Page rendering checks
Before scoring, we check whether each artifact opens as a webpage, including blank pages, truncated content, and fatal runtime errors. Artifacts that cannot render receive zero; their results are not replaced through retesting. Opening successfully is only a prerequisite for further evaluation, not proof that the requirements are met or that the artifact is deliverable.
Requirement fulfillment and design quality
Each artifact is scored out of 100. Requirement fulfillment accounts for 30 points: pages and modules, required content, interface states, and interactions are checked against the task brief and rated complete, partially complete, or incomplete. Design quality accounts for 70 points across five equally weighted dimensions: layout and readability, information hierarchy, fit between style and task, color and contrast, and image fit and content relevance. Each dimension is rated met, partially met, or unmet.
Average score and delivery rate
The average score is the arithmetic mean of all scored artifacts for a model. An artifact scoring 80 or above counts as deliverable in this evaluation. Delivery rate is the proportion of scored artifacts that meet this threshold, indicating how consistently a model produces qualifying results. Deliverable refers to this evaluation’s prototype-quality standard, not readiness for production use.
Efficiency and cost
We report measured means for completion time, cache hit rate, input and output token usage, and tool calls separately from quality scores. Each artifact’s cost is estimated from recorded token usage at standard input, output, and cache-read prices, then averaged. These estimates compare resource consumption; they are not invoices and exclude unrecorded charges. Efficiency and cost do not contribute to the evaluation score. Changing selection preferences affects recommendation rankings, not the original evaluation results.
