OpenDesignArena

Find the right AI model for design.

Compare quality, speed, and cost with real evaluations.

Best value

DeepSeek V4.1 Flash

Strong design performance at the lowest cost.

In this evaluation, DeepSeek V4.1 Flash achieved 98% of top-ranked GPT-6 Astra’s average score, at just 1% of its average cost.

Cost per artifact
$0.023vs $1.61
Evaluation average / 100
81.2vs 82.7

Highest score

Fastest delivery

Compare models

Pick two or three models and see all five dimensions on one chart. Different shapes mean different strengths.

Five-dimension profile

Farther out is better on all five axes: stronger results, faster completion, and lower cost.

RequirementDesign qualityDelivery rateDelivery speedCost efficiency

    Mean on each axis across all 13 models

    Choose by preference

    Balance evaluated quality, cost, and speed to find the model that suits your needs.

    What matters most to you?
    Quality
    50%
    Cost
    30%
    Speed
    20%
    1. 1 DeepSeek V4.1 Flash 78.8 Average score: 81.2/100; Average cost: $0.023; Average time: 5.3 min
    2. 2 DeepSeek V4 Flash 68.6 Average score: 70.6/100; Average cost: $0.031; Average time: 11.1 min
    3. 3 GPT-5.6 Sol 65.7 Average score: 77.6/100; Average cost: $0.544; Average time: 3.2 min
    4. 4 Gemini 3.8 Flash 64.7 Average score: 68.9/100; Average cost: $0.210; Average time: 3.8 min
    5. 5 DeepSeek V4 Pro 64.6 Average score: 72.9/100; Average cost: $0.061; Average time: 17.7 min
    6. 6 Muse Spark 1.3 63.2 Average score: 66.6/100; Average cost: $0.267; Average time: 3.3 min
    7. 7 GLM-5.3 Flash 61.4 Average score: 69.8/100; Average cost: $0.064; Average time: 23.5 min
    8. 8 GPT-6 Astra 57.5 Average score: 82.7/100; Average cost: $1.61; Average time: 11.1 min
    9. 9 Grok 4.6 56.5 Average score: 72.4/100; Average cost: $0.599; Average time: 11.4 min
    10. 10 Hunyuan H4 Preview 54.6 Average score: 74.6/100; Average cost: $0.418; Average time: 29.2 min
    11. 11 Qwen 3.8-Max 52.3 Average score: 72.0/100; Average cost: $0.595; Average time: 26.4 min
    12. 12 Claude Fable 5.1 52.1 Average score: 80.3/100; Average cost: $3.66; Average time: 12.8 min
    13. 13 Kimi K3 52.1 Average score: 65.7/100; Average cost: $0.614; Average time: 14.0 min

    Choose by use case

    Pick a design scenario; each of the five dimensions below recommends one model.

    Overall performance

    GPT-6 Astra

    Requirement · 26.5/30Design quality · 56.2/70Delivery rate · 60.0%Completion time · 11.1 minCost per run · $1.61 Requirement 26.5/30 Design quality 56.2/70 Delivery rate 60.0% Completion time 11.1 min Cost per run $1.61
    Avg. score
    82.7/100
    Avg. requirement
    26.5/30
    Avg. design quality
    56.2/70
    Avg. delivery rate
    60.0%
    Mean time
    11.1 min
    Mean cost / artifact
    $1.61/run
    Design quality

    Claude Fable 5.1

    Avg. design quality
    53.2/70
    Avg. score
    80.3/100
    Avg. delivery rate
    56.7%
    Avg. requirement
    27.1/30
    Mean time
    12.8 min
    Mean cost / artifact
    $3.66/run
    Requirement fulfillment

    DeepSeek V4.1 Flash

    Avg. requirement
    28.4/30
    Avg. score
    81.2/100
    Avg. delivery rate
    57.7%
    Avg. design quality
    52.8/70
    Mean time
    5.3 min
    Mean cost / artifact
    $0.023/run
    Completion speed

    GPT-5.6 Sol

    Mean time
    3.2 min
    Avg. score
    77.6/100
    Avg. delivery rate
    53.3%
    Avg. requirement
    25.8/30
    Avg. design quality
    51.8/70
    Mean cost / artifact
    $0.544/run
    Cost efficiency

    Hunyuan H4 Preview

    Mean cost / artifact
    $0.418/run
    Avg. score
    74.6/100
    Avg. delivery rate
    40.0%
    Avg. requirement
    26.8/30
    Avg. design quality
    47.8/70
    Mean time
    29.2 min

    Quality and cost

    Compare model performance and the cost of achieving it.

    Quality ranking

    82.7/100 GPT-6 Astra
    81.2/100 DeepSeek V4.1 Flash
    80.3/100 Claude Fable 5.1
    77.6/100 GPT-5.6 Sol
    74.6/100 Hunyuan H4 Preview
    72.9/100 DeepSeek V4 Pro
    72.4/100 Grok 4.6
    72.0/100 Qwen 3.8-Max
    70.6/100 DeepSeek V4 Flash
    69.8/100 GLM-5.3 Flash
    68.9/100 Gemini 3.8 Flash
    66.6/100 Muse Spark 1.3
    65.7/100 Kimi K3

    Model quality vs cost

    Average score Cost per artifact (USD) GPT-6 Astra GPT-6 Astra Average score 82.7 · $1.61 DeepSeek V4.1 Flash DeepSeek V4.1 Flash Average score 81.2 · $0.023 Claude Fable 5.1 Claude Fable 5.1 Average score 80.3 · $3.66 GPT-5.6 Sol GPT-5.6 Sol Average score 77.6 · $0.544 Hunyuan H4 Preview Hunyuan H4 Preview Average score 74.6 · $0.418 DeepSeek V4 Pro DeepSeek V4 Pro Average score 72.9 · $0.061 Grok 4.6 Grok 4.6 Average score 72.4 · $0.599 Qwen 3.8-Max Qwen 3.8-Max Average score 72.0 · $0.595 DeepSeek V4 Flash DeepSeek V4 Flash Average score 70.6 · $0.031 GLM-5.3 Flash GLM-5.3 Flash Average score 69.8 · $0.064 Gemini 3.8 Flash Gemini 3.8 Flash Average score 68.9 · $0.210 Muse Spark 1.3 Muse Spark 1.3 Average score 66.6 · $0.267 Kimi K3 Kimi K3 Average score 65.7 · $0.614

    11 models in the pink area cost more and score lower than DeepSeek V4.1 Flash.

    Scenario performance

    Web app ranking

    83.1/100 GPT-5.6 Sol
    82.6/100 Claude Fable 5.1
    82.4/100 DeepSeek V4.1 Flash
    77.1/100 GPT-6 Astra
    68.5/100 Grok 4.6
    68.4/100 Gemini 3.8 Flash
    62.0/100 Muse Spark 1.3
    60.6/100 GLM-5.3 Flash
    60.5/100 Hunyuan H4 Preview
    57.0/100 Qwen 3.8-Max
    56.4/100 DeepSeek V4 Flash
    51.8/100 DeepSeek V4 Pro
    35.9/100 Kimi K3

    Web app · Model quality vs cost

    Average score Cost per artifact (USD) GPT-5.6 Sol GPT-5.6 Sol Average score 83.1 · $0.537 Claude Fable 5.1 Claude Fable 5.1 Average score 82.6 · $5.55 DeepSeek V4.1 Flash DeepSeek V4.1 Flash Average score 82.4 · $0.030 GPT-6 Astra GPT-6 Astra Average score 77.1 · $1.87 Grok 4.6 Grok 4.6 Average score 68.5 · $0.818 Gemini 3.8 Flash Gemini 3.8 Flash Average score 68.4 · $0.185 Muse Spark 1.3 Muse Spark 1.3 Average score 62.0 · $0.293 GLM-5.3 Flash GLM-5.3 Flash Average score 60.6 · $0.093 Hunyuan H4 Preview Hunyuan H4 Preview Average score 60.5 · $0.609 Qwen 3.8-Max Qwen 3.8-Max Average score 57.0 · $0.552 DeepSeek V4 Flash DeepSeek V4 Flash Average score 56.4 · $0.041 DeepSeek V4 Pro DeepSeek V4 Pro Average score 51.8 · $0.067 Kimi K3 Kimi K3 Average score 35.9 · $0.486

    10 models in the pink area cost more and score lower than DeepSeek V4.1 Flash.

    Model efficiency

    Compare completion time, cache hit rate, and token consumption. All figures are measured means.

    Scenario
    Average scoreCompletion time
    High score · High efficiency

    Evaluation method

    Score requirement fulfillment and design quality, with efficiency and cost reported separately.

    Tasks and samples

    Models are evaluated on shared design tasks across five scenarios: web apps, mobile apps, desktop clients, dashboards and admin panels, and websites and landing pages. These results describe performance on practical prototype-generation tasks, not general model capability.

    Page rendering checks

    Before scoring, we check whether each artifact opens as a webpage, including blank pages, truncated content, and fatal runtime errors. Artifacts that cannot render receive zero; their results are not replaced through retesting. Opening successfully is only a prerequisite for further evaluation, not proof that the requirements are met or that the artifact is deliverable.

    Requirement fulfillment and design quality

    Each artifact is scored out of 100. Requirement fulfillment accounts for 30 points: pages and modules, required content, interface states, and interactions are checked against the task brief and rated complete, partially complete, or incomplete. Design quality accounts for 70 points across five equally weighted dimensions: layout and readability, information hierarchy, fit between style and task, color and contrast, and image fit and content relevance. Each dimension is rated met, partially met, or unmet.

    Average score and delivery rate

    The average score is the arithmetic mean of all scored artifacts for a model. An artifact scoring 80 or above counts as deliverable in this evaluation. Delivery rate is the proportion of scored artifacts that meet this threshold, indicating how consistently a model produces qualifying results. Deliverable refers to this evaluation’s prototype-quality standard, not readiness for production use.

    Efficiency and cost

    We report measured means for completion time, cache hit rate, input and output token usage, and tool calls separately from quality scores. Each artifact’s cost is estimated from recorded token usage at standard input, output, and cache-read prices, then averaged. These estimates compare resource consumption; they are not invoices and exclude unrecorded charges. Efficiency and cost do not contribute to the evaluation score. Changing selection preferences affects recommendation rankings, not the original evaluation results.

    OpenDesign Desktop

    One design system. Every output unmistakably your brand

    Inside the full Vibe Design Workspace, use the same brand rules across websites, slide decks, interactive prototypes, dashboards, images, and HTML video. Connect Codex, Claude Code, Cursor, and other coding agents already on your computer, then create locally for free.

    • Web, slides, prototypes, dashboards, images, and video
    • 140+ design systems, plus the full template and skill library
    • Connect local Codex and 21+ coding agents · Free to use
    Download free

    Available for macOS and Windows