The Code/X ArchiveView on X
OpenDesign

@OpenDesignHQ

We benchmarked DeepSeek V4.1 Flash by @deepseek_ai .

It reached 98% of GPT-6 Astra’s score at 1.4% of the cost on everyday design tasks based on user requests.

Every model except Astra scored lower AND cost more.

Are open models overtaking closed ones?
Full results below ↘️
Image from the post
31794710.7K3.6K
OpenDesign

@OpenDesignHQ

1) Most benchmarks test models’ limits. But we noticed that most users just want to know: which model should I use for everyday design work?

We built OpenDesign Arena to help you choose.
open-design.ai/zh/llm-arena-f…
Image from the post
31440576
OpenDesign

@OpenDesignHQ

2) Are closed models doomed?

11 models scored lower AND cost more than DeepSeek V4.1 Flash across our real-world design benchmark.

Only GPT-6 Astra scored higher.
Image from the post
11331243
OpenDesign

@OpenDesignHQ

3) Cheaper. Faster. Nearly the same score.

Average time & cost per artifact:
DeepSeek V4.1 Flash: 5.3 min / $0.023
GPT-6 Astra: 11.1 min / $1.61
Claude Fable 5.1: 12.8 min / $3.66

98% of Astra’s score. 1/70th the cost.
Image from the post
5727338
OpenDesign

@OpenDesignHQ

4) Top-2 score. Fastest among the top 3.
DeepSeek V4.1 Flash delivers in 5.3 minutes—less than half the time of Astra or Claude.
Image from the post
1413214
OpenDesign

@OpenDesignHQ

5)DeepSeek V4.1 Flash across app design tasks:

Web apps: #3 — 82.4
Mobile apps: #5 — 82.5
Desktop apps: tied #3 — 84.4
Image from the postImage from the postImage from the post
1212813
OpenDesign

@OpenDesignHQ

6) DeepSeek V4.1 Flash on websites and dashboards:

Websites & landing pages: #4 — 79.3
Dashboards & admin panels: #10 — 76.1
Ahead of Astra on landing pages. Behind on dashboards.

Choose for your task, not just the overall ranking.
Image from the postImage from the post
039714
End of thread