Key facts
- DeepSeek V4.1 Flash scored 81.2 out of 100 on design tasks, 98% of GPT-6 Astra's 82.7.
- DeepSeek V4.1 Flash costs $0.023 per finished design, compared to GPT-6 Astra's $1.61.
- DeepSeek V4.1 Flash completed tasks in 5.3 minutes, while GPT-6 Astra took 11.1 minutes.
- Claude Fable 5.1 scored 80.3, took 12.8 minutes, and cost $3.66 per design.
- DeepSeek V4.1 Flash activates 8 billion of its 552 billion parameters to read prompts.
OpenDesign Arena has benchmarked 13 AI models on their ability to perform real-world design tasks, such as building web apps, dashboards, and landing pages. OpenAI's GPT-6 Astra emerged as the top performer, scoring 82.7 out of 100. However, DeepSeek's newly released V4.1 Flash model closely followed, achieving a score of 81.2, which is 98% of GPT-6 Astra's performance. Crucially, DeepSeek's model is significantly more cost-effective and faster, costing $0.023 per finished design compared to GPT-6 Astra's $1.61, and completing tasks in 5.3 minutes versus Astra's 11.1 minutes. Other tested models, including Claude Fable 5.1, Grok 4.6, and Qwen 3.8-Max, scored lower than DeepSeek's offering and were more expensive to run. DeepSeek's technical paper attributes the V4.1 Flash model's efficiency to its Causal Encoder-Decoder design, which activates only 8 billion of its 552 billion parameters for prompt reading and 16 billion for response writing. This approach allows for faster completion times and lower operational costs. OpenDesign's scoring methodology prioritizes outputs that meet the design brief and assesses quality based on layout, hierarchy, color, and style fit, with a model's output only being scored if it renders as a working webpage.
