← all stories

DeepSeek-V4-Flash

DeepSeek-V4-Flash-0731 is an open-weights Mixture-of-Experts model released under the MIT license, with 284 billion total parameters and 13 billion active parameters. It is available for free commercial use, with paid cloud pricing listed at $0.14 per million input tokens and $0.28 per million output tokens.

Synthesized from 2 Yomimono stories · updated Aug 19

DeepSeek released the official version of DeepSeek-V4-Flash-0731 on August 7, 2026, replacing an earlier preview. The model is open under the MIT license, allowing free commercial use. It is a Mixture-of-Experts design with 284 billion total parameters, activating only 13 billion during inference, and supports a context length of 1 million tokens.

Benchmark results show the model surpasses the preview version of the higher-tier DeepSeek-V4-Pro, which is more than five times larger, across all nine agent-based benchmark items. Among open models, it outperformed GLM-5.2 from China in all eight items where scores were published, though it does not reach Anthropic's Claude Opus 4.8 in any item. Third-party tests by Artificial Analysis rate its intelligence on par with Google's Gemini 3.6 Flash, and it outperforms GLM-5.2 in coding and GPT-5.6 Luna in agent performance.

The release came hours after OpenAI cut GPT-5.6 Luna prices by 80 percent. DeepSeek has notified users of a significant price increase for the model, suggesting the low-cost strategy drew demand that strained its computing resources. Community quantized versions are available, including GGUF-format versions from Unsloth in 13 steps ranging from 1-bit at 82.5GB to 8-bit at 162GB. For paid cloud usage, DeepSeek's documentation lists the price at $0.14 per million tokens for input, $0.0028 per million tokens for cached input, and $0.28 per million tokens for output, with prices doubling during peak times.

Key facts

Release date
August 7, 2026 (official release, replacing preview)
License
MIT license, free for commercial use
Model type
Mixture-of-Experts
Parameters
284 billion total, 13 billion active
Context length
1 million tokens
Benchmark performance
Surpasses preview DeepSeek-V4-Pro on all nine agent-based items; outperforms GLM-5.2 on all eight published items; does not reach Claude Opus 4.8 on any item
Third-party intelligence rating
On par with Gemini 3.6 Flash per Artificial Analysis
Cloud pricing
$0.14 per million input tokens, $0.0028 per million cached input tokens, $0.28 per million output tokens; doubles during peak times
Quantized versions
GGUF versions from Unsloth in 13 steps, from 1-bit at 82.5GB to 8-bit at 162GB

Timeline

Synthesized by Yomimono from the cited Yomimono stories below, each itself sourced, then editorially reviewed. Every fact links the story it came from.

Facts

Noted
released the official version of its AI model DeepSeek-V4-Flash-0731, replacing the earlier preview version · 2026-08-07
Noted
weights are available under the MIT license · 2026-08-07
Noted
allowing free commercial use · 2026-08-07
Noted
model is a Mixture-of-Experts type · 2026-08-07
Noted
has 284 billion total parameters · 2026-08-07
Noted
activates only 13 billion during inference · 2026-08-07
Noted
supports a context length of 1 million tokens · 2026-08-07
Noted
surpasses the preview version of DeepSeek-V4-Pro across all nine agent-based benchmark items · 2026-08-07
Noted
outperformed GLM-5.2 in all eight items where scores were published · 2026-08-07
Noted
does not reach Anthropic's Claude Opus 4.8 in any item · 2026-08-07
Noted
community quantized versions are available · 2026-08-07
Noted
GGUF-format versions from Unsloth in 13 steps ranging from 1-bit at 82.5GB to 8-bit at 162GB · 2026-08-07
Noted
input price is $0.14 per million tokens · 2026-08-07
Noted
cached input price is $0.0028 per million tokens · 2026-08-07
Noted
output price is $0.28 per million tokens · 2026-08-07
Noted
prices double during peak times · 2026-08-07
Noted
DeepSeek has notified users of a significant price increase for the model · 2026-08-07
Noted
DeepSeek's V4 Flash model, popular for its low cost and high benchmark scores, is facing criticism after third-party tests showed it struggling on real-world agent tasks. · 2026-08-19
Noted
Composio ran 30 difficult multi-step workflows across four harnesses, with only 6 fully succeeding. · 2026-08-19
Noted
The company also raised API prices on August 6, 2026, citing unprecedented demand. · 2026-08-19

Structured graph also available as JSON at /public/entities/deepseek-v4-flash. CC BY 4.0.

All coverage

Aug 19

DeepSeek V4 Flash Stumbles on Real Agent Tasks as Prices Surge

DeepSeek's V4 Flash model, popular for its low cost and high benchmark scores, is facing criticism after third-party tests showed it struggling on real-world agent tasks. Composio ran 30 difficult multi-step workflows across four harnesses, with only 6 fully succeeding. The company also raised API prices on August 6, 2026, citing unprecedented demand.

Aug 9

DeepSeek Releases V4-Flash-0731 Open Model

DeepSeek has released DeepSeek-V4-Flash-0731 as an open model under a commercial-use license. The MoE model has 284 billion total parameters and 13 billion active parameters. Third-party tests by Artificial Analysis rate its intelligence on par with Google's Gemini 3.6 Flash, and it outperforms GLM-5.2 in coding and GPT-5.6 Luna in agent performance.

Aug 7

DeepSeek-V4-Flash Official Release Makes Weights Free for Commercial Use

DeepSeek has released the official version of its AI model DeepSeek-V4-Flash-0731, replacing the earlier preview version. The weights are available under the MIT license, allowing free commercial use. The model is a Mixture-of-Experts type with 284 billion total parameters, activating only 13 billion during inference, and supports a context length of 1 million tokens. According to benchmark scores, it surpasses the preview version of the higher-tier DeepSeek-V4-Pro, which is more than five times larger, across all nine agent-based benchmark items. Among open models, it also outperformed GLM-5.2 from China in all eight items where scores were published, though it does not reach Anthropic's Claude Opus 4.8 in any item. Community quantized versions are available, including GGUF-format versions from Unsloth in 13 steps ranging from 1-bit at 82.5GB to 8-bit at 162GB. For paid cloud usage, DeepSeek's documentation lists the price at $0.14 per million tokens for input, $0.0028 per million tokens for cached input, and $0.28 per million tokens for output, with prices doubling during peak times.