DeepSeek-V4-Flash-0731 is an open-weights Mixture-of-Experts model released under the MIT license, with 284 billion total parameters and 13 billion active parameters. It is available for free commercial use, with paid cloud pricing listed at $0.14 per million input tokens and $0.28 per million output tokens.
DeepSeek released the official version of DeepSeek-V4-Flash-0731 on August 7, 2026, replacing an earlier preview. The model is open under the MIT license, allowing free commercial use. It is a Mixture-of-Experts design with 284 billion total parameters, activating only 13 billion during inference, and supports a context length of 1 million tokens.
Benchmark results show the model surpasses the preview version of the higher-tier DeepSeek-V4-Pro, which is more than five times larger, across all nine agent-based benchmark items. Among open models, it outperformed GLM-5.2 from China in all eight items where scores were published, though it does not reach Anthropic's Claude Opus 4.8 in any item. Third-party tests by Artificial Analysis rate its intelligence on par with Google's Gemini 3.6 Flash, and it outperforms GLM-5.2 in coding and GPT-5.6 Luna in agent performance.
The release came hours after OpenAI cut GPT-5.6 Luna prices by 80 percent. DeepSeek has notified users of a significant price increase for the model, suggesting the low-cost strategy drew demand that strained its computing resources. Community quantized versions are available, including GGUF-format versions from Unsloth in 13 steps ranging from 1-bit at 82.5GB to 8-bit at 162GB. For paid cloud usage, DeepSeek's documentation lists the price at $0.14 per million tokens for input, $0.0028 per million tokens for cached input, and $0.28 per million tokens for output, with prices doubling during peak times.
Synthesized by Yomimono from the cited Yomimono stories below, each itself
sourced, then editorially reviewed. Every
fact links the story it came from.
Facts
- Noted
- released the official version of its AI model DeepSeek-V4-Flash-0731, replacing the earlier preview version · 2026-08-07
- Noted
- weights are available under the MIT license · 2026-08-07
- Noted
- allowing free commercial use · 2026-08-07
- Noted
- model is a Mixture-of-Experts type · 2026-08-07
- Noted
- has 284 billion total parameters · 2026-08-07
- Noted
- activates only 13 billion during inference · 2026-08-07
- Noted
- supports a context length of 1 million tokens · 2026-08-07
- Noted
- surpasses the preview version of DeepSeek-V4-Pro across all nine agent-based benchmark items · 2026-08-07
- Noted
- outperformed GLM-5.2 in all eight items where scores were published · 2026-08-07
- Noted
- does not reach Anthropic's Claude Opus 4.8 in any item · 2026-08-07
- Noted
- community quantized versions are available · 2026-08-07
- Noted
- GGUF-format versions from Unsloth in 13 steps ranging from 1-bit at 82.5GB to 8-bit at 162GB · 2026-08-07
- Noted
- input price is $0.14 per million tokens · 2026-08-07
- Noted
- cached input price is $0.0028 per million tokens · 2026-08-07
- Noted
- output price is $0.28 per million tokens · 2026-08-07
- Noted
- prices double during peak times · 2026-08-07
- Noted
- DeepSeek has notified users of a significant price increase for the model · 2026-08-07
- Noted
- DeepSeek's V4 Flash model, popular for its low cost and high benchmark scores, is facing criticism after third-party tests showed it struggling on real-world agent tasks. · 2026-08-19
- Noted
- Composio ran 30 difficult multi-step workflows across four harnesses, with only 6 fully succeeding. · 2026-08-19
- Noted
- The company also raised API prices on August 6, 2026, citing unprecedented demand. · 2026-08-19
Structured graph also available as JSON at /public/entities/deepseek-v4-flash.
CC BY 4.0.
Aug 19
DeepSeek's V4 Flash model, popular for its low cost and high benchmark scores, is facing criticism after third-party tests showed it struggling on real-world agent tasks. Composio ran 30 difficult multi-step workflows across four harnesses, with only 6 fully succeeding. The company also raised API prices on August 6, 2026, citing unprecedented demand.
Aug 9
DeepSeek has released DeepSeek-V4-Flash-0731 as an open model under a commercial-use license. The MoE model has 284 billion total parameters and 13 billion active parameters. Third-party tests by Artificial Analysis rate its intelligence on par with Google's Gemini 3.6 Flash, and it outperforms GLM-5.2 in coding and GPT-5.6 Luna in agent performance.
Aug 7
DeepSeek has released the official version of its AI model DeepSeek-V4-Flash-0731, replacing the earlier preview version. The weights are available under the MIT license, allowing free commercial use. The model is a Mixture-of-Experts type with 284 billion total parameters, activating only 13 billion during inference, and supports a context length of 1 million tokens. According to benchmark scores, it surpasses the preview version of the higher-tier DeepSeek-V4-Pro, which is more than five times larger, across all nine agent-based benchmark items. Among open models, it also outperformed GLM-5.2 from China in all eight items where scores were published, though it does not reach Anthropic's Claude Opus 4.8 in any item. Community quantized versions are available, including GGUF-format versions from Unsloth in 13 steps ranging from 1-bit at 82.5GB to 8-bit at 162GB. For paid cloud usage, DeepSeek's documentation lists the price at $0.14 per million tokens for input, $0.0028 per million tokens for cached input, and $0.28 per million tokens for output, with prices doubling during peak times.