Aug 19
DeepSeek's V4 Flash model, popular for its low cost and high benchmark scores, is facing criticism after third-party tests showed it struggling on real-world agent tasks. Composio ran 30 difficult multi-step workflows across four harnesses, with only 6 fully succeeding. The company also raised API prices on August 6, 2026, citing unprecedented demand.
Aug 14
Chinese AI startup DeepSeek has released its latest AI model, DeepSeek V4 Pro 0813, without a formal announcement. The release was first noticed on social media and message boards, where users and tech commentators highlighted the model's performance and cost efficiency. According to the open-source AI coding agent Cline, the model scores 15.8 percent higher on Terminal Bench, a benchmark that evaluates how accurately AI can execute real system operations and development tasks on a Linux terminal, compared to the April preview model. Cline also noted that DeepSeek V4 Pro 0813 runs at roughly one fifty-seventh the cost of Claude Fable 5 while delivering comparable performance. The model has 1.6 trillion parameters, 49 billion active parameters, and a 1 million token context window. Tech writer @ChrisGPT reported benchmark improvements across Terminal Bench 2.1, CyberGym, and DeepSWE, with scores rising from 72.1 to 87.9 percent, 52.7 to 83.3 percent, and 12.8 to 62.7 percent respectively. Pricing is set at 0.435 dollars per million input tokens and 0.87 dollars per million output tokens, with a discounted cache-hit input rate of 0.003625 dollars.
Aug 9
DeepSeek has released DeepSeek-V4-Flash-0731 as an open model under a commercial-use license. The MoE model has 284 billion total parameters and 13 billion active parameters. Third-party tests by Artificial Analysis rate its intelligence on par with Google's Gemini 3.6 Flash, and it outperforms GLM-5.2 in coding and GPT-5.6 Luna in agent performance.
Aug 7
DeepSeek has released the official version of its AI model DeepSeek-V4-Flash-0731, replacing the earlier preview version. The weights are available under the MIT license, allowing free commercial use. The model is a Mixture-of-Experts type with 284 billion total parameters, activating only 13 billion during inference, and supports a context length of 1 million tokens. According to benchmark scores, it surpasses the preview version of the higher-tier DeepSeek-V4-Pro, which is more than five times larger, across all nine agent-based benchmark items. Among open models, it also outperformed GLM-5.2 from China in all eight items where scores were published, though it does not reach Anthropic's Claude Opus 4.8 in any item. Community quantized versions are available, including GGUF-format versions from Unsloth in 13 steps ranging from 1-bit at 82.5GB to 8-bit at 162GB. For paid cloud usage, DeepSeek's documentation lists the price at $0.14 per million tokens for input, $0.0028 per million tokens for cached input, and $0.28 per million tokens for output, with prices doubling during peak times.