DeepSeek has quietly shipped a major update to its V4-Flash model — and the small, cheap model now outperforms its own much larger sibling.
The update, released on July 31 with open weights on Hugging Face (deepseek-ai/DeepSeek-V4-Flash-0731), marks an unusual moment for the efficiency-focused model. DeepSeek's own benchmarks show V4-Flash now beats V4 Pro (Preview), the 1.6-trillion-parameter flagship, on agentic coding. Independent checks by Artificial Analysis also found improvements on other benchmarks such as GPQA Diamond, with no reported regressions.
V4-Flash is a Mixture-of-Experts model with 284 billion total parameters but only 13 billion active per token, supporting a 1-million-token context window. On OpenRouter it lists at about $0.087 per million input tokens and $0.17 per million output — a small fraction of the cost of frontier rivals, which is why coverage this week framed it as beating models that cost 50x more.
The release continues DeepSeek's pattern of publishing weights immediately rather than holding them back. The company notes the update is tuned for agent workflows, coding assistants and high-throughput chat, and is integrated with tools like Claude Code, OpenClaw and OpenCode.
Observers see V4-Flash-0731 as a promising sign for the next V4 Pro update — and a reminder that the price-performance frontier is moving fast, right in the middle of a broader industry debate over AI spending.




