DeepSeek makes its cheapest model more powerful: DeepSeek-V4-Flash
DeepSeek released the official version of DeepSeek-V4-Flash-0731, which outperforms the preview version with substantially enhanced agentic capabilities. The 304 billion parameter model achieves competitive results against proprietary models despite having fewer activated parameters than the larger DeepSeek-V4-Pro. The release includes deployment instructions for vLLM and SGLang with speculative decoding support.
What is DeepSeek-V4-Flash and how does it differ from other DeepSeek-V4 versions?
DeepSeek-V4-Flash is a lighter, more cost-effective model in the DeepSeek-V4 series optimized for faster inference. Compared to Pro and Ultra versions, it has fewer activated parameters (304B) but achieves competitive benchmark results. The model includes a speculative decoding module (DSpark) to accelerate text generation.
What are the main improvements of version 0731 over the preview version?
DeepSeek-V4-Flash-0731 shows significant improvements in agentic benchmarks – for example, Terminal Bench 2.1 improved from 61.8 to 82.7, and DeepSWE from 7.3 to 54.4. The model also supports three reasoning_effort levels (low, high, max) to control the depth of deliberation.
How is the model deployed and what deployment options are available?
The model can be deployed using vLLM or SGLang with speculative decoding support enabled. For local deployment, scripts are provided in the inference folder. Recommended settings are temperature=1.0 with top_p=0.95 for agentic scenarios and top_p=1.0 otherwise, with maximum output length of 384K tokens for higher reasoning effort levels.
What is the model's license and how can it be used?
The model is licensed under the MIT License, allowing free commercial and non-commercial use. Weights and code are publicly available on Hugging Face, and the community has already created 75 quantized versions and 11 fine-tuned models.
- DeepSeek V4 Flash on a single AMD MI300X — github.com 80 % match
- DeepSeek Harness enters developer preview with open-source architecture — deepseek.com 76 % match
- LFM2.5-2.6B: small and capable local AI model — liquid.ai 75 % match
- DeepSeek-V4-Flash
- DeepSeek-V4-Flash-0731
- DeepSeek-AI
- DeepSeek-V4-Pro3
- vLLM3
- SGLang
- DSpark
- Hugging Face12