TECH FLOW Svět Androida
← Back to the stream
huggingface.co · picked by Petr Mišák · 60d ago

DeepSeek makes its cheapest model more powerful: DeepSeek-V4-Flash

Source preview: DeepSeek makes its cheapest model more powerful: DeepSeek-V4-Flash
AI summary

DeepSeek released the official version of DeepSeek-V4-Flash-0731, which outperforms the preview version with substantially enhanced agentic capabilities. The 304 billion parameter model achieves competitive results against proprietary models despite having fewer activated parameters than the larger DeepSeek-V4-Pro. The release includes deployment instructions for vLLM and SGLang with speculative decoding support.

The summary is written by AI from the source; it isn’t the newsroom’s opinion. For details, read the source.

13 people have already opened the source

AI questions & answers
What is DeepSeek-V4-Flash and how does it differ from other DeepSeek-V4 versions?

DeepSeek-V4-Flash is a lighter, more cost-effective model in the DeepSeek-V4 series optimized for faster inference. Compared to Pro and Ultra versions, it has fewer activated parameters (304B) but achieves competitive benchmark results. The model includes a speculative decoding module (DSpark) to accelerate text generation.

What are the main improvements of version 0731 over the preview version?

DeepSeek-V4-Flash-0731 shows significant improvements in agentic benchmarks – for example, Terminal Bench 2.1 improved from 61.8 to 82.7, and DeepSWE from 7.3 to 54.4. The model also supports three reasoning_effort levels (low, high, max) to control the depth of deliberation.

How is the model deployed and what deployment options are available?

The model can be deployed using vLLM or SGLang with speculative decoding support enabled. For local deployment, scripts are provided in the inference folder. Recommended settings are temperature=1.0 with top_p=0.95 for agentic scenarios and top_p=1.0 otherwise, with maximum output length of 384K tokens for higher reasoning effort levels.

What is the model's license and how can it be used?

The model is licensed under the MIT License, allowing free commercial and non-commercial use. Weights and code are publicly available on Hugging Face, and the community has already created 75 quantized versions and 11 fine-tuned models.

Questions and answers are written by AI about the topic, not taken from the source; they aren’t the newsroom’s opinion.
Related from the stream
Mentions