DeepSeek V4.1 Flash model is released. It is the smallest model in its new series of model structures, with native multi-modal visual understanding capabilities. The design of the new model structure aims to achieve higher capability limits, faster reasoning speed, greater throughput, and scalability to larger parameter models.
Input Activation Only 8B, HBM Requirement Reduced to 1/4
DeepSeek V4.1 Flash is a 552B parameter MoE model that adopts a new Causal-Encoder-Decoder structure, with asymmetric input and output—input activation is only 8B, while output activation is 16B, significantly lower in cost than similarly sized models known so far. The new model also uses a new pre-training method and has been further trained on a larger scale through reinforcement learning, successfully surpassing the intelligence level of several flagship models including DeepSeek V4 Pro in benchmark tests.
In addition, the new generation model greatly reduces the size of the KV Cache, reducing the requirement for HBM to 1/4 compared to the previous generation, and for SSD to 1/8; compared to the original model, the KV Cache is only 1/437. In Agent usage scenarios, cache hit costs often account for a large proportion, and the compression of KV Cache significantly reduces the usage cost of Agent-type tasks.

API Has Been Synchronized, V4 Pro Will Be Routed on September 14
DeepSeek V4.1 Flash has been synchronized on the DeepSeek API, natively supporting multi-modal, and can be called by changing the model name to deepseek-flash. The old version models V4 Flash and V4 Flash Vision Exp have been discontinued. For compatibility reasons, the model names deepseek-v4-flash and deepseek-v4-flash-vision-exp will be temporarily routed to V4.1 Flash.
After multiple tests, V4.1 Flash has comprehensively surpassed V4 Pro in performance, cost, speed, and total time indicators. Therefore, the official plans to orderly discontinue V4 Pro: after 12:00 Beijing Time on September 14, 2026, until the release of V4.1 Pro, requests accessing deepseek-v4-pro will be redirected to V4.1 Flash and billed at the V4.1 Flash price.

Tencent (WorkBuddy, CodeBuddy) and OpenCode, as official partners, have now fully integrated DeepSeek V4.1 Flash. Thanks to architectural innovations, the official has accordingly reduced the pricing of V4.1 Flash, while still adopting peak-valley pricing—off-peak prices are half of peak-hour prices. The new pricing takes effect starting at 12:00 on September 10, 2026. Model weights and technical reports have been synchronized and made available on Hugging Face.




