DeepSeek launches V4.1-Flash with a leaner KV cache
DeepSeek's 10 September 2026 API release note ships V4.1-Flash, a 552B MoE with 8B active input and 16B output, and a KV cache at 1/4 HBM and 1/8 SSD. After 12:00 Beijing Time on 14 September, deepseek-v4-pro routes to Flash at Flash prices until V4.1-Pro.

DeepSeek's 10 September 2026 API docs news post is a product and API release note for DeepSeek-V4.1-Flash, the smallest model in its new architecture family, with native visual understanding. The same-day changelog lists vendor scores including GPQA Diamond 90.9, Terminal-Bench 2.1 90.6, and a Codeforces rating of 3471. Those figures are DeepSeek's own table, not a third-party audit.
The model is a 552 billion parameter Mixture-of-Experts design. DeepSeek describes a new Causal Encoder-Decoder architecture that activates 8 billion parameters on input and 16 billion on output. Against the prior generation, it says the V4.1-Flash KV cache needs one quarter the HBM and one eighth the SSD storage.
It is live on the DeepSeek API as deepseek-flash, with native multimodal support. V4-Flash and V4-Flash-Vision-Exp are retired. For compatibility, deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash.
The cutoff most one-line takes skip is already on the changelog. After 12:00 Beijing Time on 14 September 2026 (04:00 UTC on the news post), and until V4.1-Pro ships, every deepseek-v4-pro request routes to V4.1-Flash and is billed at V4.1-Flash prices. DeepSeek says its testing puts V4.1-Flash ahead of V4-Pro on performance, cost, speed, and total time. That is a vendor claim.
TechTimes adds colour the docs do not headline: a global KV cache of 890 bytes per token, and DeepSWE scores that swing from 65.5% to 74.2% depending on the harness. Benchmark wins stay DeepSeek-reported. Scores on agent tasks can move with the harness wrapped around the same checkpoint.
DeepSeek posted weights on Hugging Face. TechTimes says they ship under the MIT license. Hosted API traffic still sits under Chinese intelligence and cybersecurity statutes, the same legal backdrop covered when CISA flagged China-linked distillation around DeepSeek and Alibaba. Open weights change the data-transit story for self-hosters, but they do not rewrite that legal backdrop.
The release sits next to other open-weight and hosted-model moves, including Tencent's Hy4 preview, Nvidia's agreed Hugging Face purchase, and OpenAI's private safety processing path.
If you call deepseek-v4-pro today, decide before 14 September Beijing noon whether Flash pricing and routing are acceptable for production, or migrate endpoints and eval harnesses now while V4.1-Flash is still optional.
Subscribe to Techpresso
Free daily newsletter, read in 5 minutes.
Subscribe free