DeepSeek released DeepSeek-V4.1-Flash on Sept. 10, 2026. Reuters reported the launch the same morning. The company put the model on its API as deepseek-flash and published MIT-licensed weights on Hugging Face.
DeepSeek says the model is a multimodal mixture-of-experts with 552 billion backbone parameters, native image and text input, text output, and a 1 million-token context. Company docs list a maximum output of 384,000 tokens and concurrency of 2,500 for deepseek-flash.
Company claim: a Causal Encoder-Decoder design activates about 8 billion parameters per token in prefill and 16 billion in decode, with a global KV cache of 890 bytes per token, about one-quarter of V4-Flash. The technical report is titled as a KV-cache compression push and lists Engram memory, CSA2 attention, and DSpark speculative decoding among the stack.
First-party API prices, per 1 million tokens, list cache-hit input at $0.003 off-peak and $0.006 peak, cache-miss input at $0.15 and $0.30, and output at $0.60 and $1.20. Peak windows are 01:00 to 04:00 UTC and 06:00 to 10:00 UTC, Monday through Friday. Artificial Analysis publishes the peak rates on its model page.
The Sept. 10 launch materials said deepseek-v4-pro requests would route to V4.1-Flash at Flash rates from 04:00 UTC on Sept. 14. Live pricing footnotes and the changelog, rechecked after that window on Sept. 14, say DeepSeek will keep providing V4-Pro API service with billing unchanged. The launch page still shows the superseded route line.
Independent: Artificial Analysis gave a max-effort reasoning configuration an Intelligence Index score of 40 and about 218.7 output tokens per second in its displayed class at the evidence snapshot. Those ranks move. DeepSeek's vendor tables and broad "ahead of V4-Pro" language remain company claims; the same vendor table still shows V4-Pro higher on some listed exams.
Hugging Face hosts about 510 GB of safetensors shards under an MIT license. Treat app or web chat availability as unverified from first-party launch text unless a current DeepSeek surface confirms it.