Product Signal

DeepSeek V4 Flash Becomes a Full Deployment Package

DeepSeek moved V4 Flash into API public beta with the 0731 model, new agent benchmarks, Codex support, published pricing, and MIT-licensed weights.

DeepSeek has moved the DeepSeek-V4-Flash-0731 API into public beta while retaining the deepseek-v4-flash model name. The release adds published pricing, Responses API support, Codex integration, and higher concurrency. Compared with the April Preview, it also brings stronger vendor-reported agent benchmarks and MIT-licensed weights, giving users a choice between the hosted API and self-hosting.

More Than a Release-Stage Label

V4 Flash was already available through the API during the April Preview, with a 1M-token context window and support for OpenAI ChatCompletions and the Anthropic API. The 0731 version keeps the same model name, but DeepSeek reports materially stronger agent performance: 82.7 on Terminal Bench 2.1, 76.7 on Cybergym, 70.3 on Toolathlon verified, and 54.4 on DeepSWE. The substantive change is therefore the combination of a model update and a broader deployment surface, not first-time access.

Hosted and Self-Hosted Paths

The current API supports a 1M-token context window, up to 384K output tokens, JSON Output, Tool Calls, the Responses API, and the Anthropic API. Cache-miss input and output are priced at $0.14 and $0.28 per million tokens, respectively, while the concurrency limit is 2,500 versus 500 for V4 Pro. DeepSeek also says V4 Flash is currently the only V4 model integrated with Codex. In parallel, its 284B-total, 13B-active FP4 and FP8 mixed-precision weights are downloadable under the MIT license.

Deployment Becomes the Competitive Variable

The package expands V4 Flash's competitive surface from benchmark scores to interface compatibility, concurrency, token pricing, and deployment control. Codex support can reduce migration work for existing coding-agent workflows, while open weights provide an alternative for teams that need local control. The strongest countercase is that every performance result is vendor-reported and the Preview was already usable. Until independent replications and adoption data emerge, a more complete release package should not be treated as proof of superior production performance.

What to watch next

Three observable tests come next: whether independent teams can reproduce the Terminal Bench 2.1, Cybergym, and DeepSWE results; whether V4 Pro receives Codex support as planned in early August; and what developers measure for latency, reliability, and cost per completed task at the 2,500-concurrency limit. Those results will show whether V4 Flash merely has a more complete interface or materially changes agent deployment choices.

Sources