Product Signal
DeepSeek V4 Flash Becomes a Full Deployment Package
DeepSeek moved V4 Flash into API public beta with the 0731 model, new agent benchmarks, Codex support, published pricing, and MIT-licensed weights.
DeepSeek has moved the DeepSeek-V4-Flash-0731 API into public beta while retaining the deepseek-v4-flash model name. The release adds published pricing, Responses API support, Codex integration, and higher concurrency. Compared with the April Preview, it also brings stronger vendor-reported agent benchmarks and MIT-licensed weights, giving users a choice between the hosted API and self-hosting.
More Than a Release-Stage Label
V4 Flash was already available through the API during the April Preview, with a 1M-token context window and support for OpenAI ChatCompletions and the Anthropic API. The 0731 version keeps the same model name, but DeepSeek reports materially stronger agent performance: 82.7 on Terminal Bench 2.1, 76.7 on Cybergym, 70.3 on Toolathlon verified, and 54.4 on DeepSWE. The substantive change is therefore the combination of a model update and a broader deployment surface, not first-time access.
Hosted and Self-Hosted Paths
The current API supports a 1M-token context window, up to 384K output tokens, JSON Output, Tool Calls, the Responses API, and the Anthropic API. Cache-miss input and output are priced at $0.14 and $0.28 per million tokens, respectively, while the concurrency limit is 2,500 versus 500 for V4 Pro. DeepSeek also says V4 Flash is currently the only V4 model integrated with Codex. In parallel, its 284B-total, 13B-active FP4 and FP8 mixed-precision weights are downloadable under the MIT license.
Deployment Becomes the Competitive Variable
The package expands V4 Flash's competitive surface from benchmark scores to interface compatibility, concurrency, token pricing, and deployment control. Codex support can reduce migration work for existing coding-agent workflows, while open weights provide an alternative for teams that need local control. The strongest countercase is that every performance result is vendor-reported and the Preview was already usable. Until independent replications and adoption data emerge, a more complete release package should not be treated as proof of superior production performance.
What to watch next
Three observable tests come next: whether independent teams can reproduce the Terminal Bench 2.1, Cybergym, and DeepSWE results; whether V4 Pro receives Codex support as planned in early August; and what developers measure for latency, reliability, and cost per completed task at the 2,500-concurrency limit. Those results will show whether V4 Flash merely has a more complete interface or materially changes agent deployment choices.
Sources
- DeepSeek API Changelog — DeepSeek-V4-Flash Update
- DeepSeek API Docs — DeepSeek V4 Preview Release
- DeepSeek API Docs — Models & Pricing
- DeepSeek API Docs — Integrate with Codex
- Hugging Face / DeepSeek-AI — deepseek-ai/DeepSeek-V4-Flash
- Simon Willison — deepseek-ai/DeepSeek-V4-Flash-0731
- Product Hunt — DeepSeek-V4-Flash-0731
- DeepSeek — DeepSeek-V4-Flash-0731
- Artificial Analysis — DeepSeek V4 Flash (max) - Intelligence, Performance & Price Analysis
- The Information — DeepSeek Makes a Splash with Small, Affordable V4-Flash Model
- DeepSeek — Change Log
- DeepSeek API Changelog — DeepSeek-V4-Pro Update
- DeepSeek — Models & Pricing
- National Institute of Standards and Technology — CAISI Evaluation of DeepSeek V4 Pro
- The Decoder AI — Deepseek ships improved V4 Pro, open-sources its agent software, and raises API prices
- DeepSeek — deepseek-ai/deepseek-harness
- The Information — DeepSeek’s Flagship V4-Pro Model Gets Mixed Reviews
- DeepSeek — deepseek-ai/DeepSeek-V4-Pro
- Vals AI — DeepSeek V4
- Simon Willison — DeepSeek V4 Pro 0813 (on OpenRouter)
- DeepSeek — DeepSeek-V4-Pro GA Release
- DeepSeek — Your First API Call
- The Information — DeepSeek Announces Price Hikes for V4 Models
- Engadget — DeepSeek's AI models are about to cost four times more
- 腾讯云 — 〖大模型服务平台 TokenHub〗&〖智能体开发平台〗关于腾讯云 DeepSeek-V4 正式版〖原厂直供〗模型发布计划及计费调整通知
- DeepSeek API Changelog — DeepSeek-V4-Flash-Vision-Exp Release
- Terminal-Bench — How to run Terminal-Bench 2.1
- DeepSeek — Files API
- DeepSeek — DeepSeek-V4 Preview: Entering the Era of Affordable Million-Token Context
- Bloomberg Technology — DeepSeek Unveils Test Model to Rival Anthropic’s Opus 4.8
- DeepSeek — 更新日志
- DeepSeek — 模型 & 价格
- DeepSeek — 图像理解
- The Decoder AI — Deepseek releases experimental Flash vision model that rivals Opus 4.8 on agent benchmarks
- Bloomberg Technology — DeepSeek Ends Weekend Peak Pricing for API Users From Today
- DeepSeek — 模型 & 价格 | DeepSeek API Docs
- TechNode — DeepSeek to introduce peak and off-peak pricing for its API
- 中华网 — DeepSeek再度调价 周末统一低谷价