Why it matters
  • Lead. DeepSeek released V4-Flash-0731 as a generally available model on July 31, setting a list price of $0.14 per million cache-miss input tokens and $0.28 per million output tokens — roughly 9.5 times cheaper than its V4 Pro sibling and positioning it as one of the most cost-competitive frontier-class models on the market.
  • Fact. On the Artificial Analysis Intelligence Index v4.0, V4-Flash scores 47 against Pro’s 52, a five-point gap, but outperforms Pro on Terminal Bench — an agentic coding evaluation — by 10 points, scoring 82.7 against Pro’s 72.7.
  • Stake. The release intensifies a pricing war among Chinese AI labs that is compressing Western providers’ revenue per token and pushing buyers to reassess whether frontier-level performance requires frontier-level pricing.

DeepSeek launched its V4 family on April 24, 2026 as a two-model architecture: V4 Pro at 1.6 trillion total parameters with 49 billion active, and V4 Flash at 284 billion total with just 13 billion active via mixture-of-experts routing. Both carry a one-million-token context window. Flash exited its preview period on July 31 as DeepSeek-V4-Flash-0731, signalling stability sufficient for production workloads. Cache hits are billed at $0.0028 per million tokens, making it substantially cheaper still for applications that repeatedly reference fixed context.

Where Flash Beats Pro

The Terminal Bench result — where Flash scores 82.7 to Pro’s estimated 72.7 — is notable because coding benchmarks are typically where larger, more capable models hold their clearest lead. The inversion suggests Flash’s post-training optimisations are particularly well-suited to the structured, step-by-step reasoning that software engineering tasks demand, even if its general intelligence index score lags. At a cost of approximately $0.03 per benchmark test, it becomes economically viable to run large parallel evaluation sweeps, an advantage for automated testing pipelines and agent frameworks that run thousands of inference calls per session.

The release lands alongside competing moves from other Chinese labs. Alibaba’s Qwen3.8-Max, a 2.4-trillion parameter multimodal model, is priced at $2 per million input tokens and $6 per million output tokens — nearly 15 times the Flash input cost. That gap puts Flash in a different market tier, targeting high-volume, latency-tolerant workloads rather than the reasoning-intensive tasks where Qwen3.8-Max competes. DeepSeek has been building toward hardware independence in parallel; the lab is developing a proprietary inference chip that could further reduce its operational costs once it reaches production.

What It Means for the Broader Market

Flash’s general availability arrives four days after the EU AI Act’s Article 50 transparency requirements took effect for all AI systems placed in the European market. Those rules require disclosure of AI-generated media and machine-interaction status, but do not restrict pricing or deployment of models themselves. For Western API providers, the more immediate concern is commercial: a Chinese lab offering near-frontier performance at sub-$0.30 output costs compresses the viable price band for comparable closed-source alternatives, particularly in cost-sensitive enterprise segments where inference budgets are a material line item.