DeepSeek’s Prices Jump Up to 14x While Google Halves Its Own — The Great AI Price Divergence of 2026

Introduction: The Price War Didn’t End — It Split in Two

For two years, the story of AI pricing only went one direction: down. Open-weight labs like DeepSeek undercut everyone, frontier models matched the cuts, and “cheaper per token” became the industry’s default headline. This week, that story broke in half. On August 13, 2026, DeepSeek announced that its new V4-Pro lineup arrives with a complete pricing overhaul — including peak and off-peak rates that raise some line items by more than 10x starting August 16. On the very same day, Google shipped Gemini 3.7 Flash at an introductory price half of its three-week-old predecessor, and OpenAI previewed Ultrafast, a premium tier selling speed at a premium.

The result is the great AI price divergence of 2026: budget inference is getting more expensive as demand strains capacity, while hyperscaler workhorse models get cheaper fighting for developer volume. The era of uniformly falling prices is over — and your AI bill now depends on which side of the divide your tools sit on.

What DeepSeek Announced: V4-Pro Goes GA, Pricing Gets Rewritten

Inside the launch notes for DeepSeek-V4-Pro’s general availability release are the details that hit Hacker News’ front page with 200+ points on August 14. The model news matters on its own: major agent upgrades with “strong production gains,” a flexible reasoning-effort dial (low, high, or max) for both V4-Pro and V4-Flash, native OpenAI Responses API support optimized for Codex, a new “Expert Mode” in DeepSeek’s app and web interface, and a 1M-token context with up to 384K output.

Then comes the pricing section. DeepSeek is introducing peak and off-peak billing — off-peak rates are 50% lower than peak — with the new prices taking effect at 16:00 UTC on August 16, 2026. That gives existing API customers roughly three days to re-plan their budgets.

The New Numbers: Up to 12x on a Single Line Item

Here is the before-and-after, per million tokens, from DeepSeek’s own pricing page:

Rate (per 1M tokens)Old priceNew off-peakNew peak
V4-Flash input (cache miss)$0.14$0.22$0.44
V4-Flash output$0.28$0.66$1.32
V4-Pro input (cache miss)$0.435$0.66$1.32
V4-Pro output$0.87$1.98$3.96
V4-Pro input (cache hit)$0.003625$0.022$0.044

The headline multipliers depend on where you look. V4-Flash output tokens roughly quadruple at peak ($0.28 to $1.32, a 4.7x increase). V4-Pro output rises 4.6x at peak. But the steepest climb is cached input on V4-Pro: from $0.003625 to $0.044 — a 12x jump, which is why coverage of the change describes some prices rising “more than 10x.” Reuters framed the same numbers from the other side: the new V4-Pro peaks at up to 14 times the price of V4-Flash at its old rates.

Context check: DeepSeek is still inexpensive in absolute terms — $1.32 per million output tokens at peak remains far below premium frontier pricing. What changed is the trend, not the ranking. The cheapest major API just had its first broad price increase — and flexible scheduling is now the discount.

Why the Budget King Is Raising Prices

The simplest explanation is economic: demand is straining capacity. The company that built its brand on shockingly cheap inference now has more customers than compute at the hours everyone wants to use it. Peak/off-peak billing is the classic utility response — the same structure power companies use to flatten load curves. DeepSeek’s concurrency limits tell the same story: V4-Flash allows 2,500 concurrent requests, while the heftier V4-Pro is capped at 500.

There’s also a product-cycle explanation: V4-Pro is a bigger model aimed at production agent workflows — long contexts, many tool calls, sustained traffic — exactly the workload class that burns the most GPU-hours. Charging more for the model that costs the most to serve, while discounting off-peak hours to shift batch workloads, funds capacity without a flat across-the-board hike.

Same Week, Opposite Direction: Google Halves Gemini 3.7 Flash

While DeepSeek’s prices climbed, Google launched Gemini 3.7 Flash at an introductory $0.75/$3.75 per million input/output tokens — half the launch price of Gemini 3.6 Flash, which itself debuted only three weeks earlier. The model targets coding and agent workloads (65.3% on DeepSWE v1.1) and landed in GitHub Copilot almost immediately. The cut is labeled “introductory” and reported to run through year-end — a classic land-grab for developer volume, with a built-in re-price date.

The third pole is OpenAI’s Ultrafast preview, running GPT-5.6 Sol at up to 14x standard speed on Cerebras hardware — proof that speed is now a premium feature. Budget models are getting pricier, workhorse models cheaper, and frontier tiers charge for velocity: one market, three price directions at once.

Peak/Off-Peak Billing, Explained for AI Tool Buyers

DeepSeek’s peak windows are 01:00–04:00 and 06:00–10:00 UTC; every other hour is off-peak at half price. For teams in the Americas, those windows land in the evening and overnight — meaning many US after-hours batch jobs just got cheaper while interactive usage varies more. Practical implications:

How to Shop for AI Tools When Prices Diverge

Most people don’t buy raw API tokens — they buy tools built on them. But vendors pass these costs through, and the divergence shows up in your invoices as bigger price gaps between similar-looking products. A few rules for 2026:

Compare AI Tools by Real Cost

300+ AI tools with pricing, capability breakdowns, and categories — find the option that fits your budget before the next rate change. Curated on aitrove.ai.

Browse All AI Tools →

Frequently Asked Questions

Why did DeepSeek raise its API prices?

DeepSeek introduced peak/off-peak pricing alongside the V4-Pro launch, effective 16:00 UTC on August 16, 2026. Reporting on the change points to AI demand straining capacity: peak windows concentrate usage, and the new structure charges more at busy hours while discounting off-peak workloads by 50% to flatten demand.

How much more expensive does DeepSeek get?

It depends on the line item. V4-Flash output rises from $0.28 to $1.32 per million tokens at peak (4.7x); V4-Pro output rises from $0.87 to $3.96 (4.6x); and V4-Pro cached input rises from $0.003625 to $0.044 — just over 12x, the source of the “more than 10x” coverage. Off-peak rates are half the peak rates.

When are DeepSeek’s peak hours?

Peak windows are 01:00–04:00 UTC and 06:00–10:00 UTC daily. All other hours — including the 04:00–06:00 UTC gap between the two windows — are off-peak at 50% lower rates.

Is AI getting more expensive or cheaper in 2026?

Both, which is the point of the divergence. Budget-focused open-weight APIs like DeepSeek are raising rates as demand strains capacity, Google cut Gemini 3.7 Flash to half its predecessor’s launch price, and OpenAI is selling premium speed tiers. Costs per tool now depend heavily on which model tier and pricing structure sits underneath it.

What is DeepSeek-V4-Pro?

V4-Pro is DeepSeek’s flagship general-availability model released August 13, 2026, aimed at production agent workflows. It adds flexible reasoning-effort settings (low/high/max), native OpenAI Responses API support optimized for Codex, a 1M-token context with up to 384K output, Anthropic-format API compatibility, and an “Expert Mode” in the consumer app.

Stay Ahead of AI Pricing Shifts

Your trusted directory for AI agents, coding assistants, and the models powering them — with the cost context you need to choose well.

Explore the Directory →