Skip to Content

DeepSeek V4 Pro Pricing Breakdown

Current API Prices, Peak Hours, Benchmarks, and Migration
2026-05-17 19:52:21 Updated 2026-08-20 23:52:59.601418 — min read 786 views
DeepSeek V4 Pro Pricing Breakdown
“The DeepSeek V4 Pro pricing picture changed in August 2026. The current official API page lists the V4-Pro-0813 model, peak and off-peak rates, a 1M context length, Responses API support, and a 500 concurrency limit. The May discount is historical, not the current price.

What You'll Learn

  • The current official DeepSeek V4 Pro API prices and how peak and off-peak billing works.
  • What changed between the April preview, the August general release, and the expired May promotion.
  • Which agent benchmarks and API features DeepSeek reports for V4 Pro.
  • How to compare cache status, token mix, provider margin, latency, and concurrency before choosing a route.

DeepSeek's official August 13, 2026 GA release changes the way this article should be read. V4 Pro is now available on the app, web, and API, and the API model name remains deepseek-v4-pro. The pricing page lists the current V4-Pro-0813 version, a 1M context length, and peak and off-peak rates that took effect at 16:00 UTC on August 16, 2026.

The original article was written around a May 31, 2026 promotion. It described $0.435 per million input tokens, $0.87 per million output tokens, and $0.0036 cache-hit input as if those numbers were permanent. They are not the current official API rates. A reliable DeepSeek V4 Pro pricing guide must show the date, model version, billing window, cache status, and whether a number comes from the official API or from a separate provider.

For broader context on how model pricing affects agent infrastructure, see our AI agent ROI guide. The cost of a model call is only one part of a production workflow that may also include retrieval, tools, storage, retries, and human review.

DeepSeek V4 Pro Current Status

DeepSeek V4 Pro is now a general release model in the DeepSeek app, web product, and API. The August 13 announcement says the GA release brings stronger agent capabilities, flexible reasoning effort, native OpenAI Responses API support, and a Codex-oriented setup. The change log records the same release as a V4 Pro update and says that users can keep the API model name as deepseek-v4-pro.

This status is different from the April preview and the May promotional period. The April 24 update introduced V4-Pro and V4-Flash API names, while the August release moved V4 Pro to GA. Developers maintaining an integration should confirm whether they want the current rolling model name or a versioned route, then record the behavior they tested. A model name that remains unchanged can still point to a newer serving version.

The current official pricing page also lists an OpenAI-format base URL at https://api.deepseek.com and an Anthropic-format base URL at https://api.deepseek.com/anthropic. The page says V4 Pro supports JSON Output, Tool Calls, the Responses API, the Anthropic API, and Chat Prefix Completion beta. Those capabilities make it a practical candidate for existing integrations, but compatibility should be tested rather than assumed.

Current fieldOfficial DeepSeek valueWhy it matters
Model versionDeepSeek-V4-Pro-0813Prevents May preview figures being treated as current
API model namedeepseek-v4-proExisting API calls can target the current model name
Context length1MSets the maximum input context listed on the pricing page
Concurrency limit500Defines a listed account-level capacity boundary

Current Official API Pricing

The current official pricing page lists prices per 1M tokens and bills on the total number of input and output tokens. V4 Pro has two input states. A cache hit is cheaper than a cache miss, and both states have separate peak and off-peak prices. Output tokens also have peak and off-peak prices. This is the pricing information a developer should use for a new estimate as of August 21, 2026.

V4 Pro chargeOff-peak price per 1M tokensPeak price per 1M tokens
Input cache hit$0.022$0.044
Input cache miss$0.66$1.32
Output$1.98$3.96
Context length1M1M

The official pricing page says the expense equals the number of tokens multiplied by the applicable price. It also says DeepSeek may adjust product prices and recommends checking the live page regularly. That warning is especially relevant here because the price schedule changed in the same month as the general release.

A simple example shows why the input state matters. A prompt with a cache hit and a prompt with a cache miss do not have the same input cost even when their token counts are identical. An application that repeats a long system instruction may benefit from cache behavior, while a first-time request or a changed prefix may be billed at the cache-miss rate. The billing record, not a headline rate, is the correct source for a production estimate.

The comparison with our OpenAI o3 Mini and o1 cost guide should also be made with matching assumptions. Compare the same input length, output length, reasoning mode, retry policy, and currency. A lower per-token price can still produce a higher workflow cost if a model needs more tokens or more retries.

How Token Billing Works

DeepSeek bills by token rather than by document or request. The official page defines the expense as the number of tokens multiplied by the applicable price. That makes input and output mix important. Two requests with the same document can have different totals if one generates a longer answer or if one qualifies for a cache hit.

Estimate a route with separate input, cache-hit, cache-miss, and output lines. Then add expected retries, tool calls, and any provider-specific charges. This approach is more reliable than quoting one average price because the average can hide a workload that is mostly cache misses or long outputs.

Peak and Off-Peak Billing Windows

DeepSeek defines peak hours as 01:00 to 04:00 UTC and 06:00 to 10:00 UTC. All other hours are off-peak. The official page says off-peak rates are half of peak rates. For a global team, this creates a scheduling variable. Batch work that is not latency-sensitive may be cheaper when it runs outside the peak windows, while interactive requests may need to accept the peak rate.

Scheduling should not be treated as a free optimization. Moving a job to a different hour can affect freshness, staffing, downstream dependencies, and incident response. A finance or compliance workflow may need a fixed completion window. A coding agent may have a user waiting for a response. Record both the token price and the service-level requirement when comparing the savings.

Provider or account terms can also affect the result. The official DeepSeek page lists a concurrency limit of 500 for V4 Pro. A workload that sends hundreds of simultaneous requests may encounter queuing or rate limits even when the average token price is attractive. Test the actual account with the expected request pattern and observe errors, latency, and queue behavior.

Model Features and 1M Context

The official pricing page lists a 1M context length for V4 Pro and supports both non-thinking and thinking modes. The August GA update describes low, high, and max thinking effort levels. The practical meaning is that a request can trade reasoning effort and possibly latency or token use against task complexity. The setting is not a guarantee of a particular answer quality.

V4 Pro's API feature list includes JSON Output, Tool Calls, Responses API support, Anthropic API support, and Chat Prefix Completion beta. DeepSeek's GA announcement specifically highlights native Responses API support and adaptation for Codex. Teams using an OpenAI-compatible client should still test the exact request and response shape, because a shared format does not guarantee identical error messages, tool schemas, or streaming behavior.

A 1M context limit is a capacity boundary rather than a claim that every prompt should use 1M tokens. Long prompts can increase latency, input cost, and the amount of irrelevant information the model must process. Retrieval, permissions, source citations, and document versioning remain useful even when the model can accept a large context.

For a related systems perspective, our vector database comparison explains why context management and access control remain important around any long-context model. A larger window does not remove the need to decide which data a user may see.

Agent Benchmarks in the GA Release

DeepSeek's August change log reports a set of agent-focused benchmark results for the V4 Pro GA update. It lists HLE without tools at 42.7 and with tools at 60.0. It lists Terminal Bench 2.1 at 87.9, NL2Repo at 61.5, Cybergym at 83.3, DeepSWE at 62.7, Toolathlon-Verified at 74.1, Agents' Last Exam at 25.7, AutomationBench Public at 31.8, DSBench-FullStack at 71.1, and DSBench-Hard at 67.2.

These numbers are useful as a current vendor-published release snapshot. They are not a neutral ranking of every frontier model, because the change log does not by itself provide the full harness, sampling settings, comparison set, or independent replication for each result. Benchmark names can also change versions. A responsible buyer should treat them as evidence to investigate and then run a workload test with the team's own prompts and tools.

BenchmarkReported V4 Pro GA resultInterpretation
HLE42.7 without tools, 60.0 with toolsTool access materially changes the measured task
Terminal Bench 2.187.9Terminal-agent result in the release note
NL2Repo61.5Repository-level task result
AutomationBench Public31.8Agent workflow result with a named benchmark scope

The older article used 80.6% on SWE-bench and 90.1% on GPQA as if they were current V4 Pro proof points. Those figures were not adequately sourced in the persisted article, and the official August update uses a different set of agent benchmarks. They should not be carried forward without a model card or a reproducible evaluation record.

API Model Names and Migration

The April API release used deepseek-v4-pro and deepseek-v4-flash. The August GA announcement says V4 Pro's API model name remains unchanged. This is convenient for clients, but it creates a versioning responsibility for teams. Store the date, response headers, prompts, settings, and model name in the evaluation record so that a later change can be detected.

The change log says the legacy deepseek-chat and deepseek-reasoner names were scheduled for discontinuation after July 24, 2026. On August 21, 2026, new integrations should use the current V4 model names and check the live migration documentation. Old code may still run through compatibility behavior, but relying on it can make future changes harder to diagnose.

Native Responses API support and Codex adaptation can simplify an agent migration, but the team should test tool calls with real schemas. Check whether reasoning effort settings are accepted, how partial outputs are streamed, how errors are returned, and whether a failed tool call consumes input or output tokens. Include these results in the cost model.

For comparison with another model-routing decision, see our GPT-5.5 and Grok comparison. The point is not to choose a universal winner. It is to compare models under the actual quality, speed, and reliability requirements of a route.

What the May 2026 Promotion Means Now

The old article said V4 Pro was 75% off until May 31, 2026 and recommended locking in the $0.435 input and $0.87 output rates. That deadline has passed. The historical rates can be useful for understanding the original launch economics, but they should not appear in the current summary as a live offer.

The current official schedule is materially different. V4-Pro-0813 uses cache-hit, cache-miss, and output prices that vary by peak and off-peak hours. A reader comparing an old screenshot with the current pricing page may see why the article needs a date-stamped table. The price of a model is not a fixed attribute when the provider changes versions or billing windows.

Third-party providers may still list their own V4 routes and rates. Those rates are not interchangeable with the official DeepSeek API. A provider may add infrastructure margin, expose a different model version, offer different concurrency, or apply a separate cache policy. Label the provider, model identifier, date, and billing unit whenever a comparison is published.

Pricing periodInput exampleHow to label it
May 2026 promotion$0.435 cache miss and $0.87 output in the old articleHistorical promotional rates, not current official prices
Current off-peak$0.66 cache miss and $1.98 outputOfficial V4-Pro-0813 API rate per 1M tokens
Current peak$1.32 cache miss and $3.96 outputOfficial V4-Pro-0813 API rate per 1M tokens
Third-party routeProvider-specificMust name the provider and model version

How to Compare Providers and Workloads

A fair provider comparison begins with identity. Write down the provider, exact model string, context limit, region, API format, price schedule, cache policy, concurrency limit, and data-retention terms. Do not rank Fireworks, DeepInfra, Together.ai, or any other route from a single speed figure unless the test conditions are documented.

Next, define the workload. Measure a short interactive prompt, a long document request, a tool-using agent turn, and a batch job. Record time to first token, total latency, output tokens, input cache state, failures, retries, and queue time. A fast provider can be more expensive or less reliable for the workload that matters to the buyer.

Use the same prompt and response requirements across providers. If one route uses a different system prompt, reasoning effort, tool set, or output format, the comparison is not like for like. Keep a record of the date because the official DeepSeek page says product prices may change.

Cost Optimization Without Misleading Claims

Cache-hit pricing is valuable only when a request actually qualifies for a cache hit. Keep stable prefixes stable when the provider's caching rules support that pattern, but do not make a cache assumption without reading the current documentation. A prompt with frequent changes may be billed as a miss. A cached prefix may also need invalidation when instructions or permissions change.

Peak scheduling can help batch workloads, but it should not become the only optimization. Reduce unnecessary context, avoid repeated tool calls, cap output length where appropriate, and route simple tasks to a smaller model. Keep citations and access controls intact. The cheapest request is not useful if it produces an incorrect answer that requires human rework.

Cost reporting should show at least four quantities: input tokens, cache-hit input tokens, cache-miss input tokens, and output tokens. Add retries and tool calls as separate lines. This makes the estimate auditable and prevents a low cache-hit rate from being hidden inside an average price.

For teams evaluating local alternatives, our Mac M4 Max local LLM benchmark provides a useful comparison frame. Local inference has its own hardware, electricity, maintenance, and quality costs, so it should be measured with the same workload discipline.

DeepSeek V4 Pro Adoption Checklist

Before moving a production route to V4 Pro, request the current service terms and pricing page. Confirm whether the account sees V4-Pro-0813, which thinking modes are enabled, and whether the 1M context limit is available for the selected endpoint. Confirm whether JSON Output, Tool Calls, Responses API, and Anthropic API support behave as required by the application.

  1. Record the exact model name, date, endpoint, and account region.
  2. Measure quality on representative prompts with a fixed answer key.
  3. Test tool calls, structured outputs, streaming, retries, timeouts, and partial failures.
  4. Measure cache-hit and cache-miss cases separately.
  5. Run the same workload during peak and off-peak windows when scheduling is possible.
  6. Check concurrency, queueing, privacy, retention, and deletion terms.
  7. Set a price-change alert and review the official pricing page before a large batch.

A migration is complete only when the quality and operational evidence is saved. Keep sample outputs, error logs, token counts, latency measurements, and the exact settings used. This record makes it possible to distinguish a model change from a prompt change or a provider incident.

Conclusion: Use Current Prices and Reproducible Tests

DeepSeek V4 Pro remains a significant low-cost model option, but the article's old May promotion is no longer the correct headline. As of August 21, 2026, the official API page lists V4-Pro-0813 with a 1M context, peak and off-peak pricing, cache-hit and cache-miss input rates, a 500 concurrency limit, and support for modern API features.

The August GA release also changes the capability discussion. DeepSeek reports stronger agent results, three thinking-effort levels, native Responses API support, and Codex adaptation. Those results are useful release evidence, not a universal guarantee. Compare them with your own workload and retain the benchmark version and settings.

The practical rule is simple. Do not use the expired $0.435 and $0.87 promotion as current pricing. Do not copy a third-party provider's rate into the official API table. Do not call a model the best value without defining quality, speed, reliability, and total workflow cost. Check the live official page, measure the route, and update the estimate when the provider changes the model or schedule.

Frequently Asked Questions

The official pricing page lists V4-Pro-0813 cache-hit input at $0.022 per million tokens off-peak and $0.044 peak, cache-miss input at $0.66 off-peak and $1.32 peak, and output at $1.98 off-peak and $3.96 peak.
DeepSeek defines peak hours as 01:00 to 04:00 UTC and 06:00 to 10:00 UTC. All other hours are off-peak, and the official page says off-peak rates are half the peak rates.
The current official pricing page lists the model version as DeepSeek-V4-Pro-0813. The API model name remains deepseek-v4-pro according to DeepSeek's August 2026 GA announcement.
Yes. The current official pricing page lists a 1M context length for V4 Pro. A large context limit does not guarantee lower latency or correct reasoning, so test a representative workload before relying on the full window.
No. The old article's 75% promotion ended on May 31, 2026. The current official schedule is the V4-Pro-0813 peak and off-peak price table, with separate cache-hit and cache-miss rates.
The current official page lists JSON Output, Tool Calls, the Responses API, the Anthropic API, and Chat Prefix Completion beta. The GA release also highlights native Responses API support and Codex adaptation.
Use cache hits when the request qualifies, schedule non-urgent batch work during off-peak hours, monitor input and output tokens separately, and measure retries and tool calls. Always confirm the current pricing page because DeepSeek says product prices may change.
SK Jabedul Haque
Written by

SK Jabedul Haque

Founder & Chief Editor

Building India's most trusted finance education platform — simplifying news, schemes and market trends so anyone can understand and invest confidently.

Read full bio

Never miss an update

Get our clearest explainers on schemes, markets and money — read what matters, without the noise.

Explore more articles
In this article