Key takeaways
- DeepSeek V4 Pro 0813 is listed by DeepSeek at $0.435 per 1M input tokens and $0.87 per 1M output tokens, with a release date of 2026-08-12.
- Prompt caching reduces input costs by 99 percent down to $0.004 per 1M tokens for cached context.
- The model features a 1049k token context window and a maximum output generation limit of 384k tokens.
- At $0.87 per 1M tokens, output pricing sits 82.6 percent below our tracked dataset median output rate of $5.00.
- DeepSeek V4 Pro 0813 costs more than triple DeepSeek V4 Flash 0731, which charges $0.14 input and $0.28 output per 1M tokens.
How much does DeepSeek V4 Pro 0813 cost in production?
On 2026-08-12, DeepSeek added V4 Pro 0813 to its API catalog at $0.435 per 1M input tokens and $0.87 per 1M output tokens. Understanding deepseek v4 pro 0813 pricing requires evaluating standard API charges alongside the 99 percent discount available for cached prompt inputs, which cost $0.004 per 1M tokens.
The model carries recorded dataset capabilities for open weights and long context processing. It features a context window of 1049k tokens and a maximum output limit of 384k tokens. That context window capacity sits directly behind GPT-5.6 Luna Pro, which holds the widest context in our current dataset at 1050k tokens.
When reviewing deepseek v4 pro 0813 pricing against broader market benchmarks, the output pricing is notably low. Across 34 tracked current models in our database, the median output rate is $5.00 per 1M tokens. At $0.87 per 1M output tokens, DeepSeek V4 Pro 0813 operates 82.6 percent below the dataset output median, making large generation volumes unusually affordable for a flagship option.
How does DeepSeek V4 Pro 0813 compare to alternative options?
DeepSeek V4 Pro 0813 costs significantly more than lightweight alternatives while maintaining lower rates than classical enterprise models. Compared to DeepSeek V3.2 Exp at $0.27 input and $0.41 output per 1M tokens, V4 Pro 0813 carries a 61.1 percent higher input rate and a 112.2 percent higher output rate.
Examining deepseek v4 pro 0813 vs earlier lightweight models highlights a stark price tier gap. DeepSeek V4 Flash 0731 costs $0.14 for input and $0.28 for output per 1M tokens, meaning the Pro edition costs over three times as much per token. When compared against external provider flash models, the difference widens further.
External alternatives tracked in our database offer lower entry pricing for lightweight tasks. Qwen3.7 Flash from Qwen charges $0.03 for input and $0.13 for output per 1M tokens. Nemotron 3.5 Lightning from NVIDIA charges $0.10 for input and $0.25 for output per 1M tokens. Buyers analyzing deepseek api pricing must evaluate whether long context capabilities justify paying these premiums over flash options.
| Model | Provider | Standard Input (1M) | Cached Input (1M) | Output Rate (1M) | Context Window |
|---|---|---|---|---|---|
| DeepSeek V4 Pro 0813 | DeepSeek | $0.435 | $0.004 | $0.870 | 1049k tokens |
| DeepSeek V3.2 Exp | DeepSeek | $0.270 | Not Tracked | $0.410 | Not Tracked |
| DeepSeek V4 Flash 0731 | DeepSeek | $0.140 | Not Tracked | $0.280 | Not Tracked |
| Nemotron 3.5 Lightning | NVIDIA | $0.100 | Not Tracked | $0.250 | Not Tracked |
| Qwen3.7 Flash | Qwen | $0.030 | Not Tracked | $0.130 | Not Tracked |
| GPT-5.6 Luna Pro | Not Tracked | Not Tracked | Not Tracked | Not Tracked | 1050k tokens |
What impact does prompt caching have on context cost?
Prompt caching reduces the cost per 1m tokens for repeated inputs from $0.435 to $0.004 per 1M tokens. This 99 percent reduction allows engineering teams to populate the 1049k token context window repeatedly without incurring crippling input bills.
To put this reduction in concrete financial terms, loading a full 1,000,000 token context buffer into DeepSeek V4 Pro 0813 costs $0.435 on the initial uncached call. If an application executes 100 queries against that identical cached document structure, the input cost calculation changes dramatically.
Under standard uncached rates, 100 calls with 1M input tokens each would cost $43.50 (100 multiplied by $0.435). With prompt caching enabled, the initial load costs $0.435, while the remaining 99 calls cost $0.004 each (totaling $0.396). The combined input bill for all 100 queries drops to $0.831, saving $42.669 per batch run.
Which workloads are worth paying for with DeepSeek V4 Pro 0813?
DeepSeek V4 Pro 0813 is worth paying for when workloads require massive token output, open weights deployment, or sustained context reuse. Output tasks benefit directly from the $0.87 per 1M token rate, which sits far below the $5.00 median output cost.
Consider an engineering pipeline generating massive technical documentation batches totaling 500 million output tokens per month alongside 100 million input tokens (with 80 percent prompt cache reuse). At standard rates, 20 million uncached input tokens cost $8.70, 80 million cached input tokens cost $0.32, and 500 million output tokens cost $435.00, resulting in a total monthly spend of $444.02.
If that same 500 million token output volume were run on a model charging the $5.00 market median, output costs alone would total $2,500.00 per month. The low output cost per 1m tokens makes DeepSeek V4 Pro 0813 an exceptional fit for heavy text generation, long form translation, software codebase synthesis, and recursive reasoning pipelines that emit lengthy response streams.
When is DeepSeek V4 Pro 0813 not worth paying for?
DeepSeek V4 Pro 0813 is not worth paying for on high volume, short context requests where lightweight models handle identical tasks for a fraction of the cost. Tasks like intent classification, simple entity extraction, and conversational routing do not benefit from a 1049k context window or 384k max output capability.
For example, a customer service routing system processing 10 million requests per month (averaging 1,000 input tokens and 200 output tokens per request) generates 10 billion input tokens (10,000 million) and 2 billion output tokens (2,000 million). On DeepSeek V4 Pro 0813 without caching, 10,000 million input tokens cost $4,350.00 (10,000 x $0.435) and 2,000 million output tokens cost $1,740.00 (2,000 x $0.87), creating a monthly bill of $6,090.00.
Running that exact workload on Qwen3.7 Flash ($0.03 input, $0.13 output) costs $300.00 for input (10,000 x $0.03) and $260.00 for output (2,000 x $0.13), totaling $560.00 per month. Paying a $5,530 monthly premium for deepseek v4 pro 0813 pricing on simple classification is an inefficient allocation of capital. Similarly, DeepSeek V4 Flash 0731 would cost $1,400 input (10,000 x $0.14) and $560 output (2,000 x $0.28), totaling $1,960 per month.
How does monthly API spend scale across standard production tiers?
Monthly spend scales predictably based on input reuse and output generation ratios. Calculating deepseek v4 pro 0813 api cost across standardized monthly tiers shows how cached context radically alters operational budgets.
Scenario A represents a long document processing deployment running 50 million input tokens (80 percent cached) and 5 million output tokens monthly. The 10 million uncached input tokens cost $4.35 (10 x $0.435), the 40 million cached input tokens cost $0.16 (40 x $0.004), and the 5 million output tokens cost $4.35 (5 x $0.87), resulting in an $8.86 monthly total.
Scenario B represents un-cached heavy output generation with 100 million standard input tokens and 100 million output tokens monthly. Input charges equal $43.50 (100 x $0.435) and output charges equal $87.00 (100 x $0.87), producing a monthly bill of $130.50.
The primary driver of deepseek v4 pro 0813 pricing efficiency is prompt caching. Scenario C models an enterprise context repository processing 1 billion input tokens (95 percent cached) and 50 million output tokens. Uncached input (50 million tokens) costs $21.75, cached input (950 million tokens) costs $3.80, and output (50 million tokens) costs $43.50, bringing total monthly spend to $69.05.
How do open weights impact the total cost of ownership?
Open weights availability allows organizations to host DeepSeek V4 Pro 0813 on private infrastructure rather than relying exclusively on managed API endpoints. In our database, open weights is recorded as a core capability alongside long context support for this model release.
When making architectural decisions, direct API deployment at $0.435 input and $0.87 output per 1M tokens provides immediate scalability without capital expenditure on GPU infrastructure. Hosted API pricing eliminates idle compute costs, power expenses, and cluster management overhead.
For workloads with unpredictable traffic spikes, API usage remains highly cost effective due to the low $0.87 output rate. However, enterprises requiring strict data privacy or custom hosting can utilize the open weights option to run instance deployments where regulations prevent managed API usage.
What is the financial threshold for choosing Pro over Flash?
The decision threshold between DeepSeek V4 Pro 0813 and DeepSeek V4 Flash 0731 depends on context length requirements and output volume intensity. Flash charges $0.14 input and $0.28 output per 1M tokens, while Pro charges $0.435 input and $0.87 output.
Pro becomes financially justified when a workload requires context sizes exceeding standard flash limits up to the 1049k window, or requires generation lengths up to 384k output tokens. If a system requires deep context processing across massive document collections, Flash cannot perform the task regardless of its lower unit rate.
Conversely, for requests that stay within standard prompt limits, Flash provides a 67.8 percent discount on input ($0.14 vs $0.435) and a 67.8 percent discount on output ($0.28 vs $0.87). Engineering teams should route high volume short queries to Flash or Qwen3.7 Flash ($0.03 input) while reserving DeepSeek V4 Pro 0813 for heavy context synthesis.
Frequently asked questions
- What is the input and output token cost for DeepSeek V4 Pro 0813?
- DeepSeek V4 Pro 0813 costs $0.435 per 1M input tokens and $0.87 per 1M output tokens on standard requests. When prompt caching is utilized, the cached input rate drops by 99 percent to $0.004 per 1M tokens.
- How wide is the context window on DeepSeek V4 Pro 0813?
- DeepSeek V4 Pro 0813 features a context window of 1049k tokens, which is second only to GPT-5.6 Luna Pro at 1050k tokens in our current dataset. It also supports a maximum output generation length of 384k tokens.
- Is DeepSeek V4 Pro 0813 open source or open weights?
- DeepSeek V4 Pro 0813 is recorded in our dataset as an open weights model with long context capabilities. This allows organizations the option to evaluate open architecture parameters alongside direct API access provided by DeepSeek.
- How does DeepSeek V4 Pro 0813 compare to DeepSeek V4 Flash 0731?
- DeepSeek V4 Pro 0813 is significantly more expensive than DeepSeek V4 Flash 0731. Standard input costs $0.435 compared to $0.14 for Flash, while output costs $0.87 compared to $0.28 for Flash. Flash is better suited for low context, high volume routing tasks.
- When was DeepSeek V4 Pro 0813 added to the pricing feed?
- The verified release date on record for DeepSeek V4 Pro 0813 is 2026-08-12. Our pricing intelligence system tracked its initial API rates immediately upon addition to the live feed.
About the author
AIOPLY Pricing Desk
Independent AI cost research, verified against live provider rates
How we workQuality standards
Prices were checked against provider documentation on Aug 15, 2026. Rates change without notice, so confirm current figures with the provider before committing a budget. We publish list prices only and take no payment for placement.
Report a correctionSources and further reading
More from the blog
Pricing updates
Grok API Pricing Explained: What xAI Models Actually Cost to RunA working breakdown of Grok API pricing per million tokens, including cached input, long-context billing, tool calls, and how to forecast a monthly bill.
Pricing updates
Claude Opus API Pricing: How Anthropic Bills Its Most Capable ModelA working breakdown of Claude Opus API pricing per million tokens, including cached input, batch discounts, extended thinking, and how to forecast a monthly Anthropic bill.
Comparisons
Grok vs Claude: API Pricing, Cost per Task and When Each One WinsA cost-first comparison of Grok and Claude API pricing, with worked examples, break-even token volumes, and clear guidance on which model fits which workload.
