AI Cost Intelligence, Verified Daily.
Skip to content
Model releases4 min readUpdated Aug 26, 2026

DeepSeek V4 Flash 0731 Pricing | Unit Cost and Benchmark Analysis

DeepSeek V4 Flash 0731 enters our database at $0.14 per 1M input tokens and $0.28 per 1M output tokens, combining a 1311k context window with built-in reasoning and vision.

DeepSeek V4 Flash cost chart in the AIOPLY house style
DeepSeek V4 Flash cost chart in the AIOPLY house style

Key takeaways

  • DeepSeek V4 Flash 0731 pricing starts at $0.14 per 1M input tokens and $0.28 per 1M output tokens, with a recorded release date of 2026-08-12.
  • Prompt caching reduces the input cost by 80 percent down to $0.028 per 1M tokens.
  • The model features a 1311k token context window, exceeding the 1050k token context window of GPT-5.6 Luna Pro.
  • DeepSeek V4 Flash 0731 is 88 percent cheaper on input and 92 percent cheaper on output than DeepSeek V4 Pro 0813.
  • At $0.28 per 1M output tokens, pricing sits 92 percent below the $3.90 median output rate across 35 tracked models.

What does DeepSeek V4 Flash 0731 cost to run in production?

DeepSeek V4 Flash 0731 costs $0.14 per 1M input tokens and $0.28 per 1M output tokens under standard rates. Recorded in our dataset on 2026-08-12, this baseline structure makes deepseek v4 flash 0731 pricing a competitive baseline for workloads that require Vision, Reasoning, and Long context capabilities.

When prompt caching is enabled, cached input drops to $0.028 per 1M tokens, which is an 80 percent discount compared to standard input. To calculate the operational cost per 1m tokens in a production setup, consider a system processing 500 million input tokens and 50 million output tokens each month without caching. The standard input cost equals $70.00 (500 times $0.14) and the output cost equals $14.00 (50 times $0.28), totaling $84.00 per month.

If 80 percent of those input tokens are cached, 100 million uncached input tokens cost $14.00 while the 400 million cached input tokens cost $11.20 (400 times $0.028). Combined with $14.00 for output, the total monthly API bill falls to $39.20. That yields a direct monthly saving of $44.80 through prompt design strategy alone.

How does DeepSeek V4 Flash 0731 compare against internal DeepSeek alternatives?

DeepSeek V4 Flash 0731 is significantly cheaper than both DeepSeek V4 Pro 0813 and DeepSeek V3.2 Exp across all token categories. DeepSeek V4 Pro 0813 costs $1.19 input and $3.56 output per 1M tokens, while DeepSeek V3.2 Exp costs $0.27 input and $0.41 output per 1M tokens.

Comparing deepseek v4 flash 0731 vs DeepSeek V4 Pro 0813 shows that Flash 0731 is 88.2 percent cheaper on standard input ($0.14 versus $1.19) and 92.1 percent cheaper on output ($0.28 versus $3.56). Against DeepSeek V3.2 Exp, Flash 0731 offers a 48.1 percent discount on input ($0.14 versus $0.27) and a 31.7 percent discount on output ($0.28 versus $0.41).

For an enterprise issuing 1 billion input tokens and 200 million output tokens per month, running DeepSeek V4 Pro 0813 costs $1,190.00 in input and $712.00 in output, totaling $1,902.00. The exact same volume on DeepSeek V4 Flash 0731 costs $140.00 in input and $56.00 in output, totaling $196.00. Switching to Flash 0731 cuts expenditure by $1,706.00 monthly within the same model provider.

ModelProvider / HostInput Rate (/1M)Cached Input (/1M)Output Rate (/1M)Context Window
DeepSeek V4 Flash 0731DeepSeek$0.14$0.028$0.281,311k
DeepSeek V4 Pro 0813DeepSeek$1.19N/A$3.56N/A
DeepSeek V3.2 ExpDeepSeek$0.27N/A$0.41N/A
Qwen3.7 FlashQwen$0.03N/A$0.13N/A
Nemotron 3.5 LightningNVIDIA$0.08N/A$0.20N/A
Qwen3 32BHugging Face$0.08N/A$0.28N/A
GPT-5.6 Luna ProOpenAIN/AN/AN/A1,050k
Pricing and context limits across selected AI models

Is DeepSeek V4 Flash 0731 cheaper than rival lightweight models?

DeepSeek V4 Flash 0731 is not the absolute lowest priced model in the budget tier, as Qwen3.7 Flash charges $0.03 input and $0.13 output per 1M tokens. However, Flash 0731 remains competitive against other lightweight options while offering a much larger context window and built-in vision.

When evaluating deepseek v4 flash 0731 vs other alternatives in our dataset, Nemotron 3.5 Lightning from NVIDIA charges $0.08 per 1M input tokens and $0.20 per 1M output tokens. Qwen3 32B hosted on Hugging Face charges $0.08 per 1M input tokens and matches Flash 0731 on output at $0.28 per 1M tokens.

On a monthly workload of 250 million input tokens and 25 million output tokens, standard uncached runs result in different billings across providers. Qwen3.7 Flash totals $10.75 ($7.50 input plus $3.25 output). Nemotron 3.5 Lightning totals $25.00 ($20.00 input plus $5.00 output). DeepSeek V4 Flash 0731 totals $42.00 ($35.00 input plus $7.00 output). However, if 80 percent of Flash 0731 inputs are cached, its input cost drops to $12.60 ($7.00 for 50M uncached plus $5.60 for 200M cached), bringing the total monthly bill down to $19.60 and beating Nemotron 3.5 Lightning.

How does the 1311k context window alter pipeline unit economics?

The 1311k token context window on DeepSeek V4 Flash 0731 permits extensive document analysis without context chunking overhead. This capacity exceeds the 1050k context window of GPT-5.6 Luna Pro, which was previously the widest context window recorded in our dataset.

In addition to its 1311k context window, DeepSeek V4 Flash 0731 supports a maximum output of 393k tokens. This maximum output length allows applications to generate comprehensive reports, long software files, or detailed multi-step reasoning outputs without running into generation truncations.

Evaluating cost per 1m tokens for long context processing demonstrates clear efficiency. Processing a massive 1 million token input prompt once costs $0.14. If that context remains static and benefits from the $0.028 cached input rate, every subsequent call on that prompt costs less than three cents for input context, making large retrieval setups economically viable.

Which workloads are worth paying for on DeepSeek V4 Flash 0731?

Workloads requiring multimodal vision processing, long context ingestion, and structured reasoning are fully worth paying for on DeepSeek V4 Flash 0731. The recorded dataset capabilities list Vision, Reasoning, Long context, and Cheap for this model.

Conversely, high volume pipelines handling short text transformation or simple data extraction are not worth running on Flash 0731. Paying $0.14 per 1M standard input tokens for simple tasks is inefficient when Qwen3.7 Flash handles input at $0.03 per 1M tokens without performance bottlenecks.

  • Worth paying for: Document parsing and visual data extraction where vision capabilities and long context are mandatory.
  • Worth paying for: Repository scale code reviews that utilize context sizes near the 1311k token window limit.
  • Worth paying for: Multi-step technical reasoning pipelines that leverage cached system prompts at $0.028 per 1M tokens.
  • Not worth paying for: Basic sentiment analysis or short text classification where Qwen3.7 Flash offers $0.03 per 1M input tokens.
  • Not worth paying for: Static un-cached short prompts with zero context reuse, where alternative providers offer lower baseline standard input rates.

What is the market positioning of DeepSeek api pricing across our dataset?

Current deepseek api pricing positions Flash 0731 well below broader market averages, particularly for output generation. Across 35 tracked current models in our dataset, the median output rate is $3.90 per 1M tokens.

At $0.28 per 1M output tokens, deepseek v4 flash 0731 pricing sits 92.8 percent below the market median output price. This places it alongside budget models while delivering capabilities usually restricted to higher tiers.

Incorporating deepseek v4 flash 0731 api cost into dynamic routing architectures allows development teams to offload long context vision and light reasoning tasks from higher priced tiers like DeepSeek V4 Pro 0813, lowering global API spend across production stacks.

Frequently asked questions

What is the standard deepseek v4 flash 0731 api cost?
DeepSeek V4 Flash 0731 pricing is $0.14 per 1M input tokens and $0.28 per 1M output tokens. Prompt caching lowers the input rate by 80 percent to $0.028 per 1M tokens. Released on 2026-08-12, the model provides a 1311k token context window and maximum output capacity of 393k tokens.
How does DeepSeek V4 Flash 0731 compare to DeepSeek V4 Pro 0813 on price?
DeepSeek V4 Flash 0731 is significantly cheaper than DeepSeek V4 Pro 0813. Flash 0731 charges $0.14 input and $0.28 output per 1M tokens, whereas Pro 0813 charges $1.19 input and $3.56 output per 1M tokens. Flash 0731 is 88 percent cheaper on input and 92 percent cheaper on output.
Is DeepSeek V4 Flash 0731 the cheapest model available?
No, DeepSeek V4 Flash 0731 is not the cheapest model overall. Qwen3.7 Flash offers lower standard rates at $0.03 per 1M input tokens and $0.13 per 1M output tokens. Nemotron 3.5 Lightning also charges lower standard input at $0.08 per 1M tokens, though Flash 0731 features prompt caching and a larger context window.
What is the context window size for DeepSeek V4 Flash 0731?
DeepSeek V4 Flash 0731 features a 1311k token context window and a maximum output capacity of 393k tokens. This context window exceeds the 1050k context window of GPT-5.6 Luna Pro. At $0.028 per 1M cached input tokens, processing large contexts repeatedly is cost effective.

About the author

AIOPLY Pricing Desk

Independent AI cost research, verified against live provider rates

How we work

Quality standards

Prices were checked against provider documentation on Aug 26, 2026. Rates change without notice, so confirm current figures with the provider before committing a budget. We publish list prices only and take no payment for placement.

Report a correction

Sources and further reading

More from the blog