Lowest input rate
Qwen3.7 Flash
$0.030 / 1M in
Pricing database
Every model we track, priced in USD per 1M tokens, with context limits, cached input rates, and the date each record was last verified. Public list rates only, no batch or enterprise discounts.
Lowest input rate
Qwen3.7 Flash
$0.030 / 1M in
Widest context window
Grok 4 Fast
2M tokens
Showing 43 of 43 models.
Qwen · Verified Aug 28, 2026
OpenAI · Verified Aug 28, 2026
DeepSeek · Verified Aug 28, 2026
Hugging Face · Verified Aug 28, 2026
NVIDIA · Verified Aug 28, 2026
Qwen · Verified Aug 28, 2026
OpenAI · Verified Aug 28, 2026
OpenAI · Verified Aug 28, 2026
xAI · Verified Aug 28, 2026
Google · Verified Aug 28, 2026
OpenAI · Verified Aug 28, 2026
DeepSeek · Verified Aug 28, 2026
Meta · Verified Aug 28, 2026
Google · Verified Aug 28, 2026
Google · Verified Aug 28, 2026
Google · Verified Aug 28, 2026
Qwen · Verified Aug 28, 2026
Qwen · Verified Aug 28, 2026
NVIDIA · Verified Aug 28, 2026
DeepSeek · Verified Aug 28, 2026
Google · Verified Aug 28, 2026
Qwen · Verified Aug 28, 2026
Anthropic · Verified Aug 28, 2026
Perplexity · Verified Aug 28, 2026
OpenAI · Verified Aug 28, 2026
Anthropic · Verified Aug 28, 2026
Google · Verified Aug 28, 2026
Mistral · Verified Aug 28, 2026
OpenAI · Verified Aug 28, 2026
OpenAI · Verified Aug 28, 2026
OpenAI · Verified Aug 28, 2026
OpenAI · Verified Aug 28, 2026
Qwen · Verified Aug 28, 2026
Qwen · Verified Aug 28, 2026
xAI · Verified Aug 28, 2026
xAI · Verified Aug 28, 2026
OpenAI · Verified Aug 28, 2026
Anthropic · Verified Aug 28, 2026
Perplexity · Verified Aug 28, 2026
xAI · Verified Aug 28, 2026
Anthropic · Verified Aug 28, 2026
Anthropic · Verified Aug 28, 2026
Anthropic · Verified Aug 28, 2026
Rates are published list prices in USD per 1M tokens and exclude batch, committed-spend, and enterprise agreements. Run your own numbers in the cost calculators or put models head to head on the compare page.
Claude family, long context and dependable instruction following.
Aggressively priced reasoning models with cache discounts.
Gemini family, very large context windows and low-cost tiers.
Inference providers routing to open-weight models.
Llama open-weight models served by many inference providers.
European lab with efficient small and mid-size models.
Nemotron models served through NVIDIA NIM endpoints.
GPT family models, strong general reasoning and tool use.
Aggregator routing one API across many model providers.
Sonar models with built-in web search grounding.
Alibaba's Qwen line, open weights and low-cost hosted tiers.
Grok family, fast general models with large context.
It lists every model AIOPLY tracks with input price, output price, cached input price per 1M tokens, context window, max output, release date, and the date each record was last verified.
Records are reconciled daily against a live pricing feed. Each row shows its own last verified date so you can see exactly how fresh a number is before you use it in a budget.
No. All figures are published provider list rates in USD per 1M tokens. Batch processing, committed spend, and negotiated enterprise agreements usually price lower than the rates shown here.