Cheapest LLM API Finder
Enter your workload and get the 10 cheapest LLM APIs ranked by real monthly cost.
Pricing dataset updated Oct 2, 2026
- 1$2.34per month
Ling 3.0 Flash VL
inclusionAI · 262K context · $0.021 / $0.062 per 1M
- 2$2.36per month
Ling 3.0 Flash
inclusionAI · 262K context · $0.021 / $0.063 per 1M
- 3$2.98per month
Granite 4.0 Micro
IBM · 131K context · $0.017 / $0.112 per 1M
- 4$3.25per month
Nex-N2.5-Mini
Nex AGI · 262K context · $0.025 / $0.100 per 1M
- 5$3.98per month
DeepSeek V4 Flash 0423
DeepSeek · 1.048576M context · $0.042 / $0.084 per 1M
- 6$4.08per month
Qwen3.7 Flash
Qwen · 1M context · $0.030 / $0.130 per 1M
- 7$5.02per month
Mercury 2.5
Inception · 260K context · $0.040 / $0.150 per 1M
- 8$6.50per month
Nemotron 3 Nano 30B A3B
NVIDIA · 262K context · $0.050 / $0.200 per 1M
- 9$6.50per month
Solar Mini 4
Upstage · 524K context · $0.050 / $0.200 per 1M
- 10$6.54per month
Nemotron 3.5 Lightning
NVIDIA · 262K context · $0.059 / $0.170 per 1M
Ranked from live list pricing per 1M tokens, refreshed every 6 hours. Free tiers may carry rate limits.
How it works
The cheapest model per token is not always the cheapest for your workload. A model with low input pricing but expensive output can cost more than a balanced one if your responses are long.
This finder multiplies each model's live input and output price by your own token mix and request volume, then ranks the results. Legacy models are excluded, and you can set a minimum context window so only models that fit your prompts are shown.
Before switching, test the top two or three options on a sample of real prompts. Quality differences often matter more than a few dollars a month.
Worked example
A chatbot at 50,000 requests a month
With 1,200 input and 350 output tokens per request, the cheapest capable models usually land under $10 a month, while frontier models cost several hundred dollars for the same traffic.
