LLM Cost Calculator
Get a monthly number you can defend. Compare 151 models at your traffic and see the real usage cost at list price — no discounts, no assumptions — plus the effective cost per task.
Compare live LLM API prices
| Model | Price | Per day | Monthly | Yearly |
|---|---|---|---|---|
| Loading live prices… | ||||
Monthly = per day × 30. Each price is the model's published list price on its cheapest route, refreshed weekly. Only models with a flat per-unit price are listed; image models billed per token are left out rather than estimated.
Advanced settings
Paste a prompt to count tokens
| Model | Per request | Monthly | Yearly | Cost per task* |
|---|---|---|---|---|
| Loading live prices… | ||||
Click a model to see the cost breakdown. * Cost per task = tokens a model used per task in a public benchmark × our live prices. The benchmark follows your use case. Dash when not published. Sources
Routing mix
Split traffic across 2–3 of your selected models and see the blended bill.
What fits my budget?
Enter a monthly budget and we rank every model that stays under it at your current settings.
Top up once and we match it. Run the estimate above on real traffic.
At this volume, talk to our team about routing strategy and volume pricing.
LLM API prices per 1M tokens
All 151 available text models, sorted by input price. Click a column to sort; open a model to see routes, benchmarks and alternatives. Cost per task = public benchmark token counts × our live prices; hover a cell for the token counts. Covered: general 47, coding 53, agent 21. Prices refresh weekly from the live catalogue; provider fees and discounts last verified 2026-09-28.
| Llama 3.1 8B Instruct | $0.020 | $0.050 | — | 128K | — | — | — |
| Amazon Nova Micro | $0.035 | $0.14 | $0.009 | 128K | — | — | — |
| GLM-4.6V FlashX | $0.040 | $0.40 | $0.004 | 128K | — | — | — |
| Amazon Nova 2 Lite | $0.040 | $0.16 | $0.010 | 1M | — | — | — |
| GPT-5 Nano | $0.050 | $0.40 | $0.005 | 400K | — | — | — |
| Qwen3 VL Flash | $0.050 | $0.40 | $0.010 | 262K | — | — | — |
| Qwen Flash | $0.050 | $0.40 | $0.010 | 1M | — | — | — |
| Qwen Turbo | $0.050 | $0.20 | $0.005 | 1M | — | — | — |
| Amazon Nova Lite | $0.060 | $0.24 | $0.015 | 300K | — | — | — |
| Qwen3 Coder 30B A3B Instruct | $0.070 | $0.27 | $0.007 | 262K | — | — | — |
| GPT OSS 20B | $0.070 | $0.30 | $0.007 | 131K | — | $0.0011 | — |
| GLM-4.7 FlashX | $0.070 | $0.40 | $0.010 | 200K | — | — | — |
| Qwen3 235B A22B Instruct 2507 | $0.090 | $0.58 | $0.009 | 262K | — | — | — |
| GPT-4.1 Nano | $0.10 | $0.40 | $0.010 | 1M | — | — | — |
| GPT-6 Luna | $0.10 | $0.50 | $0.010 | 1.05M | $0.0253 | — | — |
| Gemini 2.5 Flash Lite | $0.10 | $0.40 | $0.010 | 1.05M | — | — | — |
| Qwen 3.5 9B | $0.10 | $0.15 | $0.010 | 262K | — | — | — |
| Qwen3 30B A3B Instruct 2507 | $0.10 | $0.30 | $0.010 | 262K | — | — | — |
| Mistral Small 3.2 | $0.10 | $0.30 | — | 128K | — | — | — |
| Ministral 3 3B 2512 | $0.10 | $0.10 | — | 131K | — | — | — |
| GLM-4 32B (0414-128k) | $0.10 | $0.10 | — | 128K | — | — | — |
| Qwen3 Coder Next | $0.11 | $0.68 | $0.060 | 262K | $0.0102 | — | — |
| MiniMax M2.1 Lightning | $0.12 | $0.48 | $0.024 | 197K | — | — | — |
| GLM-4.5 Air | $0.13 | $0.85 | $0.030 | 128K | — | — | — |
| Gemma 4 26B A4B | $0.13 | $0.40 | $0.013 | 262K | — | — | — |
| Llama 3.3 70B Instruct | $0.14 | $0.40 | — | 131K | — | — | — |
| MiMo-V2.6-Flash | $0.14 | $0.28 | $0.003 | 1.05M | $0.0217 | — | — |
| Gemma 4 31B | $0.14 | $0.40 | $0.014 | 262K | — | $0.0047 | — |
| GPT-4o Mini | $0.15 | $0.60 | $0.075 | 128K | — | — | — |
| Qwen3 Next 80B A3B Instruct | $0.15 | $1.50 | $0.015 | 131K | — | — | — |
| GPT OSS 120B | $0.15 | $0.60 | $0.015 | 131K | $0.0162 | $0.0017 | — |
| Ministral 3 8B 2512 | $0.15 | $0.15 | — | 262K | — | — | — |
| GLM-5.3-Flash | $0.15 | $0.50 | $0.030 | 1M | $0.0343 | $0.0109 | $0.0364 |
| Llama 4 Scout 17B Instruct | $0.17 | $0.66 | — | 131K | — | — | — |
| DeepSeek V4 Flash | $0.19 | $0.51 | $0.028 | 1.05M | $0.0316 | $0.0098 | — |
| Qwen3 VL Plus | $0.20 | $1.60 | $0.040 | 262K | — | — | — |
| Qwen3 VL 30B A3B Instruct | $0.20 | $0.70 | $0.020 | 131K | — | — | — |
| Qwen3 235B A22B FP8 | $0.20 | $0.80 | $0.020 | 41K | — | — | — |
| Ministral 3 14B 2512 | $0.20 | $0.20 | — | 262K | — | — | — |
| MiniMax Text 01 | $0.20 | $1.10 | $0.040 | 1M | — | — | — |
| MiniMax M2 | $0.20 | $1.00 | $0.030 | 197K | — | — | — |
| MiniMax M2.5 | $0.20 | $1.20 | $0.030 | 205K | — | — | — |
| GPT-5.6 Luna | $0.20 | $1.20 | $0.020 | 1.05M | $0.0495 | $0.0722 | $0.0309 |
| GPT-5.4 Nano | $0.20 | $1.25 | $0.020 | 400K | $0.0832 | — | — |
| Qwen Omni Turbo | $0.20 | $0.80 | $0.020 | 33K | — | — | — |
| Qwen3.8 Flash Next | $0.20 | $0.50 | $0.050 | 262K | $0.0539 | — | — |
| Qwen VL Plus | $0.21 | $0.64 | $0.021 | 131K | — | — | — |
| Llama 4 Maverick 17B Instruct | $0.24 | $0.97 | — | 1.05M | — | — | — |
| GPT-5 Mini | $0.25 | $2.00 | $0.025 | 400K | — | — | — |
| Gemini 3.1 Flash Lite | $0.25 | $1.50 | $0.025 | 1.05M | $0.0179 | $0.0252 | — |
| GPT-5.1 Codex Mini | $0.25 | $2.00 | $0.025 | 400K | — | — | — |
| MiniMax M2.1 | $0.27 | $1.10 | $0.054 | 205K | — | — | — |
| Gemini 3.5 Flash Lite | $0.30 | $2.50 | $0.030 | 1.05M | $0.0438 | $0.0457 | — |
| Gemini 2.5 Flash | $0.30 | $2.50 | $0.030 | 1.05M | — | $0.0682 | — |
| Qwen3 Coder Flash | $0.30 | $1.50 | $0.060 | 1M | — | — | — |
| DeepSeek V4.1 Flash | $0.30 | $1.20 | $0.006 | 1M | $0.1063 | — | — |
| Qwen3 VL 235B A22B Instruct | $0.30 | $1.50 | $0.030 | 131K | — | — | — |
| MiniMax M3 | $0.30 | $1.20 | $0.060 | 1M | $0.0578 | — | — |
| GLM-4.6V | $0.30 | $0.90 | $0.055 | 131K | — | — | — |
| MiniMax M2.7 | $0.30 | $1.20 | $0.060 | 205K | $0.0252 | — | — |
| Codestral | $0.30 | $0.90 | — | 256K | — | $0.0028 | — |
| Kimi K2.5 | $0.36 | $1.98 | $0.23 | 262K | — | $0.0318 | — |
| GLM-4.7 | $0.38 | $1.98 | $0.11 | 205K | — | $0.0387 | — |
| GPT-4.1 Mini | $0.40 | $1.60 | $0.040 | 1M | — | — | — |
| Devstral 2 2512 | $0.40 | $2.00 | — | 262K | — | — | — |
| Qwen Plus Latest | $0.40 | $1.20 | $0.080 | 1M | — | — | — |
| Qwen Plus | $0.40 | $1.20 | $0.080 | 131K | — | — | — |
| GLM-4.6 | $0.43 | $1.74 | $0.080 | 205K | — | $0.0150 | — |
| MiMo-V2.6-Pro | $0.43 | $0.87 | $0.004 | 1.05M | $0.0559 | — | — |
| MiMo V2.5 Pro | $0.43 | $0.87 | $0.004 | 1.05M | $0.0260 | $0.0310 | — |
| DeepSeek V4 Flash 0731 | $0.44 | $1.32 | $0.028 | 1M | $0.0819 | $0.0916 | — |
| GLM 5.2 | $0.48 | $1.50 | $0.090 | 1M | $0.0965 | — | — |
| Gemini 3 Flash (Preview) | $0.50 | $3.00 | $0.050 | 1.05M | — | $0.0916 | — |
| Qwen3.7 Max | $0.50 | $3.00 | $0.050 | 1M | $0.0966 | $0.1522 | — |
| Qwen 3.7 Plus | $0.50 | $3.00 | $0.050 | 1M | $0.1090 | — | — |
| Qwen 3.6 Plus | $0.50 | $3.00 | $0.050 | 1.05M | — | $0.0571 | — |
| Qwen3 VL 235B A22B Thinking | $0.50 | $2.00 | $0.050 | 131K | — | — | — |
| Qwen3 Next 80B A3B Thinking | $0.50 | $6.00 | $0.050 | 131K | — | — | — |
| Mistral Large 3 2512 | $0.50 | $1.50 | — | 262K | — | $0.0051 | — |
| GPT-3.5 Turbo | $0.50 | $1.50 | $0.050 | 16K | — | — | — |
| MiniMax M2.5 Highspeed | $0.60 | $2.40 | $0.030 | 205K | — | — | — |
| GLM-4.5V | $0.60 | $1.80 | $0.11 | 128K | — | — | — |
| GLM-4.5 | $0.60 | $2.20 | $0.11 | 128K | — | $0.0315 | — |
| GLM-5 | $0.60 | $2.20 | $0.20 | 203K | — | $0.0506 | — |
| Kimi K2.6 | $0.66 | $3.41 | $0.14 | 262K | $0.1614 | $0.1690 | — |
| Llama 3.1 70B Instruct | $0.72 | $0.72 | — | 128K | — | — | — |
| Kimi K2.7-Code | $0.74 | $3.50 | $0.15 | 262K | $0.1046 | $0.1614 | $0.2075 |
| Gemini 3.8 Flash | $0.75 | $3.75 | $0.075 | 1M | $0.2663 | $0.1721 | $0.5372 |
| Gemini 3.7 Flash | $0.75 | $3.75 | $0.075 | 1.05M | — | $0.1219 | $0.4022 |
| Gemini 3.6 Flash | $0.75 | $3.75 | $0.075 | 1.05M | $0.1560 | $0.1222 | $0.3594 |
| GPT-5.4 Mini | $0.75 | $4.50 | $0.070 | 400K | $0.2325 | — | — |
| QwQ Plus | $0.80 | $2.40 | $0.080 | 131K | — | — | — |
| Amazon Nova Pro | $0.80 | $3.20 | $0.20 | 300K | — | — | — |
| Qwen VL Max | $0.80 | $3.20 | $0.080 | 131K | — | — | — |
| Qwen3 Max 2026-01-23 | $0.84 | $3.38 | $0.60 | 262K | — | — | — |
| Qwen Coder Plus | $1.00 | $5.00 | $0.10 | 131K | — | — | — |
| o3 Mini | $1.10 | $4.40 | $0.11 | 200K | — | — | — |
| GLM-4.5 AirX | $1.10 | $4.50 | $0.22 | 128K | — | — | — |
| GLM-5-Turbo | $1.20 | $4.00 | $0.24 | 200K | — | — | — |
| GLM-5V-Turbo | $1.20 | $4.00 | $0.24 | 200K | — | — | — |
| Gemini 2.5 Pro | $1.25 | $10.0 | $0.13 | 1.05M | $0.1055 | — | — |
| GPT-5 Codex | $1.25 | $10.0 | $0.13 | 400K | — | — | — |
| Grok 4.3 | $1.25 | $2.50 | $0.20 | 1M | $0.0442 | $0.0564 | — |
| GPT-5.1 Codex | $1.25 | $10.0 | $0.13 | 400K | — | $0.5140 | — |
| GPT-5.1 | $1.25 | $10.0 | $0.13 | 400K | — | — | — |
| GPT-5 | $1.25 | $10.0 | $0.13 | 400K | — | $0.1315 | — |
| DeepSeek V4 Pro | $1.32 | $3.96 | $0.044 | 1.05M | $0.2184 | $0.1717 | — |
| GLM-5.1 | $1.40 | $4.40 | $0.26 | 203K | $0.1817 | $0.1315 | — |
| Gemini 3.5 Flash | $1.50 | $9.00 | $0.15 | 1.05M | — | $0.4425 | $0.6816 |
| Qwen Max | $1.60 | $6.40 | $0.16 | 131K | — | — | — |
| GPT-5.3 Codex | $1.75 | $14.0 | $0.17 | 400K | — | $0.5618 | — |
| GPT-5.2 Codex | $1.75 | $14.0 | $0.17 | 400K | — | $0.5916 | — |
| GPT-5.2 | $1.75 | $14.0 | $0.17 | 400K | — | — | — |
| Kimi K2.7 Code HighSpeed | $1.90 | $8.00 | $0.38 | 262K | — | — | — |
| Gemini 3.1 Pro (Preview) | $2.00 | $12.0 | $0.20 | 1.05M | $0.2173 | $0.3884 | $0.3404 |
| GPT-6 Sol | $2.00 | $10.0 | $0.20 | 1.05M | $0.3124 | — | — |
| Qwen 3.8 Max | $2.00 | $6.00 | $0.17 | 1M | $0.6464 | — | $0.5705 |
| o3 | $2.00 | $8.00 | $0.20 | 200K | — | $0.1008 | — |
| Grok 4.6 | $2.00 | $6.00 | $0.50 | 500K | $0.2150 | $0.2440 | $0.3670 |
| Grok 4.5 | $2.00 | $6.00 | $0.30 | 500K | $0.1613 | $0.1802 | $0.2132 |
| GPT-4.1 | $2.00 | $8.00 | $0.20 | 1M | — | — | — |
| Claude Sonnet 5 | $2.00 | $10.0 | $0.20 | 1M | $1.1779 | $0.6403 | $0.8730 |
| GPT-5.6 Terra | $2.00 | $12.0 | $0.20 | 1.05M | $0.4668 | $1.0079 | $0.2582 |
| GLM-4.5 X | $2.20 | $8.90 | $0.45 | 128K | — | — | — |
| GPT-5.4 | $2.50 | $15.0 | $0.25 | 1.05M | — | — | — |
| GPT-4o | $2.50 | $10.0 | $1.25 | 128K | — | — | — |
| Kimi K3 | $2.66 | $13.3 | $0.28 | 1M | $0.6445 | $1.1379 | $1.0839 |
| Claude Sonnet 4.6 | $3.00 | $15.0 | $0.30 | 1M | $1.1238 | $1.7708 | $1.1424 |
| Claude Sonnet 4 (2025-05-14) | $3.00 | $15.0 | $0.30 | 200K | — | $0.0750 | — |
| Claude Opus 5.5 | $4.00 | $20.0 | $0.20 | 1M | $2.3833 | — | — |
| Mistral Large Latest | $4.00 | $12.0 | — | 128K | — | — | — |
| GPT-5.6 Sol | $4.00 | $20.0 | $0.40 | 1.05M | $0.5862 | $1.0352 | $0.5690 |
| MiMo-V2.6-Pro-UltraSpeed | $4.35 | $8.70 | $0.036 | 1.05M | — | — | — |
| Claude Opus 4.7 | $5.00 | $25.0 | $0.50 | 1M | — | $0.2245 | — |
| Claude Opus 4.6 | $5.00 | $25.0 | $0.50 | 1M | — | $0.0617 | — |
| Claude Opus 5 | $5.00 | $25.0 | $0.50 | 1M | $1.8128 | $1.8905 | $1.6052 |
| Claude Opus 4.8 | $5.00 | $25.0 | $0.50 | 1M | $1.7637 | $0.9639 | $1.2516 |
| Claude Opus 4.5 | $5.00 | $25.0 | $0.50 | 200K | — | — | — |
| GPT-5.5 | $5.00 | $30.0 | $0.50 | 1.05M | $0.7097 | $0.3367 | $0.9348 |
| Qwen3 Coder Plus | $6.00 | $60.0 | $0.60 | 1M | — | $0.1622 | — |
| Grok 4.7 | $6.00 | $18.0 | $1.50 | 500K | $1.4501 | — | — |
| GPT-6 Astra | $10.0 | $50.0 | $1.00 | 1.05M | $1.3603 | $2.8860 | $1.3253 |
| Claude Fable 5.1 | $10.0 | $50.0 | $0.25 | 1M | $3.9055 | $2.2948 | — |
| Claude Fable 5 | $10.0 | $50.0 | $1.00 | 1M | $3.3424 | $2.1088 | $2.8643 |
| GPT-4 Turbo | $10.0 | $30.0 | $1.00 | 128K | — | — | — |
| GPT-5 Pro | $15.0 | $120.0 | $1.50 | 400K | — | — | — |
| Claude Opus 4.1 | $15.0 | $75.0 | $1.50 | 200K | — | — | — |
| Video O1 | $15.0 | $60.0 | $1.50 | 200K | — | — | — |
| GPT-5.2 Pro | $21.0 | $168.0 | $2.10 | 400K | — | — | — |
| GPT-5.4 Pro | $30.0 | $180.0 | $3.00 | 1.05M | — | — | — |
| GPT-4 | $30.0 | $60.0 | $3.00 | 8K | — | — | — |
What real workloads cost
Three common workloads priced with live rates on the top-rated model, a mid-priced one, and the cheapest well-rated one.
Customer support chatbot
- Gemini 3.5 Flash Lite$720/mo
- Qwen 3.6 Plus$990/mo
- Amazon Nova Micro$56.70/mo
Coding assistant
- Gemini 3.5 Flash Lite$923/mo
- Qwen 3.6 Plus$1,275/mo
- Amazon Nova Micro$73.50/mo
Document summarization
- Gemini 3.5 Flash Lite$306/mo
- Qwen 3.6 Plus$468/mo
- Amazon Nova Micro$30.24/mo
A 60-second explainer
Models bill by tokens — chunks of text, roughly ¾ of an English word. You pay for input tokens (your prompt, system instructions and context) and output tokens (the answer). Output is usually several times more expensive.
Monthly cost = requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000
Reasoning models also bill the tokens they "think" with, which is why two models with the same list price can cost very differently per task.
Prompt caching stores the part of your prompt that repeats — a long system prompt, a document, tool definitions — so the next request reads it at a fraction of the input price.
Only the repeated prefix is cached; new tokens and all output tokens are billed normally.
Mate Academy reduced AI costs by 70%
Smart routing sent simpler tasks to cheaper models and kept premium models for advanced workflows, with real-time per-model cost visibility — a 70% cost reduction in 10 business days.
Read the case study →Going direct vs using LLM API
| Direct to each provider | LLM API | |
|---|---|---|
| Price per token | Provider's published price | Provider's published price — no markup |
| Accounts & keys | One account, key and invoice per provider | One key, one balance, one invoice |
| Switching models | New SDK or integration per provider | Change the model name; OpenAI-compatible API |
| Routing cheaper models | Build and maintain it yourself | Route across models from one endpoint |
| Cost visibility | Separate dashboards per provider | Per-model spend in one place |
| Promos & volume discounts | Negotiated per provider | Passed through at provider terms |
Where cost per task comes from
We import only token counts from public benchmarks and multiply them by our live prices. Their own dollar figures use old price snapshots, so we never copy them.
- Coding task — input + output tokens per problem in ALE-Bench, via the Epoch AI Benchmarking Hub (CC BY 4.0). Older models: prompt + completion tokens per exercise in the Aider Polyglot leaderboard (225 exercises, Apache 2.0).
- Agent task — mean output tokens per task in DeepSWE (mini-swe-agent harness), via Epoch AI. Input tokens aren't published, so this is a lower bound.
- General task — output tokens per Intelligence Index task from Artificial Analysis.
Epoch AI, ‘Capabilities & benchmarking’, epoch.ai/benchmarks. Refreshed weekly.
Read these numbers with care
- Token use depends on the test setup (harness, prompts, reasoning effort). Compare models within one column, never across columns.
- Where a model was tested at several reasoning levels we use "high"; the level is shown on hover.
- Reasoning tokens are counted in output, so they're included. Caching isn't reported, so no cache discount is applied.
- Benchmarks mostly test frontier models; cheap and older models often show a dash.
- Failed attempts cost money too — these are costs per attempt, not per successful task.
LLM cost calculator FAQ
Pricing, tokens and what the estimate does and does not include.
How is the LLM cost calculated?
Each request is split into input, cached input, cache writes, output, reasoning and tool fees, each priced at the model's published rate, then multiplied by your monthly volume. Retries add volume, and the batch share gets the provider's published batch discount.
Where do the prices come from?
From the live LLM API model catalogue, which powers every model page, refreshed weekly. Batch discounts, max output and promos come from each provider's official pricing page, with the date we last verified them.
How much does prompt caching save?
Cached input is usually billed at about a tenth of the normal input price. The calculator applies your cache-hit % to the repeated part of each request — system prompt, retrieved context and conversation history — and bills misses at the cache-write price where the provider charges one.
What is batch pricing?
Some providers (OpenAI, Anthropic, Google, Mistral, Alibaba) price asynchronous batch requests 50% lower. Set a batch share for jobs that can wait; models without a published batch price get no discount.
Why do reasoning models cost more than their price suggests?
They bill the tokens they think with, usually at the output rate. Add reasoning tokens in the advanced settings and choose an effort level; the cost-per-task column also shows the effect from benchmark runs.
What is long-context pricing?
Some models charge a higher rate once a single prompt passes a threshold, such as 200K or 272K tokens. The whole call is then billed at the higher tier. The calculator applies it per call when your scenario crosses the line.
Why does a multi-turn conversation cost so much more?
Every turn re-sends the whole conversation so far, so input grows with each turn. Caching the repeated history is the main way to keep it down.
How accurate is the paste-a-prompt token count?
It is exact for OpenAI models, which publish their tokenizer. For other models it is an estimate of about 4 characters per token and is labelled approx.
Why do some cells show a dash?
Because the provider hasn't published that figure — a cache price, a batch discount or output tokens per task. We never estimate a price, so the line stays empty.
Why is output priced higher than input?
Generating tokens is the expensive part of running a model: each output token needs a full forward pass, while input tokens are processed in parallel. Most providers charge three to five times more per output token.
How often are the prices refreshed?
Weekly, from the live LLM API model catalogue — the same data that powers every model page, so prices never drift between pages. Provider fees and discounts carry the date we last verified them against the official pricing page.
Can I share or export my estimate?
Yes. Copy share link puts your whole setup — models, volume, advanced settings — into the URL, and Download CSV exports the current results table.
Is the calculator free to use?
Yes. No sign-up is needed. When you are ready, create an API key and your first deposit is doubled.