Bonus: Top up now and we'll double your first deposit — get x2 credits instantly.
Free tool · live prices

LLM Cost Calculator

Get a monthly number you can defend. Compare 151 models at your traffic and see the real usage cost at list price — no discounts, no assumptions — plus the effective cost per task.

Compare live LLM API prices

Presets are editable example assumptions, not measured figures.
requests / day
input (user message)
output
1 word ≈ 1.33 tokens · 1 page ≈ 500 words ≈ 665 tokens.
Advanced settings
Currency: USD. Context grows each turn: every call re-sends the conversation so far.
Paste a prompt to count tokens
0 tokens
Counted in your browser — the text never leaves your device.
ModelPer requestMonthlyYearlyCost per task*
Loading live prices…

Click a model to see the cost breakdown. * Cost per task = tokens a model used per task in a public benchmark × our live prices. The benchmark follows your use case. Dash when not published. Sources

Routing mix

Split traffic across 2–3 of your selected models and see the blended bill.

What fits my budget?

Enter a monthly budget and we rank every model that stays under it at your current settings.

Double your first deposit — x2, no cap

Top up once and we match it. Run the estimate above on real traffic.

Get your API key →
Live price table

LLM API prices per 1M tokens

All 151 available text models, sorted by input price. Click a column to sort; open a model to see routes, benchmarks and alternatives. Cost per task = public benchmark token counts × our live prices; hover a cell for the token counts. Covered: general 47, coding 53, agent 21. Prices refresh weekly from the live catalogue; provider fees and discounts last verified 2026-09-28.

Llama 3.1 8B Instruct$0.020$0.050—128K———
Amazon Nova Micro$0.035$0.14$0.009128K———
GLM-4.6V FlashX$0.040$0.40$0.004128K———
Amazon Nova 2 Lite$0.040$0.16$0.0101M———
GPT-5 Nano$0.050$0.40$0.005400K———
Qwen3 VL Flash$0.050$0.40$0.010262K———
Qwen Flash$0.050$0.40$0.0101M———
Qwen Turbo$0.050$0.20$0.0051M———
Amazon Nova Lite$0.060$0.24$0.015300K———
Qwen3 Coder 30B A3B Instruct$0.070$0.27$0.007262K———
GPT OSS 20B$0.070$0.30$0.007131K—$0.0011—
GLM-4.7 FlashX$0.070$0.40$0.010200K———
Qwen3 235B A22B Instruct 2507$0.090$0.58$0.009262K———
GPT-4.1 Nano$0.10$0.40$0.0101M———
GPT-6 Luna$0.10$0.50$0.0101.05M$0.0253——
Gemini 2.5 Flash Lite$0.10$0.40$0.0101.05M———
Qwen 3.5 9B$0.10$0.15$0.010262K———
Qwen3 30B A3B Instruct 2507$0.10$0.30$0.010262K———
Mistral Small 3.2$0.10$0.30—128K———
Ministral 3 3B 2512$0.10$0.10—131K———
GLM-4 32B (0414-128k)$0.10$0.10—128K———
Qwen3 Coder Next$0.11$0.68$0.060262K$0.0102——
MiniMax M2.1 Lightning$0.12$0.48$0.024197K———
GLM-4.5 Air$0.13$0.85$0.030128K———
Gemma 4 26B A4B$0.13$0.40$0.013262K———
Llama 3.3 70B Instruct$0.14$0.40—131K———
MiMo-V2.6-Flash$0.14$0.28$0.0031.05M$0.0217——
Gemma 4 31B$0.14$0.40$0.014262K—$0.0047—
GPT-4o Mini$0.15$0.60$0.075128K———
Qwen3 Next 80B A3B Instruct$0.15$1.50$0.015131K———
GPT OSS 120B$0.15$0.60$0.015131K$0.0162$0.0017—
Ministral 3 8B 2512$0.15$0.15—262K———
GLM-5.3-Flash$0.15$0.50$0.0301M$0.0343$0.0109$0.0364
Llama 4 Scout 17B Instruct$0.17$0.66—131K———
DeepSeek V4 Flash$0.19$0.51$0.0281.05M$0.0316$0.0098—
Qwen3 VL Plus$0.20$1.60$0.040262K———
Qwen3 VL 30B A3B Instruct$0.20$0.70$0.020131K———
Qwen3 235B A22B FP8$0.20$0.80$0.02041K———
Ministral 3 14B 2512$0.20$0.20—262K———
MiniMax Text 01$0.20$1.10$0.0401M———
MiniMax M2$0.20$1.00$0.030197K———
MiniMax M2.5$0.20$1.20$0.030205K———
GPT-5.6 Luna$0.20$1.20$0.0201.05M$0.0495$0.0722$0.0309
GPT-5.4 Nano$0.20$1.25$0.020400K$0.0832——
Qwen Omni Turbo$0.20$0.80$0.02033K———
Qwen3.8 Flash Next$0.20$0.50$0.050262K$0.0539——
Qwen VL Plus$0.21$0.64$0.021131K———
Llama 4 Maverick 17B Instruct$0.24$0.97—1.05M———
GPT-5 Mini$0.25$2.00$0.025400K———
Gemini 3.1 Flash Lite$0.25$1.50$0.0251.05M$0.0179$0.0252—
GPT-5.1 Codex Mini$0.25$2.00$0.025400K———
MiniMax M2.1$0.27$1.10$0.054205K———
Gemini 3.5 Flash Lite$0.30$2.50$0.0301.05M$0.0438$0.0457—
Gemini 2.5 Flash$0.30$2.50$0.0301.05M—$0.0682—
Qwen3 Coder Flash$0.30$1.50$0.0601M———
DeepSeek V4.1 Flash$0.30$1.20$0.0061M$0.1063——
Qwen3 VL 235B A22B Instruct$0.30$1.50$0.030131K———
MiniMax M3$0.30$1.20$0.0601M$0.0578——
GLM-4.6V$0.30$0.90$0.055131K———
MiniMax M2.7$0.30$1.20$0.060205K$0.0252——
Codestral$0.30$0.90—256K—$0.0028—
Kimi K2.5$0.36$1.98$0.23262K—$0.0318—
GLM-4.7$0.38$1.98$0.11205K—$0.0387—
GPT-4.1 Mini$0.40$1.60$0.0401M———
Devstral 2 2512$0.40$2.00—262K———
Qwen Plus Latest$0.40$1.20$0.0801M———
Qwen Plus$0.40$1.20$0.080131K———
GLM-4.6$0.43$1.74$0.080205K—$0.0150—
MiMo-V2.6-Pro$0.43$0.87$0.0041.05M$0.0559——
MiMo V2.5 Pro$0.43$0.87$0.0041.05M$0.0260$0.0310—
DeepSeek V4 Flash 0731$0.44$1.32$0.0281M$0.0819$0.0916—
GLM 5.2$0.48$1.50$0.0901M$0.0965——
Gemini 3 Flash (Preview)$0.50$3.00$0.0501.05M—$0.0916—
Qwen3.7 Max$0.50$3.00$0.0501M$0.0966$0.1522—
Qwen 3.7 Plus$0.50$3.00$0.0501M$0.1090——
Qwen 3.6 Plus$0.50$3.00$0.0501.05M—$0.0571—
Qwen3 VL 235B A22B Thinking$0.50$2.00$0.050131K———
Qwen3 Next 80B A3B Thinking$0.50$6.00$0.050131K———
Mistral Large 3 2512$0.50$1.50—262K—$0.0051—
GPT-3.5 Turbo$0.50$1.50$0.05016K———
MiniMax M2.5 Highspeed$0.60$2.40$0.030205K———
GLM-4.5V$0.60$1.80$0.11128K———
GLM-4.5$0.60$2.20$0.11128K—$0.0315—
GLM-5$0.60$2.20$0.20203K—$0.0506—
Kimi K2.6$0.66$3.41$0.14262K$0.1614$0.1690—
Llama 3.1 70B Instruct$0.72$0.72—128K———
Kimi K2.7-Code$0.74$3.50$0.15262K$0.1046$0.1614$0.2075
Gemini 3.8 Flash$0.75$3.75$0.0751M$0.2663$0.1721$0.5372
Gemini 3.7 Flash$0.75$3.75$0.0751.05M—$0.1219$0.4022
Gemini 3.6 Flash$0.75$3.75$0.0751.05M$0.1560$0.1222$0.3594
GPT-5.4 Mini$0.75$4.50$0.070400K$0.2325——
QwQ Plus$0.80$2.40$0.080131K———
Amazon Nova Pro$0.80$3.20$0.20300K———
Qwen VL Max$0.80$3.20$0.080131K———
Qwen3 Max 2026-01-23$0.84$3.38$0.60262K———
Qwen Coder Plus$1.00$5.00$0.10131K———
o3 Mini$1.10$4.40$0.11200K———
GLM-4.5 AirX$1.10$4.50$0.22128K———
GLM-5-Turbo$1.20$4.00$0.24200K———
GLM-5V-Turbo$1.20$4.00$0.24200K———
Gemini 2.5 Pro$1.25$10.0$0.131.05M$0.1055——
GPT-5 Codex$1.25$10.0$0.13400K———
Grok 4.3$1.25$2.50$0.201M$0.0442$0.0564—
GPT-5.1 Codex$1.25$10.0$0.13400K—$0.5140—
GPT-5.1$1.25$10.0$0.13400K———
GPT-5$1.25$10.0$0.13400K—$0.1315—
DeepSeek V4 Pro$1.32$3.96$0.0441.05M$0.2184$0.1717—
GLM-5.1$1.40$4.40$0.26203K$0.1817$0.1315—
Gemini 3.5 Flash$1.50$9.00$0.151.05M—$0.4425$0.6816
Qwen Max$1.60$6.40$0.16131K———
GPT-5.3 Codex$1.75$14.0$0.17400K—$0.5618—
GPT-5.2 Codex$1.75$14.0$0.17400K—$0.5916—
GPT-5.2$1.75$14.0$0.17400K———
Kimi K2.7 Code HighSpeed$1.90$8.00$0.38262K———
Gemini 3.1 Pro (Preview)$2.00$12.0$0.201.05M$0.2173$0.3884$0.3404
GPT-6 Sol$2.00$10.0$0.201.05M$0.3124——
Qwen 3.8 Max$2.00$6.00$0.171M$0.6464—$0.5705
o3$2.00$8.00$0.20200K—$0.1008—
Grok 4.6$2.00$6.00$0.50500K$0.2150$0.2440$0.3670
Grok 4.5$2.00$6.00$0.30500K$0.1613$0.1802$0.2132
GPT-4.1$2.00$8.00$0.201M———
Claude Sonnet 5$2.00$10.0$0.201M$1.1779$0.6403$0.8730
GPT-5.6 Terra$2.00$12.0$0.201.05M$0.4668$1.0079$0.2582
GLM-4.5 X$2.20$8.90$0.45128K———
GPT-5.4$2.50$15.0$0.251.05M———
GPT-4o$2.50$10.0$1.25128K———
Kimi K3$2.66$13.3$0.281M$0.6445$1.1379$1.0839
Claude Sonnet 4.6$3.00$15.0$0.301M$1.1238$1.7708$1.1424
Claude Sonnet 4 (2025-05-14)$3.00$15.0$0.30200K—$0.0750—
Claude Opus 5.5$4.00$20.0$0.201M$2.3833——
Mistral Large Latest$4.00$12.0—128K———
GPT-5.6 Sol$4.00$20.0$0.401.05M$0.5862$1.0352$0.5690
MiMo-V2.6-Pro-UltraSpeed$4.35$8.70$0.0361.05M———
Claude Opus 4.7$5.00$25.0$0.501M—$0.2245—
Claude Opus 4.6$5.00$25.0$0.501M—$0.0617—
Claude Opus 5$5.00$25.0$0.501M$1.8128$1.8905$1.6052
Claude Opus 4.8$5.00$25.0$0.501M$1.7637$0.9639$1.2516
Claude Opus 4.5$5.00$25.0$0.50200K———
GPT-5.5$5.00$30.0$0.501.05M$0.7097$0.3367$0.9348
Qwen3 Coder Plus$6.00$60.0$0.601M—$0.1622—
Grok 4.7$6.00$18.0$1.50500K$1.4501——
GPT-6 Astra$10.0$50.0$1.001.05M$1.3603$2.8860$1.3253
Claude Fable 5.1$10.0$50.0$0.251M$3.9055$2.2948—
Claude Fable 5$10.0$50.0$1.001M$3.3424$2.1088$2.8643
GPT-4 Turbo$10.0$30.0$1.00128K———
GPT-5 Pro$15.0$120.0$1.50400K———
Claude Opus 4.1$15.0$75.0$1.50200K———
Video O1$15.0$60.0$1.50200K———
GPT-5.2 Pro$21.0$168.0$2.10400K———
GPT-5.4 Pro$30.0$180.0$3.001.05M———
GPT-4$30.0$60.0$3.008K———

Worked examples

What real workloads cost

Three common workloads priced with live rates on the top-rated model, a mid-priced one, and the cheapest well-rated one.

Customer support chatbot

20,000 conversations a day, long system prompt, short answers.
20,000 req/day · 1,500 in / 300 out tokens

  • Gemini 3.5 Flash Lite$720/mo
  • Qwen 3.6 Plus$990/mo
  • Amazon Nova Micro$56.70/mo

Coding assistant

5,000 requests a day with large code context and longer replies.
5,000 req/day · 8,000 in / 1,500 out tokens

  • Gemini 3.5 Flash Lite$923/mo
  • Qwen 3.6 Plus$1,275/mo
  • Amazon Nova Micro$73.50/mo

Document summarization

2,000 documents a day, heavy input, short summaries.
2,000 req/day · 12,000 in / 600 out tokens

  • Gemini 3.5 Flash Lite$306/mo
  • Qwen 3.6 Plus$468/mo
  • Amazon Nova Micro$30.24/mo
Tokens & caching

A 60-second explainer

Models bill by tokens — chunks of text, roughly ¾ of an English word. You pay for input tokens (your prompt, system instructions and context) and output tokens (the answer). Output is usually several times more expensive.

Monthly cost = requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000

Reasoning models also bill the tokens they "think" with, which is why two models with the same list price can cost very differently per task.

Prompt caching stores the part of your prompt that repeats — a long system prompt, a document, tool definitions — so the next request reads it at a fraction of the input price.

Request 12,000 infull input price
Request 2+1,800 cachedcache-read price
+ new200 infull input price

Only the repeated prefix is cached; new tokens and all output tokens are billed normally.

Customer story

Mate Academy reduced AI costs by 70%

Smart routing sent simpler tasks to cheaper models and kept premium models for advanced workflows, with real-time per-model cost visibility — a 70% cost reduction in 10 business days.

Read the case study →
70%Cost reduction
$36KSaved annually
$60k → $24kAnnual AI spend
10 daysTo results
Compare

Going direct vs using LLM API

Direct to each providerLLM API
Price per tokenProvider's published priceProvider's published price — no markup
Accounts & keysOne account, key and invoice per providerOne key, one balance, one invoice
Switching modelsNew SDK or integration per providerChange the model name; OpenAI-compatible API
Routing cheaper modelsBuild and maintain it yourselfRoute across models from one endpoint
Cost visibilitySeparate dashboards per providerPer-model spend in one place
Promos & volume discountsNegotiated per providerPassed through at provider terms
Sources

Where cost per task comes from

We import only token counts from public benchmarks and multiply them by our live prices. Their own dollar figures use old price snapshots, so we never copy them.

  • Coding task — input + output tokens per problem in ALE-Bench, via the Epoch AI Benchmarking Hub (CC BY 4.0). Older models: prompt + completion tokens per exercise in the Aider Polyglot leaderboard (225 exercises, Apache 2.0).
  • Agent task — mean output tokens per task in DeepSWE (mini-swe-agent harness), via Epoch AI. Input tokens aren't published, so this is a lower bound.
  • General task — output tokens per Intelligence Index task from Artificial Analysis.

Epoch AI, ‘Capabilities & benchmarking’, epoch.ai/benchmarks. Refreshed weekly.

Read these numbers with care

  • Token use depends on the test setup (harness, prompts, reasoning effort). Compare models within one column, never across columns.
  • Where a model was tested at several reasoning levels we use "high"; the level is shown on hover.
  • Reasoning tokens are counted in output, so they're included. Caching isn't reported, so no cache discount is applied.
  • Benchmarks mostly test frontier models; cheap and older models often show a dash.
  • Failed attempts cost money too — these are costs per attempt, not per successful task.
FAQ

LLM cost calculator FAQ

Pricing, tokens and what the estimate does and does not include.

How is the LLM cost calculated?

Each request is split into input, cached input, cache writes, output, reasoning and tool fees, each priced at the model's published rate, then multiplied by your monthly volume. Retries add volume, and the batch share gets the provider's published batch discount.

Where do the prices come from?

From the live LLM API model catalogue, which powers every model page, refreshed weekly. Batch discounts, max output and promos come from each provider's official pricing page, with the date we last verified them.

How much does prompt caching save?

Cached input is usually billed at about a tenth of the normal input price. The calculator applies your cache-hit % to the repeated part of each request — system prompt, retrieved context and conversation history — and bills misses at the cache-write price where the provider charges one.

What is batch pricing?

Some providers (OpenAI, Anthropic, Google, Mistral, Alibaba) price asynchronous batch requests 50% lower. Set a batch share for jobs that can wait; models without a published batch price get no discount.

Why do reasoning models cost more than their price suggests?

They bill the tokens they think with, usually at the output rate. Add reasoning tokens in the advanced settings and choose an effort level; the cost-per-task column also shows the effect from benchmark runs.

What is long-context pricing?

Some models charge a higher rate once a single prompt passes a threshold, such as 200K or 272K tokens. The whole call is then billed at the higher tier. The calculator applies it per call when your scenario crosses the line.

Why does a multi-turn conversation cost so much more?

Every turn re-sends the whole conversation so far, so input grows with each turn. Caching the repeated history is the main way to keep it down.

How accurate is the paste-a-prompt token count?

It is exact for OpenAI models, which publish their tokenizer. For other models it is an estimate of about 4 characters per token and is labelled approx.

Why do some cells show a dash?

Because the provider hasn't published that figure — a cache price, a batch discount or output tokens per task. We never estimate a price, so the line stays empty.

Why is output priced higher than input?

Generating tokens is the expensive part of running a model: each output token needs a full forward pass, while input tokens are processed in parallel. Most providers charge three to five times more per output token.

How often are the prices refreshed?

Weekly, from the live LLM API model catalogue — the same data that powers every model page, so prices never drift between pages. Provider fees and discounts carry the date we last verified them against the official pricing page.

Can I share or export my estimate?

Yes. Copy share link puts your whole setup — models, volume, advanced settings — into the URL, and Download CSV exports the current results table.

Is the calculator free to use?

Yes. No sign-up is needed. When you are ready, create an API key and your first deposit is doubled.