Price & performance
Cheapest per Token providers
Every model provider on LLMAPI tagged “Cheapest per Token” — 11 of 50 providers. All of them run through one API key and one OpenAI-compatible endpoint.
Providers listed11
Models reachable53
Filters available84
IntegrationOne OpenAI-compatible endpoint
11 providers matching this filter.
| Provider | Models | Regions | Best for | |
|---|---|---|---|---|
Google Gemini models via AI Studio and Vertex AI. |
7 | Global | Multimodal workloads and very long context at aggressive price points | View → |
Mistral Efficient open and commercial models from Paris. |
10 | EU, US | Price-performance and EU data-residency requirements | View → |
DeepSeek Open reasoning models at aggressive price points. |
5 | Global | Reasoning and coding at the lowest per-token prices | View → |
Zai (Zhipu AI) GLM series models from Zhipu AI. |
5 | APAC | Chinese-English bilingual applications | View → |
Multiverse Computing Compressed models via quantum-inspired techniques. |
3 | EU | Compressed models for on-device and cost-sensitive deployments | View → |
Parasail Serverless inference marketplace. |
4 | US | Serverless open-model inference at spot-style pricing | View → |
Crusoe Energy-first GPU cloud. |
1 | US | Cost-effective, sustainably powered GPU compute | View → |
DeepInfra Low-cost inference for popular open models. |
8 | US | Popular open models at the lowest per-token cost | View → |
Novita Affordable GPU cloud and model APIs. |
4 | APAC, US | Budget-friendly serverless GPUs and model APIs | View → |
Wafer Distributed inference network. |
3 | Global | Affordable open-model inference on distributed GPUs | View → |
Makora Model hosting and inference APIs. |
3 | Global | Pay-as-you-go hosting for open models | View → |
One key, every provider on this page
Consolidated billing, spend limits and instant fallback between providers.