LLM API
Text, vision, speech, video and OCR behind one OpenAI-compatible key, with uptime monitoring, spend controls and an invoice instead of credits.
We compared 20 routers, gateways and inference clouds on cost, control, latency and compliance. Answer six questions and get a shortlist built for your stack — not a generic top ten.
A platform margin on top of provider prices is invisible at prototype volume and a real line item at production volume. Routers that bill provider rates, or let you bring your own keys, close that gap.
Once legal reviews the data path, a US-only proxy in front of your prompts stops being acceptable. Regional hosting, self-hosted gateways and cloud platforms keep traffic inside a perimeter you can name.
Voice agents and live UX spend their whole budget in the first second. A direct inference provider with no extra hop usually beats any router on speed, whatever the catalogue says.
Per-team budgets, spend caps, traces, caching and one invoice stop being nice-to-have the moment a second team starts calling models on the company card.
When to stay on OpenRouter: you are prototyping, you want one bill across hundreds of models, and a small fee is cheaper than running infrastructure. The planner says so too when that is the honest answer.
The live board re-ranks every alternative as you answer. No email, no gate, and you can share the result with the link in your address bar.
If you already know which wall you hit, start here.
Text, vision, speech, video and OCR behind one OpenAI-compatible key, with uptime monitoring, spend controls and an invoice instead of credits.
An open-source proxy for 100+ providers with virtual keys, budgets and fallbacks — the closest thing to OpenRouter you can run inside your own perimeter.
The same one-key, one-bill shape with cost-aware routing and caching on top. Migration is a base URL and a key, nothing else.
Purpose-built inference hardware for open-weight models. The pick when voice agents or real-time UX own the latency budget.
Aggressive per-token pricing on Llama, Qwen, DeepSeek and friends, on an OpenAI-compatible API with no platform fee layered on top.
Frontier and open models under the AWS contract, IAM boundary and compliance paperwork your procurement team already signed off.
The same data the planner runs on. Filter by what you need, or click a column heading to sort.
| Alternative | Model access | Modalities | Hosting | Routing & failover | Pricing shape |
|---|---|---|---|---|---|
| 400+ text models | Text, some vision | Managed cloud | Automatic provider fallback | Provider price + platform fee, prepaid credits | |
| 400+ models, 30+ providers | Text, vision, speech, video, OCR | Managed cloud, EU region | Failover + live uptime monitoring | Pay per token with volume discounts, invoiced | |
| 200+ open models | Text, image, embeddings | Managed cloud | Single provider, no cross-vendor failover | Per token, plus hourly dedicated GPUs | |
| 100+ open models | Text, vision, audio, image | Managed cloud | Single provider | Per token, plus on-demand GPU hours | |
| ~20 curated models | Text, speech-to-text | Managed cloud | Single provider | Per token, free developer tier | |
| 100+ open models | Text, image, speech, embeddings | Managed cloud | Single provider | Per token, pay as you go | |
| 200+ models | Text, image, video, speech | Managed cloud | Single provider | Per token or per generation | |
| Thousands of community models | Image, video, audio, text | Managed cloud | None | Per second of compute | |
| ~50 models from selected vendors | Text, image, embeddings | Managed cloud, many regions incl. EU | Cross-region inference, no cross-vendor routing | Per token, on your AWS bill; provisioned throughput available | |
| OpenAI + 1000s of catalogue models | Text, image, speech, vision | Managed cloud, EU data boundary | Regional deployments, no cross-vendor routing | Per token or provisioned units, on the Azure bill | |
| Gemini + model garden | Text, image, video, speech | Managed cloud, EU regions | Regional endpoints, no cross-vendor routing | Per token or per media unit, on the GCP bill | |
| Bring your own provider keys | Whatever your providers support | Managed cloud or self-hosted | Conditional routing, retries, fallbacks | Per request tiers, free developer plan, enterprise self-host | |
| 100+ providers via your own keys | Whatever your providers support | Self-hosted (managed tier available) | Load balancing, retries, fallbacks | Free and open source, paid enterprise tier | |
| Bring your own provider keys | Whatever your providers support | Self-hosted or hybrid | Policy-based routing and rate limits | Open source core, enterprise licence | |
| Bring your own keys + Workers AI | Text, image, speech via Workers AI | Managed edge, EU options | Fallbacks, caching, rate limits | Free tier, usage-based extras | |
| Your own model deployments | Whatever you deploy | Your cloud account / VPC | You build it | Platform fee on top of your own compute | |
| Open models via partner providers | Text, image, audio, embeddings | Managed cloud | Provider selection per model | Free credits, then pass-through provider pricing | |
| 150+ models | Text, vision | Managed cloud | Cost-aware routing and fallbacks | Usage fee on top of provider prices | |
| 100+ AI services across vendors | Text, OCR, speech, image, translation | Managed cloud, EU company | Provider fallback per task | Per call, with a monthly plan | |
| Major providers and open models | Text | Managed cloud | Live benchmark-based routing | Provider price plus routing fee |
Facts checked against each vendor's own documentation and pricing pages on 21 September 2026. Model counts move weekly — treat them as orders of magnitude, not guarantees.
What each one actually is, who it suits, and the part their marketing page leaves out.
A marketplace-style router in front of hundreds of hosted models, with a single OpenAI-compatible endpoint, credits-based billing and automatic fallbacks between upstream providers.
A managed router covering 400+ models from 30+ providers behind one key, with per-model uptime monitoring, spend controls, invoices instead of credits, and a white-label router other companies resell under their own brand.
An inference cloud running open-weight models on its own GPU fleet, plus fine-tuning and dedicated capacity for steady workloads.
A performance-focused inference platform for open-weight models, with speculative decoding, LoRA tuning and on-demand deployments.
Custom LPU hardware serving a curated set of open models at very high throughput through an OpenAI-compatible API.
A low-cost inference provider with an OpenAI-compatible endpoint across open text, embedding, image and speech models.
An inference cloud spanning open LLMs plus image, video and speech generation, also available as GPU instances.
A hosting layer for thousands of community-published models, including image, video and audio pipelines, with custom model deployment.
Amazon's managed model service with Anthropic, Meta, Mistral, Amazon and other families, IAM controls, VPC access and regional isolation.
Microsoft's AI platform hosting OpenAI models alongside open and partner models, with private networking, content filters and EU data boundary options.
Google's managed platform for Gemini plus a model garden of partner and open models, with grounding, tuning and regional endpoints.
A gateway that sits between your app and any provider key you bring, adding routing rules, caching, retries, budgets, logs and guardrails. Available managed or self-hosted.
An MIT-licensed Python SDK and proxy that normalises 100+ provider APIs into the OpenAI format, with virtual keys, budgets and logging. Run it in your own cluster.
Kong's gateway extended with AI plugins for multi-provider routing, prompt guards, token rate limiting and semantic caching, deployable on-prem.
A thin edge layer that proxies calls to providers, adding analytics, caching, retries and rate limits, with Workers AI for models hosted by Cloudflare.
A managed Ray platform for serving and fine-tuning models on your own infrastructure, including deployment into your VPC.
Inference providers on the Hub route calls to partners such as Together, Fireworks and Novita behind one token, alongside dedicated endpoints.
A multi-provider routing layer with caching, fallbacks, per-user budgets and analytics aimed at cutting model spend.
An aggregator covering OCR, translation, speech, moderation and generative models from many vendors under a single key and invoice, with an EU base.
A router that scores providers on live latency, throughput and cost, then sends each request to the endpoint that best matches your constraints.
If you scrolled past the planner, start here.
Stay on OpenRouter, or use Hugging Face Inference. The per-request fee is irrelevant at small volume and the catalogue is the widest there is.
LLM API covers all four modalities behind one key with uptime monitoring and a single invoice. Novita AI is the budget alternative if you can accept lighter guarantees.
LiteLLM self-hosted, or Kong AI Gateway if you already run Kong. You keep provider prices with no markup and own the uptime.
AWS Bedrock, Azure AI Foundry or Google Vertex AI — the model spend lands on a cloud contract you already have, with residency and compliance attached.
Ten criteria, applied the same way to every entry, from public documentation and hands-on use.
Latency and uptime vary by region, model and hour, so we did not publish a single speed ranking — any number we quoted would be stale by the time you read it. Support quality is reported, not tested. Where a vendor does not publish a fact, the table says so instead of guessing.
Read the full methodology write-upStill unsure after the planner? These come up in almost every migration call.
For self-hosting, LiteLLM is open source and free to run — you pay only your upstream providers. For hosted, Cloudflare AI Gateway's core features are free and sit in front of your own provider keys.
Usually yes. Most alternatives on this page expose an OpenAI-compatible endpoint, so you change the base URL, the API key and sometimes the model-ID format. The table's pricing and routing columns flag the ones that need more than that.
A gateway routes your requests to providers you already pay and adds control: budgets, caching, logs, failover. An inference provider runs the models itself. Plenty of teams use both, a gateway in front and two or three providers behind it.
Bring Your Own Key: you plug in your own OpenAI, Anthropic or Google keys, so those providers bill you directly and the router charges only for its own layer — or for nothing at all.
When you are prototyping, want the widest text catalogue with no contract, and the per-request fee is smaller than the cost of running anything yourself. The planner says so too when that is the honest answer.
No. LLM API is our own product and is labelled as such wherever it appears. Every other entry is here on merit, scored on the same criteria from public documentation and hands-on use.
LLM API gives you 400+ text, vision, speech and video models behind a single OpenAI-compatible key, with uptime monitoring, spend controls and one invoice. Migration is a base URL and a key.
Start free