We publish a picker at llmapi.ai/openrouter-alternatives that narrows 20 routers, gateways and inference clouds down to the three that fit your workload. This post is the working behind it: what we looked at, how we scored, and where the comparison stops being reliable.
Why people leave OpenRouter in the first place
OpenRouter is a good product and a fair baseline. In practice teams move for four reasons, and they are rarely about the model list.
- Modality. The catalogue is deep in text and thin in speech, video and OCR. Once a product grows a voice feature, a second vendor appears in the stack.
- Billing. Prepaid credits and a per-request fee are fine for a side project and awkward for a finance team that wants an invoice, a PO and a cost centre.
- Governance. Data residency, audit logs, per-team budgets and an SLA with a name on it tend to arrive with the first enterprise customer.
- Control. Some security teams will not accept a third party in the request path at all, which points at a self-hosted gateway rather than a hosted one.
The ten criteria
Every entry was read against the same ten questions, using the vendor's own documentation, pricing page and status page, plus hands-on use where we had an account.
- Model breadth — how many models, and how quickly new releases land.
- Modality coverage — text, vision, speech, video, OCR, embeddings.
- Routing and failover — what happens when an upstream provider degrades.
- Pricing shape — markup, credits, invoices, commitments, hidden minimums.
- Spend controls — budgets, per-key limits, per-team reporting.
- Data residency — regional hosting, EU options, retention policy.
- Self-hosting — whether you can run it yourself, and under which licence.
- White label — whether you can resell it under your own brand.
- Enterprise readiness — SLAs, SOC 2, DPAs, procurement paperwork.
- Migration effort — how close it is to a drop-in base URL swap.
How the picker scores
We did not publish a single overall ranking, because there isn't one. A gateway that is perfect for a bank is the wrong answer for a two-person team shipping a demo. Instead each alternative carries tags on four axes — modality, hosting model, primary priority and workload size — and the picker adds a weight each time one of your answers matches a tag.
Hosting carries the heaviest weight, because it is the only answer that is usually non-negotiable: if your security team said self-hosted, a managed router is not a compromise, it is a rejection. Modality and your stated priority come next, workload size last, since most platforms stretch across sizes. The result is three names with the matched reasons spelled out, plus the two worst fits for your answers, so you can see what the scoring rejected and why.
What we could not measure
Three things are missing on purpose.
- Latency and throughput. These vary by model, region, hour and upstream provider. Any single number would be stale before you read it, and vendor-published benchmarks are marketing. Where speed is the deciding factor, run your own prompt on your own region for a day.
- Real-world uptime. We monitor our own per-model availability and publish it, but we have no comparable instrumentation on other platforms, so we report what routing features exist rather than how often they save you.
- Support quality. Reported from public plan documents, not tested. Community Discord and a named account manager are different products.
Where a vendor does not publish a fact, the comparison says so rather than estimating. Model counts move weekly; treat them as orders of magnitude.
The shortlist at a glance
| Platform | Modalities | Hosting | Pricing shape |
|---|---|---|---|
| OpenRouter | Text, some vision | Managed cloud | Provider price + platform fee, prepaid credits |
| LLM API | Text, vision, speech, video, OCR | Managed cloud, EU region | Pay per token with volume discounts, invoiced |
| Together AI | Text, image, embeddings | Managed cloud | Per token, plus hourly dedicated GPUs |
| Fireworks AI | Text, vision, audio, image | Managed cloud | Per token, plus on-demand GPU hours |
| Groq | Text, speech-to-text | Managed cloud | Per token, free developer tier |
| DeepInfra | Text, image, speech, embeddings | Managed cloud | Per token, pay as you go |
| Novita AI | Text, image, video, speech | Managed cloud | Per token or per generation |
| Replicate | Image, video, audio, text | Managed cloud | Per second of compute |
| AWS Bedrock | Text, image, embeddings | Managed cloud, many regions incl. EU | Per token, on your AWS bill; provisioned throughput available |
| Azure AI Foundry | Text, image, speech, vision | Managed cloud, EU data boundary | Per token or provisioned units, on the Azure bill |
| Google Vertex AI | Text, image, video, speech | Managed cloud, EU regions | Per token or per media unit, on the GCP bill |
| Portkey | Whatever your providers support | Managed cloud or self-hosted | Per request tiers, free developer plan, enterprise self-host |
| LiteLLM | Whatever your providers support | Self-hosted (managed tier available) | Free and open source, paid enterprise tier |
| Kong AI Gateway | Whatever your providers support | Self-hosted or hybrid | Open source core, enterprise licence |
| Cloudflare AI Gateway | Text, image, speech via Workers AI | Managed edge, EU options | Free tier, usage-based extras |
| Anyscale | Whatever you deploy | Your cloud account / VPC | Platform fee on top of your own compute |
| Hugging Face Inference | Text, image, audio, embeddings | Managed cloud | Free credits, then pass-through provider pricing |
| Requesty | Text, vision | Managed cloud | Usage fee on top of provider prices |
| Eden AI | Text, OCR, speech, image, translation | Managed cloud, EU company | Per call, with a monthly plan |
| Unify | Text | Managed cloud | Provider price plus routing fee |
Four situations, four answers
Prototyping with maximum choice. Stay where you are, or add Hugging Face Inference. At small volume the per-request fee is noise and the catalogue breadth is the whole point.
Production across text, voice and video. You want one key that covers all of it, with uptime monitoring and a single invoice, which is the case LLM API was built for. Novita AI is the cheaper alternative if lighter guarantees are acceptable.
Nothing third-party in the data path. Self-hosted LiteLLM, or Kong AI Gateway if your traffic already runs through Kong. You keep provider prices with no markup and you own the uptime.
Procurement has to sign. AWS Bedrock, Azure AI Foundry or Google Vertex AI put the spend on a contract that already exists, with residency and compliance attached. You trade catalogue breadth for paperwork that is already done.
Where we are biased, stated plainly
We build one of the twenty. LLM API appears in the comparison on the same criteria as everyone else, it wins the recommendations where the answers point at mixed modalities, unified billing, failover or white-label reselling, and it loses them where you asked for self-hosting, a free hobby tier or a single-vendor cloud contract. If the honest answer to your situation is LiteLLM on your own cluster, the picker will say LiteLLM.
Take the four-question picker and see which three come back for your workload.