Bonus: Top up now and we'll double your first deposit — get x2 credits instantly.
OpenRouter alternatives · updated 21 September 2026

OpenRouter got you started. Here is where to route next.

We compared 20 routers, gateways and inference clouds on cost, control, latency and compliance. Answer six questions and get a shortlist built for your stack — not a generic top ten.

20alternatives compared
6questions in the planner
10criteria applied to each
0paid placements
Why teams look elsewhere

Four reasons people outgrow a single router

~5%

A fee on every token

A platform margin on top of provider prices is invisible at prototype volume and a real line item at production volume. Routers that bill provider rates, or let you bring your own keys, close that gap.

EU

Prompts that cannot travel

Once legal reviews the data path, a US-only proxy in front of your prompts stops being acceptable. Regional hosting, self-hosted gateways and cloud platforms keep traffic inside a perimeter you can name.

p95

Latency budgets

Voice agents and live UX spend their whole budget in the first second. A direct inference provider with no extra hop usually beats any router on speed, whatever the catalogue says.

360°

Control and visibility

Per-team budgets, spend caps, traces, caching and one invoice stop being nice-to-have the moment a second team starts calling models on the company card.

When to stay on OpenRouter: you are prototyping, you want one bill across hundreds of models, and a small fee is cheaper than running infrastructure. The planner says so too when that is the honest answer.

Route planner

Six questions to your shortlist

The live board re-ranks every alternative as you answer. No email, no gate, and you can share the result with the link in your address bar.

Question 1 of 6
What are you building?

This sets how much operational weight you can reasonably carry.

Editor's picks

The short answer, by need

If you already know which wall you hit, start here.

Best for mixed modalities

LLM API

Text, vision, speech, video and OCR behind one OpenAI-compatible key, with uptime monitoring, spend controls and an invoice instead of credits.

Read the profileRunner-up: Google Vertex AI
Best self-hosted

LiteLLM

An open-source proxy for 100+ providers with virtual keys, budgets and fallbacks — the closest thing to OpenRouter you can run inside your own perimeter.

Read the profileRunner-up: Kong AI Gateway
Best drop-in swap

Requesty

The same one-key, one-bill shape with cost-aware routing and caching on top. Migration is a base URL and a key, nothing else.

Read the profileRunner-up: Unify
Lowest latency

Groq

Purpose-built inference hardware for open-weight models. The pick when voice agents or real-time UX own the latency budget.

Read the profileRunner-up: Fireworks AI
Cheapest open models

DeepInfra

Aggressive per-token pricing on Llama, Qwen, DeepSeek and friends, on an OpenAI-compatible API with no platform fee layered on top.

Read the profileRunner-up: Together AI
Best for enterprise

AWS Bedrock

Frontier and open models under the AWS contract, IAM boundary and compliance paperwork your procurement team already signed off.

Read the profileRunner-up: Azure AI Foundry

Every alternative, side by side

The same data the planner runs on. Filter by what you need, or click a column heading to sort.

AlternativeModel accessModalitiesHostingRouting & failoverPricing shape
OpenRouterbaseline 400+ text models Text, some vision Managed cloud Automatic provider fallback Provider price + platform fee, prepaid credits
LLM APIthat’s us 400+ models, 30+ providers Text, vision, speech, video, OCR Managed cloud, EU region Failover + live uptime monitoring Pay per token with volume discounts, invoiced
Together AI 200+ open models Text, image, embeddings Managed cloud Single provider, no cross-vendor failover Per token, plus hourly dedicated GPUs
Fireworks AI 100+ open models Text, vision, audio, image Managed cloud Single provider Per token, plus on-demand GPU hours
Groq ~20 curated models Text, speech-to-text Managed cloud Single provider Per token, free developer tier
DeepInfra 100+ open models Text, image, speech, embeddings Managed cloud Single provider Per token, pay as you go
Novita AI 200+ models Text, image, video, speech Managed cloud Single provider Per token or per generation
Replicate Thousands of community models Image, video, audio, text Managed cloud None Per second of compute
AWS Bedrock ~50 models from selected vendors Text, image, embeddings Managed cloud, many regions incl. EU Cross-region inference, no cross-vendor routing Per token, on your AWS bill; provisioned throughput available
Azure AI Foundry OpenAI + 1000s of catalogue models Text, image, speech, vision Managed cloud, EU data boundary Regional deployments, no cross-vendor routing Per token or provisioned units, on the Azure bill
Google Vertex AI Gemini + model garden Text, image, video, speech Managed cloud, EU regions Regional endpoints, no cross-vendor routing Per token or per media unit, on the GCP bill
Portkey Bring your own provider keys Whatever your providers support Managed cloud or self-hosted Conditional routing, retries, fallbacks Per request tiers, free developer plan, enterprise self-host
LiteLLM 100+ providers via your own keys Whatever your providers support Self-hosted (managed tier available) Load balancing, retries, fallbacks Free and open source, paid enterprise tier
Kong AI Gateway Bring your own provider keys Whatever your providers support Self-hosted or hybrid Policy-based routing and rate limits Open source core, enterprise licence
Cloudflare AI Gateway Bring your own keys + Workers AI Text, image, speech via Workers AI Managed edge, EU options Fallbacks, caching, rate limits Free tier, usage-based extras
Anyscale Your own model deployments Whatever you deploy Your cloud account / VPC You build it Platform fee on top of your own compute
Hugging Face Inference Open models via partner providers Text, image, audio, embeddings Managed cloud Provider selection per model Free credits, then pass-through provider pricing
Requesty 150+ models Text, vision Managed cloud Cost-aware routing and fallbacks Usage fee on top of provider prices
Eden AI 100+ AI services across vendors Text, OCR, speech, image, translation Managed cloud, EU company Provider fallback per task Per call, with a monthly plan
Unify Major providers and open models Text Managed cloud Live benchmark-based routing Provider price plus routing fee

Facts checked against each vendor's own documentation and pricing pages on 21 September 2026. Model counts move weekly — treat them as orders of magnitude, not guarantees.

The 20 profiles

What each one actually is, who it suits, and the part their marketing page leaves out.

OpenRouter

A marketplace-style router in front of hundreds of hosted models, with a single OpenAI-compatible endpoint, credits-based billing and automatic fallbacks between upstream providers.

Best for
Trying many models quickly, side projects, teams that want the widest text-model catalogue without contracts.
Watch out for
A per-request fee sits on top of provider prices, support is community-first, and speech, video and OCR coverage is thin compared with the text catalogue.
Pricing shape
Provider price + platform fee, prepaid credits

LLM API

A managed router covering 400+ models from 30+ providers behind one key, with per-model uptime monitoring, spend controls, invoices instead of credits, and a white-label router other companies resell under their own brand.

Best for
Teams running mixed modalities in production, agencies and platforms that need one invoice, and companies reselling AI access.
Watch out for
No self-hosted edition — the gateway runs as a managed service.
Pricing shape
Pay per token with volume discounts, invoiced

Together AI

An inference cloud running open-weight models on its own GPU fleet, plus fine-tuning and dedicated capacity for steady workloads.

Best for
Open-model workloads at volume where per-token price and throughput matter more than closed-model access.
Watch out for
No access to closed frontier models from OpenAI or Anthropic through the same endpoint.
Pricing shape
Per token, plus hourly dedicated GPUs

Fireworks AI

A performance-focused inference platform for open-weight models, with speculative decoding, LoRA tuning and on-demand deployments.

Best for
Latency-sensitive production apps built on open models.
Watch out for
Catalogue is narrower than a marketplace router and closed frontier models are not included.
Pricing shape
Per token, plus on-demand GPU hours

Groq

Custom LPU hardware serving a curated set of open models at very high throughput through an OpenAI-compatible API.

Best for
Voice agents, streaming chat and anything where response speed is the product.
Watch out for
Small catalogue, capacity limits at peak, and no cross-provider routing.
Pricing shape
Per token, free developer tier

DeepInfra

A low-cost inference provider with an OpenAI-compatible endpoint across open text, embedding, image and speech models.

Best for
Cost-sensitive batch and background workloads.
Watch out for
Fewer enterprise controls, lighter SLAs, no closed frontier models.
Pricing shape
Per token, pay as you go

Novita AI

An inference cloud spanning open LLMs plus image, video and speech generation, also available as GPU instances.

Best for
Media-heavy products that also need text models at low cost.
Watch out for
Quality of service varies by model; enterprise paperwork is lighter than a hyperscaler's.
Pricing shape
Per token or per generation

Replicate

A hosting layer for thousands of community-published models, including image, video and audio pipelines, with custom model deployment.

Best for
Generative media, experimentation and running your own containerised model.
Watch out for
Cold starts, per-second billing that is hard to forecast, and inconsistent model maintenance.
Pricing shape
Per second of compute

AWS Bedrock

Amazon's managed model service with Anthropic, Meta, Mistral, Amazon and other families, IAM controls, VPC access and regional isolation.

Best for
Enterprises already on AWS that need procurement, data residency and compliance more than catalogue breadth.
Watch out for
Model availability differs by region, quotas need raising, and the developer experience is heavier than a router.
Pricing shape
Per token, on your AWS bill; provisioned throughput available

Azure AI Foundry

Microsoft's AI platform hosting OpenAI models alongside open and partner models, with private networking, content filters and EU data boundary options.

Best for
Microsoft-committed enterprises needing OpenAI models with corporate compliance.
Watch out for
Deployment quotas per region, and the catalogue outside OpenAI is uneven.
Pricing shape
Per token or provisioned units, on the Azure bill

Google Vertex AI

Google's managed platform for Gemini plus a model garden of partner and open models, with grounding, tuning and regional endpoints.

Best for
Teams standardising on Gemini, or needing long-context and multimodal models under GCP contracts.
Watch out for
Authentication and project setup are heavier than an API key; pricing differs per modality.
Pricing shape
Per token or per media unit, on the GCP bill

Portkey

A gateway that sits between your app and any provider key you bring, adding routing rules, caching, retries, budgets, logs and guardrails. Available managed or self-hosted.

Best for
Platform teams that already hold provider contracts and need control, tracing and spend limits.
Watch out for
You still buy the models elsewhere — it is a control plane, not a catalogue.
Pricing shape
Per request tiers, free developer plan, enterprise self-host

LiteLLM

An MIT-licensed Python SDK and proxy that normalises 100+ provider APIs into the OpenAI format, with virtual keys, budgets and logging. Run it in your own cluster.

Best for
Engineering teams that want full control, no vendor in the data path, and provider prices with no markup.
Watch out for
You own uptime, upgrades and scaling; the enterprise features sit behind a paid tier.
Pricing shape
Free and open source, paid enterprise tier

Kong AI Gateway

Kong's gateway extended with AI plugins for multi-provider routing, prompt guards, token rate limiting and semantic caching, deployable on-prem.

Best for
Enterprises that already standardise API traffic on Kong and want AI under the same policy layer.
Watch out for
Infrastructure-team territory: no model catalogue, no consumer-grade onboarding.
Pricing shape
Open source core, enterprise licence

Cloudflare AI Gateway

A thin edge layer that proxies calls to providers, adding analytics, caching, retries and rate limits, with Workers AI for models hosted by Cloudflare.

Best for
Teams on Cloudflare who want cheap observability and caching without changing providers.
Watch out for
Not a catalogue or a billing layer; governance features are lighter than dedicated gateways.
Pricing shape
Free tier, usage-based extras

Anyscale

A managed Ray platform for serving and fine-tuning models on your own infrastructure, including deployment into your VPC.

Best for
ML teams with custom models and existing cloud commitments.
Watch out for
Heaviest setup on this list; no ready-made catalogue or router.
Pricing shape
Platform fee on top of your own compute

Hugging Face Inference

Inference providers on the Hub route calls to partners such as Together, Fireworks and Novita behind one token, alongside dedicated endpoints.

Best for
Open-model research, prototypes and anything already anchored on the Hub.
Watch out for
No closed frontier models, and production guarantees depend on the upstream partner.
Pricing shape
Free credits, then pass-through provider pricing

Requesty

A multi-provider routing layer with caching, fallbacks, per-user budgets and analytics aimed at cutting model spend.

Best for
Teams whose main pain is an unpredictable model bill.
Watch out for
Younger platform with a smaller catalogue and shorter operating history.
Pricing shape
Usage fee on top of provider prices

Eden AI

An aggregator covering OCR, translation, speech, moderation and generative models from many vendors under a single key and invoice, with an EU base.

Best for
Products that need document, speech and translation AI alongside LLMs.
Watch out for
Task-oriented rather than model-oriented; per-call pricing can exceed direct provider rates.
Pricing shape
Per call, with a monthly plan

Unify

A router that scores providers on live latency, throughput and cost, then sends each request to the endpoint that best matches your constraints.

Best for
Engineers who want routing decisions made on measurements rather than a fixed preference list.
Watch out for
Smaller catalogue and community; best suited to text workloads.
Pricing shape
Provider price plus routing fee

Short version, by situation

If you scrolled past the planner, start here.

You are prototyping and want maximum choice

Stay on OpenRouter, or use Hugging Face Inference. The per-request fee is irrelevant at small volume and the catalogue is the widest there is.

You are in production across text, voice and video

LLM API covers all four modalities behind one key with uptime monitoring and a single invoice. Novita AI is the budget alternative if you can accept lighter guarantees.

Your security team wants nothing in the data path

LiteLLM self-hosted, or Kong AI Gateway if you already run Kong. You keep provider prices with no markup and own the uptime.

Procurement has to sign off

AWS Bedrock, Azure AI Foundry or Google Vertex AI — the model spend lands on a cloud contract you already have, with residency and compliance attached.

How we compared them

Ten criteria, applied the same way to every entry, from public documentation and hands-on use.

  • Model breadth and how fast new releases land
  • Modality coverage: text, vision, speech, video, OCR
  • Routing, retries and failover behaviour
  • Pricing shape: markup, credits, invoices, commitments
  • Spend controls, budgets and per-team reporting
  • Data residency and regional hosting
  • Self-hosting and the licence that comes with it
  • White-label and reseller support
  • Enterprise readiness: SLAs, SOC 2, procurement
  • Migration effort from an OpenAI-compatible endpoint

What we could not measure

Latency and uptime vary by region, model and hour, so we did not publish a single speed ranking — any number we quoted would be stale by the time you read it. Support quality is reported, not tested. Where a vendor does not publish a fact, the table says so instead of guessing.

Read the full methodology write-up
FAQ

Questions we get a lot

Still unsure after the planner? These come up in almost every migration call.

What is the best free OpenRouter alternative?

For self-hosting, LiteLLM is open source and free to run — you pay only your upstream providers. For hosted, Cloudflare AI Gateway's core features are free and sit in front of your own provider keys.

Can I switch without rewriting code?

Usually yes. Most alternatives on this page expose an OpenAI-compatible endpoint, so you change the base URL, the API key and sometimes the model-ID format. The table's pricing and routing columns flag the ones that need more than that.

Gateway or inference provider — what is the difference?

A gateway routes your requests to providers you already pay and adds control: budgets, caching, logs, failover. An inference provider runs the models itself. Plenty of teams use both, a gateway in front and two or three providers behind it.

What does BYOK mean?

Bring Your Own Key: you plug in your own OpenAI, Anthropic or Google keys, so those providers bill you directly and the router charges only for its own layer — or for nothing at all.

When should I just stay on OpenRouter?

When you are prototyping, want the widest text catalogue with no contract, and the per-request fee is smaller than the cost of running anything yourself. The planner says so too when that is the honest answer.

Is anyone paying to be on this list?

No. LLM API is our own product and is labelled as such wherever it appears. Every other entry is here on merit, scored on the same criteria from public documentation and hands-on use.

Try the one built for mixed workloads

LLM API gives you 400+ text, vision, speech and video models behind a single OpenAI-compatible key, with uptime monitoring, spend controls and one invoice. Migration is a base URL and a key.

Start free