Portkey
The same routing, retries, caching and virtual keys, with an open-source core you can still self-host and a control plane if you would rather not.
We compared 20 gateways, proxies and control planes on who runs them, what they do in the request path and what the layer costs. Answer ten questions and get a shortlist built for your deployment — not a generic top ten.
The live board re-ranks every alternative as you answer. No email, no gate, and you can share the result with the link in your address bar.
If you already know which wall you hit, start here.
The same routing, retries, caching and virtual keys, with an open-source core you can still self-host and a control plane if you would rather not.
A Go gateway built for the concurrency where a Python proxy starts costing you p99, with fallbacks, budgets and a UI in the box.
AI routing declared as Gateway API resources next to the rest of your mesh, so the cluster owns it rather than a service team.
Analytics, caching, retries and fallbacks in front of the provider keys you already hold, with nothing at all to deploy.
Text, vision, speech, video and OCR behind one managed key with unified billing, for stacks that outgrew chat completions.
A commercial gateway that installs into your own VPC with RBAC, audit logs, SSO and a support contract attached.
Pick a name and see the close-up: who runs it, what it does in the request path, what the layer costs and the capabilities where the two genuinely part ways. The other fourteen are in the full table.
One managed OpenAI-compatible endpoint across text, vision, speech and video, with unified billing.
LiteLLM
The baseline you are running today
LLM API
The alternative
Where they differ
Why people pick it One managed OpenAI-compatible endpoint across text, vision, speech and video, with unified billing.
Watch out for No self-hosted edition. If your security review forbids a third party in the request path, this is the wrong row.
The closest like-for-like replacement: an AI gateway with an open-source core and a managed control plane.
LiteLLM
The baseline you are running today
Portkey
The alternative
Where they differ
Why people pick it The closest like-for-like replacement: an AI gateway with an open-source core and a managed control plane.
Watch out for The features most people migrate for — logs, governance, guardrails — live in the paid control plane, not the open-source gateway.
LLM routing as plugins on a gateway your platform team probably already runs.
LiteLLM
The baseline you are running today
Kong AI Gateway
The alternative
Where they differ
Why people pick it LLM routing as plugins on a gateway your platform team probably already runs.
Watch out for Heavier to stand up than a single proxy container, and the interesting AI plugins are Kong Enterprise, not the OSS build.
A gateway with no servers at all, running at the edge in front of your existing keys.
LiteLLM
The baseline you are running today
Cloudflare AI Gateway
The alternative
Where they differ
Why people pick it A gateway with no servers at all, running at the edge in front of your existing keys.
Watch out for Governance features are lighter than a dedicated control plane, and log retention on the free tier runs out fast.
The hosted marketplace people land on when they stop wanting infrastructure.
LiteLLM
The baseline you are running today
OpenRouter
The alternative
Where they differ
Why people pick it The hosted marketplace people land on when they stop wanting infrastructure.
Watch out for A fee on credit purchases, community-first support, thin speech and video coverage, and nothing to self-host.
A Go rewrite of the same idea, aimed squarely at LiteLLM's overhead at high concurrency.
LiteLLM
The baseline you are running today
Bifrost
The alternative
Where they differ
On the twelve capabilities we check, these two cover the same ground. The difference is who runs it and what the layer costs.
Why people pick it A Go rewrite of the same idea, aimed squarely at LiteLLM's overhead at high concurrency.
Watch out for Smaller community and shorter track record than the incumbent, so fewer answered questions when something odd happens at 2am.
The same data the planner runs on. Filter by what you need, or click a column heading to sort.
| Alternative | What it is | Who runs it | Model access | Routing & control | Pricing shape |
|---|---|---|---|---|---|
| The baseline: an open-source Python proxy in front of 100+ providers. | Self-hosted (managed tier available) | 100+ providers | Retries, fallbacks, load balancing | Free and open source; enterprise licence priced on annual request capacity | |
| One managed OpenAI-compatible endpoint across text, vision, speech and video, with unified billing. | Managed cloud, EU region | 400+ models across modalities | Failover, latency routing, caching | Per token with volume discounts, invoiced; no platform fee on tokens | |
| The closest like-for-like replacement: an AI gateway with an open-source core and a managed control plane. | Self-hosted or managed cloud | 250+ models via your keys | Conditional routing, fallbacks, caching, guardrails | Open-source gateway free; hosted control plane from a flat monthly fee | |
| LLM routing as plugins on a gateway your platform team probably already runs. | Self-hosted, hybrid or Konnect cloud | Any provider you configure | Load balancing, retries, semantic caching | Open-source core; enterprise licence for the full AI plugin set | |
| An Apache-2.0 API gateway with AI proxy plugins, no commercial licence anywhere in the path. | Self-hosted | Any provider you configure | Multi-upstream fallback, retries, rate limits | Free and open source, Apache-2.0 | |
| LLM routing as a Kubernetes-native Envoy control plane, built on Gateway API. | Self-hosted on Kubernetes | Any provider you configure | Upstream failover, token rate limiting | Free and open source, CNCF project | |
| A Go rewrite of the same idea, aimed squarely at LiteLLM's overhead at high concurrency. | Self-hosted or managed | Multi-provider via your keys | Automatic fallback, load balancing, caching | Free and open source; paid enterprise tier | |
| Observability first, gateway second — and both can run on your own infrastructure. | Managed cloud or self-hosted | Any provider via proxy or async logging | Caching, rate limits, retries | Free tier, then usage-based; self-host free | |
| Traces, evals and prompt management for teams that only used the proxy as a logging point. | Self-hosted or managed cloud | Instrumented in your own code | None — observability layer | Free self-host; usage-based cloud plans | |
| A gateway with no servers at all, running at the edge in front of your existing keys. | Managed edge, regional options | Your provider keys, plus Workers AI | Fallbacks, retries, caching | Core gateway free; 5% on credits if you use unified billing | |
| One key, hundreds of models, zero infrastructure — if your app already lives on Vercel. | Managed cloud | Hundreds of models, one key | Automatic failover, spend limits | Pass-through model pricing on a unified balance | |
| An enterprise AI gateway that installs into your own VPC or Kubernetes cluster. | Your VPC / Kubernetes, or managed | Any provider, plus your own models | Fallbacks, load balancing, rate limits | Commercial, priced per deployment or usage | |
| Gateway features attached to the lakehouse your data already sits in. | Managed inside your Databricks workspace | Foundation models and your served models | Fallbacks, rate limits, guardrails | Consumption on your Databricks contract | |
| A hosted router with cost-aware model selection and a one-line swap. | Managed cloud | Many providers, one key | Cost-aware routing, caching, fallbacks | Roughly 5% on token spend | |
| Routing decided by live quality, latency and cost benchmarks rather than a static config. | Managed cloud | Many providers and endpoints | Benchmark-driven dynamic routing | Provider price plus a routing fee | |
| The hosted marketplace people land on when they stop wanting infrastructure. | Managed cloud | 500+ models, mostly text | Automatic provider fallback | Provider price plus a fee on credit top-ups | |
| One API across specialist AI providers, not just chat models. | Managed cloud, EU company | Specialist providers across modalities | Provider switching per feature | Per call, with monthly plans and a checkout fee | |
| Frontier and open models inside the AWS account and contract you already have. | Managed cloud, many regions including EU | Frontier and open models on AWS | Prompt routing, guardrails, retries | Per token or provisioned capacity, on your AWS bill | |
| OpenAI and open models under the Microsoft contract, with an EU data boundary. | Managed cloud, EU data boundary | OpenAI plus open-model catalogue | Deployment-level routing, quotas, filters | Per token or provisioned units, on the Azure bill | |
| Gemini, open models and media generation on the Google Cloud contract. | Managed cloud, EU regions | Gemini, open and media models | Regional endpoints, quotas, retries | Per token or per media unit, on the GCP bill |
Facts read from each vendor's own documentation and pricing pages on 22 September 2026. Where something is not published, the comparison says so rather than guessing.
What each one actually is, who it suits, and the part their marketing page leaves out.
An open-source proxy and SDK that normalises 100+ providers onto the OpenAI schema, with virtual keys, budgets, retries and fallbacks. You run it — usually a container, a Postgres and a Redis.
A managed gateway: one key and one balance across text, vision, speech, video and OCR models, with uptime monitoring, spend controls, an invoice instead of credits, and a white-label option.
A gateway written in TypeScript with an Apache-2.0 core you can self-host and a hosted control plane on top: routing, retries, caching, guardrails, budgets, traces and prompt management.
AI plugins on top of Kong Gateway: multi-provider proxying, prompt templates, token rate limiting, semantic caching and guardrails, deployed the same way as the rest of your API traffic.
A high-performance OpenResty gateway with ai-proxy, ai-proxy-multi, token rate limiting and prompt guard plugins, so LLM calls become another upstream under the same config.
A CNCF project that extends Envoy Gateway with an OpenAI-compatible frontend, provider credential handling, token-aware rate limiting and upstream failover, all expressed as Kubernetes resources.
An open-source gateway written in Go with an OpenAI-compatible API, provider fallbacks, key rotation, budgets, MCP support and a built-in UI — marketed on how little latency the hop adds under load.
An open-source LLM observability platform with a gateway in front: logging, tracing, evals, caching, rate limits and cost tracking, available as a hosted service or a self-hosted stack.
An open-source LLM engineering platform: tracing, prompt management, evaluations and cost analytics, self-hostable under a permissive core licence and commonly paired with a thin gateway.
A managed edge gateway: analytics, caching, rate limiting, retries and fallbacks in front of provider keys you already hold, with optional unified billing across providers.
A managed gateway exposing hundreds of models behind one key and one balance, with automatic provider failover, spend limits and first-class support in the AI SDK.
A commercial gateway and MLOps control plane deployed in your cloud account: routing, rate limits, budgets per team, RBAC, audit logs and SSO, with support attached.
A governance layer over model serving endpoints: permissions, rate limits, payload logging, guardrails and usage tracking, managed inside the Databricks workspace.
A managed routing layer over many providers with caching, cost optimisation, fallbacks and spend analytics, reached through an OpenAI-compatible endpoint.
A hosted router that benchmarks endpoints continuously and sends each request to whichever one currently best fits the quality, speed and price targets you set.
A managed router in front of hundreds of hosted models with one OpenAI-compatible endpoint, credit billing and automatic fallback between upstream providers.
A European aggregator covering OCR, document parsing, speech, translation and generative models from dozens of specialist providers behind a single API and one bill.
A managed model service with IAM-scoped access, guardrails, provisioned throughput and, with Bedrock's routing features, model selection per request — all billed on the AWS invoice.
A managed platform for model deployment and routing with content filters, private networking, quota management and an EU data boundary, billed on the Azure subscription.
A managed AI platform with model garden access, regional endpoints, provisioned throughput, safety filters and per-project quotas, billed on the GCP invoice.
Upgrades, a Postgres, a Redis, config drift across environments and a Python process in the hot path of every request. None of it ships product, and all of it pages somebody.
A proxy that is invisible at ten requests a second is measurable at a thousand. Teams hit this and start looking at Go and Rust gateways, or at taking the hop out entirely.
Per-team budgets exist, but audit logs, RBAC, SSO, guardrails and an evidence trail an auditor accepts are a product, not a config file.
Speech, video and OCR arrive, and a proxy designed around chat completions ends up with a second vendor bolted alongside it.
When to stay on LiteLLM: you need provider abstraction inside your own perimeter, your traffic is moderate, and the team that runs it is the team that wanted it. The planner says so too when that is the honest answer.
Self-hosting has no licence fee and a very real floor: servers, a database, a cache and the person who keeps them up. Drag the slider to your monthly token spend and compare the layer, not the tokens.
Layer cost only — tokens are billed on top and cost about the same wherever you route. The self-hosted rows assume roughly $180 a month of infrastructure for a two-replica gateway, a small managed Postgres and a cache; your figure will differ, and neither row prices the engineering time, which is usually the larger number. Vendor fee shapes taken from each vendor's own pricing page, checked 22 September 2026 — the full working is in our comparison write-up.
Every alternative against the twelve capabilities the planner scores on. Filled means the platform covers it as a first-class capability; empty means it does not — no half-credit for “on the roadmap”.
| Alternative | Self-host | Managed | Kubernetes | Edge | BYOK | Unified billing | Failover | Caching | Budgets | Observability | Enterprise | Drop-in swap |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| covered | covered | covered | not covered | covered | not covered | covered | covered | covered | covered | covered | covered | |
| not covered | covered | not covered | not covered | not covered | covered | covered | covered | covered | covered | covered | covered | |
| covered | covered | covered | covered | covered | not covered | covered | covered | covered | covered | covered | covered | |
| covered | covered | covered | not covered | covered | not covered | covered | covered | covered | covered | covered | not covered | |
| covered | not covered | covered | not covered | covered | not covered | covered | covered | not covered | not covered | covered | not covered | |
| covered | not covered | covered | not covered | covered | not covered | covered | not covered | covered | covered | covered | covered | |
| covered | covered | covered | not covered | covered | not covered | covered | covered | covered | covered | covered | covered | |
| covered | covered | covered | covered | covered | not covered | covered | covered | covered | covered | covered | covered | |
| covered | covered | covered | not covered | covered | not covered | not covered | not covered | covered | covered | covered | not covered | |
| not covered | covered | not covered | covered | covered | covered | covered | covered | covered | covered | covered | covered | |
| not covered | covered | not covered | covered | covered | covered | covered | covered | covered | covered | not covered | covered | |
| covered | covered | covered | not covered | covered | not covered | covered | covered | covered | covered | covered | covered | |
| not covered | covered | not covered | not covered | covered | not covered | covered | not covered | covered | covered | covered | not covered | |
| not covered | covered | not covered | not covered | covered | covered | covered | covered | covered | covered | not covered | covered | |
| not covered | covered | not covered | not covered | covered | covered | covered | not covered | covered | covered | not covered | covered | |
| not covered | covered | not covered | not covered | covered | covered | covered | covered | covered | covered | not covered | covered | |
| not covered | covered | not covered | not covered | covered | covered | covered | not covered | covered | covered | covered | not covered | |
| not covered | covered | not covered | not covered | not covered | not covered | covered | covered | covered | covered | covered | not covered | |
| not covered | covered | not covered | not covered | not covered | not covered | covered | covered | covered | covered | covered | covered | |
| not covered | covered | not covered | not covered | not covered | not covered | covered | covered | covered | covered | covered | not covered |
If you scrolled past the planner, start here.
Portkey does the closest like-for-like job, self-hosted or hosted. Bifrost is the pick if the complaint is specifically Python overhead.
Kong AI Gateway or Envoy AI Gateway put AI traffic under the policy, auth and telemetry you already operate, instead of beside it.
Cloudflare AI Gateway keeps your own provider keys with nothing to deploy. LLM API is the answer when speech, video or one invoice also matter.
TrueFoundry in your own VPC, or AWS Bedrock, Azure AI Foundry or Google Vertex AI if the spend belongs on a cloud contract you already hold.
Twelve capabilities and ten planner axes, applied the same way to every entry, from public documentation and hands-on use.
Gateway overhead varies by deployment, hardware and traffic shape, so we did not publish a single latency ranking — benchmark it on your own cluster with your own prompts. Support quality is reported from plan documents, not tested. Where a vendor does not publish a fact, the comparison says so instead of guessing.
Read the full methodology write-upStill unsure after the planner? These come up in almost every migration call.
Portkey's gateway core and Bifrost are the closest open-source like-for-like replacements. If the gateway belongs to your platform team rather than an application team, Kong AI Gateway and Envoy AI Gateway do the same routing job as part of infrastructure you already run, and Apache APISIX keeps everything in the path permissively licensed.
Usually yes. Most alternatives here expose an OpenAI-compatible endpoint, so you change a base URL, a key and sometimes the model-ID format. The exceptions are the hyperscaler platforms and Databricks, where you adopt their SDK and IAM model instead.
There is no licence fee, but there is a floor: a couple of gateway replicas, a Postgres, a cache and somebody on call. Around $150–$250 a month of infrastructure is typical before any engineering time. Below roughly $1,000 a month of token spend, a flat-fee managed control plane is usually cheaper than the servers alone.
The Python proxy adds real latency and CPU at high concurrency, which is why Go-based gateways like Bifrost and Envoy-based routing exist. At modest request rates it is irrelevant; measure yours before it drives the decision.
When you need provider abstraction inside your own perimeter, your traffic is moderate, and the team running it is the team that wanted it. The planner says so too when that is the honest answer.
No. LLM API is our own product and is labelled as such wherever it appears. Every other entry is here on merit, scored on the same criteria from public documentation and hands-on use.
LLM API gives you 400+ text, vision, speech and video models behind a single OpenAI-compatible key, with uptime monitoring, spend controls and one invoice. Migration from a self-hosted proxy is a base URL and a key.
Start free