Bonus: Top up now and we'll double your first deposit — get x2 credits instantly.
LiteLLM alternatives · updated 22 September 2026

LiteLLM got you this far. Here is what replaces it.

We compared 20 gateways, proxies and control planes on who runs them, what they do in the request path and what the layer costs. Answer ten questions and get a shortlist built for your deployment — not a generic top ten.

20alternatives compared
10questions in the planner
12capabilities checked on each
0paid placements
Route planner

Ten questions to your shortlist

The live board re-ranks every alternative as you answer. No email, no gate, and you can share the result with the link in your address bar.

Question 1 of 10
Who should run the gateway?

The single biggest fork on this page.

Editor's picks

The short answer, by need

If you already know which wall you hit, start here.

Closest like-for-like

Portkey

The same routing, retries, caching and virtual keys, with an open-source core you can still self-host and a control plane if you would rather not.

Read the profileRunner-up: Bifrost
Best if you keep self-hosting

Bifrost

A Go gateway built for the concurrency where a Python proxy starts costing you p99, with fallbacks, budgets and a UI in the box.

Read the profileRunner-up: Envoy AI Gateway
Best for Kubernetes platforms

Envoy AI Gateway

AI routing declared as Gateway API resources next to the rest of your mesh, so the cluster owns it rather than a service team.

Read the profileRunner-up: Kong AI Gateway
Best zero-ops swap

Cloudflare AI Gateway

Analytics, caching, retries and fallbacks in front of the provider keys you already hold, with nothing at all to deploy.

Read the profileRunner-up: Vercel AI Gateway
Best when modalities arrive

LLM API

Text, vision, speech, video and OCR behind one managed key with unified billing, for stacks that outgrew chat completions.

Read the profileRunner-up: Eden AI
Best for enterprise governance

TrueFoundry

A commercial gateway that installs into your own VPC with RBAC, audit logs, SSO and a support contract attached.

Read the profileRunner-up: AWS Bedrock
Head to head

LiteLLM vs the six people actually weigh it against

Pick a name and see the close-up: who runs it, what it does in the request path, what the layer costs and the capabilities where the two genuinely part ways. The other fourteen are in the full table.

LiteLLM vs LLM API

One managed OpenAI-compatible endpoint across text, vision, speech and video, with unified billing.

LiteLLM

The baseline you are running today

Who runs it
Self-hosted (managed tier available)
Model access
100+ providers
In the request path
Retries, fallbacks, load balancing
Pricing shape
Free and open source; enterprise licence priced on annual request capacity

LLM API

The alternative

Who runs it
Managed cloud, EU region
Model access
400+ models across modalities
In the request path
Failover, latency routing, caching
Pricing shape
Per token with volume discounts, invoiced; no platform fee on tokens

Where they differ

  • CapabilityLiteLLMLLM API
  • Self-hostyesno
  • Kubernetesyesno
  • BYOKyesno
  • Unified billingnoyes

Why people pick it One managed OpenAI-compatible endpoint across text, vision, speech and video, with unified billing.

Watch out for No self-hosted edition. If your security review forbids a third party in the request path, this is the wrong row.

Every alternative, side by side

The same data the planner runs on. Filter by what you need, or click a column heading to sort.

AlternativeWhat it isWho runs itModel accessRouting & controlPricing shape
LiteLLMbaseline The baseline: an open-source Python proxy in front of 100+ providers. Self-hosted (managed tier available) 100+ providers Retries, fallbacks, load balancing Free and open source; enterprise licence priced on annual request capacity
LLM APIours One managed OpenAI-compatible endpoint across text, vision, speech and video, with unified billing. Managed cloud, EU region 400+ models across modalities Failover, latency routing, caching Per token with volume discounts, invoiced; no platform fee on tokens
Portkey The closest like-for-like replacement: an AI gateway with an open-source core and a managed control plane. Self-hosted or managed cloud 250+ models via your keys Conditional routing, fallbacks, caching, guardrails Open-source gateway free; hosted control plane from a flat monthly fee
Kong AI Gateway LLM routing as plugins on a gateway your platform team probably already runs. Self-hosted, hybrid or Konnect cloud Any provider you configure Load balancing, retries, semantic caching Open-source core; enterprise licence for the full AI plugin set
Apache APISIX An Apache-2.0 API gateway with AI proxy plugins, no commercial licence anywhere in the path. Self-hosted Any provider you configure Multi-upstream fallback, retries, rate limits Free and open source, Apache-2.0
Envoy AI Gateway LLM routing as a Kubernetes-native Envoy control plane, built on Gateway API. Self-hosted on Kubernetes Any provider you configure Upstream failover, token rate limiting Free and open source, CNCF project
Bifrost A Go rewrite of the same idea, aimed squarely at LiteLLM's overhead at high concurrency. Self-hosted or managed Multi-provider via your keys Automatic fallback, load balancing, caching Free and open source; paid enterprise tier
Helicone Observability first, gateway second — and both can run on your own infrastructure. Managed cloud or self-hosted Any provider via proxy or async logging Caching, rate limits, retries Free tier, then usage-based; self-host free
Langfuse Traces, evals and prompt management for teams that only used the proxy as a logging point. Self-hosted or managed cloud Instrumented in your own code None — observability layer Free self-host; usage-based cloud plans
Cloudflare AI Gateway A gateway with no servers at all, running at the edge in front of your existing keys. Managed edge, regional options Your provider keys, plus Workers AI Fallbacks, retries, caching Core gateway free; 5% on credits if you use unified billing
Vercel AI Gateway One key, hundreds of models, zero infrastructure — if your app already lives on Vercel. Managed cloud Hundreds of models, one key Automatic failover, spend limits Pass-through model pricing on a unified balance
TrueFoundry An enterprise AI gateway that installs into your own VPC or Kubernetes cluster. Your VPC / Kubernetes, or managed Any provider, plus your own models Fallbacks, load balancing, rate limits Commercial, priced per deployment or usage
Databricks Mosaic AI Gateway Gateway features attached to the lakehouse your data already sits in. Managed inside your Databricks workspace Foundation models and your served models Fallbacks, rate limits, guardrails Consumption on your Databricks contract
Requesty A hosted router with cost-aware model selection and a one-line swap. Managed cloud Many providers, one key Cost-aware routing, caching, fallbacks Roughly 5% on token spend
Unify Routing decided by live quality, latency and cost benchmarks rather than a static config. Managed cloud Many providers and endpoints Benchmark-driven dynamic routing Provider price plus a routing fee
OpenRouter The hosted marketplace people land on when they stop wanting infrastructure. Managed cloud 500+ models, mostly text Automatic provider fallback Provider price plus a fee on credit top-ups
Eden AI One API across specialist AI providers, not just chat models. Managed cloud, EU company Specialist providers across modalities Provider switching per feature Per call, with monthly plans and a checkout fee
AWS Bedrock Frontier and open models inside the AWS account and contract you already have. Managed cloud, many regions including EU Frontier and open models on AWS Prompt routing, guardrails, retries Per token or provisioned capacity, on your AWS bill
Azure AI Foundry OpenAI and open models under the Microsoft contract, with an EU data boundary. Managed cloud, EU data boundary OpenAI plus open-model catalogue Deployment-level routing, quotas, filters Per token or provisioned units, on the Azure bill
Google Vertex AI Gemini, open models and media generation on the Google Cloud contract. Managed cloud, EU regions Gemini, open and media models Regional endpoints, quotas, retries Per token or per media unit, on the GCP bill

Facts read from each vendor's own documentation and pricing pages on 22 September 2026. Where something is not published, the comparison says so rather than guessing.

The 20 profiles

What each one actually is, who it suits, and the part their marketing page leaves out.

LiteLLM

An open-source proxy and SDK that normalises 100+ providers onto the OpenAI schema, with virtual keys, budgets, retries and fallbacks. You run it — usually a container, a Postgres and a Redis.

Best for
Teams that want provider abstraction inside their own perimeter and are happy to own the deployment.
Watch out for
You own uptime, upgrades, the database and the config sprawl. Python proxy overhead shows up under heavy concurrency, and the enterprise features sit behind a licence.
Pricing shape
Free and open source; enterprise licence priced on annual request capacity

LLM API

A managed gateway: one key and one balance across text, vision, speech, video and OCR models, with uptime monitoring, spend controls, an invoice instead of credits, and a white-label option.

Best for
Teams tired of running the proxy who want the same abstraction plus modalities LiteLLM has no endpoint for.
Watch out for
No self-hosted edition. If your security review forbids a third party in the request path, this is the wrong row.
Pricing shape
Per token with volume discounts, invoiced; no platform fee on tokens

Portkey

A gateway written in TypeScript with an Apache-2.0 core you can self-host and a hosted control plane on top: routing, retries, caching, guardrails, budgets, traces and prompt management.

Best for
Teams that want LiteLLM's job done by someone else without giving up the option to run it themselves.
Watch out for
The features most people migrate for — logs, governance, guardrails — live in the paid control plane, not the open-source gateway.
Pricing shape
Open-source gateway free; hosted control plane from a flat monthly fee

Kong AI Gateway

AI plugins on top of Kong Gateway: multi-provider proxying, prompt templates, token rate limiting, semantic caching and guardrails, deployed the same way as the rest of your API traffic.

Best for
Platform teams that already operate Kong and want AI traffic under the same policy, auth and observability.
Watch out for
Heavier to stand up than a single proxy container, and the interesting AI plugins are Kong Enterprise, not the OSS build.
Pricing shape
Open-source core; enterprise licence for the full AI plugin set

Apache APISIX

A high-performance OpenResty gateway with ai-proxy, ai-proxy-multi, token rate limiting and prompt guard plugins, so LLM calls become another upstream under the same config.

Best for
Teams with a hard requirement that everything in the request path is permissively licensed and self-run.
Watch out for
No product-grade LLM console: budgets, traces and spend reporting are yours to build on top of the plugin metrics.
Pricing shape
Free and open source, Apache-2.0

Envoy AI Gateway

A CNCF project that extends Envoy Gateway with an OpenAI-compatible frontend, provider credential handling, token-aware rate limiting and upstream failover, all expressed as Kubernetes resources.

Best for
Platform teams already running Envoy or a service mesh who want AI routing declared in the cluster, not in a Python process.
Watch out for
Kubernetes only, young, and everything is YAML — there is no dashboard for a product manager to look at.
Pricing shape
Free and open source, CNCF project

Bifrost

An open-source gateway written in Go with an OpenAI-compatible API, provider fallbacks, key rotation, budgets, MCP support and a built-in UI — marketed on how little latency the hop adds under load.

Best for
Teams whose only complaint about LiteLLM is what the Python proxy costs them in p99 and CPU at high request rates.
Watch out for
Smaller community and shorter track record than the incumbent, so fewer answered questions when something odd happens at 2am.
Pricing shape
Free and open source; paid enterprise tier

Helicone

An open-source LLM observability platform with a gateway in front: logging, tracing, evals, caching, rate limits and cost tracking, available as a hosted service or a self-hosted stack.

Best for
Teams who moved to LiteLLM mostly for logs and cost attribution and never really wanted to run a router.
Watch out for
Routing logic is thinner than a dedicated gateway's, and self-hosting the full stack is more moving parts than one proxy container.
Pricing shape
Free tier, then usage-based; self-host free

Langfuse

An open-source LLM engineering platform: tracing, prompt management, evaluations and cost analytics, self-hostable under a permissive core licence and commonly paired with a thin gateway.

Best for
Teams replacing the observability half of LiteLLM and keeping direct provider SDKs for the calls themselves.
Watch out for
Not a gateway. There is no request routing, no failover and no virtual keys — you need something else in the path.
Pricing shape
Free self-host; usage-based cloud plans

Cloudflare AI Gateway

A managed edge gateway: analytics, caching, rate limiting, retries and fallbacks in front of provider keys you already hold, with optional unified billing across providers.

Best for
Teams whose real goal is to stop operating a proxy and who are happy for the hop to be someone else's edge.
Watch out for
Governance features are lighter than a dedicated control plane, and log retention on the free tier runs out fast.
Pricing shape
Core gateway free; 5% on credits if you use unified billing

Vercel AI Gateway

A managed gateway exposing hundreds of models behind one key and one balance, with automatic provider failover, spend limits and first-class support in the AI SDK.

Best for
Product teams shipping on Next.js who want the abstraction without a deployment of their own.
Watch out for
Strongly tied to the Vercel account and ecosystem, and there is nothing to self-host if the data path becomes a compliance question.
Pricing shape
Pass-through model pricing on a unified balance

TrueFoundry

A commercial gateway and MLOps control plane deployed in your cloud account: routing, rate limits, budgets per team, RBAC, audit logs and SSO, with support attached.

Best for
Enterprises that need the proxy inside their perimeter but cannot justify owning the code that runs it.
Watch out for
Commercial pricing and a sales process; heavier than a single container if all you wanted was routing.
Pricing shape
Commercial, priced per deployment or usage

Databricks Mosaic AI Gateway

A governance layer over model serving endpoints: permissions, rate limits, payload logging, guardrails and usage tracking, managed inside the Databricks workspace.

Best for
Teams whose prompts and evaluation data already live in Databricks and who want one governance boundary.
Watch out for
Only makes sense if you are a Databricks customer; outside that context it is not a general-purpose gateway.
Pricing shape
Consumption on your Databricks contract

Requesty

A managed routing layer over many providers with caching, cost optimisation, fallbacks and spend analytics, reached through an OpenAI-compatible endpoint.

Best for
Small teams that want the proxy gone today and will accept a percentage fee for it.
Watch out for
A fee on spend, no self-hosted edition, and governance depth well short of an enterprise control plane.
Pricing shape
Roughly 5% on token spend

Unify

A hosted router that benchmarks endpoints continuously and sends each request to whichever one currently best fits the quality, speed and price targets you set.

Best for
Teams whose fallback config in LiteLLM has become a hand-maintained spreadsheet of endpoints.
Watch out for
Text-centric, hosted only, and dynamic routing makes cost and behaviour harder to predict request by request.
Pricing shape
Provider price plus a routing fee

OpenRouter

A managed router in front of hundreds of hosted models with one OpenAI-compatible endpoint, credit billing and automatic fallback between upstream providers.

Best for
Prototypes and side projects that need the widest text catalogue and no servers at all.
Watch out for
A fee on credit purchases, community-first support, thin speech and video coverage, and nothing to self-host.
Pricing shape
Provider price plus a fee on credit top-ups

Eden AI

A European aggregator covering OCR, document parsing, speech, translation and generative models from dozens of specialist providers behind a single API and one bill.

Best for
Products whose AI surface is document and speech processing rather than chat completions.
Watch out for
Per-call pricing with a checkout fee, and it is not a gateway you can put in front of your own provider keys.
Pricing shape
Per call, with monthly plans and a checkout fee

AWS Bedrock

A managed model service with IAM-scoped access, guardrails, provisioned throughput and, with Bedrock's routing features, model selection per request — all billed on the AWS invoice.

Best for
Organisations whose gateway requirement is really a procurement, residency and IAM requirement.
Watch out for
Catalogue is narrower than a router's, cross-cloud routing is not the point, and the console is nobody's idea of a developer gateway.
Pricing shape
Per token or provisioned capacity, on your AWS bill

Azure AI Foundry

A managed platform for model deployment and routing with content filters, private networking, quota management and an EU data boundary, billed on the Azure subscription.

Best for
Microsoft-committed enterprises where the data path has to stay inside an existing agreement.
Watch out for
Quota and region juggling is real work, and multi-cloud routing is explicitly not what this is for.
Pricing shape
Per token or provisioned units, on the Azure bill

Google Vertex AI

A managed AI platform with model garden access, regional endpoints, provisioned throughput, safety filters and per-project quotas, billed on the GCP invoice.

Best for
Teams standardised on Google Cloud, especially where video and speech generation matter.
Watch out for
Vertex-specific SDKs and IAM mean this is an integration, not a base-URL swap, and it routes within Google only.
Pricing shape
Per token or per media unit, on the GCP bill
Why teams move off it

Four reasons a self-run proxy stops paying for itself

2am

You are the on-call for the gateway

Upgrades, a Postgres, a Redis, config drift across environments and a Python process in the hot path of every request. None of it ships product, and all of it pages somebody.

p99

Overhead at concurrency

A proxy that is invisible at ten requests a second is measurable at a thousand. Teams hit this and start looking at Go and Rust gateways, or at taking the hop out entirely.

SOC 2

Governance you have to build

Per-team budgets exist, but audit logs, RBAC, SSO, guardrails and an evidence trail an auditor accepts are a product, not a config file.

The stack outgrows text

Speech, video and OCR arrive, and a proxy designed around chat completions ends up with a second vendor bolted alongside it.

When to stay on LiteLLM: you need provider abstraction inside your own perimeter, your traffic is moderate, and the team that runs it is the team that wanted it. The planner says so too when that is the honest answer.

What the gateway layer costs you

Self-hosting has no licence fee and a very real floor: servers, a database, a cache and the person who keeps them up. Drag the slider to your monthly token spend and compare the layer, not the tokens.

$1,000 / month$12,000 a year in tokens
$100$1k$10k$100k

    Layer cost only — tokens are billed on top and cost about the same wherever you route. The self-hosted rows assume roughly $180 a month of infrastructure for a two-replica gateway, a small managed Postgres and a cache; your figure will differ, and neither row prices the engineering time, which is usually the larger number. Vendor fee shapes taken from each vendor's own pricing page, checked 22 September 2026 — the full working is in our comparison write-up.

    Coverage at a glance

    Every alternative against the twelve capabilities the planner scores on. Filled means the platform covers it as a first-class capability; empty means it does not — no half-credit for “on the roadmap”.

    AlternativeSelf-hostManagedKubernetesEdgeBYOKUnified billingFailoverCachingBudgetsObservabilityEnterpriseDrop-in swap
    LiteLLMcoveredcoveredcoverednot coveredcoverednot coveredcoveredcoveredcoveredcoveredcoveredcovered
    LLM APInot coveredcoverednot coverednot coverednot coveredcoveredcoveredcoveredcoveredcoveredcoveredcovered
    Portkeycoveredcoveredcoveredcoveredcoverednot coveredcoveredcoveredcoveredcoveredcoveredcovered
    Kong AI Gatewaycoveredcoveredcoverednot coveredcoverednot coveredcoveredcoveredcoveredcoveredcoverednot covered
    Apache APISIXcoverednot coveredcoverednot coveredcoverednot coveredcoveredcoverednot coverednot coveredcoverednot covered
    Envoy AI Gatewaycoverednot coveredcoverednot coveredcoverednot coveredcoverednot coveredcoveredcoveredcoveredcovered
    Bifrostcoveredcoveredcoverednot coveredcoverednot coveredcoveredcoveredcoveredcoveredcoveredcovered
    Heliconecoveredcoveredcoveredcoveredcoverednot coveredcoveredcoveredcoveredcoveredcoveredcovered
    Langfusecoveredcoveredcoverednot coveredcoverednot coverednot coverednot coveredcoveredcoveredcoverednot covered
    Cloudflare AI Gatewaynot coveredcoverednot coveredcoveredcoveredcoveredcoveredcoveredcoveredcoveredcoveredcovered
    Vercel AI Gatewaynot coveredcoverednot coveredcoveredcoveredcoveredcoveredcoveredcoveredcoverednot coveredcovered
    TrueFoundrycoveredcoveredcoverednot coveredcoverednot coveredcoveredcoveredcoveredcoveredcoveredcovered
    Databricks Mosaic AI Gatewaynot coveredcoverednot coverednot coveredcoverednot coveredcoverednot coveredcoveredcoveredcoverednot covered
    Requestynot coveredcoverednot coverednot coveredcoveredcoveredcoveredcoveredcoveredcoverednot coveredcovered
    Unifynot coveredcoverednot coverednot coveredcoveredcoveredcoverednot coveredcoveredcoverednot coveredcovered
    OpenRouternot coveredcoverednot coverednot coveredcoveredcoveredcoveredcoveredcoveredcoverednot coveredcovered
    Eden AInot coveredcoverednot coverednot coveredcoveredcoveredcoverednot coveredcoveredcoveredcoverednot covered
    AWS Bedrocknot coveredcoverednot coverednot coverednot coverednot coveredcoveredcoveredcoveredcoveredcoverednot covered
    Azure AI Foundrynot coveredcoverednot coverednot coverednot coverednot coveredcoveredcoveredcoveredcoveredcoveredcovered
    Google Vertex AInot coveredcoverednot coverednot coverednot coverednot coveredcoveredcoveredcoveredcoveredcoverednot covered
    First-class capabilityNot offered

    Short version, by situation

    If you scrolled past the planner, start here.

    You want the same thing, just maintained

    Portkey does the closest like-for-like job, self-hosted or hosted. Bifrost is the pick if the complaint is specifically Python overhead.

    Your platform team already runs a gateway

    Kong AI Gateway or Envoy AI Gateway put AI traffic under the policy, auth and telemetry you already operate, instead of beside it.

    You want out of the operations business

    Cloudflare AI Gateway keeps your own provider keys with nothing to deploy. LLM API is the answer when speech, video or one invoice also matter.

    Procurement and audit have to sign

    TrueFoundry in your own VPC, or AWS Bedrock, Azure AI Foundry or Google Vertex AI if the spend belongs on a cloud contract you already hold.

    How we compared them

    Twelve capabilities and ten planner axes, applied the same way to every entry, from public documentation and hands-on use.

    • Who can run it: self-hosted, managed, or both
    • Deployment shape: Kubernetes, VM, edge, fully hosted
    • Licence of the part that sits in the request path
    • Routing, retries and cross-provider failover
    • Caching and spend guardrails
    • Budgets, virtual keys and per-team reporting
    • Observability: built-in traces or export to your stack
    • Keys and billing: BYOK, unified balance, cloud commitment
    • Data residency, retention and compliance paperwork
    • Enterprise readiness: SLAs, RBAC, SSO, support
    • Model ecosystem reachable through it
    • Migration effort from an OpenAI-compatible proxy

    What we could not measure

    Gateway overhead varies by deployment, hardware and traffic shape, so we did not publish a single latency ranking — benchmark it on your own cluster with your own prompts. Support quality is reported from plan documents, not tested. Where a vendor does not publish a fact, the comparison says so instead of guessing.

    Read the full methodology write-up
    FAQ

    Questions we get a lot

    Still unsure after the planner? These come up in almost every migration call.

    What is the best open-source LiteLLM alternative?

    Portkey's gateway core and Bifrost are the closest open-source like-for-like replacements. If the gateway belongs to your platform team rather than an application team, Kong AI Gateway and Envoy AI Gateway do the same routing job as part of infrastructure you already run, and Apache APISIX keeps everything in the path permissively licensed.

    Can I switch without rewriting application code?

    Usually yes. Most alternatives here expose an OpenAI-compatible endpoint, so you change a base URL, a key and sometimes the model-ID format. The exceptions are the hyperscaler platforms and Databricks, where you adopt their SDK and IAM model instead.

    Is self-hosting actually cheaper?

    There is no licence fee, but there is a floor: a couple of gateway replicas, a Postgres, a cache and somebody on call. Around $150–$250 a month of infrastructure is typical before any engineering time. Below roughly $1,000 a month of token spend, a flat-fee managed control plane is usually cheaper than the servers alone.

    What about the proxy overhead people complain about?

    The Python proxy adds real latency and CPU at high concurrency, which is why Go-based gateways like Bifrost and Envoy-based routing exist. At modest request rates it is irrelevant; measure yours before it drives the decision.

    When should I just stay on LiteLLM?

    When you need provider abstraction inside your own perimeter, your traffic is moderate, and the team running it is the team that wanted it. The planner says so too when that is the honest answer.

    Is anyone paying to be on this list?

    No. LLM API is our own product and is labelled as such wherever it appears. Every other entry is here on merit, scored on the same criteria from public documentation and hands-on use.

    Or stop running a gateway altogether

    LLM API gives you 400+ text, vision, speech and video models behind a single OpenAI-compatible key, with uptime monitoring, spend controls and one invoice. Migration from a self-hosted proxy is a base URL and a key.

    Start free