Bonus: Top up now and we'll double your first deposit — get x2 credits instantly.
Alternatives · updated weekly

Pixtral Large Latest Alternatives, Compared

28 models on LLM.API accept the same requests as Pixtral Large Latest. They are rated 1.0–5.0 from our own catalogue data, benchmarked where independent scores exist, and priced against Pixtral Large Latest, so you can see what a swap costs before you make it.

ReplacingPixtral Large Latest
★ 2.8
$4.00 / 1M in128K contextAPI only
28 drop-in options
28alternatives
99%max input saving
11benchmarked

Why are you switching?

Pick your main pain point — the picks, reviews and table below narrow to models that fix it.

Top picks

Chosen automatically from rating, published price and capability match, updated weekly. Every point comes from published data, including where Pixtral Large Latest is still the better choice.

Gemini 3.5 Flash Lite★ 4.1

Closest match
Where it beats Pixtral Large Latest
  • 93% cheaper input ($0.30 vs $4.00 per 1M).
  • Cheaper output: $2.50 vs $12.00 per 1M.
  • 8.2× the context: 1M vs 128K.
Where Pixtral Large Latest still wins
  • Nothing in the published data — the difference will show up in your own evals and prompt tuning.
Best suited for
high-volume and batch workloads, long documents and RAG, agents and tool calling.

Gemini 3.1 Flash Lite★ 4.0

Cheapest swap
Where it beats Pixtral Large Latest
  • 94% cheaper input ($0.25 vs $4.00 per 1M).
  • Cheaper output: $1.50 vs $12.00 per 1M.
  • 8.2× the context: 1M vs 128K.
Where Pixtral Large Latest still wins
  • Nothing in the published data — the difference will show up in your own evals and prompt tuning.
Best suited for
high-volume and batch workloads, long documents and RAG, agents and tool calling.

DeepSeek V4.1 Flash★ 3.9

Open weights
Where it beats Pixtral Large Latest
  • 93% cheaper input ($0.30 vs $4.00 per 1M).
  • Cheaper output: $1.20 vs $12.00 per 1M.
  • 7.8× the context: 1M vs 128K.
Where Pixtral Large Latest still wins
  • Nothing in the published data — the difference will show up in your own evals and prompt tuning.
Best suited for
high-volume and batch workloads, long documents and RAG, agents and tool calling.

GLM-5.3-Flash★ 3.9

Longest context
Where it beats Pixtral Large Latest
  • 96% cheaper input ($0.15 vs $4.00 per 1M).
  • Cheaper output: $0.50 vs $12.00 per 1M.
  • 7.8× the context: 1M vs 128K.
Where Pixtral Large Latest still wins
  • Nothing in the published data — the difference will show up in your own evals and prompt tuning.
Best suited for
high-volume and batch workloads, long documents and RAG, agents and tool calling.

What would switching save you?

Pick the stage closest to yours. Costs use each model's published per-token list price and assume 75% of tokens are input, which is typical for chat and RAG. Your own ratio will move the numbers, not the ranking much.

Pixtral Large Latest at this volume: $300 a month75% input / 25% output tokens, list prices
ModelMonthly costPer monthPer yearvs Pixtral Large Latest / year
Ministral 3 3B 2512$5.00$60.00Save $3,540
GLM-5.3-Flash$11.88$143Save $3,458
DeepSeek V4.1 Flash$26.25$315Save $3,285
Gemini 3.1 Flash Lite$28.13$338Save $3,263
Gemini 3.5 Flash Lite$42.50$510Save $3,090
Pixtral Large LatestToday$300$3,600

Pixtral Large Latest vs the alternatives

Prices per 1M tokens. Green is cheaper than Pixtral Large Latest's input price, red is dearer. Index and speed come from Artificial Analysis.

ModelRatingInputOutputContextIndexSpeedOpen weights
Pixtral Large LatestYou are here★ 2.8$4.00$12.00128KNo
Gemini 3.5 Flash LiteClosest match★ 4.1$0.30 −93%$2.501M23348 tok/sNo
Gemini 3.1 Flash LiteCheapest★ 4.0$0.25 −94%$1.501MNo
DeepSeek V4.1 FlashOpen weights★ 3.9$0.30 −93%$1.201M40199 tok/sYes
GLM-5.3-Flash★ 3.9$0.15 −96%$0.501MYes
Qwen3.7 Max★ 3.9$0.50 −88%$3.001MNo
Qwen 3.7 Plus★ 3.9$0.50 −88%$3.001M2665 tok/sNo
GPT-5.6 Luna★ 3.8$0.20 −95%$1.201.1M38110 tok/sNo
GPT-5.4 Nano★ 3.8$0.20 −95%$1.25400KNo
GLM-4.6V FlashX★ 3.8$0.04 −99%$0.40128KNo
Qwen 3.6 Plus★ 3.8$0.50 −88%$3.001MNo
MiniMax M3★ 3.7$0.60 −85%$2.401M3092 tok/sYes
Qwen3.8 Flash Next★ 3.6$0.20 −95%$0.50262K4053 tok/sYes

Showing the top 12 of 28. Pick a filter to see every match.

Benchmarks: quality and speed, not just price

Independent scores from Artificial Analysis: the intelligence index blends reasoning, knowledge, maths and coding tests; output speed is measured tokens per second. "Input $ per index point" shows how much quality each dollar buys.

ModelIntelligence indexOutput speedInput $ per index point
Grok 4.6 44 55 tok/s$0.0455
DeepSeek V4.1 Flash 40 199 tok/s$0.0075
Qwen3.8 Flash Next 40 53 tok/s$0.00502
Grok 4.5 39 51 tok/s$0.0513
GPT-5.6 Luna 38 110 tok/s$0.00526
MiniMax M3 30 92 tok/s$0.02
Qwen 3.7 Plus 26 65 tok/s$0.0192
Kimi K2.7-Code 26 50 tok/s$0.0285
Gemini 3.5 Flash Lite 23 348 tok/s$0.013
Amazon Nova Lite 7 169 tok/s$0.00857
Pixtral Large LatestYou are hereNot publishedNot published

Source: Artificial Analysis. 11 of 28 alternatives have published scores; the rest are shown without one rather than estimated. Always confirm on your own prompts.

Which one for your workload?

Different pipelines stress different numbers. Here is what actually drives the choice for each one, which model wins on that number, and how it stacks up against Pixtral Large Latest.

If you needDocument extraction and RAGGPT-5.6 Luna
If you needAgents and tool callingDeepSeek V4.1 Flash
If you needAir-gapped or compliance-boundDeepSeek V4.1 Flash
If you needBudget batch jobsGLM-4.6V FlashX
If you needCoding assistantsKimi K2.7-Code

Document extraction and RAG

What matters: Long inputs, lots of them: context size and input price matter most.

Retrieval pipelines send far more tokens in than they get back — often 20 to 50 input tokens for every output token. That makes input price and context size the two numbers that decide your bill. A bigger window also means fewer chunks, simpler retrieval logic and less risk of cutting off the passage that holds the answer.

Our pick: GPT-5.6 Luna ★ 3.8
1.1M context at $0.20 / 1M in.
Context window (tokens)
Pixtral Large Latest128K
GPT-5.6 Luna1.1M

8.2× Pixtral Large Latest.

Agents and tool calling

What matters: Needs reliable function calls and JSON output across many steps.

Agents fail in quiet ways: a malformed JSON argument or a skipped tool call breaks the loop several steps later. Native tool calling plus JSON mode is the baseline, and general reasoning strength matters because the model plans each next step itself. The intelligence index is the best public proxy for that.

Our pick: DeepSeek V4.1 Flash ★ 3.9
Tool calling and JSON mode, rated ★ 3.9, intelligence index 40.
Key facts

Tool calling and JSON mode, rated ★ 3.9, intelligence index 40.

Air-gapped or compliance-bound

What matters: Data cannot leave your cloud or region.

If data residency or an air-gapped network is a hard requirement, no hosted-only model qualifies, however good it is. Open weights give you an exit: prototype through the API today, then run the same model inside your own cloud when the compliance review lands, without rewriting prompts.

Our pick: DeepSeek V4.1 Flash ★ 3.9
Published weights — self-host it and keep every token in your own environment.
LLM API rating (1–5)
Pixtral Large Latest2.8
DeepSeek V4.1 Flash3.9

1.4× Pixtral Large Latest.

Budget batch jobs

What matters: Classification, tagging, summaries at volume where cost per token dominates.

Tagging, classification and summaries at volume rarely need a frontier model. Here cost per token dominates and a small quality drop is invisible after aggregation. We only consider models rated 3.0 or higher, so the cheapest pick is still a reliable one.

Our pick: GLM-4.6V FlashX ★ 3.8
$0.04 per 1M input — 99% under Pixtral Large Latest.
Input price per 1M tokens (lower is better)
Pixtral Large Latest$4.00
GLM-4.6V FlashX$0.04

99% cheaper than Pixtral Large Latest.

Coding assistants

What matters: Code generation, review and refactoring inside editors or CI.

Code work punishes small mistakes: one wrong import and the suggestion is useless. Models built or tuned for code usually beat general models of the same size, and a high intelligence index (which includes coding tests) is the next best signal.

Our pick: Kimi K2.7-Code ★ 3.6
Built or tuned for code.
Key facts

Built or tuned for code.

Switching takes one line

Point your existing OpenAI client at LLM.API and change the model name. Nothing else in your code moves.

cURL
# Same OpenAI-compatible request — only the model name changes.
curl https://api.llmapi.ai/v1/chat/completions \
  -H "Authorization: Bearer $LLMAPI_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-3.5-flash-lite",
    "messages": [{"role":"user","content":"Hello!"}]
  }'

Should you actually switch?

Good reasons to move off Pixtral Large Latest

1Token spend is your biggest line item and a cheaper model passes your evals.
2You need published weights so the model can move in-house or into a private region later.
3Your prompts or agent traces no longer fit the context window.
4You want a second model behind the same key as a fallback when one provider has a bad day.

Good reasons to stay

1Your prompts are tuned to Pixtral Large Latest and the quality gap shows up in your own evals.
2You rely on features the alternatives do not list on their model page.
3The saving is small next to the cost of re-testing a production pipeline.
4Running Pixtral Large Latest through LLM.API already cuts the list price, so the gap is narrower than it looks.

Frequently asked questions

Are these Pixtral Large Latest alternatives drop-in replacements?

Every model listed here is served on LLM.API through the same OpenAI-compatible /v1/chat/completions endpoint, and each one accepts the same input types as Pixtral Large Latest. In most cases you change the "model" string and nothing else. Provider-specific parameters are listed on each model page, so check those before you switch a production prompt.

Do I need a separate API key for each alternative?

No. One LLM.API key covers every model on this page, with one invoice and one set of usage logs, so you can A/B two models against the same traffic without opening new accounts.

How is the rating calculated?

It comes from our own catalogue data, not paid reviews: 40% capability (price tier and how recent the model is), 35% value for money compared with similar models, and 25% context size plus support for tools, JSON and streaming. Each model has one rating, and it is the same on every page it appears on. It updates weekly.

Where do the benchmark scores come from?

The intelligence index and output speed are published by Artificial Analysis, an independent benchmarking site. We only show a score where they publish one; models without a score show a dash rather than an estimate.

Will I lose quality by moving off Pixtral Large Latest?

It depends on the workload. Swap the top-rated alternative in for a sample of real traffic, compare the outputs side by side, and keep whichever wins. Because both models sit behind the same key, running that comparison costs you only the tokens.

Are the prices on this page current?

Prices come straight from the LLM.API catalogue and refresh every Sunday. Where a provider has not published a rate, the table shows a dash instead of an estimate.

28 alternatives live now

Swap Pixtral Large Latest in one line.
Keep the one that wins.

One API key for every model on this page. Run Pixtral Large Latest and any alternative side by side on your real traffic, then keep whichever is cheaper, faster or better.

1key & invoice
28+drop-in models
0code rewrites
# your existing OpenAI client
client = OpenAI(
  base_url="https://api.llmapi.ai/v1",
)
client.chat.completions.create(
  model="pixtral-large-latest"
  model="gemini-3.5-flash-lite",
  messages=msgs,
)