Model provider7 modelsSelf-hosted

Meta

Open-weight Llama models for self-hosting and fine-tuning.

StatusOperational Live status
LLMAPI discountUp to −25%
Models on LLMAPI7
From$0.06 / 1M blended
RegionsSelf-hosted
Best uptime (24h)100.00%

Verdict

Meta releases the Llama family of open-weight models, the de facto standard for teams that need to self-host, fine-tune, or distill LLMs. Through LLMAPI you reach 7 Meta models on the same OpenAI-compatible endpoint as every other provider, so switching between them is a one-string change.

Good fit for

  • Text Generation
  • Open-Weight Models
  • Best for Local Deployment
  • Self-Hosted / On-Premise
  • Fine-Tunable Models
  • Most Transparent

Consider another provider for

  • No code generation models in this catalog
  • No image generation models in this catalog
  • No vision models in this catalog
  • No video generation models in this catalog
  • Not eligible for LLMAPI volume discounts

Key facts

Where a provider does not document something we say so rather than guess. Everything LLMAPI records about Meta, in one table.

Overview

ProviderMeta
Websitemeta.com
Best forSelf-hosting, fine-tuning, and keeping full control of model weights
Models on LLMAPI7
Popular modelsLlama 4 Maverick, Llama 4 Scout, Llama 3.3 70B, Llama 3.3 8B, Llama 3.1 405B, Llama 3.1 70B

Access & routing

API compatibilityOpenAI-compatible — /v1/chat/completions
Model ID prefixmeta/<model>
RegionsSelf-hosted
FallbackConfigurable — on 429 or 5xx the router switches to any other model or provider on your list
BillingOne invoice across every provider, with per-key spend limits
Volume discountNot eligible

Performance (median across listed models)

Blended price / 1M$0.15
Uptime (24h)99.95% (median across listed models)

Modality

Tags

Deployment

Tags

Trust & safety

Tags

Audience

Tags

Privacy & data

Inference locationSelf-hosted
Data residencyNot documented by the provider
Default retentionNot documented by the provider
Zero-data retentionNot documented by the provider
Training on API dataNot documented by the provider
Human reviewNot documented by the provider

Compliance & lifecycle

CertificationsISO 27001
Deprecation noticeNot documented by the provider
Operational statusOperational Live status

Prices and context from the OpenRouter public catalogue; uptime and output speed from OpenRouter provider endpoints; quality scores from LiveBench. Refreshed nightly — last updated 2026-09-29. — means no figure is published for that model.

Best Meta models

A quick view of how Meta 's 7 models compare on intelligence, output speed and price — so you can pick the right one for your use case.

Most intelligent

#1Llama 4 Scout42
#2Llama 3.1 405B41
#3Llama 3.3 70B34
#4Llama 4 Maverick30

Intelligence index · 7 models

Fastest

#1Llama 3.1 405B425 t/s
#2Llama 3.1 8B425 t/s
#3Llama 3.1 70B418 t/s
#4Llama 3.3 70B386 t/s

Output tokens / second · 7 models

Lowest price

#1Llama 3.1 8B$0.56
#2Llama 3.1 405B$3.49
#3Llama 4 Maverick$7.91
#4Llama 4 Scout$9.89

Blended price per 1M tokens · 7 models

Compared with other hosts

How Meta sits against the rest of the catalog, using the same measurements shown on every provider page.

  • Price. Median blended price of $9.89 per 1M tokens — 61% above the $6.13 median across all 50 providers in the catalog.
  • Throughput. Median output speed of 386 tokens/s, 11% faster than the 348 tok/s catalog median.
  • Catalog depth. 7 models listed, against a catalog average of 5 per provider.
  • Integration. Identical to every other provider here — same endpoint, same SDK, same keys, so a switch costs one string change.

Median blended price per 1M tokens

lower is better
Meta$9.89
DigitalOcean$9.25
xAI (SpaceXAI)$10.60
Parasail$8.93
InclusionAI$8.76
Nebius$8.66

Median output speed

higher is better
Meta386 tok/s
DigitalOcean348 tok/s
xAI (SpaceXAI)242 tok/s
Parasail822 tok/s
InclusionAI244 tok/s
Nebius580 tok/s

Figures are medians across the models listed on each provider page and are refreshed with the catalog.

Models & pricing

Every Meta model reachable with your LLMAPI key. Click a column heading to sort.

ModelAPI model nameUptime 24hPrice / 1MContext
Llama 4 Maverickmeta/llama-4-maverick99.95%$ 0.31049 k
Llama 4 Scoutmeta/llama-4-scout99.94%$ 0.151311 k
Llama 3.3 70Bmeta/llama-3-3-70b99.80%$ 0.15131 k
Llama 3.3 8Bmeta/llama-3-3-8b———
Llama 3.1 405Bmeta/llama-3-1-405b———
Llama 3.1 70Bmeta/llama-3-1-70b99.99%$ 0.4131 k
Llama 3.1 8Bmeta/llama-3-1-8b100.00%$ 0.06131 k

Price is a blended per-1M-token figure; latency is time to first token and throughput is output tokens per second, measured on the LLMAPI edge.

Cost calculator

Pick a model, enter your traffic, and see the monthly bill at LLMAPI rates.

Estimates use the blended per-1M-token price shown in the table above. Word counts differ by language — Cyrillic and CJK text uses 2–3× more tokens per word.

Estimated monthly cost

$—
Input tokens$—
Output tokens$—
Per request$—
Get API key

Meta via LLMAPI

Same models, same features, one key. Prefix the model ID with meta/ and point your OpenAI client at our base URL.

# Python · OpenAI SDK
from openai import OpenAI

client = OpenAI(
    base_url="https://api.llmapi.ai/v1",
    api_key="LLMAPI_KEY",
)

r = client.chat.completions.create(
    model="meta/llama-4-maverick",
    messages=[{"role": "user", "content": "Hello"}],
)
PricingBilled at the rates in the table above, on one invoice with every other provider.
FeaturesFull pass-through — tool calling, JSON schema, streaming, vision and reasoning behave exactly as on Meta’s own API.
FallbackConfigurable. On 429 or 5xx the router switches to any model or provider on your fallback list.
RegionsSelf-hosted
Your dataLLMAPI stores request metadata only — model, token counts, latency and status. No prompt or completion content.

Capabilities & use cases

Each tag below is its own catalog page listing every provider that shares it.

Developer resources

Where to go next.

Join thousands of developers building on one API key

Every provider in the catalog, one key, one invoice.

Questions

How do I call Meta models through LLMAPI?

Point any OpenAI-compatible client at https://api.llmapi.ai/v1 and use the model ID meta/llama-4-maverick. No other change is needed.

Does Meta cost more through LLMAPI?

No. You pay the rates shown in the table above, on a single invoice with every other provider.

Which regions are used?

Self-hosted

What happens if the provider returns an error?

On 429 or 5xx responses the router falls back to any other model or provider on your configured list, so requests keep succeeding.

How many Meta models are available?

7 at the moment, all listed in the models table above and updated as the provider ships new ones.