Deployment
Fine-Tunable Models providers
Every model provider on LLMAPI tagged “Fine-Tunable Models” — 24 of 50 providers. All of them run through one API key and one OpenAI-compatible endpoint.
24 providers matching this filter.
| Provider | Models | Regions | Best for | |
|---|---|---|---|---|
Meta Open-weight Llama models for self-hosting and fine-tuning. |
7 | Self-hosted | Self-hosting, fine-tuning, and keeping full control of model weights | View → |
DeepSeek Open reasoning models at aggressive price points. |
5 | Global | Reasoning and coding at the lowest per-token prices | View → |
Alibaba Cloud Qwen models plus a full Model Studio platform. |
10 | APAC, Global | APAC-focused apps and multilingual workloads with Qwen models | View → |
Xiaomi MiMo open models focused on reasoning. |
3 | Self-hosted | On-prem math and reasoning workloads with open weights | View → |
Arcee AI Small language models and model merging. |
4 | US | Small, fine-tuned models and continual pre-training | View → |
InclusionAI Ling open models from Ant Group. |
4 | APAC | Open MoE models for APAC-focused applications | View → |
Thinking Machines Frontier research lab building custom-model tooling. |
1 | US | Custom and collaboratively fine-tuned models | View → |
Baseten High-performance inference for open models. |
8 | US | Production inference for open models with autoscaling | View → |
Parasail Serverless inference marketplace. |
4 | US | Serverless open-model inference at spot-style pricing | View → |
Modal Serverless GPUs for custom model code. |
1 | US, EU | Running custom inference code on serverless GPUs | View → |
CoreWeave Specialized GPU cloud at hyperscale. |
1 | US, EU | Large-scale training and inference on dedicated GPU clusters | View → |
Cerebras Wafer-scale chips for the fastest inference. |
8 | US | The fastest token-per-second inference for open models | View → |
Fireworks Fast generative AI inference platform. |
10 | US, EU | High-speed OpenAI-compatible inference for open models | View → |
Together AI Open-model inference and fine-tuning cloud. |
8 | US | Open-model inference plus fine-tuning in one cloud | View → |
SiliconFlow High-throughput inference for open models. |
7 | APAC, Global | High-throughput open-model APIs with APAC availability | View → |
SambaNova Dataflow chips with fast open-model APIs. |
4 | US | High-throughput open-model serving on custom hardware | View → |
Replicate Run thousands of community models via API. |
5 | US | Prototyping with thousands of community models | View → |
Databricks Model serving inside the data lakehouse. |
4 | Global | Serving and fine-tuning models where your data already lives | View → |
DeepInfra Low-cost inference for popular open models. |
8 | US | Popular open models at the lowest per-token cost | View → |
Wafer Distributed inference network. |
3 | Global | Affordable open-model inference on distributed GPUs | View → |
Groq LPU hardware for ultra-low-latency inference. |
9 | US, EU | Ultra-low-latency inference for real-time apps | View → |
Celeris Fast inference for open-source models. |
3 | US | Fast, simple endpoints for open-source models | View → |
Makora Model hosting and inference APIs. |
3 | Global | Pay-as-you-go hosting for open models | View → |
Modular MAX serving stack and Mojo ecosystem. |
1 | US | Custom serving stacks and AI systems engineering | View → |
One key, every provider on this page
Consolidated billing, spend limits and instant fallback between providers.