Modality
Text Generation providers
Every model provider on LLMAPI tagged “Text Generation” — 50 of 50 providers. All of them run through one API key and one OpenAI-compatible endpoint.
50 providers matching this filter.
| Provider | Models | Regions | Best for | |
|---|---|---|---|---|
Anthropic Claude models with a focus on safety and long context. |
6 | US, EU | Long-document analysis, coding assistants, and agents that need careful reasoning | View → |
OpenAI GPT series, o-series reasoning models, and multimodal APIs. |
15 | US, EU | General-purpose chat, multimodal apps, and the widest tooling ecosystem | View → |
Google Gemini models via AI Studio and Vertex AI. |
7 | Global | Multimodal workloads and very long context at aggressive price points | View → |
Meta Open-weight Llama models for self-hosting and fine-tuning. |
7 | Self-hosted | Self-hosting, fine-tuning, and keeping full control of model weights | View → |
Mistral Efficient open and commercial models from Paris. |
10 | EU, US | Price-performance and EU data-residency requirements | View → |
xAI (SpaceXAI) Grok models built by xAI. |
5 | US | Real-time, X-integrated assistants and fast reasoning workloads | View → |
DeepSeek Open reasoning models at aggressive price points. |
5 | Global | Reasoning and coding at the lowest per-token prices | View → |
Alibaba Cloud Qwen models plus a full Model Studio platform. |
10 | APAC, Global | APAC-focused apps and multilingual workloads with Qwen models | View → |
Zai (Zhipu AI) GLM series models from Zhipu AI. |
5 | APAC | Chinese-English bilingual applications | View → |
MiniMax Long-context models and multimodal APIs. |
6 | APAC, Global | Very long context and text-to-speech/video generation | View → |
Kimi (Moonshot AI) Moonshot AI's long-context Kimi models. |
3 | APAC, Global | Long-document Q&A and agentic tool use | View → |
Xiaomi MiMo open models focused on reasoning. |
3 | Self-hosted | On-prem math and reasoning workloads with open weights | View → |
StepFun Step series multimodal models. |
4 | APAC | Multimodal language, vision, and audio generation in APAC | View → |
Cohere Enterprise-focused Command models and embeddings. |
7 | US, EU | Enterprise RAG, search, and reranking pipelines | View → |
Liquid AI Liquid foundation models beyond transformers. |
6 | US | Low-latency inference at the edge and on-device | View → |
Upstage Solar models and document AI. |
3 | APAC, US | Document parsing and lightweight enterprise chat | View → |
Reka AI Multimodal models for video, image, and text. |
3 | US | Apps that natively understand video, images, and audio | View → |
Inception Diffusion-based language models for speed. |
3 | US | Ultra-fast text generation for latency-critical products | View → |
Arcee AI Small language models and model merging. |
4 | US | Small, fine-tuned models and continual pre-training | View → |
Multiverse Computing Compressed models via quantum-inspired techniques. |
3 | EU | Compressed models for on-device and cost-sensitive deployments | View → |
InclusionAI Ling open models from Ant Group. |
4 | APAC | Open MoE models for APAC-focused applications | View → |
Agnes AI Agentic models for research workflows. |
2 | APAC | Agentic deep-research and multi-step task automation | View → |
Thinking Machines Frontier research lab building custom-model tooling. |
1 | US | Custom and collaboratively fine-tuned models | View → |
Baseten High-performance inference for open models. |
8 | US | Production inference for open models with autoscaling | View → |
DigitalOcean GPU droplets and a managed genAI platform. |
4 | US, EU, APAC | Startups hosting agents and GPUs on a developer-friendly cloud | View → |
Parasail Serverless inference marketplace. |
4 | US | Serverless open-model inference at spot-style pricing | View → |
Modal Serverless GPUs for custom model code. |
1 | US, EU | Running custom inference code on serverless GPUs | View → |
CoreWeave Specialized GPU cloud at hyperscale. |
1 | US, EU | Large-scale training and inference on dedicated GPU clusters | View → |
Cerebras Wafer-scale chips for the fastest inference. |
8 | US | The fastest token-per-second inference for open models | View → |
Fireworks Fast generative AI inference platform. |
10 | US, EU | High-speed OpenAI-compatible inference for open models | View → |
Together AI Open-model inference and fine-tuning cloud. |
8 | US | Open-model inference plus fine-tuning in one cloud | View → |
Nebius Full-stack AI cloud built in Europe. |
5 | EU, US | EU-based GPU clusters and managed open-model inference | View → |
SiliconFlow High-throughput inference for open models. |
7 | APAC, Global | High-throughput open-model APIs with APAC availability | View → |
Crusoe Energy-first GPU cloud. |
1 | US | Cost-effective, sustainably powered GPU compute | View → |
SambaNova Dataflow chips with fast open-model APIs. |
4 | US | High-throughput open-model serving on custom hardware | View → |
Replicate Run thousands of community models via API. |
5 | US | Prototyping with thousands of community models | View → |
Amazon Bedrock AWS managed access to many foundation models. |
12 | Global (AWS) | AWS-native teams needing many models behind one API | View → |
Microsoft Azure Azure OpenAI Service and model catalog. |
9 | Global (Azure) | Enterprises needing OpenAI models with Azure compliance | View → |
Databricks Model serving inside the data lakehouse. |
4 | Global | Serving and fine-tuning models where your data already lives | View → |
DeepInfra Low-cost inference for popular open models. |
8 | US | Popular open models at the lowest per-token cost | View → |
Novita Affordable GPU cloud and model APIs. |
4 | APAC, US | Budget-friendly serverless GPUs and model APIs | View → |
Bitdeer AI GPU cloud from a datacenter operator. |
1 | APAC, US | GPU cloud capacity from a datacenter operator | View → |
FriendliAI Optimized serving engine for LLMs. |
3 | US, APAC | High-efficiency LLM serving with dedicated endpoints | View → |
Wafer Distributed inference network. |
3 | Global | Affordable open-model inference on distributed GPUs | View → |
Scaleway European cloud with GPU and inference APIs. |
4 | EU | EU data residency for GPUs and generative APIs | View → |
Groq LPU hardware for ultra-low-latency inference. |
9 | US, EU | Ultra-low-latency inference for real-time apps | View → |
Celeris Fast inference for open-source models. |
3 | US | Fast, simple endpoints for open-source models | View → |
Makora Model hosting and inference APIs. |
3 | Global | Pay-as-you-go hosting for open models | View → |
Inco Confidential-computing model inference. |
2 | US | Privacy-sensitive workloads in confidential enclaves | View → |
Modular MAX serving stack and Mojo ecosystem. |
1 | US | Custom serving stacks and AI systems engineering | View → |
One key, every provider on this page
Consolidated billing, spend limits and instant fallback between providers.