Verdict
Alibaba Cloud develops the Qwen family of open and commercial models and offers them through its Model Studio platform alongside broader cloud services. Through LLMAPI you reach 10 Alibaba Cloud models on the same OpenAI-compatible endpoint as every other provider, so switching between them is a one-string change.
Good fit for
- Text Generation
- Code Generation
- Multilingual
- Best for High Volume Production
- Open-Weight Models
- Best for Local Deployment
Consider another provider for
- No image generation models in this catalog
- No vision models in this catalog
- No video generation models in this catalog
- No video understanding models in this catalog
Key facts
Where a provider does not document something we say so rather than guess. Everything LLMAPI records about Alibaba Cloud, in one table.
Overview | |
| Provider | Alibaba Cloud |
| Website | alibabacloud.com |
| Best for | APAC-focused apps and multilingual workloads with Qwen models |
| Models on LLMAPI | 25 |
| Popular models | GLM-5.3, Qwen 3.8 Max, DeepSeek V4 Flash 0731, Kimi K3, Kimi K2.7-Code, Qwen3.7 Max |
Access & routing | |
| API compatibility | OpenAI-compatible — /v1/chat/completions |
| Model ID prefix | alibaba-cloud/<model> |
| Regions | APAC, Global |
| Fallback | Configurable — on 429 or 5xx the router switches to any other model or provider on your list |
| Billing | One invoice across every provider, with per-key spend limits |
| Volume discount | Eligible |
Performance (median across listed models) | |
| Blended price / 1M | $1.13 |
| Uptime (24h) | 99.76% (median across listed models) |
Modality | |
| Tags | |
Capability | |
| Tags | |
Price & performance | |
| Tags | |
Deployment | |
| Tags | |
Industry | |
| Tags | |
Workflow | |
| Tags | |
Trust & safety | |
| Tags | |
Audience | |
| Tags | |
Privacy & data | |
| Inference location | APAC, Global |
| Data residency | Not documented by the provider |
| Default retention | Not documented by the provider |
| Zero-data retention | Not documented by the provider |
| Training on API data | Not documented by the provider |
| Human review | Not documented by the provider |
Compliance & lifecycle | |
| Certifications | SOC 2 · ISO 27001 · GDPR |
| Deprecation notice | Not documented by the provider |
| Operational status | Operational Live status |
Prices and context from the OpenRouter public catalogue; uptime and output speed from OpenRouter provider endpoints; quality scores from LiveBench. Refreshed nightly — last updated 2026-09-29. — means no figure is published for that model.
Best Alibaba Cloud models
A quick view of how Alibaba Cloud 's 10 models compare on intelligence, output speed and price — so you can pick the right one for your use case.
Most intelligent
Intelligence index · 10 models
Fastest
Output tokens / second · 10 models
Lowest price
Blended price per 1M tokens · 10 models
Compared with other hosts
How Alibaba Cloud sits against the rest of the catalog, using the same measurements shown on every provider page.
- Price. Median blended price of $6.23 per 1M tokens — 2% above the $6.13 median across all 50 providers in the catalog.
- Throughput. Median output speed of 150 tokens/s, 57% slower than the 348 tok/s catalog median.
- Catalog depth. 10 models listed, against a catalog average of 5 per provider.
- Integration. Identical to every other provider here — same endpoint, same SDK, same keys, so a switch costs one string change.
Median blended price per 1M tokens
lower is better| Alibaba Cloud | $6.23 | |
| Wafer | $6.22 | |
| SambaNova | $6.20 | |
| $6.31 | ||
| Arcee AI | $6.07 | |
| Anthropic | $6.63 |
Median output speed
higher is better| Alibaba Cloud | 150 tok/s | |
| Wafer | 303 tok/s | |
| SambaNova | 954 tok/s | |
| 239 tok/s | ||
| Arcee AI | 369 tok/s | |
| Anthropic | 365 tok/s |
Figures are medians across the models listed on each provider page and are refreshed with the catalog.
Models & pricing
Every Alibaba Cloud model reachable with your LLMAPI key. Click a column heading to sort.
| Model | API model name | Uptime 24h | Price / 1M | Context |
|---|---|---|---|---|
| GLM-5.3 | glm-5.3 | 99.91% | $ 2.15 | 1000 k |
| Qwen 3.8 Max | qwen3.8-max | — | $ 3 | 1000 k |
| DeepSeek V4 Flash 0731 | deepseek-v4-flash-0731 | 99.68% | $ 0.66 | 1000 k |
| Kimi K3 | kimi-k3 | 93.83% | $ 6 | 1049 k |
| Kimi K2.7-Code | kimi-k2.7-code | 97.36% | $ 1.71 | 262 k |
| Qwen3.7 Max | qwen3.7-max | 99.98% | $ 1.13 | 1000 k |
| Qwen 3.7 Plus | qwen3.7-plus | 99.96% | $ 1.13 | 1000 k |
| Qwen3 Max 2026-01-23 | qwen3-max-2026-01-23 | — | $ 2.4 | 262 k |
| Qwen 3.6 Plus | qwen3.6-plus | 99.96% | $ 1.13 | 1049 k |
| Qwen3 VL 235B A22B Instruct | qwen3-vl-235b-a22b-instruct | 98.66% | $ 0.88 | 131 k |
| Qwen3 VL 235B A22B Thinking | qwen3-vl-235b-a22b-thinking | 87.60% | $ 0.88 | 131 k |
| Qwen3 Next 80B A3B Thinking | qwen3-next-80b-a3b-thinking | 99.25% | $ 1.88 | 131 k |
| Qwen3 Next 80B A3B Instruct | qwen3-next-80b-a3b-instruct | 99.74% | $ 0.88 | 131 k |
| Qwen3 Max | qwen3-max | 99.78% | $ 6 | 262 k |
| Qwen3 Coder Plus | qwen3-coder-plus | 100.00% | $ 19.5 | 1000 k |
| Qwen3 Coder Flash | qwen3-coder-flash | 99.98% | $ 0.6 | 1000 k |
| Qwen3 VL Plus | qwen3-vl-plus | — | $ 0.55 | 262 k |
| Qwen3 VL Flash | qwen3-vl-flash | — | $ 0.14 | 262 k |
| QwQ Plus | qwq-plus | — | $ 1.2 | 131 k |
| Qwen Coder Plus | qwen-coder-plus | — | $ 2 | 131 k |
| Qwen Omni Turbo | qwen-omni-turbo | — | $ 0.35 | 33 k |
| Qwen2.5 VL 32B Instruct | qwen2-5-vl-32b-instruct | — | $ 2.1 | 131 k |
| Qwen Flash | qwen-flash | — | $ 0.14 | 1000 k |
| Qwen VL Plus | qwen-vl-plus | — | $ 0.32 | 131 k |
| Qwen VL Max | qwen-vl-max | — | $ 1.4 | 131 k |
Price is a blended per-1M-token figure; latency is time to first token and throughput is output tokens per second, measured on the LLMAPI edge.
Cost calculator
Pick a model, enter your traffic, and see the monthly bill at LLMAPI rates.
Estimates use the blended per-1M-token price shown in the table above. Word counts differ by language — Cyrillic and CJK text uses 2–3× more tokens per word.
Alibaba Cloud via LLMAPI
Same models, same features, one key. Prefix the model ID with alibaba-cloud/ and point your OpenAI client at our base URL.
# Python · OpenAI SDK from openai import OpenAI client = OpenAI( base_url="https://api.llmapi.ai/v1", api_key="LLMAPI_KEY", ) r = client.chat.completions.create( model="alibaba-cloud/qwen3-max", messages=[{"role": "user", "content": "Hello"}], )
| Pricing | Eligible for LLMAPI volume discounts; billed on one invoice with every other provider. |
| Features | Full pass-through — tool calling, JSON schema, streaming, vision and reasoning behave exactly as on Alibaba Cloud’s own API. |
| Fallback | Configurable. On 429 or 5xx the router switches to any model or provider on your fallback list. |
| Regions | APAC, Global |
| Your data | LLMAPI stores request metadata only — model, token counts, latency and status. No prompt or completion content. |
Capabilities & use cases
Each tag below is its own catalog page listing every provider that shares it.
Developer resources
Where to go next.
Join thousands of developers building on one API key
Every provider in the catalog, one key, one invoice.
Questions
How do I call Alibaba Cloud models through LLMAPI?
Point any OpenAI-compatible client at https://api.llmapi.ai/v1 and use the model ID alibaba-cloud/qwen3-max. No other change is needed.
Does Alibaba Cloud cost more through LLMAPI?
No. This provider is eligible for LLMAPI volume discounts, so heavy usage lands below list price.
Which regions are used?
APAC, Global
What happens if the provider returns an error?
On 429 or 5xx responses the router falls back to any other model or provider on your configured list, so requests keep succeeding.
How many Alibaba Cloud models are available?
10 at the moment, all listed in the models table above and updated as the provider ships new ones.