Kimi K2 0905 API: Pricing, Context & Providers
Kimi K2 0905 is the September update of Kimi K2 0711.
What is Kimi K2 0905?
Kimi K2 0905 is the September update of Kimi K2 0711.
Developer: Moonshot AI. Released 4 Sep 2025. Context window 262,144 tokens, up to 98,304 output tokens.
Can you self-host Kimi K2 0905?
No. We found no public weights for Kimi K2 0905 on Hugging Face, so there is nothing to download, quantize or serve on your own GPUs: no VRAM budget, no vLLM or GGUF build to plan for. The only way to run it is over an API.
Not on LLM.API yet
We do not route this model through our API at the moment. Browse the models you can call today — most workloads have a close match already live.
Providers
Companies that host this model today, with their public list prices. LLM API does not route this model yet.
List price by provider ($ / 1M tokens)
InputOutputProvider list prices from OpenRouter's public catalogue.
| Provider | Input /M | Output /M | Cache read /M | Context | Precision | Uptime (24h) |
|---|---|---|---|---|---|---|
| Novita | $0.6 | $2.5 | — | 262,144 | fp8 | 100.0% |
Source: OpenRouter public catalogue. Last updated 23 Sep 2026.
Not available on LLM API yet
We do not route this model through our API at the moment, so there is no endpoint or code snippet for it yet. Browse the models you can call today — most workloads have a close match already live.
Browse available modelsWhy Build on LLM.API?
One unified API. Every major model. Built-in reliability, cost control, and observability.
-
Unified AI Routing
Intelligently route each request across models and providers based on latency, cost, or quality. One integration that always picks the best path for you.
Smart multi-model routing -
Cost-Aware Orchestration
Define budget and quality targets, then let LLM.API choose the optimal models. Automatically downgrade, upgrade, or mix providers to keep spend under control.
Optimize every token -
Automatic Fallbacks
Configure policy-based failover across regions and providers. When a model errors or times out, LLM.API seamlessly retries on backups without changing your code.
Resilience by default -
Deep Observability
Centralize logs, traces, metrics, and cost for every provider in one place. Quickly debug prompts, spot regressions, and understand real-world model performance.
See every request -
Task-Level Abstractions
Describe tasks—chat, scoring, extraction—once and let LLM.API match them to the right models and prompts. Ship features faster with consistent, reusable interfaces.
From models to tasks -
High-Throughput Batching
Send thousands of requests in a single batch with built-in rate control and retries. Maximize throughput while staying within provider limits and budgets.
Scale without throttling
5 Core Capabilities
Agentic coding intelligence
Tuned for coding agents: 69.2 on SWE-bench Verified and 44.5 on Terminal-Bench, both up on the previous K2 release.
256K context window
Context doubled from 128K to 256K tokens, enough to review a whole codebase in one pass.
Front-end generation
The update targeted the look and practicality of front-end code, including web pages and charts.
Native tool calling
Built for multi-step planning with tool use, including file editing and terminal operations in coding harnesses.
1T-parameter open MoE
One trillion total parameters with 32B active per token across 384 experts, published under a modified MIT licence.
6 Most Valuable Use Cases
- Autonomous coding agents that edit files and run terminal commands
- Whole-repository review and bug tracing inside a 256K context
- Front-end and UI generation where output has to look presentable
- Multilingual software maintenance across mixed-language codebases
- Long-horizon agent tasks that span many tool calls
- Self-hosted deployments needing open weights with FP8 quantisation
When to Use — When NOT to Use
Use it if...
- Your main workload is agentic coding with tool calls
- You need a context window beyond 128K tokens
- You want open weights competitive with proprietary coding models
- You run a coding harness and want smooth tool-use integration
Avoid if...
- You need multimodal input — this model is text in, text out
- You want a reasoning-first model with visible thinking traces
- You need the cheapest possible general chat model
- You want the current Kimi generation, since later K2 releases supersede this checkpoint
BENCHMARKS
Kimi K2 0905 benchmark scores
Intelligence index
Scale: 0-100 index points
Reference price per 1M tokens
Bars compare published reference list prices, not LLM.API pricing.
Scores as published by Artificial Analysis and the official model card (source). Reference prices are provider list prices, not LLM.API pricing. Figures with no published value are omitted.
COMMUNITY
What developers say about Kimi K2 0905
Summarised from publicly published developer write-ups and the model's own documentation. Opinions are the sources’, not LLM.API’s.
- Reviewers describe 0905 as a coding-and-agents upgrade to the July K2 release rather than a new model family.
- The doubled 256K context is the most cited practical change, letting agents work across a repository without losing track.
- Testers report smoother integration with coding harnesses, with file editing and terminal steps working without the earlier friction.
- Its SWE-bench Verified score puts it level with proprietary coding models of the period while keeping open weights.
SOURCES
Frequently Asked Questions
What is Kimi K2 0905?
Kimi K2 0905 is the September update of Kimi K2 0711.
Who makes Kimi K2 0905?
Kimi K2 0905 is developed by Moonshot AI. It was released on 4 Sep 2025.
What is the context window of Kimi K2 0905?
262,144 tokens, with up to 98,304 output tokens per response.
How much does Kimi K2 0905 cost?
The reference list price is $0.6 per 1M input tokens and $2.5 per 1M output tokens. The cheapest host right now is Novita at $0.6 / $2.5 per 1M tokens.
Which providers host Kimi K2 0905?
Novita.
What modalities does Kimi K2 0905 support?
Text input and text output. It supports tool calling, structured outputs, JSON mode.
Can I use Kimi K2 0905 through LLM API?
Not yet. LLM API does not route this model at the moment. Browse the models page for close alternatives you can call today with one API key.
COMPARE
Competitive Models
-
Kimi K2.6 (free)
Up to 30%
Kimi K2.6 (free) is MoonshotAI’s open-source, multimodal Mixture-of-Experts model optimized for long-horizon coding, autonomous agents, and large-context reasoning. The free variant provides access to these capabilities via selected platforms and endpoints without direct model licensing costs.
-
Kimi K2.5
Kimi K2.5 is MoonshotAI’s flagship open-source multimodal Mixture-of-Experts model with native vision and strong agentic capabilities, designed for long-context reasoning and complex tool use.
-
Kimi K2.6
Up to 30%
Kimi K2.6 is MoonshotAI’s open-source, 1-trillion-parameter Mixture-of-Experts multimodal model optimized for long-horizon coding, agentic tool use, and image/video understanding. It is notable for its large ~262K-token context window and strong performance on complex software engineering and tool-using benchmarks.
Get one key to every model
Swap your API key. Keep your code.