Bonus: Top up now and we'll double your first deposit — get x2 credits instantly.

Kimi K2 0905 API: Pricing, Context & Providers

Kimi K2 0905 is the September update of Kimi K2 0711.

What is Kimi K2 0905?

Kimi K2 0905 is the September update of Kimi K2 0711.

Developer: Moonshot AI. Released 4 Sep 2025. Context window 262,144 tokens, up to 98,304 output tokens.


Can you self-host Kimi K2 0905?

No. We found no public weights for Kimi K2 0905 on Hugging Face, so there is nothing to download, quantize or serve on your own GPUs: no VRAM budget, no vLLM or GGUF build to plan for. The only way to run it is over an API.

Not on LLM.API yet

We do not route this model through our API at the moment. Browse the models you can call today — most workloads have a close match already live.

Providers

Companies that host this model today, with their public list prices. LLM API does not route this model yet.

List price by provider ($ / 1M tokens)

InputOutput
Novita$0.6 in
$2.5 out

Provider list prices from OpenRouter's public catalogue.

ProviderInput /MOutput /MCache read /MContextPrecisionUptime (24h)
Novita$0.6$2.5262,144fp8100.0%

Source: OpenRouter public catalogue. Last updated 23 Sep 2026.

Not available on LLM API yet

We do not route this model through our API at the moment, so there is no endpoint or code snippet for it yet. Browse the models you can call today — most workloads have a close match already live.

Browse available models

Why Build on LLM.API?

One unified API. Every major model. Built-in reliability, cost control, and observability.

  • Unified AI Routing

    Intelligently route each request across models and providers based on latency, cost, or quality. One integration that always picks the best path for you.

    Smart multi-model routing
  • Cost-Aware Orchestration

    Define budget and quality targets, then let LLM.API choose the optimal models. Automatically downgrade, upgrade, or mix providers to keep spend under control.

    Optimize every token
  • Automatic Fallbacks

    Configure policy-based failover across regions and providers. When a model errors or times out, LLM.API seamlessly retries on backups without changing your code.

    Resilience by default
  • Deep Observability

    Centralize logs, traces, metrics, and cost for every provider in one place. Quickly debug prompts, spot regressions, and understand real-world model performance.

    See every request
  • Task-Level Abstractions

    Describe tasks—chat, scoring, extraction—once and let LLM.API match them to the right models and prompts. Ship features faster with consistent, reusable interfaces.

    From models to tasks
  • High-Throughput Batching

    Send thousands of requests in a single batch with built-in rate control and retries. Maximize throughput while staying within provider limits and budgets.

    Scale without throttling

5 Core Capabilities

  • Agentic coding intelligence

    Tuned for coding agents: 69.2 on SWE-bench Verified and 44.5 on Terminal-Bench, both up on the previous K2 release.

  • 256K context window

    Context doubled from 128K to 256K tokens, enough to review a whole codebase in one pass.

  • Front-end generation

    The update targeted the look and practicality of front-end code, including web pages and charts.

  • Native tool calling

    Built for multi-step planning with tool use, including file editing and terminal operations in coding harnesses.

  • 1T-parameter open MoE

    One trillion total parameters with 32B active per token across 384 experts, published under a modified MIT licence.

6 Most Valuable Use Cases

  • Autonomous coding agents that edit files and run terminal commands
  • Whole-repository review and bug tracing inside a 256K context
  • Front-end and UI generation where output has to look presentable
  • Multilingual software maintenance across mixed-language codebases
  • Long-horizon agent tasks that span many tool calls
  • Self-hosted deployments needing open weights with FP8 quantisation

When to Use — When NOT to Use

Use it if...

  • Your main workload is agentic coding with tool calls
  • You need a context window beyond 128K tokens
  • You want open weights competitive with proprietary coding models
  • You run a coding harness and want smooth tool-use integration

Avoid if...

  • You need multimodal input — this model is text in, text out
  • You want a reasoning-first model with visible thinking traces
  • You need the cheapest possible general chat model
  • You want the current Kimi generation, since later K2 releases supersede this checkpoint

Kimi K2 0905 benchmark scores

Intelligence index

This model15
Tracked median12

Scale: 0-100 index points

Reference price per 1M tokens

Input$0.60
Output$2.50

Bars compare published reference list prices, not LLM.API pricing.

Artificial Analysis Intelligence Index15
Median index across all tracked models12
SWE-bench Verified (official model card)69.2
SWE-bench Multilingual (official model card)55.9
Terminal-Bench (official model card)44.5
Reference input price$0.60 / 1M tokens
Reference output price$2.50 / 1M tokens

Scores as published by Artificial Analysis and the official model card (source). Reference prices are provider list prices, not LLM.API pricing. Figures with no published value are omitted.

What developers say about Kimi K2 0905

Summarised from publicly published developer write-ups and the model's own documentation. Opinions are the sources’, not LLM.API’s.

  • Reviewers describe 0905 as a coding-and-agents upgrade to the July K2 release rather than a new model family.
  • The doubled 256K context is the most cited practical change, letting agents work across a repository without losing track.
  • Testers report smoother integration with coding harnesses, with file editing and terminal steps working without the earlier friction.
  • Its SWE-bench Verified score puts it level with proprietary coding models of the period while keeping open weights.

Frequently Asked Questions

  • What is Kimi K2 0905?

    Kimi K2 0905 is the September update of Kimi K2 0711.

  • Who makes Kimi K2 0905?

    Kimi K2 0905 is developed by Moonshot AI. It was released on 4 Sep 2025.

  • What is the context window of Kimi K2 0905?

    262,144 tokens, with up to 98,304 output tokens per response.

  • How much does Kimi K2 0905 cost?

    The reference list price is $0.6 per 1M input tokens and $2.5 per 1M output tokens. The cheapest host right now is Novita at $0.6 / $2.5 per 1M tokens.

  • Which providers host Kimi K2 0905?

    Novita.

  • What modalities does Kimi K2 0905 support?

    Text input and text output. It supports tool calling, structured outputs, JSON mode.

  • Can I use Kimi K2 0905 through LLM API?

    Not yet. LLM API does not route this model at the moment. Browse the models page for close alternatives you can call today with one API key.

Get one key to every model

Swap your API key. Keep your code.