This model is deprecated. We only continue performance benchmarking for the default 10k input token workload. Results for other workloads are historical and no longer updated.

Kimi has launched a newer model, Kimi K2.6. We suggest considering it instead.

For more information, see comparison of Kimi K2.6 to other models and API provider benchmarks for Kimi K2.6.

Kimi K2.5

Open weights model

Released January 2026

Kimi K2.5 (Reasoning) Intelligence, Performance & Price Analysis

Model summary

IntelligenceUpdated

28
Artificial Analysis Intelligence Index
3 out of 4 units for Intelligence.

Speed

N/A
Output tokens per second
Unknown out of 4 units for Speed.
In $0.60Out $2.75Cache Discount 42%
N/A
Cost per Intelligence Index task
Unknown out of 4 units for Cost.

Verbosity

N/A
Output tokens from Intelligence Index
Unknown out of 4 units for Verbosity.

Kimi K2.5 (Reasoning) is above average in intelligence, but somewhat expensive when comparing to other open weight models of similar size. The model supports text, image, and video input, outputs text, and has a 256k tokens context window.

Kimi K2.5 (Reasoning) scores 28 on the Artificial Analysis Intelligence Index, placing it above average among comparable models (median: 22).

Pricing for Kimi K2.5 (Reasoning) is $0.60 per 1M input tokens (somewhat expensive, median: $0.30) and $2.75 per 1M output tokens (somewhat expensive, median: $1.17).

ReasoningYes

This page shows the reasoning version of this model.

A non-reasoning variant may also exist.

Input modality

Supports: text, image, and video

Output modality

Supports: text

Context window256k
~384 A4 pages of size 12 Arial font
Total parameters1T
Active parameters32B
Number of parameters active per token during inference
LicenseModified MIT License
Model weightsHugging Face

Metrics are compared against models of the same class:

  • Non-reasoning models → compared only with other non-reasoning models
  • Reasoning models → compared across both reasoning and non-reasoning
  • Open weights models → compared only with other open weights models of the same size class:
    • Tiny: ≤4B parameters
    • Small: 4B–40B parameters
    • Medium: 40B–150B parameters
    • Large: >150B parameters
  • Proprietary models → compared across proprietary and open weights models of the same price range, using a blended 3:1 input/output price ratio:
    • <$0.15 per 1M tokens
    • $0.15–$1 per 1M tokens
    • >$1 per 1M tokens

Highlights

Updated
Artificial Analysis Intelligence Index · Higher is better

Speed

Output tokens per second · Higher is better
Weighted average cost (USD) per Intelligence Index task · Lower is better

IntelligenceUpdated

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.2 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)

Artificial Analysis Intelligence Index v4.2 includes: AA-Briefcase, GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Artificial Analysis Intelligence Index by Open Weights / Proprietary

Artificial Analysis Intelligence Index v4.2 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)

Artificial Analysis Intelligence Index v4.2 includes: AA-Briefcase, GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Indicates whether the model weights are available. Models are labelled as 'Commercial Use Restricted' if commercial use is limited by conditions, and as 'Non-commercial' if the license prohibits commercial use.

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
See more

Agentic knowledge work, (Elo-500)/2000

Agentic real-world work tasks, (Elo-500)/2000

Agentic tool use

Agentic coding & terminal use

Coding

Reasoning & knowledge

Professional document reasoning, All-pass

Physics reasoning

Long context reasoning

Agentic SaaS workflows

Legal agentic work, criterion pass rate

Agentic business operations

Scientific reasoning

Quantitative analysis on spreadsheets & documents

Instruction following

Long-horizon agentic tasks

Kubernetes incident root-cause analysis

Visual reasoning

While model intelligence generally translates across use cases, specific evaluations may be more relevant for certain use cases.

Artificial Analysis Intelligence Index v4.2 includes: AA-Briefcase, GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

AA-Omniscience

AA-Omniscience Index

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.

Openness Index

Artificial Analysis Openness Index: Score

Openness Index assesses model openness on a 0 to 100 normalized scale (higher is more open)

Intelligence Index Comparisons

Intelligence Index vs. Cost per Intelligence Index Task

Artificial Analysis Intelligence Index · Weighted average cost (USD) per Artificial Analysis Intelligence Index task
Most attractive quadrant
Pareto line

Weighted average cost per Intelligence Index task. Each evaluation’s cost is calculated from input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight.

Artificial Analysis Intelligence Index v4.2 includes: AA-Briefcase, GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

Cost

Cost per Intelligence Index Task

Weighted average cost (USD) per Artificial Analysis Intelligence Index task, segmented by token type. Lower is better

Weighted average cost per Intelligence Index task. Each evaluation’s cost is calculated from input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight.

Cost to Run Artificial Analysis Intelligence Index

Cost (USD) to run all evaluations in the Artificial Analysis Intelligence Index

The cost to run the evaluations in the Artificial Analysis Intelligence Index, calculated using the model's input, cache hit, cache write, reasoning, and answer token prices and the number of tokens used across evaluations (excluding repeats).

Pricing: Cache Hit, Input, and Output

Price (USD per M Tokens)

Price per token for cached prompts (previously processed), typically offering a significant discount compared to regular input price, represented as USD per million tokens. The values shown here are the cache hit price; cache write and cache storage are billed separately and vary by provider — see "Cache pricing by provider" for detail.

Context Window

Context Window

Context window: tokens limit · Higher is better

Larger context windows are relevant to RAG (Retrieval Augmented Generation) LLM workflows which typically involve reasoning and information retrieval of large amounts of data.

Maximum number of combined input & output tokens. Output tokens commonly have a significantly lower limit (varied by model).

Model Size (Open Weights Models Only)

Model Size: Total and Active Parameters

Comparison between total model parameters and parameters active during inference

The total number of trainable weights and biases in the model, expressed in billions. These parameters are learned during training and determine the model's ability to process and generate responses.

The number of parameters actually executed during each inference forward pass, expressed in billions. For Mixture of Experts (MoE) models, a routing mechanism selects a subset of experts per token, resulting in fewer active than total parameters. Dense models use all parameters, so active equals total.

Frequently Asked Questions

Common questions about Kimi K2.5 (Reasoning)

Kimi K2.5 (Reasoning) was released on January 27, 2026.

Kimi K2.5 (Reasoning) was created by Kimi.

Kimi K2.5 (Reasoning) scores 28 (estimated) on the Artificial Analysis Intelligence Index, placing it above average among other open weight models of similar size (median: 22).

Kimi K2.5 (Reasoning) costs $0.60 per 1M input tokens (somewhat higher than average, median: $0.50) and $2.75 per 1M output tokens (somewhat higher than average, median: $2.20), based on the median across providers serving the model.

Kimi K2.5 (Reasoning) costs $0.60 per 1M input tokens and $2.75 per 1M output tokens (based on the median across providers serving the model). For a blended rate (7:2:1 cache hit/input/output ratio), this is $0.64 per 1M tokens. Pricing may vary by provider. Compare provider pricing

Yes, Kimi K2.5 (Reasoning) is a reasoning model. It uses extended thinking or chain-of-thought reasoning to work through complex problems before providing an answer.

Kimi K2.5 (Reasoning) supports text, image, and video input.

Kimi K2.5 (Reasoning) supports text output.

Yes, Kimi K2.5 (Reasoning) supports image input and can analyze, describe, and answer questions about images.

Yes, Kimi K2.5 (Reasoning) is multimodal. It can process text, image, and video input and generate text output.

Kimi K2.5 (Reasoning) has a context window of 260k tokens. This determines how much text and conversation history the model can process in a single request.

Yes, Kimi K2.5 (Reasoning) is open weights. The model weights are publicly available and can be downloaded for self-hosting.

Kimi K2.5 (Reasoning) has 1 trillion parameters (32 billion active).

Kimi K2.5 (Reasoning) is a Mixture of Experts (MoE) model with 1 trillion total parameters, but only 32 billion active parameters are used during inference.

Kimi K2.5 (Reasoning) is released under the Modified MIT License license. Commercial use is allowed with restrictions. View license

Kimi K2.5 (Reasoning) achieves a score of 28 on the Artificial Analysis Intelligence Index. This composite benchmark evaluates models across reasoning, knowledge, mathematics, and coding.

Yes, Kimi K2.5 (Reasoning) is available via API through 6 providers. Compare API providers

Kimi K2.5 (Reasoning) is available through 6 API providers. Compare providers