GMI:模型智能、性能与价格

GMI
GMI

本分析旨在帮助你根据使用场景,选择 GMI 提供的最佳模型。

最智能

#1
MiMo-V2.6-Pro (BF16)MiMo-V2.6-Pro (BF16)
46
#2
GLM-5.3 (max)GLM-5.3 (max)
45
#3
GLM-5.3-FlashGLM-5.3-Flash
42
#4
DeepSeek V4.1 Flash (max)DeepSeek V4.1 Flash (max)
39
#5
MiMo-V2.6-Flash (BF16)MiMo-V2.6-Flash (BF16)
38

Intelligence Index

共 12 个模型

速度最快

#1
DeepSeek V4 Flash Vision (max)DeepSeek V4 Flash Vision (max)
221 t/s
#2
DeepSeek V4.1 Flash (max)DeepSeek V4.1 Flash (max)
160 t/s
#3
DeepSeek V4 Flash 0731 (max)DeepSeek V4 Flash 0731 (max)
107 t/s
#4
MiniMax-M2.5 FP8MiniMax-M2.5 FP8
91 t/s
#5
Hy3Hy3
89 t/s

输出速度

共 12 个模型

价格最低

#1
MiMo-V2.6-Flash (BF16)MiMo-V2.6-Flash (BF16)
$0.06
#2
GLM-5.3-FlashGLM-5.3-Flash
$0.06
#3
Hy3Hy3
$0.11
#4
DeepSeek V4.1 Flash (max)DeepSeek V4.1 Flash (max)
$0.11
#5
DeepSeek V4 Flash 0731 (max)DeepSeek V4 Flash 0731 (max)
$0.15

每 100 万 token 的混合价格

共 12 个模型

GMI 提供 12 个模型,每个模型的智能、性能和价格特征各不相同。 下方对比了各模型的关键指标。

  • 智能方面,GMI 上表现最好的模型是 MiMo-V2.6-Pro (BF16)(46)、GLM-5.3 (max)(45)和GLM-5.3-Flash(42)。
  • 输出速度方面,最快的模型是 DeepSeek V4 Flash Vision (max)(221 t/s)、DeepSeek V4.1 Flash (max)(160 t/s)和DeepSeek V4 Flash 0731 (max)(107 t/s)。 各模型之间的速度差异显著,最快与最慢相差 147%。
  • 延迟方面,Kimi K2.5 (non-reasoning)(2.98 秒)、DeepSeek V4 Flash Vision (max)(11.47 秒)和DeepSeek V4.1 Flash (max)(14.63 秒) 的首个答案 Token 延迟最低。
  • 价格方面,MiMo-V2.6-Flash (BF16)($0.06)、GLM-5.3-Flash($0.06)和Hy3($0.11) 每 100 万 token 的混合价格最低。 各模型价格最多相差 2.6 倍。
  • 上下文窗口方面,DeepSeek V4 Pro 0813 (max)(1M)、DeepSeek V4 Flash 0731 (max)(1M)和MiMo-V2.6-Pro (BF16)(1M) 支持 GMI 上最大的上下文窗口。
Artificial Analysis Intelligence Index · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens (blended) · Lower is better

智能评测

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
查看更多

Agentic knowledge work, (Elo-500)/2000

Agentic real-world work tasks, (Elo-500)/2000

Agentic SaaS workflows

Agentic coding & terminal use

SciCodeUnder review

Coding

Reasoning & knowledge

Professional document reasoning, All-pass

CritPtUnder review

Physics reasoning

Long context reasoning

Legal agentic work, Hallucination-Gated All-Pass Rate

Agentic business operations

Agentic scientific research workflows in a terminal

Quantitative analysis on spreadsheets & documents

No data available

Kubernetes incident root-cause analysis

Visual reasoning

Medical long context reasoning

Intelligence Index 与价格

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

上下文窗口

上下文窗口

Context window: tokens limit · Higher is better

价格

Intelligence Index 与价格

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

性能摘要

输出速度与价格

Output speed: output tokens per second · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

速度

按输出速度(每秒 token 数)衡量

输出速度

Output tokens per second · Higher is better

延迟

按首 Token 延迟(秒)衡量

延迟: 首个答案 Token 延迟

收到首个回答 token 的秒数 · 包含推理模型的“思考”时间

缓存行为

缓存命中率

可缓存输入 token 中从缓存提供的比例 · 最近四周的中位数,更新于 2026年9月26日

每任务成本与缓存命中率

Weighted average cost (USD) per Intelligence Index task · Cache hit rate: median of the last four weeks, updated Sep 26, 2026
Most attractive quadrant

端到端响应时间

Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed

端到端响应时间与价格

End-to-end response time: end-to-end seconds to output 500 tokens · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

进一步分析
Xiaomi 标志
MiMo-V2.6-Pro (BF16)
1M
开放
46
$0.12
44
6.03
63.35
45.86
Z AI 标志
GLM-5.3 (max)
1M
开放
45
$1.42
78
3.80
36.03
25.78
Z AI 标志
GLM-5.3-Flash
1M
开放
42
$0.31
55
3.92
49.77
36.68
DeepSeek 标志
DeepSeek V4.1 Flash (max)
1M
开放
39
$0.38
160
2.18
17.85
12.54
Xiaomi 标志
MiMo-V2.6-Flash (BF16)
524k
开放
38
$0.06
50
3.78
54.09
40.24
DeepSeek 标志
DeepSeek V4 Pro 0813 (max)
1.05M
开放
36
$1.76
66
3.60
41.39
30.23
DeepSeek 标志
DeepSeek V4 Flash Vision (max)
1M
专有
35
$0.28
220
2.41
13.76
9.08
DeepSeek 标志
DeepSeek V4 Flash 0731 (max)
1.05M
开放
34
--
108
2.85
26.04
18.55
Z AI 标志
GLM-5.2 (max) (FP8)
1.05M
开放
34
--
--
--
--
--
DeepSeek 标志
DeepSeek V4 Pro (max)
1.05M
开放
30
--
--
--
--
--
MiniMax 标志
MiniMax-M3
1.05M
开放
29
--
--
--
--
--
Kimi 标志
Kimi K2.6 FP8
262k
开放
27
--
--
--
--
--
Z AI 标志
GLM-5.1 (FP8)
203k
开放
26
--
--
--
--
--
Xiaomi 标志
MiMo-V2.5-Pro
1.05M
开放
26
--
--
--
--
--
Kimi 标志
Kimi K2.7 Code (FP8)
65.5k
开放
26
$1.37
83
5.22
38.11
26.86
Tencent 标志
Hy3
262k
开放
25
$0.07
89
3.34
31.30
22.37
Xiaomi 标志
MiMo-V2.5
1.05M
开放
25*
--
--
--
--
--
DeepSeek 标志
DeepSeek V4 Flash (high)
1.05M
开放
24
--
--
--
--
--
DeepSeek 标志
DeepSeek V4 Flash (max)
1.05M
开放
24
--
--
--
--
--
NVIDIA 标志
Nemotron 3 Ultra
262k
开放
23
--
--
--
--
--
Alibaba 标志
Qwen3.5 27B (FP8)
262k
开放
23*
--
--
--
--
--
MiniMax 标志
MiniMax-M2.5 FP8
197k
开放
23*
--
90
2.04
29.68
22.11
MiniMax 标志
MiniMax-M2.7 (FP8)
197k
开放
23
--
--
--
--
--
Tencent 标志
Hy3-preview
262k
开放
23*
--
--
--
--
--
Kimi 标志
Kimi K2.5 (non-reasoning)
262k
开放
19*
--
55
3.15
12.26
--
Alibaba 标志
Qwen3.5 35B A3B (FP8)
262k
开放
19*
--
--
--
--
--
DeepSeek 标志
DeepSeek V4 Flash (non-reasoning)
1.05M
开放
19*
--
--
--
--
--
Alibaba 标志
Qwen3.5 397B A17B (FP8)
262k
开放
18
$0.47
--
--
--
--
Xiaomi 标志
MiMo-V2.5-Pro (non-reasoning)
1.05M
开放
18*
--
--
--
--
--
Alibaba 标志
Qwen3.6 35B A3B FP8
262k
开放
18
$0.31
--
--
--
--
Tencent 标志
Hy3-preview (non-reasoning)
262k
开放
17*
--
--
--
--
--
Google 标志
Gemma 4 26B A4B (FP8)
1.05M
开放
17*
--
--
--
--
--
DeepSeek 标志
DeepSeek V3.2 (non-reasoning)
164k
开放
16*
--
--
--
--
--
Alibaba 标志
Qwen3.5 122B A10B (FP8)
262k
开放
16
$0.32
--
--
--
--
Alibaba 标志
Qwen3.6 35B A3B (non-reasoning) FP8
262k
开放
15*
--
--
--
--
--
Google 标志
Gemma 4 31B (FP8)
262k
开放
15
$0.10
--
--
--
--
Google 标志
Gemma 4 26B A4B (non-reasoning) (FP8)
1.05M
开放
13*
--
--
--
--
--
Alibaba 标志
Qwen3 Next 80B A3B
262k
开放
11*
--
--
--
--
--

关键定义

常见问题

关于 GMI 的常见问题