GMI: Models Intelligence, Performance & Price

GMI
GMI

This analysis is intended to support you in choosing the best model provided by GMI for your use-case.

Most Intelligent

Updated
#1
GLM-5.3-FlashGLM-5.3-Flash
42
#2
DeepSeek V4 Pro 0813 (max)DeepSeek V4 Pro 0813 (max)
36
#3
DeepSeek V4 Flash 0731 (max)DeepSeek V4 Flash 0731 (max)
35
#4
Kimi K2.7 Code (FP8)Kimi K2.7 Code (FP8)
26
#5
Hy3Hy3
26

Intelligence index

Total 9 models

Fastest

#1
Nemotron 3.5 Lightning (BF16)Nemotron 3.5 Lightning (BF16)
298 t/s
#2
DeepSeek V4 Flash 0731 (max)DeepSeek V4 Flash 0731 (max)
159 t/s
#3
GLM-5.3-FlashGLM-5.3-Flash
105 t/s
#4
Hy3Hy3
96 t/s
#5
MiniMax-M2.5 FP8MiniMax-M2.5 FP8
92 t/s

Output speed

Total 9 models

Lowest Price

#1
GLM-5.3-FlashGLM-5.3-Flash
$0.10
#2
Hy3Hy3
$0.11
#3
DeepSeek V3.2 (Non-reasoning)DeepSeek V3.2 (Non-reasoning)
$0.22
#4
DeepSeek V4 Flash 0731 (max)DeepSeek V4 Flash 0731 (max)
$0.23
#5
MiniMax-M2.5 FP8MiniMax-M2.5 FP8
$0.39

Blended price (per 1M tokens)

Total 9 models

GMI offers 9 models, each with different intelligence, performance, and pricing characteristics. Below is a comparison of the key metrics across models.

  • For intelligence, the top models on GMI are GLM-5.3-Flash (42), DeepSeek V4 Pro 0813 (max) (36), and DeepSeek V4 Flash 0731 (max) (35).
  • For output speed, the fastest models are Nemotron 3.5 Lightning (BF16) (298 t/s), DeepSeek V4 Flash 0731 (max) (159 t/s), and GLM-5.3-Flash (105 t/s). Speed varies significantly across models, with a 223% difference between the fastest and slowest.
  • For latency, DeepSeek V3.2 (Non-reasoning) (2.65s), Kimi K2.5 (Non-reasoning) (3.30s), and Nemotron 3.5 Lightning (BF16) (8.44s) offer the lowest time to first answer token.
  • For pricing, GLM-5.3-Flash ($0.10), Hy3 ($0.11), and DeepSeek V3.2 (Non-reasoning) ($0.22) offer the lowest blended prices per 1M tokens. Prices vary up to 3.9x across models.
  • For context window size, DeepSeek V4 Pro 0813 (max) (1M), DeepSeek V4 Flash 0731 (max) (1M), and GLM-5.3-Flash (1M) support the largest context windows on GMI.
  • GLM-5.3-Flash provides the best balance of intelligence and cost-effectiveness. For the fastest output, Nemotron 3.5 Lightning (BF16) is the top choice.

Highlights

Updated
Artificial Analysis Intelligence Index · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens (blended) · Lower is better

Intelligence Evaluations

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
See more

Agentic knowledge work, (Elo-500)/2000

Agentic real-world work tasks, (Elo-500)/2000

Agentic SaaS workflows

Agentic coding & terminal use

Coding

Reasoning & knowledge

Professional document reasoning, All-pass

Physics reasoning

Long context reasoning

Legal agentic work, criterion pass rate

Agentic business operations

Quantitative analysis on spreadsheets & documents

No data available

Instruction following

Agentic tool use

Long-horizon agentic tasks

No data available

Kubernetes incident root-cause analysis

No data available

Visual reasoning

Intelligence Index vs. Price

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Context Window

Context Window

Context window: tokens limit · Higher is better

Pricing

Intelligence Index vs. Price

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Performance Summary

Output Speed vs. Price

Output speed: output tokens per second · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Speed

Measured by Output Speed (tokens per second)

Output Speed

Output tokens per second · Higher is better

Latency

Measured by Time (seconds) to First Token

Latency: Time To First Answer Token

Seconds to first answer token received · Accounts for reasoning model 'thinking' time

End-to-End Response Time

Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed

End-to-End Response Time vs. Price

End-to-end response time: end-to-end seconds to output 500 tokens · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Further Analysis
Z AI logo
GLM-5.3-Flash
1M
Open
42
$0.42
105
2.17
26.04
19.10
DeepSeek logo
DeepSeek V4 Pro 0813 (max)
1.05M
Open
36
$1.61
74
3.10
36.99
27.11
DeepSeek logo
DeepSeek V4 Flash 0731 (max)
1.05M
Open
35
$0.58
159
2.29
17.99
12.56
Z AI logo
GLM-5.2 (max) (FP8)
1.05M
Open
34
--
--
--
--
--
DeepSeek logo
DeepSeek V4 Pro (max)
1.05M
Open
31
--
--
--
--
--
MiniMax logo
MiniMax-M3
1.05M
Open
30
--
--
--
--
--
Kimi logo
Kimi K2.6 FP8
262k
Open
27
--
--
--
--
--
Z AI logo
GLM-5.1 (FP8)
203k
Open
26
--
--
--
--
--
Xiaomi logo
MiMo-V2.5-Pro
1.05M
Open
26
--
--
--
--
--
Kimi logo
Kimi K2.7 Code (FP8)
65.5k
Open
26
$1.20
57
5.16
53.06
39.12
Tencent logo
Hy3
262k
Open
26
$0.07
96
3.10
29.23
20.90
DeepSeek logo
DeepSeek V4 Flash (high)
1.05M
Open
25
--
--
--
--
--
DeepSeek logo
DeepSeek V4 Flash (max)
1.05M
Open
25
--
--
--
--
--
NVIDIA logo
Nemotron 3 Ultra
262k
Open
23
--
--
--
--
--
MiniMax logo
MiniMax-M2.7 (FP8)
197k
Open
23
--
--
--
--
--
Alibaba logo
Qwen3.5 27B (FP8)
262k
Open
23*
--
--
--
--
--
MiniMax logo
MiniMax-M2.5 FP8
197k
Open
23*
--
92
1.86
28.95
21.68
Tencent logo
Hy3-preview
262k
Open
23*
--
--
--
--
--
Xiaomi logo
MiMo-V2.5
1.05M
Open
22
$0.04
--
--
--
--
Kimi logo
Kimi K2.5 (Non-reasoning)
262k
Open
19*
--
40
3.30
15.73
--
Alibaba logo
Qwen3.5 35B A3B (FP8)
262k
Open
19*
--
--
--
--
--
Alibaba logo
Qwen3.5 397B A17B (FP8)
262k
Open
19
$0.47
--
--
--
--
DeepSeek logo
DeepSeek V4 Flash (Non-reasoning)
1.05M
Open
19*
--
--
--
--
--
Alibaba logo
Qwen3.6 35B A3B FP8
262k
Open
19
$0.31
--
--
--
--
Xiaomi logo
MiMo-V2.5-Pro (Non-reasoning)
1.05M
Open
18*
--
--
--
--
--
Tencent logo
Hy3-preview (Non-reasoning)
262k
Open
17*
--
--
--
--
--
Google logo
Gemma 4 26B A4B (FP8)
1.05M
Open
17*
--
--
--
--
--
Alibaba logo
Qwen3.5 122B A10B (FP8)
262k
Open
16
$0.32
--
--
--
--
DeepSeek logo
DeepSeek V3.2 (Non-reasoning)
164k
Open
16*
--
50
2.65
12.58
--
Google logo
Gemma 4 31B (FP8)
262k
Open
15
$0.10
--
--
--
--
Alibaba logo
Qwen3.6 35B A3B (Non-reasoning) FP8
262k
Open
15*
--
--
--
--
--
NVIDIA logo
Nemotron 3.5 Lightning (BF16)
65.5k
Open
14
$0.00
298
1.72
10.12
6.72
Google logo
Gemma 4 26B A4B (Non-reasoning) (FP8)
1.05M
Open
13*
--
--
--
--
--
Alibaba logo
Qwen3 Next 80B A3B (Reasoning)
262k
Open
11*
--
--
--
--
--
Alibaba logo
Qwen3 Next 80B A3B
262k
Open
10*
--
--
--
--
--

Key definitions

Frequently Asked Questions

Common questions about GMI