Microsoft Azure: Models Intelligence, Performance & Price

Microsoft Azure
Microsoft Azure

This analysis is intended to support you in choosing the best model provided by Microsoft Azure for your use-case.

Most Intelligent

Updated
#1
GPT-6 Astra (max)GPT-6 Astra (max)
53
#2
Claude Fable 5 (with fallback)Claude Fable 5 (with fallback)
50
#3
GPT-6 Sol (max)GPT-6 Sol (max)
48
#4
GPT-5.6 Sol (max)GPT-5.6 Sol (max)
47
#5
GPT-6 Astra (low)GPT-6 Astra (low)
46

Intelligence index

Total 82 models

Fastest

#1
Llama 4 Maverick (FP8)Llama 4 Maverick (FP8)
434 t/s
#2
gpt-oss-120b (low)gpt-oss-120b (low)
338 t/s
#3
gpt-oss-120b (high)gpt-oss-120b (high)
324 t/s
#4
GPT-5.4 mini (xhigh)GPT-5.4 mini (xhigh)
292 t/s
#5
GPT-4.1 nanoGPT-4.1 nano
286 t/s

Output speed

Total 82 models

Lowest Price

#1
GPT-5 nano (high)GPT-5 nano (high)
$0.05
#2
GPT-5 nano (medium)GPT-5 nano (medium)
$0.05
#3
GPT-6 Luna (max)GPT-6 Luna (max)
$0.08
#4
GPT-4.1 nanoGPT-4.1 nano
$0.08
#5
GPT-4o miniGPT-4o mini
$0.14

Blended price (per 1M tokens)

Total 82 models

Azure offers 82 models, each with different intelligence, performance, and pricing characteristics. Below is a comparison of the key metrics across models.

  • For intelligence, the top models on Azure are GPT-6 Astra (max) (53), Claude Fable 5 (with fallback) (50), and GPT-6 Sol (max) (48).
  • For output speed, the fastest models are Llama 4 Maverick (FP8) (434 t/s), gpt-oss-120b (low) (338 t/s), and gpt-oss-120b (high) (324 t/s). Speed varies significantly across models, with a 52% difference between the fastest and slowest.
  • For latency, Phi-4 Mini (0.83s), Llama 4 Scout (0.85s), and GPT-5.4 mini (non-reasoning) (1.03s) offer the lowest time to first answer token.
  • For pricing, GPT-5 nano (high) ($0.05), GPT-5 nano (medium) ($0.05), and GPT-6 Luna (max) ($0.08) offer the lowest blended prices per 1M tokens. Prices vary up to 2.7x across models.
  • For context window size, GPT-6 Sol (max) (1M), GPT-5.4 (xhigh) (1M), and Claude Fable 5 (with fallback) (1M) support the largest context windows on Azure.

Highlights

Updated
Artificial Analysis Intelligence Index · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens (blended) · Lower is better

Intelligence Evaluations

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
See more

Agentic knowledge work, (Elo-500)/2000

Agentic real-world work tasks, (Elo-500)/2000

Agentic SaaS workflows

Agentic coding & terminal use

SciCodeUnder review

Coding

Reasoning & knowledge

Professional document reasoning, All-pass

CritPtUnder review

Physics reasoning

Long context reasoning

Legal agentic work, criterion pass rate

Agentic business operations

Agentic scientific research workflows in a terminal

Quantitative analysis on spreadsheets & documents

Kubernetes incident root-cause analysis

Visual reasoning

Medical long context reasoning

Intelligence Index vs. Price

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Context Window

Context Window

Context window: tokens limit · Higher is better

Pricing

Intelligence Index vs. Price

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Performance Summary

Output Speed vs. Price

Output speed: output tokens per second · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Speed

Measured by Output Speed (tokens per second)

Output Speed

Output tokens per second · Higher is better

Latency

Measured by Time (seconds) to First Token

Latency: Time To First Answer Token

Seconds to first answer token received · Accounts for reasoning model 'thinking' time

End-to-End Response Time

Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed

End-to-End Response Time vs. Price

End-to-end response time: end-to-end seconds to output 500 tokens · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Further Analysis
OpenAI logo
GPT-6 Astra (max)
524k
Proprietary
53
$3.19
69
354.77
362.00
--
Anthropic logo
Claude Fable 5 (with fallback)
1M
Proprietary
50
$8.50
66
100.61
108.24
--
OpenAI logo
GPT-6 Sol (max)
1.05M
Proprietary
48
--
149
99.81
103.16
--
OpenAI logo
GPT-5.6 Sol (max)
524k
Proprietary
47
$1.95
103
127.37
132.24
--
OpenAI logo
GPT-6 Astra (low)
524k
Proprietary
46
$0.79
56
3.10
12.11
--
OpenAI logo
GPT-5.6 Sol (xhigh)
524k
Proprietary
44
$1.26
98
66.87
71.95
--
OpenAI logo
GPT-5.6 Sol (high)
524k
Proprietary
42
$0.85
88
17.16
22.86
--
Anthropic logo
Claude Opus 4.7 (max)
1M
Proprietary
41*
--
46
20.67
31.46
--
OpenAI logo
GPT-5.4 (xhigh)
1.05M
Proprietary
39*
--
100
181.33
186.34
--
Anthropic logo
Claude Sonnet 5 (max)
1M
Proprietary
38
$4.75
85
194.75
200.63
--
OpenAI logo
GPT-6 Luna (max)
524k
Proprietary
38
--
206
70.92
73.35
--
OpenAI logo
GPT-5.6 Terra (xhigh)
524k
Proprietary
38
$0.67
116
45.63
49.93
--
OpenAI logo
GPT-5.5 (high)
524k
Proprietary
37
$1.13
83
35.33
41.33
--
OpenAI logo
GPT-5.6 Luna (xhigh)
524k
Proprietary
35
$0.09
162
46.38
49.47
--
OpenAI logo
GPT-5.6 Terra (high)
524k
Proprietary
34
$0.35
109
3.09
7.67
--
OpenAI logo
GPT-5.6 Luna (high)
524k
Proprietary
32
$0.05
150
13.55
16.89
--
Anthropic logo
Claude Opus 4.6 (max)
1M
Proprietary
32*
--
41
24.77
37.10
--
DeepSeek logo
DeepSeek V4 Pro (max)
1M
Open
30
$9.65
81
1.72
62.11
54.20
OpenAI logo
GPT-5.2 (xhigh)
400k
Proprietary
30*
--
89
110.20
115.85
--
DeepSeek logo
DeepSeek V4 Pro (high)
1M
Open
30*
--
82
1.85
32.08
24.17
Anthropic logo
Claude Sonnet 4.6 (max)
200k
Proprietary
30
$2.70
56
130.59
139.48
--
Anthropic logo
Claude Opus 4.5
200k
Proprietary
29*
--
48
18.72
29.21
--
OpenAI logo
GPT-5.2 Codex (xhigh)
400k
Proprietary
29*
--
115
86.52
90.87
--
Kimi logo
Kimi K2.6
262k
Open
27
$1.91
227
1.53
23.34
19.61
OpenAI logo
GPT-5.2 (medium)
400k
Proprietary
27*
--
88
5.18
10.87
--
Anthropic logo
Claude Opus 4.6 (non-reasoning, high)
1M
Proprietary
26*
--
38
1.88
15.12
--
SpaceXAI logo
Grok 4.20 0309 v2
262k
Proprietary
26*
--
216
10.56
12.87
--
OpenAI logo
GPT-5 Codex (high)
400k
Proprietary
25*
--
156
10.45
13.65
--
SpaceXAI logo
Grok 4.3 (high)
200k
Proprietary
25
$0.53
157
20.96
24.14
--
SpaceXAI logo
Grok 4.3 (medium)
200k
Proprietary
25*
--
157
14.72
17.91
--
OpenAI logo
GPT-5.1 (high)
272k
Proprietary
25*
--
115
41.64
45.99
--
Anthropic logo
Claude Sonnet 4.6 (non-reasoning, high)
200k
Proprietary
25*
--
43
1.63
13.15
--
SpaceXAI logo
Grok 4.3 (low)
200k
Proprietary
24*
--
139
6.02
9.61
--
OpenAI logo
GPT-5.4 mini (xhigh)
400k
Proprietary
24
$0.43
254
127.11
129.08
--
OpenAI logo
GPT-5.1 Codex (high)
400k
Proprietary
24*
--
120
9.35
13.50
--
Anthropic logo
Claude Opus 4.5 (non-reasoning)
200k
Proprietary
24*
--
47
1.51
12.13
--
Kimi logo
Kimi K2.6 (non-reasoning)
262k
Open
24*
--
194
1.60
4.18
--
Kimi logo
Kimi K2.5
262k
Open
23*
--
131
1.40
27.93
22.70
OpenAI logo
GPT-5 (high)
400k
Proprietary
23*
--
97
68.03
73.21
--
OpenAI logo
GPT-5 (medium)
400k
Proprietary
23*
--
91
35.09
40.55
--
Anthropic logo
Claude 4.1 Opus
200k
Proprietary
23*
--
--
--
--
--
SpaceXAI logo
Grok 4
256k
Proprietary
22*
--
71
10.93
17.93
--
Kimi logo
Kimi K2 Thinking
256k
Open
22*
--
99
1.91
27.12
20.17
OpenAI logo
o3-pro
200k
Proprietary
22*
--
--
--
--
--
DeepSeek logo
DeepSeek V4 Pro (non-reasoning)
1M
Open
21*
--
89
1.58
7.21
--
OpenAI logo
GPT-5 (low)
400k
Proprietary
21*
--
92
11.44
16.85
--
Anthropic logo
Claude 4.5 Sonnet
200k
Proprietary
21
$0.54
44
14.87
26.12
--
OpenAI logo
GPT-5 mini (medium)
400k
Proprietary
21*
--
132
10.36
14.15
--
OpenAI logo
GPT-5.1 Codex mini (high)
400k
Proprietary
20*
--
163
15.60
18.67
--
OpenAI logo
o3
200k
Proprietary
20*
--
97
24.29
29.43
--
OpenAI logo
GPT-5.4 mini (medium)
400k
Proprietary
20*
--
211
9.14
11.51
--
Kimi logo
Kimi K2.5 (non-reasoning)
262k
Open
19*
--
122
1.35
5.46
--
Anthropic logo
Claude 4.5 Sonnet (non-reasoning)
200k
Proprietary
19*
--
42
1.46
13.31
--
Anthropic logo
Claude 4.1 Opus (non-reasoning)
200k
Proprietary
19*
--
--
--
--
--
SpaceXAI logo
Grok 4 Fast
2M
Proprietary
18*
--
--
--
--
--
OpenAI logo
GPT-5.2 (non-reasoning)
400k
Proprietary
17*
--
86
1.48
7.31
--
Anthropic logo
Claude 4.5 Haiku
200k
Proprietary
17
$0.22
98
15.36
20.47
--
OpenAI logo
GPT-5 mini (high)
400k
Proprietary
17
$0.05
131
53.20
57.02
--
OpenAI logo
o4-mini (high)
200k
Proprietary
17*
--
138
22.09
25.71
--
DeepSeek logo
DeepSeek V3.2 (non-reasoning)
128k
Open
16*
--
166
1.90
4.91
--
Anthropic logo
Claude 4.5 Haiku (non-reasoning)
200k
Proprietary
15*
--
90
1.11
6.68
--
OpenAI logo
o1
200k
Proprietary
15*
--
--
--
--
--
SpaceXAI logo
Grok 3 mini Reasoning (high)
32k
Proprietary
15*
--
--
--
--
--
OpenAI logo
GPT-5.1 (non-reasoning)
400k
Proprietary
13*
--
110
1.24
5.80
--
OpenAI logo
GPT-5 nano (high)
400k
Proprietary
13*
--
203
69.34
71.80
--
OpenAI logo
GPT-4.1
1M
Proprietary
13*
--
114
1.64
6.02
--
OpenAI logo
GPT-5 nano (medium)
400k
Proprietary
12*
--
205
32.76
35.20
--
OpenAI logo
o3-mini
200k
Proprietary
12*
--
218
7.00
9.29
--
SpaceXAI logo
Grok 3
16k
Proprietary
12*
--
--
--
--
--
OpenAI logo
gpt-oss-120b (high)
131k
Open
12
$0.11
315
0.72
8.65
6.34
OpenAI logo
GPT-5 (minimal)
400k
Proprietary
11*
--
93
1.78
7.14
--
OpenAI logo
o1-preview
128k
Proprietary
11*
--
--
--
--
--
OpenAI logo
GPT-5.4 mini (non-reasoning)
400k
Proprietary
11*
--
187
1.05
3.73
--
SpaceXAI logo
Grok 4 Fast (non-reasoning)
2M
Proprietary
11*
--
--
--
--
--
OpenAI logo
o3-mini (high)
200k
Proprietary
11
--
228
20.21
22.40
--
OpenAI logo
gpt-oss-120b (low)
131k
Open
10*
--
334
0.84
8.32
5.99
OpenAI logo
GPT-4.1 mini
1M
Proprietary
10*
--
103
1.54
6.38
--
Meta logo
Llama 4 Maverick (FP8)
128k
Open
10*
--
455
1.28
2.38
--
OpenAI logo
GPT-5 mini (minimal) East US 2 - Global Standard
400k
Proprietary
10*
--
--
--
--
--
Mistral logo
Mistral Large 3
256k
Open
9
$0.10
56
1.86
10.78
--
Mistral logo
Mistral Medium 3
128k
Proprietary
9*
--
47
2.27
12.91
--
OpenAI logo
GPT-4o (Nov)
128k
Proprietary
8*
--
125
1.90
5.91
--
Meta logo
Llama 4 Scout
128k
Open
8*
--
134
0.85
4.57
--
OpenAI logo
GPT-4.1 nano
1M
Proprietary
8*
--
276
1.47
3.27
--
OpenAI logo
GPT-4o (Aug)
128k
Proprietary
8*
--
141
1.38
4.93
--
Meta logo
Llama 3.3 70B
128k
Open
8*
--
124
2.18
6.20
--
OpenAI logo
GPT-4o (May)
128k
Proprietary
7*
--
143
1.55
5.03
--
OpenAI logo
GPT-5 nano (minimal) East US 2 - Global Standard
400k
Proprietary
7*
--
--
--
--
--
OpenAI logo
GPT-4 Turbo
128k
Proprietary
7*
--
108
1.78
6.42
--
Cohere logo
Command A
256k
Open
7*
--
42
3.04
14.82
--
OpenAI logo
GPT-4o mini
128k
Proprietary
7*
--
78
1.87
8.24
--
Microsoft logo
Phi-4 Mini
128k
Open
6*
--
44
0.82
12.09
--
Microsoft logo
Phi-4
16.4k
Open
6*
--
36
2.79
16.59
--
Microsoft logo
Phi-4 Multimodal
4.1k
Open
6*
--
--
--
--
--

Key definitions

Frequently Asked Questions

Common questions about Microsoft Azure