Microsoft Azure: Models Intelligence, Performance & Price

Microsoft Azure
Microsoft Azure

This analysis is intended to support you in choosing the best model provided by Microsoft Azure for your use-case.

Most Intelligent

Updated
#1
GPT-6 Astra (max)GPT-6 Astra (max)
53
#2
Claude Fable 5 (with fallback)Claude Fable 5 (with fallback)
50
#3
GPT-6 Sol (max)GPT-6 Sol (max)
48
#4
GPT-5.6 Sol (max)GPT-5.6 Sol (max)
47
#5
GPT-6 Astra (low)GPT-6 Astra (low)
46

Intelligence index

Total 82 models

Fastest

#1
Llama 4 Maverick (FP8)Llama 4 Maverick (FP8)
455 t/s
#2
gpt-oss-120b (low)gpt-oss-120b (low)
328 t/s
#3
gpt-oss-120b (high)gpt-oss-120b (high)
310 t/s
#4
GPT-4.1 nanoGPT-4.1 nano
264 t/s
#5
GPT-5.4 mini (xhigh)GPT-5.4 mini (xhigh)
254 t/s

Output speed

Total 82 models

Lowest Price

#1
GPT-5 nano (high)GPT-5 nano (high)
$0.05
#2
GPT-5 nano (medium)GPT-5 nano (medium)
$0.05
#3
GPT-6 Luna (max)GPT-6 Luna (max)
$0.08
#4
GPT-4.1 nanoGPT-4.1 nano
$0.08
#5
GPT-4o miniGPT-4o mini
$0.14

Blended price (per 1M tokens)

Total 82 models

Azure offers 82 models, each with different intelligence, performance, and pricing characteristics. Below is a comparison of the key metrics across models.

  • For intelligence, the top models on Azure are GPT-6 Astra (max) (53), Claude Fable 5 (with fallback) (50), and GPT-6 Sol (max) (48).
  • For output speed, the fastest models are Llama 4 Maverick (FP8) (455 t/s), gpt-oss-120b (low) (328 t/s), and gpt-oss-120b (high) (310 t/s). Speed varies significantly across models, with a 79% difference between the fastest and slowest.
  • For latency, Phi-4 Mini (0.82s), Llama 4 Scout (0.88s), and Claude 4.5 Haiku (non-reasoning) (1.05s) offer the lowest time to first answer token.
  • For pricing, GPT-5 nano (high) ($0.05), GPT-5 nano (medium) ($0.05), and GPT-6 Luna (max) ($0.08) offer the lowest blended prices per 1M tokens. Prices vary up to 2.7x across models.
  • For context window size, GPT-5.4 (xhigh) (1M), Claude Fable 5 (with fallback) (1M), and Claude Opus 4.7 (max) (1M) support the largest context windows on Azure.

Highlights

Updated
Artificial Analysis Intelligence Index · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens (blended) · Lower is better

Intelligence Evaluations

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
See more

Agentic knowledge work, (Elo-500)/2000

Agentic real-world work tasks, (Elo-500)/2000

Agentic SaaS workflows

Agentic coding & terminal use

SciCodeUnder review

Coding

Reasoning & knowledge

Professional document reasoning, All-pass

CritPtUnder review

Physics reasoning

Long context reasoning

Legal agentic work, criterion pass rate

Agentic business operations

Agentic scientific research workflows in a terminal

Quantitative analysis on spreadsheets & documents

Kubernetes incident root-cause analysis

Visual reasoning

Medical long context reasoning

Intelligence Index vs. Price

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Context Window

Context Window

Context window: tokens limit · Higher is better

Pricing

Intelligence Index vs. Price

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Performance Summary

Output Speed vs. Price

Output speed: output tokens per second · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Speed

Measured by Output Speed (tokens per second)

Output Speed

Output tokens per second · Higher is better

Latency

Measured by Time (seconds) to First Token

Latency: Time To First Answer Token

Seconds to first answer token received · Accounts for reasoning model 'thinking' time

End-to-End Response Time

Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed

End-to-End Response Time vs. Price

End-to-end response time: end-to-end seconds to output 500 tokens · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Further Analysis
OpenAI logo
GPT-6 Astra (max)
524k
Proprietary
53
$3.19
67
355.86
363.37
--
Anthropic logo
Claude Fable 5 (with fallback)
1M
Proprietary
50
$8.50
64
104.07
111.85
--
OpenAI logo
GPT-6 Sol (max)
524k
Proprietary
48
--
148
92.61
95.99
--
OpenAI logo
GPT-5.6 Sol (max)
524k
Proprietary
47
$1.95
91
123.98
129.48
--
OpenAI logo
GPT-6 Astra (low)
524k
Proprietary
46
$0.79
51
3.14
12.94
--
OpenAI logo
GPT-5.6 Sol (xhigh)
524k
Proprietary
44
$1.26
94
46.66
51.96
--
OpenAI logo
GPT-5.6 Sol (high)
524k
Proprietary
42
$0.85
85
16.57
22.46
--
Anthropic logo
Claude Opus 4.7 (max)
1M
Proprietary
41*
--
46
19.20
29.99
--
OpenAI logo
GPT-5.4 (xhigh)
1.05M
Proprietary
39*
--
100
151.93
156.95
--
Anthropic logo
Claude Sonnet 5 (max)
1M
Proprietary
38
$4.75
80
178.29
184.58
--
OpenAI logo
GPT-6 Luna (max)
524k
Proprietary
38
--
205
70.06
72.49
--
OpenAI logo
GPT-5.6 Terra (xhigh)
524k
Proprietary
38
$0.67
111
39.37
43.87
--
OpenAI logo
GPT-5.5 (high)
524k
Proprietary
37
$1.13
87
24.72
30.44
--
OpenAI logo
GPT-5.6 Luna (xhigh)
524k
Proprietary
35
$0.09
143
46.07
49.56
--
OpenAI logo
GPT-5.6 Terra (high)
524k
Proprietary
34
$0.35
105
2.91
7.66
--
OpenAI logo
GPT-5.6 Luna (high)
524k
Proprietary
32
$0.05
135
12.24
15.96
--
Anthropic logo
Claude Opus 4.6 (max)
1M
Proprietary
32*
--
39
23.51
36.39
--
DeepSeek logo
DeepSeek V4 Pro (max)
1M
Open
30
$9.65
82
1.67
61.01
53.25
OpenAI logo
GPT-5.2 (xhigh)
400k
Proprietary
30*
--
88
100.64
106.33
--
DeepSeek logo
DeepSeek V4 Pro (high)
1M
Open
30*
--
87
1.65
30.22
22.84
Anthropic logo
Claude Sonnet 4.6 (max)
200k
Proprietary
30
$2.70
54
130.59
139.78
--
Anthropic logo
Claude Opus 4.5
200k
Proprietary
29*
--
47
18.72
29.29
--
OpenAI logo
GPT-5.2 Codex (xhigh)
400k
Proprietary
29*
--
114
78.58
82.94
--
Kimi logo
Kimi K2.6
262k
Open
27
$1.91
211
1.53
24.95
21.06
OpenAI logo
GPT-5.2 (medium)
400k
Proprietary
27*
--
81
4.31
10.49
--
Anthropic logo
Claude Opus 4.6 (non-reasoning, high)
1M
Proprietary
26*
--
37
2.10
15.55
--
SpaceXAI logo
Grok 4.20 0309 v2
262k
Proprietary
26*
--
212
10.25
12.60
--
OpenAI logo
GPT-5 Codex (high)
400k
Proprietary
25*
--
155
7.64
10.87
--
SpaceXAI logo
Grok 4.3 (high)
200k
Proprietary
25
$0.53
144
18.80
22.27
--
SpaceXAI logo
Grok 4.3 (medium)
200k
Proprietary
25*
--
148
13.99
17.36
--
OpenAI logo
GPT-5.1 (high)
272k
Proprietary
25*
--
115
40.47
44.82
--
Anthropic logo
Claude Sonnet 4.6 (non-reasoning, high)
200k
Proprietary
25*
--
43
1.75
13.48
--
SpaceXAI logo
Grok 4.3 (low)
200k
Proprietary
24*
--
133
7.38
11.15
--
OpenAI logo
GPT-5.4 mini (xhigh)
400k
Proprietary
24
$0.43
254
114.62
116.59
--
OpenAI logo
GPT-5.1 Codex (high)
400k
Proprietary
24*
--
124
11.03
15.06
--
Anthropic logo
Claude Opus 4.5 (non-reasoning)
200k
Proprietary
24*
--
47
1.62
12.22
--
Kimi logo
Kimi K2.6 (non-reasoning)
262k
Open
24*
--
197
1.92
4.46
--
Kimi logo
Kimi K2.5
262k
Open
23*
--
126
1.45
29.00
23.58
OpenAI logo
GPT-5 (high)
400k
Proprietary
23*
--
91
53.20
58.70
--
OpenAI logo
GPT-5 (medium)
400k
Proprietary
23*
--
90
34.37
39.91
--
Anthropic logo
Claude 4.1 Opus
200k
Proprietary
23*
--
--
--
--
--
SpaceXAI logo
Grok 4
256k
Proprietary
22*
--
72
9.82
16.78
--
Kimi logo
Kimi K2 Thinking
256k
Open
22*
--
109
1.75
24.64
18.31
OpenAI logo
o3-pro
200k
Proprietary
22*
--
--
--
--
--
DeepSeek logo
DeepSeek V4 Pro (non-reasoning)
1M
Open
21*
--
89
1.64
7.27
--
OpenAI logo
GPT-5 (low)
400k
Proprietary
21*
--
91
7.34
12.86
--
Anthropic logo
Claude 4.5 Sonnet
200k
Proprietary
21
$0.54
42
14.02
25.81
--
OpenAI logo
GPT-5 mini (medium)
400k
Proprietary
21*
--
134
10.25
13.98
--
OpenAI logo
GPT-5.1 Codex mini (high)
400k
Proprietary
20*
--
159
18.38
21.52
--
OpenAI logo
o3
200k
Proprietary
20*
--
96
26.92
32.11
--
OpenAI logo
GPT-5.4 mini (medium)
400k
Proprietary
20*
--
209
9.67
12.06
--
Kimi logo
Kimi K2.5 (non-reasoning)
262k
Open
19*
--
119
1.31
5.51
--
Anthropic logo
Claude 4.5 Sonnet (non-reasoning)
200k
Proprietary
19*
--
40
1.53
14.03
--
Anthropic logo
Claude 4.1 Opus (non-reasoning)
200k
Proprietary
19*
--
--
--
--
--
SpaceXAI logo
Grok 4 Fast
2M
Proprietary
18*
--
--
--
--
--
OpenAI logo
GPT-5.2 (non-reasoning)
400k
Proprietary
17*
--
84
1.55
7.53
--
Anthropic logo
Claude 4.5 Haiku
200k
Proprietary
17
$0.22
94
14.94
20.28
--
OpenAI logo
GPT-5 mini (high)
400k
Proprietary
17
$0.05
137
53.20
56.84
--
OpenAI logo
o4-mini (high)
200k
Proprietary
17*
--
147
19.75
23.16
--
DeepSeek logo
DeepSeek V3.2 (non-reasoning)
128k
Open
16*
--
158
1.84
5.01
--
Anthropic logo
Claude 4.5 Haiku (non-reasoning)
200k
Proprietary
15*
--
85
1.05
6.91
--
OpenAI logo
o1
200k
Proprietary
15*
--
--
--
--
--
SpaceXAI logo
Grok 3 mini Reasoning (high)
32k
Proprietary
15*
--
--
--
--
--
OpenAI logo
GPT-5.1 (non-reasoning)
400k
Proprietary
13*
--
106
1.32
6.03
--
OpenAI logo
GPT-5 nano (high)
400k
Proprietary
13*
--
200
64.15
66.65
--
OpenAI logo
GPT-4.1
1M
Proprietary
13*
--
116
1.64
5.96
--
OpenAI logo
GPT-5 nano (medium)
400k
Proprietary
12*
--
196
35.02
37.57
--
OpenAI logo
o3-mini
200k
Proprietary
12*
--
215
6.35
8.68
--
SpaceXAI logo
Grok 3
16k
Proprietary
12*
--
--
--
--
--
OpenAI logo
gpt-oss-120b (high)
131k
Open
12
$0.11
310
0.73
8.79
6.45
OpenAI logo
GPT-5 (minimal)
400k
Proprietary
11*
--
93
1.78
7.16
--
OpenAI logo
o1-preview
128k
Proprietary
11*
--
--
--
--
--
OpenAI logo
GPT-5.4 mini (non-reasoning)
400k
Proprietary
11*
--
183
1.30
4.03
--
SpaceXAI logo
Grok 4 Fast (non-reasoning)
2M
Proprietary
11*
--
--
--
--
--
OpenAI logo
o3-mini (high)
200k
Proprietary
11
--
231
20.49
22.66
--
OpenAI logo
gpt-oss-120b (low)
131k
Open
10*
--
328
0.87
8.49
6.10
OpenAI logo
GPT-4.1 mini
1M
Proprietary
10*
--
103
1.48
6.32
--
Meta logo
Llama 4 Maverick (FP8)
128k
Open
10*
--
455
1.25
2.35
--
OpenAI logo
GPT-5 mini (minimal) East US 2 - Global Standard
400k
Proprietary
10*
--
--
--
--
--
Mistral logo
Mistral Large 3
256k
Open
9
$0.10
66
1.62
9.22
--
Mistral logo
Mistral Medium 3
128k
Proprietary
9*
--
47
2.08
12.80
--
OpenAI logo
GPT-4o (Nov)
128k
Proprietary
8*
--
123
1.96
6.01
--
Meta logo
Llama 4 Scout
128k
Open
8*
--
135
0.88
4.60
--
OpenAI logo
GPT-4.1 nano
1M
Proprietary
8*
--
264
1.52
3.41
--
OpenAI logo
GPT-4o (Aug)
128k
Proprietary
8*
--
130
1.46
5.32
--
Meta logo
Llama 3.3 70B
128k
Open
8*
--
116
2.14
6.43
--
OpenAI logo
GPT-4o (May)
128k
Proprietary
7*
--
142
1.59
5.12
--
OpenAI logo
GPT-5 nano (minimal) East US 2 - Global Standard
400k
Proprietary
7*
--
--
--
--
--
OpenAI logo
GPT-4 Turbo
128k
Proprietary
7*
--
106
1.78
6.48
--
Cohere logo
Command A
256k
Open
7*
--
42
3.14
14.96
--
OpenAI logo
GPT-4o mini
128k
Proprietary
7*
--
78
1.87
8.32
--
Microsoft logo
Phi-4 Mini
128k
Open
6*
--
45
0.82
11.91
--
Microsoft logo
Phi-4
16.4k
Open
6*
--
40
2.64
15.27
--
Microsoft logo
Phi-4 Multimodal
4.1k
Open
6*
--
--
--
--
--

Key definitions

Frequently Asked Questions

Common questions about Microsoft Azure