DeepInfra: inteligência, desempenho e preço dos modelos

DeepInfra
DeepInfra

Esta análise ajuda você a escolher o melhor modelo oferecido por DeepInfra para seu caso de uso.

Mais inteligente

Updated
#1
GLM-5.3 (max)GLM-5.3 (max)
45
#2
GLM-5.3-FlashGLM-5.3-Flash
42
#3
Qwen3.8 2.4T A95BQwen3.8 2.4T A95B
40
#4
DeepSeek V4 Pro 0813 (max)DeepSeek V4 Pro 0813 (max)
36
#5
DeepSeek V4 Flash 0731 (max)DeepSeek V4 Flash 0731 (max)
35

Intelligence Index

108 modelos no total

Mais rápido

#1
Inkling SmallInkling Small
244 t/s
#2
gpt-oss-120b (high) (Turbo)gpt-oss-120b (high) (Turbo)
230 t/s
#3
Granite 4.2 3BGranite 4.2 3B
214 t/s
#4
Inkling (FP8)Inkling (FP8)
204 t/s
#5
Step 3.7 FlashStep 3.7 Flash
168 t/s

Velocidade de saída

108 modelos no total

Menor preço

#1
Llama 3.1 8B (Turbo, FP8)Llama 3.1 8B (Turbo, FP8)
$0.02
#2
Llama 3.1 8BLlama 3.1 8B
$0.02
#3
Granite 4.2 3BGranite 4.2 3B
$0.02
#4
Gemma 4 E4BGemma 4 E4B
$0.03
#5
Gemma 4 E4B (Non-reasoning)Gemma 4 E4B (Non-reasoning)
$0.03

Preço combinado (por 1M de tokens)

108 modelos no total

DeepInfra oferece 108 modelos, cada um com características diferentes de inteligência, desempenho e preço. Veja abaixo uma comparação das principais métricas entre os modelos.

  • Em inteligência, os principais modelos oferecidos por DeepInfra são GLM-5.3 (max) (45), GLM-5.3-Flash (42) e Qwen3.8 2.4T A95B (40).
  • Em velocidade de saída, os modelos mais rápidos são Inkling Small (244 t/s), gpt-oss-120b (high) (Turbo) (230 t/s) e Granite 4.2 3B (214 t/s).
  • Em latência, Llama 4 Maverick (FP8) (0.50s), Qwen3 30B (Non-reasoning) (FP8) (0.51s) e Qwen3 Coder 480B (Turbo, FP4) (0.64s) têm o menor tempo até o primeiro token da resposta final.
  • Em preço, Llama 3.1 8B (Turbo, FP8) ($0.02), Llama 3.1 8B ($0.02) e Granite 4.2 3B ($0.02) têm os menores preços combinados por 1M de tokens.
  • Em tamanho da janela de contexto, GLM-5.2 (max) (FP4) (1M), DeepSeek V4 Pro (max) (FP4) (1M) e DeepSeek V4 Flash (high) (FP4) (1M) aceitam as maiores janelas entre os modelos oferecidos por DeepInfra.

Destaques

Updated
Artificial Analysis Intelligence Index · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens (blended) · Lower is better

Avaliações de inteligência

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
See more

Agentic knowledge work, (Elo-500)/2000

Agentic real-world work tasks, (Elo-500)/2000

Agentic SaaS workflows

Agentic coding & terminal use

Coding

Reasoning & knowledge

Professional document reasoning, All-pass

Physics reasoning

Long context reasoning

Legal agentic work, criterion pass rate

Agentic business operations

Quantitative analysis on spreadsheets & documents

Instruction following

Agentic tool use

Long-horizon agentic tasks

Kubernetes incident root-cause analysis

Visual reasoning

Intelligence Index vs. Price

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Janela de contexto

Context Window

Context window: tokens limit · Higher is better

Preços

Intelligence Index vs. Price

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Resumo do desempenho

Output Speed vs. Price

Output speed: output tokens per second · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Velocidade

Medida pela velocidade de saída (tokens por segundo)

Output Speed

Output tokens per second · Higher is better

Latência

Medida pelo tempo (segundos) até o primeiro token

Latency: Time To First Answer Token

Seconds to first answer token received · Accounts for reasoning model 'thinking' time

Tempo de resposta de ponta a ponta

Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed

End-to-End Response Time vs. Price

End-to-end response time: end-to-end seconds to output 500 tokens · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Análises adicionais
Logo: Z AI
GLM-5.3 (max)
1.05M
Aberto
45
$2.16
141
1.04
18.82
14.22
Logo: Z AI
GLM-5.3-Flash
1M
Aberto
42
$0.43
31
1.36
81.71
64.28
Logo: Alibaba
Qwen3.8 2.4T A95B
262k
Aberto
40
$2.89
89
1.48
29.69
22.56
Logo: DeepSeek
DeepSeek V4 Pro 0813 (max)
1M
Aberto
36
$1.00
114
0.80
22.66
17.49
Logo: DeepSeek
DeepSeek V4 Flash 0731 (max)
1M
Aberto
35
$0.10
57
0.94
45.09
35.32
Logo: Z AI
GLM-5.2 (max) (FP4)
1.05M
Aberto
34
$0.65
63
1.05
40.91
31.89
Logo: Alibaba
Qwen3.8 27B (xhigh)
262k
Aberto
34
$0.65
36
1.02
70.84
55.86
Logo: DeepSeek
DeepSeek V4 Pro (max) (FP4)
1.05M
Aberto
31
$1.19
52
1.09
94.64
83.95
Logo: DeepSeek
DeepSeek V4 Pro (high) (FP4)
65.5k
Aberto
30*
--
56
1.50
45.91
35.50
Logo: MiniMax
MiniMax-M3
524k
Aberto
30
$0.44
24
0.96
106.39
84.35
Logo: Z AI
GLM-5 (FP4)
203k
Aberto
28*
--
70
0.98
52.34
44.24
Logo: Kimi
Kimi K2.6 (FP4)
262k
Aberto
27
$0.51
45
1.43
110.34
97.91
Logo: Z AI
GLM-5.1 (FP4)
203k
Aberto
26
$0.69
31
0.83
137.60
120.83
Logo: Xiaomi
MiMo-V2.5-Pro
65.5k
Aberto
26
$0.43
37
1.11
68.77
54.13
Logo: Kimi
Kimi K2.7 Code
262k
Aberto
26
$0.39
41
0.95
67.92
54.70
Logo: Thinking Machines
Inkling Small
524k
Aberto
26
$0.12
244
0.98
11.24
8.21
Logo: Tencent
Hy3 (FP8)
262k
Aberto
26
$0.07
60
1.28
42.76
33.19
Logo: Thinking Machines
Inkling (FP8)
131k
Aberto
26
$0.81
204
0.56
12.82
9.81
Logo: DeepSeek
DeepSeek V4 Flash (high) (FP4)
1.05M
Aberto
25
$0.10
53
0.89
34.02
23.61
Logo: DeepSeek
DeepSeek V4 Flash (max) (FP4)
1.05M
Aberto
25
$0.09
42
0.95
145.74
132.95
Logo: Z AI
GLM-5.1 (Non-reasoning) (FP4)
203k
Aberto
24*
--
34
1.05
15.84
--
Logo: Kimi
Kimi K2.6 (Non-reasoning) (FP4)
262k
Aberto
24*
--
36
1.49
15.39
--
Logo: Kimi
Kimi K2.5
262k
Aberto
23*
--
47
1.34
75.25
63.25
Logo: NVIDIA
Nemotron 3 Ultra
262k
Aberto
23
$0.47
143
3.48
22.89
15.92
Logo: NVIDIA
Nemotron 3 Ultra BF16
262k
Aberto
23
$0.92
144
3.26
22.50
15.77
Logo: Alibaba
Qwen3.5 27B (FP8)
262k
Aberto
23*
--
70
0.90
36.81
28.73
Logo: MiniMax
MiniMax-M2.5 (FP8)
197k
Aberto
23*
--
22
3.30
116.26
90.36
Logo: Xiaomi
MiMo-V2.5
262k
Aberto
22
$0.02
29
1.57
87.29
68.58
Logo: Z AI
GLM-4.7 (FP4)
203k
Aberto
22*
--
18
1.52
141.69
112.14
Logo: Alibaba
Qwen3.6 27B FP8
262k
Aberto
22
$0.37
67
0.94
93.57
85.14
Logo: Z AI
GLM-5 (Non-reasoning) (FP8)
203k
Aberto
22*
--
54
1.33
10.64
--
Logo: DeepSeek
DeepSeek V3.2 (FP4)
164k
Aberto
21*
--
27
1.67
92.95
73.03
Logo: Alibaba
Qwen3.5 397B A17B (Non-reasoning) (FP8)
262k
Aberto
21*
--
41
0.97
13.07
--
Logo: InclusionAI
Ling 3.0 Flash
131k
Aberto
21
$0.03
61
0.63
41.39
32.61
Logo: Alibaba
Qwen3.6 27B (Non-reasoning) FP8
262k
Aberto
20*
--
61
0.91
9.07
--
Logo: StepFun
Step 3.7 Flash
256k
Aberto
19*
--
168
0.66
15.52
11.89
Logo: Alibaba
Qwen3.5 27B (Non-reasoning) FP8
262k
Aberto
19*
--
67
0.92
8.41
--
Logo: Alibaba
Qwen3.5 35B A3B (FP8)
262k
Aberto
19*
--
159
0.80
16.48
12.54
Logo: Alibaba
Qwen3.5 397B A17B (FP8)
262k
Aberto
19
$0.23
39
0.97
94.63
80.95
Logo: Alibaba
Qwen3.6 35B A3B (FP8)
262k
Aberto
19
$0.14
97
0.70
61.76
55.88
Logo: Z AI
GLM-4.6 (FP4)
203k
Aberto
19*
--
41
0.75
62.48
49.38
Logo: Xiaomi
MiMo-V2.5-Pro (Non-reasoning)
65.5k
Aberto
18*
--
32
1.16
16.83
--
Logo: Meta
Muse Glimmer (high)
131k
Aberto
18
$0.05
145
0.94
18.14
13.76
Logo: Alibaba
Qwen3.5 122B A10B (Non-reasoning) (FP4)
262k
Aberto
18*
--
143
0.92
4.42
--
Logo: Z AI
GLM-4.7 (Non-reasoning) (FP4)
203k
Aberto
17*
--
18
1.10
28.40
--
Logo: Google
Gemma 4 26B A4B (FP8)
262k
Aberto
17*
--
30
0.87
84.86
67.19
Logo: Alibaba
Qwen3.5 122B A10B (FP4)
262k
Aberto
16
$0.14
152
0.79
17.27
13.18
Logo: DeepSeek
DeepSeek V3.2 (Non-reasoning)
164k
Aberto
16*
--
32
1.35
16.89
--
Logo: Google
Gemma 4 31B
262k
Aberto
15
$0.02
15
2.14
149.38
114.32
Logo: Alibaba
Qwen3.6 35B A3B (Non-reasoning) (FP8)
262k
Aberto
15*
--
108
0.69
5.33
--
Logo: Alibaba
Qwen3.5 35B A3B (Non-reasoning) FP8
262k
Aberto
15*
--
154
0.72
3.97
--
Logo: Z AI
GLM-4.7-Flash
203k
Aberto
15*
--
35
1.56
72.92
57.09
Logo: IBM
Granite 4.2 30B
131k
Aberto
15*
--
77
0.81
33.19
25.90
Logo: DeepSeek
DeepSeek V3.1 Terminus (Non-reasoning) (FP4)
164k
Aberto
14*
--
60
0.97
9.35
--
Logo: Google
Gemma 4 31B (Non-reasoning) (FP8)
262k
Aberto
14*
--
25
1.45
21.60
--
Logo: DeepSeek
DeepSeek V3.1 (Non-reasoning) (FP4)
164k
Aberto
14*
--
12
1.00
42.55
--
Logo: NVIDIA
Nemotron 3.5 Lightning (NVFP4)
262k
Aberto
14
$0.12
--
--
--
--
Logo: NVIDIA
Nemotron 3 Super
262k
Aberto
14
$0.48
87
14.99
43.62
22.91
Logo: Google
Gemma 4 26B A4B (Non-reasoning) (FP8)
262k
Aberto
13*
--
23
0.87
22.47
--
Logo: Alibaba
Qwen3.5 4B (FP8)
262k
Aberto
13*
--
29
0.70
85.83
68.10
Logo: DeepSeek
DeepSeek R1 0528
164k
Aberto
13*
--
24
0.97
105.07
83.28
Logo: Alibaba
Qwen3 235B A22B 2507 (FP8)
262k
Aberto
13
--
114
0.72
22.72
17.60
Logo: OpenAI
gpt-oss-120b (high) (Turbo)
131k
Aberto
12
$0.11
230
0.72
11.58
8.69
Logo: OpenAI
gpt-oss-120b (high)
131k
Aberto
12
$0.03
53
0.61
47.75
37.71
Logo: Alibaba
Qwen3 235B 2507 (Non-reasoning) (FP8)
262k
Aberto
12*
--
16
0.67
32.20
--
Logo: Alibaba
Qwen3 Coder 480B (Turbo, FP4)
262k
Aberto
12*
--
79
0.64
6.93
--
Logo: IBM
Granite 4.2 8B
131k
Aberto
12
$0.02
76
0.62
33.54
26.34
Logo: Alibaba
Qwen3.5 4B (Non-reasoning) FP8
262k
Aberto
11*
--
23
0.69
21.98
--
Logo: OpenAI
gpt-oss-120b (low)
131k
Aberto
10*
--
55
0.63
46.38
36.60
Logo: DeepSeek
DeepSeek V3 0324 (FP4)
164k
Aberto
10
$0.02
46
1.16
12.14
--
Logo: Alibaba
Qwen3 Next 80B A3B
262k
Aberto
10*
--
160
0.68
3.80
--
Logo: Meta
Llama 4 Maverick (FP8)
1.05M
Aberto
9
$0.03
74
0.50
7.25
--
Logo: IBM
Granite 4.2 3B
131k
Aberto
9
$0.01
214
0.42
12.11
9.35
Logo: OpenAI
gpt-oss-20b (high)
131k
Aberto
9
$0.01
117
0.44
21.85
17.13
Logo: NVIDIA
Llama Nemotron Super 49B v1.5
131k
Aberto
9*
--
95
4.25
30.66
21.13
Logo: Google
Gemma 4 E4B
262k
Aberto
9*
--
44
0.82
57.83
45.61
Logo: NVIDIA
Nemotron 3 Nano
262k
Aberto
9
$0.02
97
5.80
31.45
20.52
Logo: DeepSeek
DeepSeek V3 (Dec)
164k
Aberto
8
$0.02
29
0.70
17.85
--
Logo: DeepSeek
DeepSeek R1 Distill Llama 70B
131k
Aberto
8*
--
25
1.00
101.29
80.23
Logo: Alibaba
Qwen2.5 72B
32.8k
Aberto
8*
--
29
2.33
19.47
--
Logo: Meta
Llama 3.3 70B (Turbo, FP8)
131k
Aberto
8*
--
16
2.05
34.11
--
Logo: Alibaba
Qwen3 30B (FP8)
41k
Aberto
8*
--
157
0.51
16.46
12.76
Logo: NVIDIA
NVIDIA Nemotron Nano 12B v2 VL (FP8)
131k
Aberto
7*
--
137
3.82
22.01
14.55
Logo: Google
Gemma 4 E4B (Non-reasoning)
262k
Aberto
7*
--
45
0.78
11.85
--
Logo: NVIDIA
NVIDIA Nemotron Nano 9B V2
131k
Aberto
7*
--
95
7.88
34.25
21.10
Logo: Mistral
Mistral Small 3.1
128k
Aberto
7
$0.02
36
0.94
14.67
--
Logo: NVIDIA
Llama Nemotron Super 49B v1.5 (Non-reasoning)
131k
Aberto
7*
--
144
3.02
6.49
--
Logo: Alibaba
Qwen3 32B (Non-reasoning) (FP8)
41k
Aberto
7*
--
27
1.43
19.72
--
Logo: Alibaba
Qwen3 32B (FP8)
41k
Aberto
7
--
29
1.47
88.80
69.86
Logo: Mistral
Mistral Small 3.2 (FP8)
128k
Aberto
7
$0.11
37
0.92
14.36
--
Logo: NVIDIA
Llama 3.1 Nemotron 70B
131k
Aberto
7*
--
126
3.22
7.18
--
Logo: Meta
Llama 3.1 8B (Turbo, FP8)
131k
Aberto
7*
--
23
0.95
22.25
--
Logo: Meta
Llama 3.1 8B
131k
Aberto
7*
--
20
0.97
25.69
--
Logo: NVIDIA
Nemotron 3 Nano (Non-reasoning)
262k
Aberto
7*
--
103
0.64
5.51
--
Logo: NVIDIA
NVIDIA Nemotron Nano 9B V2 (Non-reasoning)
131k
Aberto
7*
--
102
7.39
12.29
--
Logo: Alibaba
Qwen3 14B (Non-reasoning) (FP8)
41k
Aberto
7*
--
64
0.80
8.59
--
Logo: Mistral
Mistral Small 3
32.8k
Aberto
7*
--
53
0.83
10.27
--
Logo: Alibaba
Qwen3 30B (Non-reasoning) (FP8)
41k
Aberto
7*
--
164
0.51
3.56
--
Logo: Meta
Llama 3.1 70B
131k
Aberto
7*
--
35
1.90
16.11
--
Logo: Meta
Llama 3.1 70B (Turbo, FP8)
131k
Aberto
7*
--
36
1.94
15.80
--
Logo: Meta
Llama 4 Scout
328k
Aberto
6
$0.06
22
0.91
23.55
--
Logo: Alibaba
Qwen3 14B (FP8)
32.8k
Aberto
6
--
57
0.99
45.03
35.23
Logo: Nous Research
Hermes 3 - Llama-3.1 70B
131k
Aberto
6*
--
31
2.13
18.17
--
Logo: Microsoft
Phi-4
16.4k
Aberto
6*
--
56
0.93
9.82
--
Logo: NVIDIA
NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) (FP8)
131k
Aberto
6*
--
128
4.58
8.48
--
Logo: Meta
Llama 3.2 11B (Vision)
131k
Aberto
5*
--
22
1.27
23.75
--
Logo: Google
Gemma 3 27B
131k
Aberto
5
$0.16
25
1.37
21.27
--
Logo: Google
Gemma 3 4B
131k
Aberto
5*
--
22
1.52
23.96
--
Logo: Meta
Llama 3 8B
8.19k
Aberto
5*
--
--
--
--
--
Logo: Google
Gemma 3 12B
131k
Aberto
4
$0.13
39
1.04
13.74
--

Principais definições

Perguntas frequentes

Perguntas comuns sobre DeepInfra