DeepInfra: inteligência, desempenho e preço dos modelos

DeepInfra
DeepInfra

Esta análise ajuda você a escolher o melhor modelo oferecido por DeepInfra para seu caso de uso.

Mais inteligente

Updated
#1
MiMo-V2.6-ProMiMo-V2.6-Pro
46
#2
GLM-5.3 (max)GLM-5.3 (max)
45
#3
GLM-5.3-FlashGLM-5.3-Flash
42
#4
Qwen3.8 2.4T A95BQwen3.8 2.4T A95B
40
#5
DeepSeek V4.1 Flash (max)DeepSeek V4.1 Flash (max)
39

Intelligence Index

107 modelos no total

Mais rápido

#1
Nemotron 3.5 Lightning (NVFP4)Nemotron 3.5 Lightning (NVFP4)
295 t/s
#2
gpt-oss-120b (high) (Turbo)gpt-oss-120b (high) (Turbo)
240 t/s
#3
Granite 4.2 3BGranite 4.2 3B
220 t/s
#4
Qwen3.6 35B A3B (non-reasoning) (FP8)Qwen3.6 35B A3B (non-reasoning) (FP8)
179 t/s
#5
Qwen3.6 35B A3B (FP8)Qwen3.6 35B A3B (FP8)
166 t/s

Velocidade de saída

107 modelos no total

Menor preço

#1
Llama 3.1 8B (Turbo, FP8)Llama 3.1 8B (Turbo, FP8)
$0.02
#2
Llama 3.1 8BLlama 3.1 8B
$0.02
#3
Granite 4.2 3BGranite 4.2 3B
$0.02
#4
Gemma 4 E4BGemma 4 E4B
$0.03
#5
Gemma 4 E4B (non-reasoning)Gemma 4 E4B (non-reasoning)
$0.03

Preço combinado (por 1M de tokens)

107 modelos no total

DeepInfra oferece 107 modelos, cada um com características diferentes de inteligência, desempenho e preço. Veja abaixo uma comparação das principais métricas entre os modelos.

  • Em inteligência, os principais modelos oferecidos por DeepInfra são MiMo-V2.6-Pro (46), GLM-5.3 (max) (45) e GLM-5.3-Flash (42).
  • Em velocidade de saída, os modelos mais rápidos são Nemotron 3.5 Lightning (NVFP4) (295 t/s), gpt-oss-120b (high) (Turbo) (240 t/s) e Granite 4.2 3B (220 t/s). A velocidade varia significativamente entre os modelos, com uma diferença de 78% entre o mais rápido e o mais lento.
  • Em latência, Qwen3 30B (non-reasoning) (FP8) (0.57s), Qwen3.5 35B A3B (non-reasoning) FP8 (0.58s) e Qwen3 Coder 480B (Turbo, FP4) (0.67s) têm o menor tempo até o primeiro token da resposta final.
  • Em preço, Llama 3.1 8B (Turbo, FP8) ($0.02), Llama 3.1 8B ($0.02) e Granite 4.2 3B ($0.02) têm os menores preços combinados por 1M de tokens.
  • Em tamanho da janela de contexto, GLM-5.2 (max) (FP4) (1M), DeepSeek V4 Pro (max) (FP4) (1M) e DeepSeek V4 Flash (high) (FP4) (1M) aceitam as maiores janelas entre os modelos oferecidos por DeepInfra.
Updated
Artificial Analysis Intelligence Index · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens (blended) · Lower is better

Avaliações de inteligência

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
Ver mais

Agentic knowledge work, (Elo-500)/2000

Agentic real-world work tasks, (Elo-500)/2000

Agentic SaaS workflows

Agentic coding & terminal use

SciCodeUnder review

Coding

Reasoning & knowledge

Professional document reasoning, All-pass

CritPtUnder review

Physics reasoning

Long context reasoning

Legal agentic work, criterion pass rate

Agentic business operations

Agentic scientific research workflows in a terminal

Quantitative analysis on spreadsheets & documents

Kubernetes incident root-cause analysis

Visual reasoning

Medical long context reasoning

Intelligence Index vs. preço

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Janela de contexto

Janela de contexto

Context window: tokens limit · Higher is better

Preços

Intelligence Index vs. preço

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Resumo do desempenho

Velocidade de saída vs. preço

Output speed: output tokens per second · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Velocidade

Medida pela velocidade de saída (tokens por segundo)

Velocidade de saída

Output tokens per second · Higher is better

Latência

Medida pelo tempo (segundos) até o primeiro token

Latência: Tempo até o primeiro token da resposta final

Segundos até receber o primeiro token de resposta · Inclui o tempo de “pensamento” dos modelos de raciocínio

Tempo de resposta de ponta a ponta

Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed

Tempo de resposta de ponta a ponta vs. preço

End-to-end response time: end-to-end seconds to output 500 tokens · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Análises adicionais
Logo: Xiaomi
MiMo-V2.6-Pro
1M
Aberto
46
$0.12
26
1.72
99.53
78.24
Logo: Z AI
GLM-5.3 (max)
1.05M
Aberto
45
$1.39
137
1.01
19.24
14.58
Logo: Z AI
GLM-5.3-Flash
1M
Aberto
42
$0.34
24
1.46
105.74
83.42
Logo: Alibaba
Qwen3.8 2.4T A95B
262k
Aberto
40
$2.89
95
1.47
27.92
21.16
Logo: DeepSeek
DeepSeek V4.1 Flash (max)
1M
Aberto
39
$0.44
79
0.85
32.54
25.35
Logo: Xiaomi
MiMo-V2.6-Flash
1M
Aberto
38
$0.42
26
1.67
98.94
77.82
Logo: DeepSeek
DeepSeek V4 Pro 0813 (max)
1M
Aberto
36
$1.00
56
1.27
45.78
35.61
Logo: DeepSeek
DeepSeek V4 Flash Vision (max)
1M
Proprietário
35
$0.51
150
0.92
17.55
13.30
Logo: DeepSeek
DeepSeek V4 Flash 0731 (max)
1M
Aberto
34
$0.10
42
1.04
59.93
47.11
Logo: Z AI
GLM-5.2 (max) (FP4)
1.05M
Aberto
34
$0.37
90
1.12
28.93
22.25
Logo: Alibaba
Qwen3.8 27B (xhigh)
262k
Aberto
34
$0.52
29
1.54
87.35
68.65
Logo: DeepSeek
DeepSeek V4 Pro (max) (FP4)
1.05M
Aberto
30
$1.19
37
1.37
132.72
117.88
Logo: DeepSeek
DeepSeek V4 Pro (high) (FP4)
65.5k
Aberto
30*
--
35
1.66
72.15
56.35
Logo: MiniMax
MiniMax-M3
524k
Aberto
29
$0.44
25
0.91
102.44
81.22
Logo: Z AI
GLM-5 (FP4)
203k
Aberto
28*
--
78
1.28
47.68
39.96
Logo: Kimi
Kimi K2.6 (FP4)
262k
Aberto
27
$0.51
20
1.67
243.21
217.15
Logo: Z AI
GLM-5.1 (FP4)
203k
Aberto
26
$0.58
41
1.32
106.17
92.64
Logo: Xiaomi
MiMo-V2.5-Pro
65.5k
Aberto
26
$0.28
26
2.09
97.60
76.40
Logo: Kimi
Kimi K2.7 Code
262k
Aberto
26
$0.39
25
1.08
108.25
87.54
Logo: Thinking Machines
Inkling Small
524k
Aberto
26
--
238
0.60
11.11
8.41
Logo: Tencent
Hy3 (FP8)
262k
Aberto
25
$0.07
33
1.03
76.02
59.99
Logo: Xiaomi
MiMo-V2.5
262k
Aberto
25*
--
28
1.57
89.64
70.46
Logo: Thinking Machines
Inkling (xhigh) (FP8)
131k
Aberto
25
--
128
0.76
20.23
15.57
Logo: DeepSeek
DeepSeek V4 Flash (high) (FP4)
1.05M
Aberto
24
$0.08
35
1.54
50.66
35.01
Logo: Z AI
GLM-5.1 (non-reasoning) (FP4)
203k
Aberto
24*
--
48
1.19
11.54
--
Logo: DeepSeek
DeepSeek V4 Flash (max) (FP4)
1.05M
Aberto
24
$0.07
36
1.37
173.39
157.95
Logo: Kimi
Kimi K2.6 (non-reasoning) (FP4)
262k
Aberto
24*
--
18
1.54
29.23
--
Logo: Kimi
Kimi K2.5
262k
Aberto
23*
--
25
1.35
140.24
118.86
Logo: NVIDIA
Nemotron 3 Ultra
262k
Aberto
23
$0.47
66
1.27
43.10
34.30
Logo: NVIDIA
Nemotron 3 Ultra BF16
262k
Aberto
23
$0.99
85
1.68
34.48
26.89
Logo: Alibaba
Qwen3.5 27B (FP8)
262k
Aberto
23*
--
61
1.01
42.00
32.79
Logo: MiniMax
MiniMax-M2.5 (FP8)
197k
Aberto
23*
--
27
0.91
94.98
75.26
Logo: Z AI
GLM-4.7 (FP4)
203k
Aberto
22*
--
17
1.33
150.59
119.41
Logo: Z AI
GLM-5 (non-reasoning) (FP8)
203k
Aberto
22*
--
71
1.09
8.12
--
Logo: DeepSeek
DeepSeek V3.2 (FP4)
164k
Aberto
21*
--
16
1.53
156.57
124.03
Logo: Alibaba
Qwen3.5 397B A17B (non-reasoning) (FP8)
262k
Aberto
21*
--
35
1.15
15.44
--
Logo: Alibaba
Qwen3.6 27B FP8
262k
Aberto
21
$0.37
63
1.03
98.44
89.52
Logo: InclusionAI
Ling 3.0 Flash
131k
Aberto
20
$0.02
48
0.75
52.98
41.78
Logo: Alibaba
Qwen3.6 27B (non-reasoning) FP8
262k
Aberto
20*
--
61
1.01
9.17
--
Logo: StepFun
Step 3.7 Flash
256k
Aberto
19*
--
169
0.71
15.47
11.81
Logo: Alibaba
Qwen3.5 27B (non-reasoning) FP8
262k
Aberto
19*
--
60
1.00
9.27
--
Logo: Alibaba
Qwen3.5 35B A3B (FP8)
262k
Aberto
19*
--
120
0.67
21.47
16.64
Logo: Z AI
GLM-4.6 (FP4)
203k
Aberto
19*
--
19
1.71
130.18
102.77
Logo: Alibaba
Qwen3.5 397B A17B (FP8)
262k
Aberto
18
--
38
1.22
98.11
83.74
Logo: Xiaomi
MiMo-V2.5-Pro (non-reasoning)
65.5k
Aberto
18*
--
29
1.47
18.58
--
Logo: Alibaba
Qwen3.6 35B A3B (FP8)
262k
Aberto
18
$0.14
32
0.85
183.26
166.94
Logo: Alibaba
Qwen3.5 122B A10B (non-reasoning) (FP4)
262k
Aberto
18*
--
148
0.70
4.07
--
Logo: Meta
Muse Glimmer (high)
131k
Aberto
17
$0.05
143
0.86
18.33
13.98
Logo: Z AI
GLM-4.7 (non-reasoning) (FP4)
203k
Aberto
17*
--
19
1.49
27.60
--
Logo: Google
Gemma 4 26B A4B (FP8)
262k
Aberto
17*
--
35
0.80
73.17
57.89
Logo: DeepSeek
DeepSeek V3.2 (non-reasoning)
164k
Aberto
16*
--
14
1.62
38.33
--
Logo: Alibaba
Qwen3.5 122B A10B (FP4)
262k
Aberto
16
$0.24
163
0.68
15.99
12.25
Logo: Alibaba
Qwen3.6 35B A3B (non-reasoning) (FP8)
262k
Aberto
15*
--
28
1.00
18.87
--
Logo: Alibaba
Qwen3.5 35B A3B (non-reasoning) FP8
262k
Aberto
15*
--
139
0.70
4.28
--
Logo: Z AI
GLM-4.7-Flash
203k
Aberto
15*
--
20
1.32
127.66
101.07
Logo: IBM
Granite 4.2 30B
131k
Aberto
15*
--
77
0.88
33.27
25.92
Logo: Google
Gemma 4 31B
262k
Aberto
15
$0.10
15
2.72
148.94
113.52
Logo: DeepSeek
DeepSeek V3.1 Terminus (non-reasoning) (FP4)
164k
Aberto
14*
--
40
1.17
13.58
--
Logo: Google
Gemma 4 31B (non-reasoning) (FP8)
262k
Aberto
14*
--
15
2.23
35.13
--
Logo: DeepSeek
DeepSeek V3.1 (non-reasoning) (FP4)
164k
Aberto
14*
--
9
1.46
59.76
--
Logo: Google
Gemma 4 26B A4B (non-reasoning) (FP8)
262k
Aberto
13*
--
31
0.72
16.91
--
Logo: Alibaba
Qwen3.5 4B (FP8)
262k
Aberto
13*
--
31
0.87
82.79
65.53
Logo: DeepSeek
DeepSeek R1 0528
164k
Aberto
13*
--
31
0.89
82.75
65.49
Logo: NVIDIA
Nemotron 3.5 Lightning (NVFP4)
262k
Aberto
13
$0.12
337
0.52
7.94
5.93
Logo: NVIDIA
Nemotron 3 Super
262k
Aberto
13
$0.48
69
77.50
113.81
29.04
Logo: Alibaba
Qwen3 235B A22B 2507 (FP8)
262k
Aberto
13
--
35
0.82
72.18
57.09
Logo: Alibaba
Qwen3 235B 2507 (FP8)
262k
Aberto
12*
--
15
1.37
35.59
--
Logo: Alibaba
Qwen3 Coder 480B (Turbo, FP4)
262k
Aberto
12*
--
39
0.94
13.70
--
Logo: OpenAI
gpt-oss-120b (high) (Turbo)
131k
Aberto
12
$0.11
322
0.73
8.48
6.20
Logo: OpenAI
gpt-oss-120b (high)
131k
Aberto
12
$0.03
52
0.60
48.97
38.70
Logo: IBM
Granite 4.2 8B
131k
Aberto
11
$0.02
65
0.73
39.01
30.62
Logo: Alibaba
Qwen3.5 4B (non-reasoning) FP8
262k
Aberto
11*
--
16
0.94
31.67
--
Logo: OpenAI
gpt-oss-120b (low)
131k
Aberto
10*
--
44
0.72
56.96
44.99
Logo: Meta
Llama 4 Maverick (FP8)
1.05M
Aberto
10*
--
68
0.58
7.99
--
Logo: DeepSeek
DeepSeek V3 0324 (FP4)
164k
Aberto
10
$0.02
48
1.12
11.52
--
Logo: Alibaba
Qwen3 Next 80B A3B
262k
Aberto
10*
--
147
0.67
4.08
--
Logo: IBM
Granite 4.2 3B
131k
Aberto
9
$0.01
223
0.52
11.71
8.95
Logo: OpenAI
gpt-oss-20b (high)
131k
Aberto
9
$0.01
111
0.50
23.09
18.07
Logo: Google
Gemma 4 E4B
262k
Aberto
9*
--
70
0.84
36.80
28.77
Logo: NVIDIA
Nemotron 3 Nano
262k
Aberto
9
$0.02
105
6.23
30.00
19.02
Logo: Alibaba
Qwen3 32B (FP8)
41k
Aberto
9*
--
31
1.50
81.16
63.73
Logo: DeepSeek
DeepSeek V3 (Dec)
164k
Aberto
8
$0.02
19
2.12
27.88
--
Logo: Mistral
Mistral Small 3.2 (FP8)
128k
Aberto
8*
--
37
0.96
14.41
--
Logo: Alibaba
Qwen3 14B (FP8)
32.8k
Aberto
8*
--
27
1.56
92.73
72.93
Logo: Meta
Llama 4 Scout
328k
Aberto
8*
--
33
0.83
15.82
--
Logo: DeepSeek
DeepSeek R1 Distill Llama 70B
131k
Aberto
8*
--
--
--
--
--
Logo: Alibaba
Qwen2.5 72B
32.8k
Aberto
8*
--
26
2.48
21.42
--
Logo: Meta
Llama 3.3 70B (Turbo, FP8)
131k
Aberto
8*
--
17
2.26
31.79
--
Logo: Alibaba
Qwen3 30B (FP8)
41k
Aberto
8*
--
126
0.58
20.36
15.83
Logo: Google
Gemma 4 E4B (non-reasoning)
262k
Aberto
7*
--
72
0.79
7.73
--
Logo: NVIDIA
NVIDIA Nemotron Nano 9B V2
131k
Aberto
7*
--
104
5.12
29.19
19.26
Logo: Alibaba
Qwen3 32B (non-reasoning) (FP8)
41k
Aberto
7*
--
30
1.34
17.90
--
Logo: Mistral
Mistral Small 3.1
128k
Aberto
7
$0.02
39
0.95
13.82
--
Logo: Meta
Llama 3.1 8B (Turbo, FP8)
131k
Aberto
7*
--
18
1.16
28.67
--
Logo: Meta
Llama 3.1 8B
131k
Aberto
7*
--
19
1.13
27.39
--
Logo: NVIDIA
Nemotron 3 Nano (non-reasoning)
262k
Aberto
7*
--
105
0.69
5.44
--
Logo: NVIDIA
NVIDIA Nemotron Nano 9B V2 (non-reasoning)
131k
Aberto
7*
--
99
6.30
11.33
--
Logo: Alibaba
Qwen3 14B (non-reasoning) (FP8)
41k
Aberto
7*
--
29
1.21
18.45
--
Logo: Mistral
Mistral Small 3
32.8k
Aberto
7*
--
51
0.96
10.67
--
Logo: Alibaba
Qwen3 30B (non-reasoning) (FP8)
41k
Aberto
7*
--
148
0.58
3.95
--
Logo: Meta
Llama 3.1 70B
131k
Aberto
7*
--
40
1.96
14.39
--
Logo: Meta
Llama 3.1 70B (Turbo, FP8)
131k
Aberto
7*
--
39
2.02
14.97
--
Logo: Nous Research
Hermes 3 - Llama-3.1 70B
131k
Aberto
6*
--
28
2.28
19.88
--
Logo: Microsoft
Phi-4
16.4k
Aberto
6*
--
70
0.96
8.06
--
Logo: Meta
Llama 3.2 11B (Vision)
131k
Aberto
5*
--
15
3.03
35.78
--
Logo: Google
Gemma 3 27B
131k
Aberto
5
$0.16
25
1.25
21.45
--
Logo: Google
Gemma 3 4B
131k
Aberto
5*
--
21
1.55
25.74
--
Logo: Meta
Llama 3 8B
8.19k
Aberto
5*
--
--
--
--
--
Logo: Google
Gemma 3 12B
131k
Aberto
4
$0.13
43
1.09
12.77
--

Principais definições

Perguntas frequentes

Perguntas comuns sobre DeepInfra