DeepInfra: 모델 지능, 성능 및 가격

DeepInfra
DeepInfra

이 분석은 사용 사례에 가장 적합한 DeepInfra 제공 모델을 선택하는 데 도움을 드립니다.

가장 높은 지능

Updated
#1
GLM-5.3 (max)GLM-5.3 (max)
45
#2
GLM-5.3-FlashGLM-5.3-Flash
42
#3
Qwen3.8 2.4T A95BQwen3.8 2.4T A95B
40
#4
DeepSeek V4 Pro 0813 (max)DeepSeek V4 Pro 0813 (max)
36
#5
DeepSeek V4 Flash 0731 (max)DeepSeek V4 Flash 0731 (max)
35

Intelligence Index

총 모델 108개

가장 빠름

#1
Inkling SmallInkling Small
250 t/s
#2
gpt-oss-120b (high) (Turbo)gpt-oss-120b (high) (Turbo)
230 t/s
#3
Granite 4.2 3BGranite 4.2 3B
218 t/s
#4
Inkling (FP8)Inkling (FP8)
208 t/s
#5
Qwen3 30B (FP8)Qwen3 30B (FP8)
171 t/s

출력 속도

총 모델 108개

가장 낮은 가격

#1
Llama 3.1 8B (Turbo, FP8)Llama 3.1 8B (Turbo, FP8)
$0.02
#2
Llama 3.1 8BLlama 3.1 8B
$0.02
#3
Granite 4.2 3BGranite 4.2 3B
$0.02
#4
Gemma 4 E4BGemma 4 E4B
$0.03
#5
Gemma 4 E4B (Non-reasoning)Gemma 4 E4B (Non-reasoning)
$0.03

혼합 가격(토큰 100만 개당)

총 모델 108개

DeepInfra에서 지능, 성능, 가격 특성이 서로 다른 모델 108개를 제공합니다. 아래에서 모델별 주요 지표를 비교합니다.

  • DeepInfra에서 지능이 가장 높은 모델은 GLM-5.3 (max)(45), GLM-5.3-Flash(42) 및 Qwen3.8 2.4T A95B(40)입니다.
  • 출력 속도가 가장 빠른 모델은 Inkling Small(250 t/s), gpt-oss-120b (high) (Turbo)(230 t/s) 및 Granite 4.2 3B(218 t/s)입니다.
  • 첫 답변 토큰까지 걸린 시간이 가장 짧은 모델은 Qwen3 30B (Non-reasoning) (FP8)(0.46초), Llama 4 Maverick (FP8)(0.51초) 및 Qwen3 Coder 480B (Turbo, FP4)(0.64초)입니다.
  • 토큰 100만 개당 혼합 가격이 가장 낮은 모델은 Llama 3.1 8B (Turbo, FP8)($0.02), Llama 3.1 8B($0.02) 및 Granite 4.2 3B($0.02)입니다.
  • DeepInfra에서 가장 큰 컨텍스트 창을 지원하는 모델은 GLM-5.2 (max) (FP4)(1M), DeepSeek V4 Pro (max) (FP4)(1M) 및 DeepSeek V4 Flash (high) (FP4)(1M)입니다.

주요 내용

Updated
Artificial Analysis Intelligence Index · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens (blended) · Lower is better

지능 평가

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
See more

Agentic knowledge work, (Elo-500)/2000

Agentic real-world work tasks, (Elo-500)/2000

Agentic SaaS workflows

Agentic coding & terminal use

Coding

Reasoning & knowledge

Professional document reasoning, All-pass

Physics reasoning

Long context reasoning

Legal agentic work, criterion pass rate

Agentic business operations

Quantitative analysis on spreadsheets & documents

Instruction following

Agentic tool use

Long-horizon agentic tasks

Kubernetes incident root-cause analysis

Visual reasoning

Intelligence Index vs. Price

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

컨텍스트 창

Context Window

Context window: tokens limit · Higher is better

가격

Intelligence Index vs. Price

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

성능 요약

Output Speed vs. Price

Output speed: output tokens per second · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

속도

출력 속도(초당 토큰 수)로 측정

Output Speed

Output tokens per second · Higher is better

지연 시간

첫 토큰까지 걸린 시간(초)으로 측정

Latency: Time To First Answer Token

Seconds to first answer token received · Accounts for reasoning model 'thinking' time

종단 간 응답 시간

Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed

End-to-End Response Time vs. Price

End-to-end response time: end-to-end seconds to output 500 tokens · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

추가 분석
Z AI 로고
GLM-5.3 (max)
1.05M
오픈
45
$2.16
136
1.03
19.47
14.75
Z AI 로고
GLM-5.3-Flash
1M
오픈
42
$0.43
30
1.28
84.56
66.62
Alibaba 로고
Qwen3.8 2.4T A95B
262k
오픈
40
$2.89
78
1.48
33.57
25.67
DeepSeek 로고
DeepSeek V4 Pro 0813 (max)
1M
오픈
36
$1.00
113
0.80
22.97
17.74
DeepSeek 로고
DeepSeek V4 Flash 0731 (max)
1M
오픈
35
$0.10
57
0.87
45.01
35.32
Z AI 로고
GLM-5.2 (max) (FP4)
1.05M
오픈
34
$0.65
60
1.06
42.74
33.34
Alibaba 로고
Qwen3.8 27B (xhigh)
262k
오픈
34
$0.65
35
1.13
71.77
56.51
DeepSeek 로고
DeepSeek V4 Pro (max) (FP4)
1.05M
오픈
31
$1.19
52
1.20
94.74
83.95
DeepSeek 로고
DeepSeek V4 Pro (high) (FP4)
65.5k
오픈
30*
--
56
1.50
45.78
35.40
MiniMax 로고
MiniMax-M3
524k
오픈
30
$0.44
24
1.00
106.63
84.50
Z AI 로고
GLM-5 (FP4)
203k
오픈
28*
--
68
0.95
53.77
45.50
Kimi 로고
Kimi K2.6 (FP4)
262k
오픈
27
$0.51
44
1.37
114.58
101.77
Z AI 로고
GLM-5.1 (FP4)
203k
오픈
26
$0.69
31
0.85
137.62
120.83
Xiaomi 로고
MiMo-V2.5-Pro
65.5k
오픈
26
$0.43
37
1.07
68.73
54.13
Kimi 로고
Kimi K2.7 Code
262k
오픈
26
$0.39
41
0.91
67.89
54.70
Thinking Machines 로고
Inkling Small
524k
오픈
26
$0.12
250
0.98
10.98
8.00
Tencent 로고
Hy3 (FP8)
262k
오픈
26
$0.07
59
1.15
43.50
33.88
Thinking Machines 로고
Inkling (FP8)
131k
오픈
26
$0.81
210
0.56
12.48
9.53
DeepSeek 로고
DeepSeek V4 Flash (high) (FP4)
1.05M
오픈
25
$0.10
45
0.91
39.59
27.56
DeepSeek 로고
DeepSeek V4 Flash (max) (FP4)
1.05M
오픈
25
$0.09
44
0.95
141.41
128.98
Z AI 로고
GLM-5.1 (Non-reasoning) (FP4)
203k
오픈
24*
--
29
1.05
18.12
--
Kimi 로고
Kimi K2.6 (Non-reasoning) (FP4)
262k
오픈
24*
--
34
1.49
16.08
--
Kimi 로고
Kimi K2.5
262k
오픈
23*
--
43
1.38
81.31
68.41
NVIDIA 로고
Nemotron 3 Ultra
262k
오픈
23
$0.47
127
3.35
25.16
17.88
NVIDIA 로고
Nemotron 3 Ultra BF16
262k
오픈
23
$0.92
153
2.65
20.78
14.87
Alibaba 로고
Qwen3.5 27B (FP8)
262k
오픈
23*
--
70
0.90
36.81
28.73
MiniMax 로고
MiniMax-M2.5 (FP8)
197k
오픈
23*
--
22
5.49
118.62
90.51
Xiaomi 로고
MiMo-V2.5
262k
오픈
22
$0.02
29
1.74
87.46
68.58
Z AI 로고
GLM-4.7 (FP4)
203k
오픈
22*
--
18
1.74
143.59
113.48
Alibaba 로고
Qwen3.6 27B FP8
262k
오픈
22
$0.37
65
0.94
95.81
87.19
Z AI 로고
GLM-5 (Non-reasoning) (FP8)
203k
오픈
22*
--
56
1.39
10.30
--
DeepSeek 로고
DeepSeek V3.2 (FP4)
164k
오픈
21*
--
28
1.67
91.22
71.64
Alibaba 로고
Qwen3.5 397B A17B (Non-reasoning) (FP8)
262k
오픈
21*
--
41
0.94
13.25
--
InclusionAI 로고
Ling 3.0 Flash
131k
오픈
21
$0.03
57
0.67
44.16
34.79
Alibaba 로고
Qwen3.6 27B (Non-reasoning) FP8
262k
오픈
20*
--
60
0.92
9.25
--
StepFun 로고
Step 3.7 Flash
256k
오픈
19*
--
168
0.66
15.54
11.90
Alibaba 로고
Qwen3.5 27B (Non-reasoning) FP8
262k
오픈
19*
--
66
0.92
8.45
--
Alibaba 로고
Qwen3.5 35B A3B (FP8)
262k
오픈
19*
--
159
0.80
16.50
12.56
Alibaba 로고
Qwen3.5 397B A17B (FP8)
262k
오픈
19
$0.23
39
0.97
95.92
82.06
Alibaba 로고
Qwen3.6 35B A3B (FP8)
262k
오픈
19
$0.14
93
0.70
63.90
57.84
Z AI 로고
GLM-4.6 (FP4)
203k
오픈
19*
--
42
0.75
59.62
47.10
Xiaomi 로고
MiMo-V2.5-Pro (Non-reasoning)
65.5k
오픈
18*
--
30
1.16
17.64
--
Meta 로고
Muse Glimmer (high)
131k
오픈
18
$0.05
148
0.89
17.78
13.51
Alibaba 로고
Qwen3.5 122B A10B (Non-reasoning) (FP4)
262k
오픈
18*
--
143
0.92
4.42
--
Z AI 로고
GLM-4.7 (Non-reasoning) (FP4)
203k
오픈
17*
--
18
1.05
28.51
--
Google 로고
Gemma 4 26B A4B (FP8)
262k
오픈
17*
--
31
0.89
81.03
64.11
Alibaba 로고
Qwen3.5 122B A10B (FP4)
262k
오픈
16
$0.14
155
0.81
16.96
12.92
DeepSeek 로고
DeepSeek V3.2 (Non-reasoning)
164k
오픈
16*
--
33
1.34
16.44
--
Google 로고
Gemma 4 31B
262k
오픈
15
$0.02
15
2.14
149.62
114.50
Alibaba 로고
Qwen3.6 35B A3B (Non-reasoning) (FP8)
262k
오픈
15*
--
92
0.69
6.12
--
Alibaba 로고
Qwen3.5 35B A3B (Non-reasoning) FP8
262k
오픈
15*
--
154
0.72
3.97
--
Z AI 로고
GLM-4.7-Flash
203k
오픈
15*
--
35
1.56
72.92
57.09
IBM 로고
Granite 4.2 30B
131k
오픈
15*
--
77
0.80
33.16
25.89
DeepSeek 로고
DeepSeek V3.1 Terminus (Non-reasoning) (FP4)
164k
오픈
14*
--
60
0.97
9.35
--
Google 로고
Gemma 4 31B (Non-reasoning) (FP8)
262k
오픈
14*
--
20
1.45
26.12
--
DeepSeek 로고
DeepSeek V3.1 (Non-reasoning) (FP4)
164k
오픈
14*
--
12
1.00
43.44
--
NVIDIA 로고
Nemotron 3.5 Lightning (NVFP4)
262k
오픈
14
$0.12
--
--
--
--
NVIDIA 로고
Nemotron 3 Super
262k
오픈
14
$0.48
90
13.04
40.86
22.26
Google 로고
Gemma 4 26B A4B (Non-reasoning) (FP8)
262k
오픈
13*
--
23
0.82
22.42
--
Alibaba 로고
Qwen3.5 4B (FP8)
262k
오픈
13*
--
29
0.72
85.88
68.12
DeepSeek 로고
DeepSeek R1 0528
164k
오픈
13*
--
23
0.93
109.29
86.69
Alibaba 로고
Qwen3 235B A22B 2507 (FP8)
262k
오픈
13
--
92
0.78
27.85
21.65
OpenAI 로고
gpt-oss-120b (high) (Turbo)
131k
오픈
12
$0.11
230
0.76
11.61
8.68
OpenAI 로고
gpt-oss-120b (high)
131k
오픈
12
$0.03
59
0.59
43.28
34.15
Alibaba 로고
Qwen3 235B 2507 (Non-reasoning) (FP8)
262k
오픈
12*
--
16
0.67
32.20
--
Alibaba 로고
Qwen3 Coder 480B (Turbo, FP4)
262k
오픈
12*
--
73
0.65
7.54
--
IBM 로고
Granite 4.2 8B
131k
오픈
12
$0.02
72
0.64
35.43
27.84
Alibaba 로고
Qwen3.5 4B (Non-reasoning) FP8
262k
오픈
11*
--
22
0.69
23.52
--
OpenAI 로고
gpt-oss-120b (low)
131k
오픈
10*
--
55
0.63
46.38
36.60
DeepSeek 로고
DeepSeek V3 0324 (FP4)
164k
오픈
10
$0.02
46
1.09
12.01
--
Alibaba 로고
Qwen3 Next 80B A3B
262k
오픈
10*
--
160
0.70
3.83
--
Meta 로고
Llama 4 Maverick (FP8)
1.05M
오픈
9
$0.03
74
0.51
7.26
--
IBM 로고
Granite 4.2 3B
131k
오픈
9
$0.01
214
0.45
12.15
9.37
OpenAI 로고
gpt-oss-20b (high)
131k
오픈
9
$0.01
118
0.44
21.64
16.96
NVIDIA 로고
Llama Nemotron Super 49B v1.5
131k
오픈
9*
--
95
4.25
30.66
21.13
Google 로고
Gemma 4 E4B
262k
오픈
9*
--
44
0.81
57.51
45.36
NVIDIA 로고
Nemotron 3 Nano
262k
오픈
9
$0.02
104
4.46
28.43
19.18
DeepSeek 로고
DeepSeek V3 (Dec)
164k
오픈
8
$0.02
29
0.66
17.81
--
DeepSeek 로고
DeepSeek R1 Distill Llama 70B
131k
오픈
8*
--
25
0.99
100.22
79.39
Alibaba 로고
Qwen2.5 72B
32.8k
오픈
8*
--
29
2.33
19.68
--
Meta 로고
Llama 3.3 70B (Turbo, FP8)
131k
오픈
8*
--
16
2.10
34.16
--
Alibaba 로고
Qwen3 30B (FP8)
41k
오픈
8*
--
159
0.51
16.23
12.57
NVIDIA 로고
NVIDIA Nemotron Nano 12B v2 VL (FP8)
131k
오픈
7*
--
148
3.29
20.22
13.55
Google 로고
Gemma 4 E4B (Non-reasoning)
262k
오픈
7*
--
45
0.77
11.92
--
NVIDIA 로고
NVIDIA Nemotron Nano 9B V2
131k
오픈
7*
--
92
7.88
35.01
21.70
Mistral 로고
Mistral Small 3.1
128k
오픈
7
$0.02
36
0.93
15.00
--
NVIDIA 로고
Llama Nemotron Super 49B v1.5 (Non-reasoning)
131k
오픈
7*
--
161
3.02
6.13
--
Alibaba 로고
Qwen3 32B (Non-reasoning) (FP8)
41k
오픈
7*
--
27
1.51
20.32
--
Alibaba 로고
Qwen3 32B (FP8)
41k
오픈
7
--
29
1.48
88.88
69.92
Mistral 로고
Mistral Small 3.2 (FP8)
128k
오픈
7
$0.11
41
0.90
13.22
--
NVIDIA 로고
Llama 3.1 Nemotron 70B
131k
오픈
7*
--
145
2.97
6.41
--
Meta 로고
Llama 3.1 8B (Turbo, FP8)
131k
오픈
7*
--
25
0.92
20.83
--
Meta 로고
Llama 3.1 8B
131k
오픈
7*
--
20
0.98
25.70
--
NVIDIA 로고
Nemotron 3 Nano (Non-reasoning)
262k
오픈
7*
--
102
0.67
5.59
--
NVIDIA 로고
NVIDIA Nemotron Nano 9B V2 (Non-reasoning)
131k
오픈
7*
--
103
5.72
10.57
--
Alibaba 로고
Qwen3 14B (Non-reasoning) (FP8)
41k
오픈
7*
--
61
0.80
9.00
--
Mistral 로고
Mistral Small 3
32.8k
오픈
7*
--
52
0.82
10.51
--
Alibaba 로고
Qwen3 30B (Non-reasoning) (FP8)
41k
오픈
7*
--
164
0.50
3.56
--
Meta 로고
Llama 3.1 70B
131k
오픈
7*
--
35
1.93
16.22
--
Meta 로고
Llama 3.1 70B (Turbo, FP8)
131k
오픈
7*
--
35
1.97
16.06
--
Meta 로고
Llama 4 Scout
328k
오픈
6
$0.06
23
1.06
22.49
--
Alibaba 로고
Qwen3 14B (FP8)
32.8k
오픈
6
--
57
0.85
44.88
35.23
Nous Research 로고
Hermes 3 - Llama-3.1 70B
131k
오픈
6*
--
31
2.14
18.03
--
Microsoft 로고
Phi-4
16.4k
오픈
6*
--
56
0.93
9.82
--
NVIDIA 로고
NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) (FP8)
131k
오픈
6*
--
138
3.91
7.53
--
Meta 로고
Llama 3.2 11B (Vision)
131k
오픈
5*
--
19
1.23
27.14
--
Google 로고
Gemma 3 27B
131k
오픈
5
$0.16
25
1.41
21.03
--
Google 로고
Gemma 3 4B
131k
오픈
5*
--
19
1.71
28.14
--
Meta 로고
Llama 3 8B
8.19k
오픈
5*
--
--
--
--
--
Google 로고
Gemma 3 12B
131k
오픈
4
$0.13
39
1.06
13.76
--

주요 용어 정의

자주 묻는 질문

DeepInfra에 관한 일반적인 질문