SiliconFlow: inteligência, desempenho e preço dos modelos

SiliconFlow
SiliconFlow

Esta análise ajuda você a escolher o melhor modelo oferecido por SiliconFlow para seu caso de uso.

Mais inteligente

Updated
#1
GLM-5.3-FlashGLM-5.3-Flash
42
#2
DeepSeek V4 Pro 0813 (max)DeepSeek V4 Pro 0813 (max)
36
#3
DeepSeek V4 Flash 0731 (max)DeepSeek V4 Flash 0731 (max)
35
#4
GLM-5.2 (max) (FP8)GLM-5.2 (max) (FP8)
34
#5
DeepSeek V4 Pro (max) (FP8)DeepSeek V4 Pro (max) (FP8)
31

Intelligence Index

35 modelos no total

Mais rápido

#1
MiniMax-M3 (FP8)MiniMax-M3 (FP8)
134 t/s
#2
Qwen3.6 35B A3B (FP8)Qwen3.6 35B A3B (FP8)
113 t/s
#3
DeepSeek V4 Flash (max) (FP8)DeepSeek V4 Flash (max) (FP8)
111 t/s
#4
GLM-5.3-FlashGLM-5.3-Flash
110 t/s
#5
Gemma 4 12BGemma 4 12B
109 t/s

Velocidade de saída

35 modelos no total

Menor preço

#1
DeepSeek V4 Flash (high) (FP8)DeepSeek V4 Flash (high) (FP8)
$0.07
#2
DeepSeek V4 Flash (max) (FP8)DeepSeek V4 Flash (max) (FP8)
$0.07
#3
GLM-5.3-FlashGLM-5.3-Flash
$0.10
#4
Hy3 (FP8)Hy3 (FP8)
$0.10
#5
Qwen3.5 9B (FP8)Qwen3.5 9B (FP8)
$0.11

Preço combinado (por 1M de tokens)

35 modelos no total

SiliconFlow oferece 35 modelos, cada um com características diferentes de inteligência, desempenho e preço. Veja abaixo uma comparação das principais métricas entre os modelos.

  • Em inteligência, os principais modelos oferecidos por SiliconFlow são GLM-5.3-Flash (42), DeepSeek V4 Pro 0813 (max) (36) e DeepSeek V4 Flash 0731 (max) (35).
  • Em velocidade de saída, os modelos mais rápidos são MiniMax-M3 (FP8) (134 t/s), Qwen3.6 35B A3B (FP8) (113 t/s) e DeepSeek V4 Flash (max) (FP8) (111 t/s).
  • Em latência, Kimi K2.6 (Non-reasoning) (FP8) (1.64s), GLM-5.2 (Non-reasoning) (FP8) (2.08s) e Gemma 4 12B (Non-reasoning) (2.40s) têm o menor tempo até o primeiro token da resposta final.
  • Em preço, DeepSeek V4 Flash (high) (FP8) ($0.07), DeepSeek V4 Flash (max) (FP8) ($0.07) e GLM-5.3-Flash ($0.10) têm os menores preços combinados por 1M de tokens.
  • Em tamanho da janela de contexto, GLM-5.2 (max) (FP8) (1M), DeepSeek V4 Pro (max) (FP8) (1M) e DeepSeek V4 Pro (high) (FP8) (1M) aceitam as maiores janelas entre os modelos oferecidos por SiliconFlow.

Destaques

Updated
Artificial Analysis Intelligence Index · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens (blended) · Lower is better

Avaliações de inteligência

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
See more

Agentic knowledge work, (Elo-500)/2000

Agentic real-world work tasks, (Elo-500)/2000

Agentic SaaS workflows

Agentic coding & terminal use

Coding

Reasoning & knowledge

Professional document reasoning, All-pass

Physics reasoning

Long context reasoning

Legal agentic work, criterion pass rate

Agentic business operations

Quantitative analysis on spreadsheets & documents

Instruction following

Agentic tool use

Long-horizon agentic tasks

Kubernetes incident root-cause analysis

Visual reasoning

Intelligence Index vs. Price

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Janela de contexto

Context Window

Context window: tokens limit · Higher is better

Preços

Intelligence Index vs. Price

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Resumo do desempenho

Output Speed vs. Price

Output speed: output tokens per second · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Velocidade

Medida pela velocidade de saída (tokens por segundo)

Output Speed

Output tokens per second · Higher is better

Latência

Medida pelo tempo (segundos) até o primeiro token

Latency: Time To First Answer Token

Seconds to first answer token received · Accounts for reasoning model 'thinking' time

Tempo de resposta de ponta a ponta

Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed

End-to-End Response Time vs. Price

End-to-end response time: end-to-end seconds to output 500 tokens · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Análises adicionais
Logo: Z AI
GLM-5.3-Flash
1M
Aberto
42
$0.29
108
1.46
24.54
18.46
Logo: DeepSeek
DeepSeek V4 Pro 0813 (max)
1.05M
Aberto
36
$6.15
63
1.77
41.66
31.91
Logo: DeepSeek
DeepSeek V4 Flash 0731 (max)
1.05M
Aberto
35
$0.85
83
2.44
32.71
24.22
Logo: Z AI
GLM-5.2 (max) (FP8)
1.05M
Aberto
34
$0.85
100
2.19
27.19
20.00
Logo: DeepSeek
DeepSeek V4 Pro (max) (FP8)
1.05M
Aberto
31
--
56
2.08
89.69
78.63
Logo: DeepSeek
DeepSeek V4 Pro (high) (FP8)
1.05M
Aberto
30*
--
50
2.15
52.24
40.04
Logo: MiniMax
MiniMax-M3 (FP8)
1M
Aberto
30
$0.45
139
1.26
19.31
14.43
Logo: Z AI
GLM-5 (FP8)
200k
Aberto
28*
--
--
--
--
--
Logo: Kimi
Kimi K2.6 (FP8)
262k
Aberto
27
$0.54
26
1.70
193.41
172.35
Logo: Z AI
GLM-5.1 (FP8)
205k
Aberto
26
$1.66
47
2.09
92.70
80.06
Logo: Tencent
Hy3 (FP8)
256k
Aberto
26
$0.07
96
2.89
28.87
20.78
Logo: DeepSeek
DeepSeek V4 Flash (high) (FP8)
1.05M
Aberto
25
--
108
1.51
17.65
11.50
Logo: DeepSeek
DeepSeek V4 Flash (max) (FP8)
1.05M
Aberto
25
--
127
1.86
49.82
44.04
Logo: Z AI
GLM-5.1 (Non-reasoning) (FP8)
205k
Aberto
24*
--
30
2.48
19.31
--
Logo: Kimi
Kimi K2.6 (Non-reasoning) (FP8)
262k
Aberto
24*
--
26
1.63
20.73
--
Logo: Kimi
Kimi K2.5 (FP8)
262k
Aberto
23*
--
26
2.13
133.09
112.08
Logo: Alibaba
Qwen3.5 27B (FP8)
262k
Aberto
23*
--
41
3.43
64.10
48.54
Logo: MiniMax
MiniMax-M2.5 (FP8)
197k
Aberto
23*
--
--
--
--
--
Logo: Z AI
GLM-5.2 (Non-reasoning) (FP8)
1.05M
Aberto
22*
--
90
2.09
7.62
--
Logo: Alibaba
Qwen3.6 27B (FP8)
262k
Aberto
22
$0.36
43
3.45
146.44
131.42
Logo: Z AI
GLM-5 (Non-reasoning) (FP8)
205k
Aberto
22*
--
--
--
--
--
Logo: DeepSeek
DeepSeek V3.2 (FP8)
164k
Aberto
21*
--
19
2.53
132.85
104.26
Logo: LongCat
LongCat 2.0 (FP8)
262k
Aberto
20
$0.15
52
2.91
50.99
38.46
Logo: Alibaba
Qwen3.5 35B A3B (FP8)
262k
Aberto
19*
--
108
2.13
25.24
18.49
Logo: Alibaba
Qwen3.5 397B A17B (FP8)
262k
Aberto
19
$0.31
--
--
--
--
Logo: Alibaba
Qwen3.6 35B A3B (FP8)
262k
Aberto
19
$0.27
105
2.11
58.10
51.24
Logo: StepFun
Step 3.5 Flash (FP8)
262k
Aberto
17*
--
77
1.77
34.43
26.13
Logo: Alibaba
Qwen3.5 122B A10B (FP8)
262k
Aberto
16
$0.21
100
1.86
26.80
19.95
Logo: DeepSeek
DeepSeek V3.2 (Non-reasoning) (FP8)
164k
Aberto
16*
--
22
2.48
24.76
--
Logo: Google
Gemma 4 31B (FP8)
262k
Aberto
15
$0.10
59
3.68
41.56
29.41
Logo: Google
Gemma 4 12B
262k
Aberto
14*
--
108
2.44
25.65
18.57
Logo: Google
Gemma 4 31B (Non-reasoning) (FP8)
262k
Aberto
14*
--
57
3.45
12.25
--
Logo: Alibaba
Qwen3.5 9B (FP8)
262k
Aberto
14*
--
60
2.43
44.38
33.57
Logo: Google
Gemma 4 26B A4B (Non-reasoning) (FP8)
262k
Aberto
13*
--
65
2.45
10.09
--
Logo: ByteDance Seed
Seed-OSS-36B-Instruct
262k
Aberto
12*
--
37
3.01
70.75
54.19
Logo: Z AI
GLM-4.6V
128k
Aberto
11*
--
--
--
--
--
Logo: Z AI
GLM-4.5-Air
98.3k
Aberto
11*
--
63
2.50
41.97
31.58
Logo: Google
Gemma 4 12B (Non-reasoning)
262k
Aberto
9*
--
103
2.35
7.19
--
Logo: Z AI
GLM-4.6V (Non-reasoning)
128k
Aberto
8*
--
--
--
--
--
Logo: InclusionAI
Ling-flash-2.0
131k
Aberto
8*
--
6
2.74
81.40
--
Logo: Alibaba
Qwen2.5 72B (FP8)
32k
Aberto
8*
--
25
4.36
24.14
--
Logo: Baidu
ERNIE 4.5 300B A47B
131k
Aberto
8*
--
--
--
--
--
Logo: InclusionAI
Ring-flash-2.0
131k
Aberto
7*
--
--
--
--
--

Principais definições

Perguntas frequentes

Perguntas comuns sobre SiliconFlow