SiliconFlow: inteligência, desempenho e preço dos modelos

SiliconFlow
SiliconFlow

Esta análise ajuda você a escolher o melhor modelo oferecido por SiliconFlow para seu caso de uso.

Mais inteligente

#1
GLM-5.3-FlashGLM-5.3-Flash
42
#2
DeepSeek V4.1 Flash (max)DeepSeek V4.1 Flash (max)
39
#3
DeepSeek V4 Pro 0813 (max)DeepSeek V4 Pro 0813 (max)
36
#4
DeepSeek V4 Flash Vision (max)DeepSeek V4 Flash Vision (max)
35
#5
DeepSeek V4 Flash 0731 (max)DeepSeek V4 Flash 0731 (max)
34

Intelligence Index

37 modelos no total

Mais rápido

#1
DeepSeek V4 Flash Vision (max)DeepSeek V4 Flash Vision (max)
221 t/s
#2
DeepSeek V4.1 Flash (max)DeepSeek V4.1 Flash (max)
168 t/s
#3
GLM-5.2 (max) (FP8)GLM-5.2 (max) (FP8)
121 t/s
#4
Gemma 4 12B (non-reasoning)Gemma 4 12B (non-reasoning)
115 t/s
#5
Gemma 4 12BGemma 4 12B
115 t/s

Velocidade de saída

37 modelos no total

Menor preço

#1
DeepSeek V4 Flash (high) (FP8)DeepSeek V4 Flash (high) (FP8)
$0.07
#2
DeepSeek V4 Flash (max) (FP8)DeepSeek V4 Flash (max) (FP8)
$0.07
#3
DeepSeek V4.1 Flash (max)DeepSeek V4.1 Flash (max)
$0.09
#4
GLM-5.3-FlashGLM-5.3-Flash
$0.10
#5
Hy3 (FP8)Hy3 (FP8)
$0.10

Preço combinado (por 1M de tokens)

37 modelos no total

SiliconFlow oferece 37 modelos, cada um com características diferentes de inteligência, desempenho e preço. Veja abaixo uma comparação das principais métricas entre os modelos.

  • Em inteligência, os principais modelos oferecidos por SiliconFlow são GLM-5.3-Flash (42), DeepSeek V4.1 Flash (max) (39) e DeepSeek V4 Pro 0813 (max) (36).
  • Em velocidade de saída, os modelos mais rápidos são DeepSeek V4 Flash Vision (max) (221 t/s), DeepSeek V4.1 Flash (max) (168 t/s) e GLM-5.2 (max) (FP8) (121 t/s). A velocidade varia significativamente entre os modelos, com uma diferença de 92% entre o mais rápido e o mais lento.
  • Em latência, GLM-5.2 (non-reasoning) (FP8) (1.44s), Kimi K2.6 (non-reasoning) (FP8) (1.54s) e Gemma 4 26B A4B (non-reasoning) (FP8) (2.36s) têm o menor tempo até o primeiro token da resposta final.
  • Em preço, DeepSeek V4 Flash (high) (FP8) ($0.07), DeepSeek V4 Flash (max) (FP8) ($0.07) e DeepSeek V4.1 Flash (max) ($0.09) têm os menores preços combinados por 1M de tokens.
  • Em tamanho da janela de contexto, GLM-5.2 (max) (FP8) (1M), DeepSeek V4 Pro (max) (FP8) (1M) e DeepSeek V4 Pro (high) (FP8) (1M) aceitam as maiores janelas entre os modelos oferecidos por SiliconFlow.
Artificial Analysis Intelligence Index · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens (blended) · Lower is better

Avaliações de inteligência

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
Ver mais

Agentic knowledge work, (Elo-500)/2000

Agentic real-world work tasks, (Elo-500)/2000

Agentic SaaS workflows

Agentic coding & terminal use

SciCodeUnder review

Coding

Reasoning & knowledge

Professional document reasoning, All-pass

CritPtUnder review

Physics reasoning

Long context reasoning

Legal agentic work, criterion pass rate

Agentic business operations

Agentic scientific research workflows in a terminal

Quantitative analysis on spreadsheets & documents

Kubernetes incident root-cause analysis

Visual reasoning

Medical long context reasoning

Intelligence Index vs. preço

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Janela de contexto

Janela de contexto

Context window: tokens limit · Higher is better

Preços

Intelligence Index vs. preço

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Resumo do desempenho

Velocidade de saída vs. preço

Output speed: output tokens per second · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Velocidade

Medida pela velocidade de saída (tokens por segundo)

Velocidade de saída

Output tokens per second · Higher is better

Latência

Medida pelo tempo (segundos) até o primeiro token

Latência: Tempo até o primeiro token da resposta final

Segundos até receber o primeiro token de resposta · Inclui o tempo de “pensamento” dos modelos de raciocínio

Comportamento do cache

Taxa de acertos do cache

Proporção de tokens de entrada elegíveis para cache atendidos pelo cache · Mediana das últimas quatro semanas, atualizada em 26 de set. de 2026

Custo por tarefa vs. Taxa de acertos do cache

Weighted average cost (USD) per Intelligence Index task · Cache hit rate: median of the last four weeks, updated Sep 26, 2026
Most attractive quadrant
Pareto line

Tempo de resposta de ponta a ponta

Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed

Tempo de resposta de ponta a ponta vs. preço

End-to-end response time: end-to-end seconds to output 500 tokens · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Análises adicionais
Logo: Z AI
GLM-5.3-Flash
1M
Aberto
42
$0.50
55
1.55
46.86
36.25
Logo: DeepSeek
DeepSeek V4.1 Flash (max)
1M
Aberto
39
$0.86
163
2.05
17.36
12.25
Logo: DeepSeek
DeepSeek V4 Pro 0813 (max)
1.05M
Aberto
36
$1.45
92
1.70
28.75
21.64
Logo: DeepSeek
DeepSeek V4 Flash Vision (max)
1M
Proprietário
35
$2.51
217
1.36
12.86
9.20
Logo: DeepSeek
DeepSeek V4 Flash 0731 (max)
1.05M
Aberto
34
$0.15
49
2.02
53.45
41.14
Logo: Z AI
GLM-5.2 (max) (FP8)
1.05M
Aberto
34
$0.85
126
1.29
21.18
15.91
Logo: DeepSeek
DeepSeek V4 Pro (max) (FP8)
1.05M
Aberto
30
--
52
1.68
94.95
83.70
Logo: DeepSeek
DeepSeek V4 Pro (high) (FP8)
1.05M
Aberto
30*
--
50
2.19
51.70
39.58
Logo: MiniMax
MiniMax-M3 (FP8)
1M
Aberto
29
$0.45
87
2.76
31.60
23.07
Logo: Z AI
GLM-5 (FP8)
200k
Aberto
28*
--
--
--
--
--
Logo: Kimi
Kimi K2.6 (FP8)
262k
Aberto
27
$0.61
27
1.59
182.00
162.19
Logo: Z AI
GLM-5.1 (FP8)
205k
Aberto
26
$1.55
59
2.06
74.79
64.26
Logo: Tencent
Hy3 (FP8)
256k
Aberto
25
$0.07
85
2.97
32.35
23.50
Logo: DeepSeek
DeepSeek V4 Flash (high) (FP8)
1.05M
Aberto
24
$0.12
55
1.55
32.92
22.35
Logo: Z AI
GLM-5.1 (non-reasoning) (FP8)
205k
Aberto
24*
--
30
2.38
19.17
--
Logo: DeepSeek
DeepSeek V4 Flash (max) (FP8)
1.05M
Aberto
24
$0.10
57
1.52
109.31
98.97
Logo: Kimi
Kimi K2.6 (non-reasoning) (FP8)
262k
Aberto
24*
--
29
1.55
19.04
--
Logo: Kimi
Kimi K2.5 (FP8)
262k
Aberto
23*
--
27
2.22
131.02
110.23
Logo: Alibaba
Qwen3.5 27B (FP8)
262k
Aberto
23*
--
22
3.62
119.06
92.35
Logo: Z AI
GLM-5.2 (non-reasoning) (FP8)
1.05M
Aberto
22*
--
95
1.44
6.71
--
Logo: Z AI
GLM-5 (non-reasoning) (FP8)
205k
Aberto
22*
--
--
--
--
--
Logo: DeepSeek
DeepSeek V3.2 (FP8)
164k
Aberto
21*
--
22
2.60
114.82
89.78
Logo: Alibaba
Qwen3.6 27B (FP8)
262k
Aberto
21
$0.36
36
3.16
175.95
158.80
Logo: Alibaba
Qwen3.5 35B A3B (FP8)
262k
Aberto
19*
--
52
1.91
50.18
38.62
Logo: LongCat
LongCat 2.0 (FP8)
262k
Aberto
19
$0.15
72
2.46
37.27
27.84
Logo: Alibaba
Qwen3.6 35B A3B (FP8)
262k
Aberto
18
$0.32
106
1.84
57.51
50.94
Logo: StepFun
Step 3.5 Flash (FP8)
262k
Aberto
17*
--
69
1.62
37.76
28.91
Logo: DeepSeek
DeepSeek V3.2 (non-reasoning) (FP8)
164k
Aberto
16*
--
13
4.50
43.87
--
Logo: Alibaba
Qwen3.5 122B A10B (FP8)
262k
Aberto
16
$0.21
53
1.74
49.16
37.94
Logo: Google
Gemma 4 31B (FP8)
262k
Aberto
15
$0.54
42
3.58
56.43
41.03
Logo: Google
Gemma 4 12B
262k
Aberto
14*
--
113
2.42
24.50
17.67
Logo: Google
Gemma 4 31B (non-reasoning) (FP8)
262k
Aberto
14*
--
45
3.51
14.52
--
Logo: Google
Gemma 4 26B A4B (non-reasoning) (FP8)
262k
Aberto
13*
--
99
2.49
7.54
--
Logo: ByteDance Seed
Seed-OSS-36B-Instruct
262k
Aberto
12*
--
40
2.90
65.66
50.21
Logo: Z AI
GLM-4.6V
128k
Aberto
11*
--
--
--
--
--
Logo: Alibaba
Qwen3.5 9B (FP8)
262k
Aberto
11
$0.16
38
2.24
68.09
52.67
Logo: Z AI
GLM-4.5-Air
98.3k
Aberto
11*
--
68
2.54
39.10
29.25
Logo: Google
Gemma 4 12B (non-reasoning)
262k
Aberto
9*
--
115
2.40
6.73
--
Logo: Z AI
GLM-4.6V (non-reasoning)
128k
Aberto
8*
--
--
--
--
--
Logo: InclusionAI
Ling-flash-2.0
131k
Aberto
8*
--
5
2.47
99.87
--
Logo: Alibaba
Qwen2.5 72B (FP8)
32k
Aberto
8*
--
29
4.24
21.54
--
Logo: Baidu
ERNIE 4.5 300B A47B
131k
Aberto
8*
--
--
--
--
--
Logo: InclusionAI
Ring-flash-2.0
131k
Aberto
7*
--
--
--
--
--

Principais definições

Perguntas frequentes

Perguntas comuns sobre SiliconFlow