Parasail: inteligência, desempenho e preço dos modelos

Parasail
Parasail

Esta análise ajuda você a escolher o melhor modelo oferecido por Parasail para seu caso de uso.

Mais inteligente

Updated
#1
GLM-5.3 (max)GLM-5.3 (max)
45
#2
Kimi K3 (max)Kimi K3 (max)
44
#3
GLM-5.3-FlashGLM-5.3-Flash
42
#4
DeepSeek V4.1 Flash (max)DeepSeek V4.1 Flash (max)
39
#5
DeepSeek V4 Flash 0731 (max)DeepSeek V4 Flash 0731 (max)
34

Intelligence Index

29 modelos no total

Mais rápido

#1
GLM-5.2 (max) (NVFP4)GLM-5.2 (max) (NVFP4)
238 t/s
#2
GLM-5.3-FlashGLM-5.3-Flash
229 t/s
#3
DeepSeek V4.1 Flash (max)DeepSeek V4.1 Flash (max)
193 t/s
#4
Kimi K2.6Kimi K2.6
182 t/s
#5
MiniMax-M3 (MXFP8)MiniMax-M3 (MXFP8)
173 t/s

Velocidade de saída

29 modelos no total

Menor preço

#1
Gemma 3 4B (FP8)Gemma 3 4B (FP8)
$0.05
#2
Gemma 3 27BGemma 3 27B
$0.09
#3
GLM-5.3-FlashGLM-5.3-Flash
$0.10
#4
Gemma 4 26B A4BGemma 4 26B A4B
$0.10
#5
Gemma 4 26B A4B (non-reasoning)Gemma 4 26B A4B (non-reasoning)
$0.10

Preço combinado (por 1M de tokens)

29 modelos no total

Parasail oferece 29 modelos, cada um com características diferentes de inteligência, desempenho e preço. Veja abaixo uma comparação das principais métricas entre os modelos.

  • Em inteligência, os principais modelos oferecidos por Parasail são GLM-5.3 (max) (45), Kimi K3 (max) (44) e GLM-5.3-Flash (42).
  • Em velocidade de saída, os modelos mais rápidos são GLM-5.2 (max) (NVFP4) (238 t/s), GLM-5.3-Flash (229 t/s) e DeepSeek V4.1 Flash (max) (193 t/s).
  • Em latência, Gemma 3 4B (FP8) (0.98s), Qwen3.6 35B A3B (non-reasoning) (FP8) (1.05s) e Qwen3 Coder Next (FP8) (1.10s) têm o menor tempo até o primeiro token da resposta final.
  • Em preço, Gemma 3 4B (FP8) ($0.05), Gemma 3 27B ($0.09) e GLM-5.3-Flash ($0.10) têm os menores preços combinados por 1M de tokens. Os preços variam em até 2.2x entre os modelos.
  • Em tamanho da janela de contexto, Kimi K3 (max) (1M), DeepSeek V4.1 Flash (max) (1M) e DeepSeek V4 Flash 0731 (max) (1M) aceitam as maiores janelas entre os modelos oferecidos por Parasail.
Updated
Artificial Analysis Intelligence Index · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens (blended) · Lower is better

Avaliações de inteligência

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
Ver mais

Agentic knowledge work, (Elo-500)/2000

Agentic real-world work tasks, (Elo-500)/2000

Agentic SaaS workflows

Agentic coding & terminal use

SciCodeUnder review

Coding

Reasoning & knowledge

Professional document reasoning, All-pass

CritPtUnder review

Physics reasoning

Long context reasoning

Legal agentic work, criterion pass rate

Agentic business operations

Agentic scientific research workflows in a terminal

Quantitative analysis on spreadsheets & documents

Kubernetes incident root-cause analysis

Visual reasoning

Medical long context reasoning

Intelligence Index vs. preço

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Janela de contexto

Janela de contexto

Context window: tokens limit · Higher is better

Preços

Intelligence Index vs. preço

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Resumo do desempenho

Velocidade de saída vs. preço

Output speed: output tokens per second · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Velocidade

Medida pela velocidade de saída (tokens por segundo)

Velocidade de saída

Output tokens per second · Higher is better

Latência

Medida pelo tempo (segundos) até o primeiro token

Latência: Tempo até o primeiro token da resposta final

Segundos até receber o primeiro token de resposta · Inclui o tempo de “pensamento” dos modelos de raciocínio

Tempo de resposta de ponta a ponta

Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed

Tempo de resposta de ponta a ponta vs. preço

End-to-end response time: end-to-end seconds to output 500 tokens · USD per 1M tokens (blended)
Most attractive quadrant

Análises adicionais
Logo: Z AI
GLM-5.3 (max)
1M
Aberto
45
$1.97
148
1.01
17.88
13.50
Logo: Kimi
Kimi K3 (max)
1.05M
Aberto
44
$4.14
146
1.23
18.30
13.66
Logo: Z AI
GLM-5.3-Flash
1M
Aberto
42
$0.67
229
0.92
11.81
8.72
Logo: DeepSeek
DeepSeek V4.1 Flash (max)
1.05M
Aberto
39
$0.90
193
0.76
13.69
10.35
Logo: DeepSeek
DeepSeek V4 Flash 0731 (max)
1.05M
Aberto
34
--
135
1.20
19.75
14.84
Logo: Z AI
GLM-5.2 (max) (NVFP4)
1M
Aberto
34
--
238
1.11
11.62
8.40
Logo: Alibaba
Qwen3.8 27B (xhigh) (FP8)
262k
Aberto
34
$0.60
85
1.26
30.72
23.57
Logo: MiniMax
MiniMax-M3 (MXFP8)
1M
Aberto
29
$0.46
173
1.18
15.65
11.58
Logo: Kimi
Kimi K2.6
262k
Aberto
27
$0.68
182
1.29
28.52
24.48
Logo: DeepSeek
DeepSeek V4 Flash (high) (FP8)
1.05M
Aberto
24
--
85
1.79
22.31
14.62
Logo: DeepSeek
DeepSeek V4 Flash (max) (FP8)
1.05M
Aberto
24
--
91
1.66
68.96
61.79
Logo: Kimi
Kimi K2.6 (non-reasoning) (INT4)
262k
Aberto
24*
--
168
1.22
4.20
--
Logo: Alibaba
Qwen3.5 397B A17B
262k
Aberto
18
$0.29
55
1.36
68.02
57.61
Logo: Alibaba
Qwen3.6 35B A3B
262k
Aberto
18
$0.12
83
1.04
72.06
65.00
Logo: Google
Gemma 4 26B A4B
256k
Aberto
17*
--
52
1.82
50.19
38.69
Logo: Alibaba
Qwen3.6 35B A3B (non-reasoning) (FP8)
262k
Aberto
15*
--
78
1.05
7.48
--
Logo: Google
Gemma 4 31B
262k
Aberto
15
$0.07
37
3.15
63.11
46.55
Logo: Google
Gemma 4 31B (non-reasoning)
262k
Aberto
14*
--
44
1.95
13.44
--
Logo: Google
Gemma 4 26B A4B (non-reasoning)
262k
Aberto
13*
--
69
1.70
8.92
--
Logo: Alibaba
Qwen3 235B 2507
131k
Aberto
12*
--
39
1.23
13.95
--
Logo: OpenAI
gpt-oss-120b (high)
131k
Aberto
12
$0.06
168
0.80
15.69
11.91
Logo: OpenAI
gpt-oss-120b (low)
131k
Aberto
10*
--
152
0.84
17.29
13.16
Logo: Meta
Llama 4 Maverick (FP8)
1.05M
Aberto
10*
--
76
1.28
7.90
--
Logo: Alibaba
Qwen3 VL 235B A22B (FP8)
131k
Aberto
10*
--
51
1.30
11.20
--
Logo: Alibaba
Qwen3 Next 80B A3B
262k
Aberto
10*
--
152
1.16
4.45
--
Logo: Alibaba
Qwen3 Coder Next (FP8)
262k
Aberto
9
$0.12
66
1.10
8.73
--
Logo: Meta
Llama 3.3 70B (FP8)
131k
Aberto
8*
--
74
2.58
9.30
--
Logo: Allen Institute for AI
Olmo 3.1 32B Think
65.5k
Aberto
7*
--
--
--
--
--
Logo: Allen Institute for AI
Olmo 3 7B
65.5k
Aberto
5*
--
--
--
--
--
Logo: Google
Gemma 3 27B
131k
Aberto
5
$0.10
37
2.26
15.61
--
Logo: Google
Gemma 3 4B (FP8)
131k
Aberto
5*
--
169
0.98
3.95
--

Principais definições

Perguntas frequentes

Perguntas comuns sobre Parasail