Parasail: Models Intelligence, Performance & Price

Parasail
Parasail

This analysis is intended to support you in choosing the best model provided by Parasail for your use-case.

Most Intelligent

Updated
#1
GLM-5.3 (max)GLM-5.3 (max)
45
#2
Kimi K3 (max)Kimi K3 (max)
44
#3
GLM-5.3-FlashGLM-5.3-Flash
42
#4
DeepSeek V4.1 Flash (max)DeepSeek V4.1 Flash (max)
39
#5
DeepSeek V4 Flash 0731 (max)DeepSeek V4 Flash 0731 (max)
34

Intelligence index

Total 29 models

Fastest

#1
Qwen3 Next 80B A3BQwen3 Next 80B A3B
187 t/s
#2
GLM-5.2 (max) (NVFP4)GLM-5.2 (max) (NVFP4)
178 t/s
#3
gpt-oss-120b (high)gpt-oss-120b (high)
162 t/s
#4
gpt-oss-120b (low)gpt-oss-120b (low)
148 t/s
#5
GLM-5.3 (max)GLM-5.3 (max)
148 t/s

Output speed

Total 29 models

Lowest Price

#1
Gemma 3 4B (FP8)Gemma 3 4B (FP8)
$0.05
#2
Gemma 3 27BGemma 3 27B
$0.09
#3
GLM-5.3-FlashGLM-5.3-Flash
$0.10
#4
Gemma 4 26B A4BGemma 4 26B A4B
$0.10
#5
Gemma 4 26B A4B (Non-reasoning)Gemma 4 26B A4B (Non-reasoning)
$0.10

Blended price (per 1M tokens)

Total 29 models

Parasail offers 29 models, each with different intelligence, performance, and pricing characteristics. Below is a comparison of the key metrics across models.

  • For intelligence, the top models on Parasail are GLM-5.3 (max) (45), Kimi K3 (max) (44), and GLM-5.3-Flash (42).
  • For output speed, the fastest models are Qwen3 Next 80B A3B (187 t/s), GLM-5.2 (max) (NVFP4) (178 t/s), and gpt-oss-120b (high) (162 t/s).
  • For latency, Qwen3 Next 80B A3B (0.87s), Qwen3.6 35B A3B (Non-reasoning) (FP8) (0.92s), and Llama 4 Maverick (FP8) (0.95s) offer the lowest time to first answer token.
  • For pricing, Gemma 3 4B (FP8) ($0.05), Gemma 3 27B ($0.09), and GLM-5.3-Flash ($0.10) offer the lowest blended prices per 1M tokens. Prices vary up to 2.2x across models.
  • For context window size, Kimi K3 (max) (1M), DeepSeek V4.1 Flash (max) (1M), and DeepSeek V4 Flash 0731 (max) (1M) support the largest context windows on Parasail.

Highlights

Updated
Artificial Analysis Intelligence Index · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens (blended) · Lower is better

Intelligence Evaluations

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
See more

Agentic knowledge work, (Elo-500)/2000

Agentic real-world work tasks, (Elo-500)/2000

Agentic SaaS workflows

Agentic coding & terminal use

Coding

Reasoning & knowledge

Professional document reasoning, All-pass

Physics reasoning

Long context reasoning

Legal agentic work, criterion pass rate

Agentic business operations

Quantitative analysis on spreadsheets & documents

Agentic tool use

Kubernetes incident root-cause analysis

Visual reasoning

Medical long context reasoning

Intelligence Index vs. Price

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Context Window

Context Window

Context window: tokens limit · Higher is better

Pricing

Intelligence Index vs. Price

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Performance Summary

Output Speed vs. Price

Output speed: output tokens per second · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Speed

Measured by Output Speed (tokens per second)

Output Speed

Output tokens per second · Higher is better

Latency

Measured by Time (seconds) to First Token

Latency: Time To First Answer Token

Seconds to first answer token received · Accounts for reasoning model 'thinking' time

End-to-End Response Time

Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed

End-to-End Response Time vs. Price

End-to-end response time: end-to-end seconds to output 500 tokens · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Further Analysis
Z AI logo
GLM-5.3 (max)
1M
Open
45
$1.97
148
1.00
17.94
13.55
Kimi logo
Kimi K3 (max)
1.05M
Open
44
$2.42
76
1.36
34.27
26.33
Z AI logo
GLM-5.3-Flash
1M
Open
42
$0.28
100
1.03
26.07
20.03
DeepSeek logo
DeepSeek V4.1 Flash (max)
1.05M
Open
39
$0.90
110
1.33
23.99
18.12
DeepSeek logo
DeepSeek V4 Flash 0731 (max)
1.05M
Open
34
$0.41
69
1.06
37.43
29.10
Z AI logo
GLM-5.2 (max) (NVFP4)
1M
Open
34
--
178
1.18
15.19
11.21
Alibaba logo
Qwen3.8 27B (xhigh) (FP8)
262k
Open
34
$2.29
67
1.38
38.69
29.85
MiniMax logo
MiniMax-M3 (MXFP8)
1M
Open
29
$0.46
139
1.29
19.22
14.34
Kimi logo
Kimi K2.6
262k
Open
27
$0.50
84
1.30
60.17
52.93
DeepSeek logo
DeepSeek V4 Flash (high) (FP8)
1.05M
Open
26*
--
65
1.60
28.27
19.01
DeepSeek logo
DeepSeek V4 Flash (max) (FP8)
1.05M
Open
24
--
64
1.69
97.80
88.25
Kimi logo
Kimi K2.6 (Non-reasoning) (INT4)
262k
Open
24*
--
80
1.39
7.64
--
Google logo
Gemma 4 31B
262k
Open
19*
--
80
1.60
29.58
21.72
Alibaba logo
Qwen3.5 397B A17B
262k
Open
18
$0.30
49
1.01
75.81
64.66
Alibaba logo
Qwen3.6 35B A3B
262k
Open
18
$0.12
87
0.84
68.49
61.91
Google logo
Gemma 4 26B A4B
256k
Open
17*
--
45
2.36
57.93
44.46
Alibaba logo
Qwen3.6 35B A3B (Non-reasoning) (FP8)
262k
Open
15*
--
92
0.92
6.34
--
Google logo
Gemma 4 31B (Non-reasoning)
262k
Open
14*
--
75
1.52
8.19
--
Google logo
Gemma 4 26B A4B (Non-reasoning)
262k
Open
13*
--
26
1.87
21.26
--
Alibaba logo
Qwen3 235B 2507
262k
Open
12*
--
36
1.10
14.86
--
OpenAI logo
gpt-oss-120b (high)
131k
Open
12
$0.06
162
0.95
16.39
12.35
OpenAI logo
gpt-oss-120b (low)
131k
Open
10*
--
148
0.89
17.79
13.52
Meta logo
Llama 4 Maverick (FP8)
1.05M
Open
10*
--
56
0.95
9.92
--
Alibaba logo
Qwen3 VL 235B A22B (FP8)
131k
Open
10*
--
32
1.44
16.98
--
Alibaba logo
Qwen3 Next 80B A3B
262k
Open
10*
--
187
0.87
3.55
--
Alibaba logo
Qwen3 Coder Next (FP8)
262k
Open
9
$0.12
73
1.15
7.96
--
Meta logo
Llama 3.3 70B (FP8)
131k
Open
8*
--
78
2.57
8.97
--
Allen Institute for AI logo
Olmo 3.1 32B Think
65.5k
Open
7*
--
--
--
--
--
Allen Institute for AI logo
Olmo 3 7B
65.5k
Open
5*
--
--
--
--
--
Google logo
Gemma 3 27B
131k
Open
5
$0.10
40
2.18
14.60
--
Google logo
Gemma 3 4B (FP8)
131k
Open
5*
--
80
1.30
7.52
--

Key definitions

Frequently Asked Questions

Common questions about Parasail