•

Open weights model

•

Released August 2026

GLM 5.3 Flash API Provider Benchmarking & Analysis

Model Comparison

This analysis is intended to support you in choosing the best API provider of GLM 5.3 Flash for your use-case.

Fastest

#1
LithosAI (FP4, ULTRA CHAT)LithosAI (FP4, ULTRA CHAT)
730.6 t/s
#2
LithosAI (FP4)LithosAI (FP4)
532.3 t/s
#3
Inco (FAST)Inco (FAST)
462.3 t/s
#4
DatabricksDatabricks
221.3 t/s
#5
ParasailParasail
217.7 t/s

Output speed

Total 21 providers

Lowest Latency

#1
LithosAI (FP4, ULTRA CHAT)LithosAI (FP4, ULTRA CHAT)
3.23 s
#2
LithosAI (FP4)LithosAI (FP4)
5.56 s
#3
Inco (FAST)Inco (FAST)
8.26 s
#4
DatabricksDatabricks
9.73 s
#5
ParasailParasail
10.15 s

Time to first answer token

Total 21 providers

Lowest Price

#1
DeepInfraDeepInfra
$0.05
#2
Bitdeer AIBitdeer AI
$0.05
#3
ZaiZai
$0.10
#4
FireworksFireworks
$0.10
#5
BasetenBaseten
$0.10

Blended price (per 1M tokens)

Total 21 providers

GLM-5.3-Flash is available through 21 API providers, each offering different performance characteristics and pricing. Below is a comparison of the key metrics across providers.

  • For output speed, the top providers are LithosAI (FP4, ULTRA CHAT) (730.6 t/s), LithosAI (FP4) (532.3 t/s), and Inco (FAST) (462.3 t/s). Speed varies significantly across providers, with a 2833% difference between the fastest and slowest.
  • For latency, LithosAI (FP4, ULTRA CHAT) (3.23s), LithosAI (FP4) (5.56s), and Inco (FAST) (8.26s) offer the lowest time to first answer token.
  • For pricing, DeepInfra (0.05), Bitdeer AI (0.05), and Zai (0.10) offer the lowest blended prices per 1M tokens. Prices vary up to 6.0x across providers.
  • LithosAI (FP4, ULTRA CHAT) offers the best performance with both the highest speed and lowest latency. For cost optimization, DeepInfra provides the most competitive pricing.
Output tokens per second · Higher is better
Seconds to first answer token received · Lower is better
USD per 1M tokens (blended) · Lower is better

Update: Default performance benchmarking workload has updated to 10k input tokens to better reflect production use cases. You can still select different workloads above.

Pricing

Pricing: Cache Hit, Input, and Output

Price (USD per M Tokens) · Lower is better · 10,000 input tokens

Pricing: Blended Price

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended) · Lower is better

Pricing: Cache Discount

1 - (cache hit price / input price) · Higher is better

Output Speed vs. Price

Blended at 7:2:1 (cache-input-output) · Output speed: output tokens per second · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Speed

Measured by Output Speed (tokens per second)

Output Speed: GLM 5.3 Flash

Output speed: output tokens per second · 10,000 input tokens

Latency vs. Output Speed

Latency: seconds to first token received · Output speed: output tokens per second · 10,000 input tokens
Most attractive quadrant

Latency

Measured by Time (seconds) to First Token

Time to First Answer Token

Seconds to first token received · Lower is better · 10,000 input tokens

End-to-End Response Time

Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed

End-to-End Response Time

Seconds to output 500 tokens, including reasoning model 'thinking' time · Lower is better · 10,000 input tokens

Key Comparison Metrics & API Features

Modal logoModal
1M
Open
$1.03
95
0.96
27.29
21.06
Zai logoZai
1M
Open
$0.25
45
3.31
59.12
44.64
Baseten logoBaseten
1.05M
Open
$0.60
225
0.54
11.65
8.89
Novita logoNovita
1.05M
Open
$0.29
61
2.60
43.51
32.72
DeepInfra logoDeepInfra
1M
Open
$0.34
24
1.46
105.74
83.42
SiliconFlow logoSiliconFlow
1M
Open
$0.50
54
2.67
49.32
37.32
Wafer logoWafer
1M
Open
$0.45
58
1.13
44.46
34.66
Together AI logoTogether AI
1.05M
Open
$0.47
--
--
--
--
Fireworks logoFireworks
1.05M
Open
$0.37
171
1.36
16.02
11.73
Databricks logoDatabricks
1M
Open
$0.28
138
0.83
18.93
14.48
GMI logoGMI
1M
Open
$0.52
59
3.09
45.45
33.89
Parasail logoParasail
1M
Open
$0.67
193
0.93
13.91
10.38
Bitdeer AI logoBitdeer AI
1M
Open
$0.14
196
2.10
14.89
10.23
FriendliAI logoFriendliAI
1M
Open
$0.51
130
0.96
20.15
15.36
Nebius logoNebius
1.05M
Open
$1.19
204
2.72
14.95
9.78
DigitalOcean logoDigitalOcean
1.05M
Open
$0.42
44
1.13
57.81
45.35
CoreWeave logoCoreWeave
1M
Open
$0.76
128
1.70
21.20
15.60
Crusoe logoCrusoe
1.05M
Open
$0.81
125
0.91
20.93
16.01
Inco logoInco
1M
Open
$0.46
464
17.76
23.14
4.31

Frequently Asked Questions

Common questions about GLM 5.3 Flash providers