Baseten: inteligencia, rendimiento y precio de sus modelos

Baseten
Baseten

Este análisis está pensado para ayudarte a elegir el mejor modelo ofrecido por Baseten para tu caso de uso.

Más inteligente

Updated
#1
GLM-5.3 (max)GLM-5.3 (max)
45
#2
GLM-5.3 (max) (FAST)GLM-5.3 (max) (FAST)
45
#3
Kimi K3 (max)Kimi K3 (max)
44
#4
GLM-5.3-FlashGLM-5.3-Flash
42
#5
DeepSeek V4.1 Flash (max)DeepSeek V4.1 Flash (max)
39

Índice de inteligencia

23 modelos en total

Más rápido

#1
DeepSeek V4 Flash 0731 (max)DeepSeek V4 Flash 0731 (max)
376 t/s
#2
GLM-5.2 (Non-reasoning) (FAST)GLM-5.2 (Non-reasoning) (FAST)
310 t/s
#3
GLM-5.2 (max) (FAST)GLM-5.2 (max) (FAST)
284 t/s
#4
InklingInkling
272 t/s
#5
gpt-oss-120b (low)gpt-oss-120b (low)
269 t/s

Velocidad de salida

23 modelos en total

Menor precio

#1
DeepSeek V4 Flash 0731 (max)DeepSeek V4 Flash 0731 (max)
$0.07
#2
GLM-5.3-FlashGLM-5.3-Flash
$0.10
#3
gpt-oss-120b (high)gpt-oss-120b (high)
$0.14
#4
gpt-oss-120b (low)gpt-oss-120b (low)
$0.14
#5
DeepSeek V4.1 Flash (max)DeepSeek V4.1 Flash (max)
$0.20

Precio combinado (por 1M de tokens)

23 modelos en total

Baseten ofrece 23 modelos, cada uno con distintas características de inteligencia, rendimiento y precio. A continuación se comparan las métricas clave entre modelos.

  • En inteligencia, los mejores modelos en Baseten son GLM-5.3 (max) (45), GLM-5.3 (max) (FAST) (45) y Kimi K3 (max) (44).
  • En velocidad de salida, los modelos más rápidos son DeepSeek V4 Flash 0731 (max) (376 t/s), GLM-5.2 (Non-reasoning) (FAST) (310 t/s) y GLM-5.2 (max) (FAST) (284 t/s).
  • En latencia, GLM-5.2 (Non-reasoning) (FAST) (0.52s), GLM-4.7 (Non-reasoning) (0.53s) y DeepSeek V4 Pro (Non-reasoning) (0.66s) ofrecen el menor tiempo hasta el primer token de respuesta.
  • En precios, DeepSeek V4 Flash 0731 (max) ($0.07), GLM-5.3-Flash ($0.10) y gpt-oss-120b (high) ($0.14) ofrecen los menores precios combinados por 1M de tokens. Los precios varían hasta 2.8x entre modelos.
  • En tamaño de ventana de contexto, GLM-5.3 (max) (1M), Kimi K3 (max) (1M) y DeepSeek V4 Flash 0731 (max) (1M) admiten las ventanas de contexto más grandes en Baseten.
  • DeepSeek V4 Flash 0731 (max) ofrece la salida más rápida y los mejores precios, lo que lo hace atractivo para aplicaciones sensibles al rendimiento y al costo. GLM-5.3 (max) lidera en inteligencia para tareas que requieren la máxima calidad.

Aspectos destacados

Updated
Artificial Analysis Intelligence Index · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens (blended) · Lower is better

Evaluaciones de inteligencia

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
See more

Agentic knowledge work, (Elo-500)/2000

Agentic real-world work tasks, (Elo-500)/2000

Agentic SaaS workflows

Agentic coding & terminal use

Coding

Reasoning & knowledge

Professional document reasoning, All-pass

CritPtUnder review

Physics reasoning

Long context reasoning

Legal agentic work, criterion pass rate

Agentic business operations

Agentic scientific research workflows in a terminal

Quantitative analysis on spreadsheets & documents

Agentic tool use

Kubernetes incident root-cause analysis

Visual reasoning

Medical long context reasoning

Intelligence Index vs. Price

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Ventana de contexto

Context Window

Context window: tokens limit · Higher is better

Precios

Intelligence Index vs. Price

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Resumen de rendimiento

Output Speed vs. Price

Output speed: output tokens per second · USD per 1M tokens (blended)
Most attractive quadrant

Velocidad

Medida por la velocidad de salida (tokens por segundo)

Output Speed

Output tokens per second · Higher is better

Latencia

Medida por el tiempo (segundos) hasta el primer token

Latency: Time To First Answer Token

Seconds to first answer token received · Accounts for reasoning model 'thinking' time

Tiempo de respuesta de extremo a extremo

Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed

End-to-End Response Time vs. Price

End-to-end response time: end-to-end seconds to output 500 tokens · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Análisis adicional
Logo de Z AI
GLM-5.3 (max)
1.05M
Abierto
45
$1.43
142
1.44
19.09
14.12
Logo de Z AI
GLM-5.3 (max) (FAST)
1M
Abierto
45
$2.15
218
0.68
12.14
9.17
Logo de Kimi
Kimi K3 (max)
1.05M
Abierto
44
$2.42
166
0.91
16.00
12.08
Logo de Z AI
GLM-5.3-Flash
1.05M
Abierto
42
$0.60
184
0.59
14.18
10.87
Logo de DeepSeek
DeepSeek V4.1 Flash (max)
1M
Abierto
39
$1.21
242
0.48
10.80
8.25
Logo de DeepSeek
DeepSeek V4 Pro 0813 (max)
1M
Abierto
36
$5.51
157
0.89
16.79
12.72
Logo de Kimi
Kimi K3 (low)
1M
Abierto
34*
--
171
0.95
15.58
11.70
Logo de DeepSeek
DeepSeek V4 Flash 0731 (max)
1.05M
Abierto
34
$0.42
376
0.54
7.19
5.32
Logo de Z AI
GLM-5.2 (max) (FAST)
524k
Abierto
34
$1.04
284
0.64
9.44
7.04
Logo de Z AI
GLM-5.2 (max)
1.05M
Abierto
34
$0.65
136
1.95
20.34
14.72
Logo de DeepSeek
DeepSeek V4 Pro (max)
1M
Abierto
30
$2.90
154
0.53
32.21
28.43
Logo de DeepSeek
DeepSeek V4 Pro (high)
1M
Abierto
30*
--
156
0.71
16.69
12.77
Logo de Thinking Machines
Inkling Small
262k
Abierto
28*
--
250
0.59
10.59
8.00
Logo de Kimi
Kimi K2.6
262k
Abierto
27
$0.74
222
0.63
22.94
20.06
Logo de Kimi
Kimi K2.7 Code
262k
Abierto
26
$0.65
161
0.32
17.23
13.81
Logo de Thinking Machines
Inkling
262k
Abierto
25
$0.72
272
0.82
10.02
7.35
Logo de Kimi
Kimi K2.6 (Non-reasoning)
262k
Abierto
24*
--
--
--
--
--
Logo de Z AI
GLM-5.2 (Non-reasoning) (FAST)
524k
Abierto
22*
--
310
0.52
2.13
--
Logo de Z AI
GLM-5.2 (Non-reasoning)
1.05M
Abierto
22*
--
123
1.52
5.58
--
Logo de Z AI
GLM-4.7
200k
Abierto
22*
--
222
0.65
11.89
9.00
Logo de DeepSeek
DeepSeek V4 Pro (Non-reasoning)
1M
Abierto
21*
--
155
0.66
3.89
--
Logo de Z AI
GLM-4.7 (Non-reasoning)
200k
Abierto
17*
--
225
0.53
2.75
--
Logo de OpenAI
gpt-oss-120b (high)
128k
Abierto
12
$0.07
267
0.19
9.55
7.49
Logo de OpenAI
gpt-oss-120b (low)
128k
Abierto
10*
--
269
0.19
9.50
7.45

Definiciones clave

Preguntas frecuentes

Preguntas comunes sobre Baseten