Nebius: inteligencia, rendimiento y precio de sus modelos

Nebius
Nebius

Este análisis está pensado para ayudarte a elegir el mejor modelo ofrecido por Nebius para tu caso de uso.

Más inteligente

Updated
#1
GLM-5.3 (max) (FP4)GLM-5.3 (max) (FP4)
45
#2
Kimi K3 (max)Kimi K3 (max)
44
#3
GLM-5.3-FlashGLM-5.3-Flash
42
#4
DeepSeek V4.1 Flash (max)DeepSeek V4.1 Flash (max)
39
#5
DeepSeek V4 Pro 0813 (max)DeepSeek V4 Pro 0813 (max)
36

Índice de inteligencia

31 modelos en total

Más rápido

#1
GLM-5.3 (max) (FP4)GLM-5.3 (max) (FP4)
387 t/s
#2
Nemotron 3 UltraNemotron 3 Ultra
327 t/s
#3
MiniMax-M3 (FP8)MiniMax-M3 (FP8)
301 t/s
#4
Nemotron 3.5 Lightning (BF16)Nemotron 3.5 Lightning (BF16)
273 t/s
#5
GLM-5.3-FlashGLM-5.3-Flash
273 t/s

Velocidad de salida

31 modelos en total

Menor precio

#1
Nemotron 3.5 Lightning (BF16)Nemotron 3.5 Lightning (BF16)
$0.08
#2
Nemotron 3 NanoNemotron 3 Nano
$0.08
#3
Qwen3 30B A3B 2507 (Non-reasoning)Qwen3 30B A3B 2507 (Non-reasoning)
$0.12
#4
Gemma 3 27B (FP8)Gemma 3 27B (FP8)
$0.12
#5
DeepSeek V4 Flash 0731 (max)DeepSeek V4 Flash 0731 (max)
$0.15

Precio combinado (por 1M de tokens)

31 modelos en total

Nebius ofrece 31 modelos, cada uno con distintas características de inteligencia, rendimiento y precio. A continuación se comparan las métricas clave entre modelos.

  • En inteligencia, los mejores modelos en Nebius son GLM-5.3 (max) (FP4) (45), Kimi K3 (max) (44) y GLM-5.3-Flash (42).
  • En velocidad de salida, los modelos más rápidos son GLM-5.3 (max) (FP4) (387 t/s), Nemotron 3 Ultra (327 t/s) y MiniMax-M3 (FP8) (301 t/s).
  • En latencia, GLM-5.2 (Non-reasoning) (FP4) (1.00s), Qwen3 30B A3B 2507 (Non-reasoning) (1.05s) y DeepSeek V4 Pro (Non-reasoning) (1.29s) ofrecen el menor tiempo hasta el primer token de respuesta.
  • En precios, Nemotron 3.5 Lightning (BF16) ($0.08), Nemotron 3 Nano ($0.08) y Qwen3 30B A3B 2507 (Non-reasoning) ($0.12) ofrecen los menores precios combinados por 1M de tokens.
  • En tamaño de ventana de contexto, MiniMax-M3 (FP8) (1M), Kimi K3 (max) (1M) y GLM-5.3-Flash (1M) admiten las ventanas de contexto más grandes en Nebius.
  • GLM-5.3 (max) (FP4) ofrece la mejor combinación de inteligencia y velocidad. Para optimizar costos, Nemotron 3.5 Lightning (BF16) ofrece los precios más competitivos.

Aspectos destacados

Updated
Artificial Analysis Intelligence Index · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens (blended) · Lower is better

Evaluaciones de inteligencia

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
See more

Agentic knowledge work, (Elo-500)/2000

Agentic real-world work tasks, (Elo-500)/2000

Agentic SaaS workflows

Agentic coding & terminal use

Coding

Reasoning & knowledge

Professional document reasoning, All-pass

CritPtUnder review

Physics reasoning

Long context reasoning

Legal agentic work, criterion pass rate

Agentic business operations

Agentic scientific research workflows in a terminal

Quantitative analysis on spreadsheets & documents

Agentic tool use

Kubernetes incident root-cause analysis

Visual reasoning

Medical long context reasoning

Intelligence Index vs. Price

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Ventana de contexto

Context Window

Context window: tokens limit · Higher is better

Precios

Intelligence Index vs. Price

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Resumen de rendimiento

Output Speed vs. Price

Output speed: output tokens per second · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Velocidad

Medida por la velocidad de salida (tokens por segundo)

Output Speed

Output tokens per second · Higher is better

Latencia

Medida por el tiempo (segundos) hasta el primer token

Latency: Time To First Answer Token

Seconds to first answer token received · Accounts for reasoning model 'thinking' time

Tiempo de respuesta de extremo a extremo

Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed

End-to-End Response Time vs. Price

End-to-end response time: end-to-end seconds to output 500 tokens · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Análisis adicional
Logo de Z AI
GLM-5.3 (max) (FP4)
1M
Abierto
45
$8.68
389
1.80
8.22
5.14
Logo de Kimi
Kimi K3 (max)
1.05M
Abierto
44
$9.22
241
1.47
11.82
8.28
Logo de Z AI
GLM-5.3-Flash
1.05M
Abierto
42
$1.19
274
1.79
10.91
7.30
Logo de DeepSeek
DeepSeek V4.1 Flash (max)
1.05M
Abierto
39
$3.13
158
1.71
17.53
12.65
Logo de DeepSeek
DeepSeek V4 Pro 0813 (max)
1.05M
Abierto
36
$8.43
203
2.33
14.64
9.85
Logo de Kimi
Kimi K3 (low)
1M
Abierto
34*
--
205
1.46
13.62
9.73
Logo de DeepSeek
DeepSeek V4 Flash 0731 (max)
1.05M
Abierto
34
$0.65
71
2.03
37.39
28.29
Logo de Z AI
GLM-5.2 (max) (FP4)
432k
Abierto
34
$2.94
262
1.02
10.55
7.63
Logo de DeepSeek
DeepSeek V4 Pro (max)
1M
Abierto
30
$9.70
171
1.28
29.84
25.63
Logo de DeepSeek
DeepSeek V4 Pro (high)
1M
Abierto
30*
--
161
1.30
16.75
12.35
Logo de MiniMax
MiniMax-M3 (FP8)
1.05M
Abierto
29
$1.91
301
1.75
10.05
6.64
Logo de Kimi
Kimi K2.6
262k
Abierto
27
$1.91
274
1.76
19.84
16.26
Logo de Z AI
GLM-5.1 (FP8, Base)
200k
Abierto
26
$2.62
41
1.78
106.05
92.13
Logo de Kimi
Kimi K2.7 Code (FP4)
256k
Abierto
26
$1.89
257
1.69
12.29
8.66
Logo de Z AI
GLM-5.1 (Non-reasoning) (FP8, Base)
200k
Abierto
24*
--
42
1.72
13.76
--
Logo de Kimi
Kimi K2.6 (Non-reasoning)
262k
Abierto
24*
--
251
1.75
3.75
--
Logo de NVIDIA
Nemotron 3 Ultra
256k
Abierto
23
$1.16
327
2.22
10.71
6.96
Logo de Z AI
GLM-5.2 (Non-reasoning) (FP4)
432k
Abierto
22*
--
212
1.00
3.35
--
Logo de Alibaba
Qwen3.5 397B A17B (Non-reasoning) (Base, FP4)
262k
Abierto
21*
--
159
1.93
5.07
--
Logo de DeepSeek
DeepSeek V4 Pro (Non-reasoning)
1.05M
Abierto
21*
--
169
1.29
4.25
--
Logo de Alibaba
Qwen3.5 397B A17B (Base, FP4)
262k
Abierto
18
$0.47
153
1.96
26.12
20.89
Logo de NVIDIA
Nemotron 3.5 Lightning (BF16)
1.05M
Abierto
13
$0.09
275
0.99
10.08
7.27
Logo de NVIDIA
Nemotron 3 Super
256k
Abierto
13
$1.64
197
1.59
14.28
10.15
Logo de Alibaba
Qwen3 235B 2507
262k
Abierto
12*
--
55
1.39
10.55
--
Logo de OpenAI
gpt-oss-120b (high) (Base)
128k
Abierto
12
$0.11
252
0.94
10.85
7.93
Logo de OpenAI
gpt-oss-120b (low) Base
128k
Abierto
10*
--
257
0.99
10.70
7.77
Logo de NVIDIA
Nemotron 3 Nano
262k
Abierto
9
$0.02
116
1.04
22.66
17.30
Logo de Alibaba
Qwen3 30B A3B 2507 (Non-reasoning)
262k
Abierto
8*
--
42
1.05
12.99
--
Logo de Nous Research
Hermes 4 405B (FP8)
128k
Abierto
7*
--
43
2.32
59.82
45.99
Logo de Nous Research
Hermes 4 405B (Non-reasoning) (FP8)
128k
Abierto
7*
--
44
2.32
13.80
--
Logo de Google
Gemma 3 27B (FP8)
110k
Abierto
5
$0.20
46
2.80
13.70
--

Definiciones clave

Preguntas frecuentes

Preguntas comunes sobre Nebius