Replicate: inteligencia, rendimiento y precio de sus modelos

Replicate
Replicate

Este análisis está pensado para ayudarte a elegir el mejor modelo ofrecido por Replicate para tu caso de uso.

Más inteligente

Updated
#1
Granite 4.0 H SmallGranite 4.0 H Small
6

Índice de inteligencia

1 modelo en total

Más rápido

#1
Granite 4.0 H SmallGranite 4.0 H Small
15 t/s

Velocidad de salida

1 modelo en total

Menor precio

#1
Granite 4.0 H SmallGranite 4.0 H Small
$0.08

Precio combinado (por 1M de tokens)

1 modelo en total

Replicate actualmente ofrece Granite 4.0 H Small.

Aspectos destacados

Updated
Artificial Analysis Intelligence Index · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens (blended) · Lower is better

Evaluaciones de inteligencia

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
See more

Agentic knowledge work, (Elo-500)/2000

No data available

Agentic real-world work tasks, (Elo-500)/2000

No data available

Agentic SaaS workflows

No data available

Agentic coding & terminal use

No data available

Coding

No data available

Reasoning & knowledge

Professional document reasoning, All-pass

No data available
CritPtUnder review

Physics reasoning

Long context reasoning

Legal agentic work, criterion pass rate

No data available

Agentic business operations

No data available

Agentic scientific research workflows in a terminal

No data available

Quantitative analysis on spreadsheets & documents

No data available

Agentic tool use

No data available

Kubernetes incident root-cause analysis

No data available

Visual reasoning

No data available

Medical long context reasoning

No data available

Intelligence Index vs. Price

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant

Ventana de contexto

Context Window

Context window: tokens limit · Higher is better

Precios

Intelligence Index vs. Price

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant

Resumen de rendimiento

Output Speed vs. Price

Output speed: output tokens per second · USD per 1M tokens (blended)
Most attractive quadrant

Velocidad

Medida por la velocidad de salida (tokens por segundo)

Output Speed

Output tokens per second · Higher is better

Latencia

Medida por el tiempo (segundos) hasta el primer token

Latency: Time To First Token

Seconds to first token received · Lower is better

Tiempo de respuesta de extremo a extremo

Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed

End-to-End Response Time vs. Price

End-to-end response time: end-to-end seconds to output 500 tokens · USD per 1M tokens (blended)
Most attractive quadrant

Análisis adicional
Logo de DeepSeek
DeepSeek V3 0324
128k
Abierto
10
$0.18
--
--
--
--
Logo de IBM
Granite 4.0 H Small
128k
Abierto
6*
--
15
35.15
69.23
--
Logo de Meta
Llama 2 Chat 7B
4.1k
Abierto
6*
--
--
--
--
--
Logo de Meta
Llama 3 70B
8.19k
Abierto
5*
--
--
--
--
--
Logo de IBM
Granite 3.3 8B
128k
Abierto
5*
--
15
26.88
60.26
--
Logo de Meta
Llama 3 8B
8.19k
Abierto
5*
--
--
--
--
--

Definiciones clave

Preguntas frecuentes

Preguntas comunes sobre Replicate