Replicate: inteligência, desempenho e preço dos modelos

Replicate
Replicate

Esta análise ajuda você a escolher o melhor modelo oferecido por Replicate para seu caso de uso.

Mais inteligente

Updated
#1
Granite 4.0 H SmallGranite 4.0 H Small
6

Intelligence Index

1 modelo no total

Mais rápido

#1
Granite 4.0 H SmallGranite 4.0 H Small
14 t/s

Velocidade de saída

1 modelo no total

Menor preço

#1
Granite 4.0 H SmallGranite 4.0 H Small
$0.08

Preço combinado (por 1M de tokens)

1 modelo no total

Atualmente, Replicate oferece o Granite 4.0 H Small.

Updated
Artificial Analysis Intelligence Index · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens (blended) · Lower is better

Avaliações de inteligência

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
Ver mais

Agentic knowledge work, (Elo-500)/2000

No data available

Agentic real-world work tasks, (Elo-500)/2000

No data available

Agentic SaaS workflows

No data available

Agentic coding & terminal use

No data available
SciCodeUnder review

Coding

No data available

Reasoning & knowledge

Professional document reasoning, All-pass

No data available
CritPtUnder review

Physics reasoning

Long context reasoning

Legal agentic work, criterion pass rate

No data available

Agentic business operations

No data available

Agentic scientific research workflows in a terminal

No data available

Quantitative analysis on spreadsheets & documents

No data available

Kubernetes incident root-cause analysis

No data available

Visual reasoning

No data available

Medical long context reasoning

No data available

Intelligence Index vs. preço

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant

Janela de contexto

Janela de contexto

Context window: tokens limit · Higher is better

Preços

Intelligence Index vs. preço

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant

Resumo do desempenho

Velocidade de saída vs. preço

Output speed: output tokens per second · USD per 1M tokens (blended)
Most attractive quadrant

Velocidade

Medida pela velocidade de saída (tokens por segundo)

Velocidade de saída

Output tokens per second · Higher is better

Latência

Medida pelo tempo (segundos) até o primeiro token

Latência: Tempo até o primeiro token

Seconds to first token received · Lower is better

Tempo de resposta de ponta a ponta

Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed

Tempo de resposta de ponta a ponta vs. preço

End-to-end response time: end-to-end seconds to output 500 tokens · USD per 1M tokens (blended)
Most attractive quadrant

Análises adicionais
Logo: IBM
Granite 4.0 H Small
128k
Aberto
6*
--
14
35.79
70.92
--
Logo: Meta
Llama 2 Chat 7B
4.1k
Aberto
6*
--
--
--
--
--
Logo: Meta
Llama 3 70B
8.19k
Aberto
5*
--
--
--
--
--
Logo: IBM
Granite 3.3 8B (non-reasoning)
128k
Aberto
5*
--
--
--
--
--
Logo: Meta
Llama 3 8B
8.19k
Aberto
5*
--
--
--
--
--

Principais definições

Perguntas frequentes

Perguntas comuns sobre Replicate