Replicate: Models Intelligence, Performance & Price

Replicate
Replicate

This analysis is intended to support you in choosing the best model provided by Replicate for your use-case.

Most Intelligent

Updated
#1
Granite 4.0 H SmallGranite 4.0 H Small
6

Intelligence index

Total 1 model

Fastest

#1
Granite 4.0 H SmallGranite 4.0 H Small
15 t/s

Output speed

Total 1 model

Lowest Price

#1
Granite 4.0 H SmallGranite 4.0 H Small
$0.08

Blended price (per 1M tokens)

Total 1 model

Replicate currently offers Granite 4.0 H Small.

Highlights

Updated
Artificial Analysis Intelligence Index · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens (blended) · Lower is better

Intelligence Evaluations

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
See more

Agentic knowledge work, (Elo-500)/2000

No data available

Agentic real-world work tasks, (Elo-500)/2000

No data available

Agentic SaaS workflows

No data available

Agentic coding & terminal use

No data available

Coding

No data available

Reasoning & knowledge

Professional document reasoning, All-pass

No data available
CritPtUnder review

Physics reasoning

Long context reasoning

Legal agentic work, criterion pass rate

No data available

Agentic business operations

No data available

Agentic scientific research workflows in a terminal

No data available

Quantitative analysis on spreadsheets & documents

No data available

Agentic tool use

No data available

Kubernetes incident root-cause analysis

No data available

Visual reasoning

No data available

Medical long context reasoning

No data available

Intelligence Index vs. Price

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant

Context Window

Context Window

Context window: tokens limit · Higher is better

Pricing

Intelligence Index vs. Price

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant

Performance Summary

Output Speed vs. Price

Output speed: output tokens per second · USD per 1M tokens (blended)
Most attractive quadrant

Speed

Measured by Output Speed (tokens per second)

Output Speed

Output tokens per second · Higher is better

Latency

Measured by Time (seconds) to First Token

Latency: Time To First Token

Seconds to first token received · Lower is better

End-to-End Response Time

Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed

End-to-End Response Time vs. Price

End-to-end response time: end-to-end seconds to output 500 tokens · USD per 1M tokens (blended)
Most attractive quadrant

Further Analysis
IBM logo
Granite 4.0 H Small
128k
Open
6*
--
15
34.42
68.16
--
Meta logo
Llama 2 Chat 7B
4.1k
Open
6*
--
--
--
--
--
Meta logo
Llama 3 70B
8.19k
Open
5*
--
--
--
--
--
IBM logo
Granite 3.3 8B
128k
Open
5*
--
--
--
--
--
Meta logo
Llama 3 8B
8.19k
Open
5*
--
--
--
--
--

Key definitions

Frequently Asked Questions

Common questions about Replicate