Replicate: Intelligenz, Leistung und Preise der Modelle

Replicate
Replicate

Diese Analyse soll Sie bei der Auswahl des besten von Replicate angebotenen Modells für Ihren Anwendungsfall unterstützen.

Höchste Intelligenz

Updated
#1
Granite 4.0 H SmallGranite 4.0 H Small
6

Intelligence Index

Insgesamt 1 Modell

Am schnellsten

#1
Granite 4.0 H SmallGranite 4.0 H Small
14 t/s

Ausgabegeschwindigkeit

Insgesamt 1 Modell

Niedrigster Preis

#1
Granite 4.0 H SmallGranite 4.0 H Small
$0.08

Mischpreis (pro 1 Mio. Tokens)

Insgesamt 1 Modell

Replicate bietet derzeit Granite 4.0 H Small an.

Updated
Artificial Analysis Intelligence Index · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens (blended) · Lower is better

Intelligenzevaluationen

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
Mehr anzeigen

Agentic knowledge work, (Elo-500)/2000

No data available

Agentic real-world work tasks, (Elo-500)/2000

No data available

Agentic SaaS workflows

No data available

Agentic coding & terminal use

No data available
SciCodeUnder review

Coding

No data available

Reasoning & knowledge

Professional document reasoning, All-pass

No data available
CritPtUnder review

Physics reasoning

Long context reasoning

Legal agentic work, criterion pass rate

No data available

Agentic business operations

No data available

Agentic scientific research workflows in a terminal

No data available

Quantitative analysis on spreadsheets & documents

No data available

Kubernetes incident root-cause analysis

No data available

Visual reasoning

No data available

Medical long context reasoning

No data available

Intelligence Index vs. Preis

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant

Kontextfenster

Kontextfenster

Context window: tokens limit · Higher is better

Preise

Intelligence Index vs. Preis

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant

Leistungsübersicht

Ausgabegeschwindigkeit vs. Preis

Output speed: output tokens per second · USD per 1M tokens (blended)
Most attractive quadrant

Geschwindigkeit

Gemessen anhand der Ausgabegeschwindigkeit (Tokens pro Sekunde)

Ausgabegeschwindigkeit

Output tokens per second · Higher is better

Latenz

Gemessen anhand der Zeit (Sekunden) bis zum ersten Token

Latenz: Zeit bis zum ersten Token

Seconds to first token received · Lower is better

Ende-zu-Ende-Antwortzeit

Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed

Ende-zu-Ende-Antwortzeit vs. Preis

End-to-end response time: end-to-end seconds to output 500 tokens · USD per 1M tokens (blended)
Most attractive quadrant

Weitere Analyse
Logo von IBM
Granite 4.0 H Small
128k
Offen
6*
--
14
35.79
70.92
--
Logo von Meta
Llama 2 Chat 7B
4.1k
Offen
6*
--
--
--
--
--
Logo von Meta
Llama 3 70B
8.19k
Offen
5*
--
--
--
--
--
Logo von IBM
Granite 3.3 8B (non-reasoning)
128k
Offen
5*
--
--
--
--
--
Logo von Meta
Llama 3 8B
8.19k
Offen
5*
--
--
--
--
--

Wichtige Definitionen

Häufig gestellte Fragen

Häufige Fragen zu Replicate