Replicate: Intelligenz, Leistung und Preise der Modelle

Replicate
Replicate

Diese Analyse soll Sie bei der Auswahl des besten von Replicate angebotenen Modells für Ihren Anwendungsfall unterstützen.

Höchste Intelligenz

Updated
#1
Granite 4.0 H SmallGranite 4.0 H Small
6
#2
Granite 3.3 8BGranite 3.3 8B
5

Intelligence Index

Insgesamt 2 Modelle

Am schnellsten

#1
Granite 3.3 8BGranite 3.3 8B
16 t/s
#2
Granite 4.0 H SmallGranite 4.0 H Small
15 t/s

Ausgabegeschwindigkeit

Insgesamt 2 Modelle

Niedrigster Preis

#1
Granite 3.3 8BGranite 3.3 8B
$0.05
#2
Granite 4.0 H SmallGranite 4.0 H Small
$0.08

Mischpreis (pro 1 Mio. Tokens)

Insgesamt 2 Modelle

Replicate bietet 2 Modelle mit jeweils unterschiedlichen Merkmalen bei Intelligenz, Leistung und Preis an. Nachfolgend werden die wichtigsten Metriken der Modelle verglichen.

  • Bei der Intelligenz sind Granite 4.0 H Small (6) und Granite 3.3 8B (5) die führenden Modelle von Replicate.
  • Bei der Ausgabegeschwindigkeit sind Granite 3.3 8B (16 t/s) und Granite 4.0 H Small (15 t/s) am schnellsten.
  • Bei der Latenz bieten Granite 3.3 8B (26.67 s) und Granite 4.0 H Small (29.59 s) die kürzeste Zeit bis zum ersten Antworttoken.
  • Beim Preis bieten Granite 3.3 8B ($0.05) und Granite 4.0 H Small ($0.08) die niedrigsten Mischpreise pro 1 Mio. Tokens.
  • Bei der Größe des Kontextfensters unterstützen Granite 4.0 H Small (128k) und Granite 3.3 8B (128k) die größten Kontextfenster von Replicate.
  • Granite 3.3 8B bietet sowohl die schnellste Ausgabe als auch die besten Preise und ist damit attraktiv für durchsatz- und kostensensible Anwendungen. Granite 4.0 H Small führt bei der Intelligenz für Aufgaben, die höchste Qualität erfordern.

Wichtigste Ergebnisse

Updated
Artificial Analysis Intelligence Index · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens (blended) · Lower is better

Intelligenzevaluationen

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
See more

Agentic knowledge work, (Elo-500)/2000

No data available

Agentic real-world work tasks, (Elo-500)/2000

No data available

Agentic SaaS workflows

No data available

Agentic coding & terminal use

No data available

Coding

No data available

Reasoning & knowledge

Professional document reasoning, All-pass

No data available

Physics reasoning

Long context reasoning

Legal agentic work, criterion pass rate

No data available

Agentic business operations

No data available

Quantitative analysis on spreadsheets & documents

No data available

Instruction following

Agentic tool use

No data available

Long-horizon agentic tasks

No data available

Kubernetes incident root-cause analysis

No data available

Visual reasoning

No data available

Intelligence Index vs. Price

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Kontextfenster

Context Window

Context window: tokens limit · Higher is better

Preise

Intelligence Index vs. Price

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

Leistungsübersicht

Output Speed vs. Price

Output speed: output tokens per second · USD per 1M tokens (blended)
Most attractive quadrant

Geschwindigkeit

Gemessen anhand der Ausgabegeschwindigkeit (Tokens pro Sekunde)

Output Speed

Output tokens per second · Higher is better

Latenz

Gemessen anhand der Zeit (Sekunden) bis zum ersten Token

Latency: Time To First Token

Seconds to first token received · Lower is better

Ende-zu-Ende-Antwortzeit

Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed

End-to-End Response Time vs. Price

End-to-end response time: end-to-end seconds to output 500 tokens · USD per 1M tokens (blended)
Most attractive quadrant

Weitere Analyse
Logo von DeepSeek
DeepSeek V3 0324
128k
Offen
10
$0.18
--
--
--
--
Logo von IBM
Granite 4.0 H Small
128k
Offen
6*
--
15
29.59
63.25
--
Logo von Meta
Llama 2 Chat 7B
4.1k
Offen
6*
--
--
--
--
--
Logo von Meta
Llama 3 70B
8.19k
Offen
5*
--
--
--
--
--
Logo von IBM
Granite 3.3 8B
128k
Offen
5*
--
16
26.67
58.45
--
Logo von Meta
Llama 3 8B
8.19k
Offen
5*
--
--
--
--
--

Wichtige Definitionen

Häufig gestellte Fragen

Häufige Fragen zu Replicate