Replicate:モデルの知能、性能、料金

Replicate
Replicate

この分析は、ユースケースに最適なReplicate提供モデルを選ぶための参考情報です。

最高の知能

Updated
#1
Granite 4.0 H SmallGranite 4.0 H Small
6
#2
Granite 3.3 8B (non-reasoning)Granite 3.3 8B (non-reasoning)
5

Intelligence Index

モデル合計:2件

最速

#1
Granite 3.3 8B (non-reasoning)Granite 3.3 8B (non-reasoning)
44 t/s
#2
Granite 4.0 H SmallGranite 4.0 H Small
14 t/s

出力速度

モデル合計:2件

最安料金

#1
Granite 3.3 8B (non-reasoning)Granite 3.3 8B (non-reasoning)
$0.05
#2
Granite 4.0 H SmallGranite 4.0 H Small
$0.08

ブレンド料金(100万トークンあたり)

モデル合計:2件

Replicateは2モデルを提供しており、知能、性能、料金の特性はそれぞれ異なります。 以下でモデル間の主要指標を比較します。

  • Replicateで知能が上位のモデルはGranite 4.0 H Small(6)、Granite 3.3 8B (non-reasoning)(5)です。
  • 出力速度が最も速いモデルはGranite 3.3 8B (non-reasoning)(44 t/s)、Granite 4.0 H Small(14 t/s)です。
  • 遅延では、Granite 3.3 8B (non-reasoning)(9.45秒)、Granite 4.0 H Small(36.00秒)の最初の回答トークンまでの時間が最短です。
  • 料金では、Granite 3.3 8B (non-reasoning)($0.05)、Granite 4.0 H Small($0.08)の100万トークンあたりのブレンド料金が最安です。
  • Replicateで最大のコンテキストウィンドウに対応するモデルはGranite 4.0 H Small(128k)、Granite 3.3 8B (non-reasoning)(128k)です。
  • Granite 3.3 8B (non-reasoning)は出力が最も速く、料金も最も優れているため、スループットと費用を重視するアプリケーションに適しています。最高品質が必要なタスクでは、Granite 4.0 H Smallが知能で首位です。
Updated
Artificial Analysis Intelligence Index · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens (blended) · Lower is better

知能評価

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
もっと見る

Agentic knowledge work, (Elo-500)/2000

No data available

Agentic real-world work tasks, (Elo-500)/2000

No data available

Agentic SaaS workflows

No data available

Agentic coding & terminal use

No data available
SciCodeUnder review

Coding

No data available

Reasoning & knowledge

Professional document reasoning, All-pass

No data available
CritPtUnder review

Physics reasoning

Long context reasoning

Legal agentic work, criterion pass rate

No data available

Agentic business operations

No data available

Agentic scientific research workflows in a terminal

No data available

Quantitative analysis on spreadsheets & documents

No data available

Kubernetes incident root-cause analysis

No data available

Visual reasoning

No data available

Medical long context reasoning

No data available

Intelligence Indexと料金

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

コンテキストウィンドウ

コンテキストウィンドウ

Context window: tokens limit · Higher is better

料金

Intelligence Indexと料金

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

性能の概要

出力速度と料金

Output speed: output tokens per second · USD per 1M tokens (blended)
Most attractive quadrant

速度

出力速度(1秒あたりのトークン数)で測定

出力速度

Output tokens per second · Higher is better

遅延

最初のトークンまでの時間(秒)で測定

遅延: 最初のトークンまでの時間

Seconds to first token received · Lower is better

エンドツーエンド応答時間

Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed

エンドツーエンド応答時間と料金

End-to-end response time: end-to-end seconds to output 500 tokens · USD per 1M tokens (blended)
Most attractive quadrant

詳細分析
IBMのロゴ
Granite 4.0 H Small
128k
オープン
6*
--
14
35.79
70.92
--
Metaのロゴ
Llama 2 Chat 7B
4.1k
オープン
6*
--
--
--
--
--
Metaのロゴ
Llama 3 70B
8.19k
オープン
5*
--
--
--
--
--
IBMのロゴ
Granite 3.3 8B (non-reasoning)
128k
オープン
5*
--
--
--
--
--
Metaのロゴ
Llama 3 8B
8.19k
オープン
5*
--
--
--
--
--

主要な定義

よくある質問

Replicateに関するよくある質問