Replicate:模型智能、性能与价格

Replicate
Replicate

本分析旨在帮助你根据使用场景,选择 Replicate 提供的最佳模型。

最智能

#1
Granite 4.0 H SmallGranite 4.0 H Small
6
#2
Granite 3.3 8B (non-reasoning)Granite 3.3 8B (non-reasoning)
5

Intelligence Index

共 2 个模型

速度最快

#1
Granite 3.3 8B (non-reasoning)Granite 3.3 8B (non-reasoning)
44 t/s
#2
Granite 4.0 H SmallGranite 4.0 H Small
9 t/s

输出速度

共 2 个模型

价格最低

#1
Granite 3.3 8B (non-reasoning)Granite 3.3 8B (non-reasoning)
$0.05
#2
Granite 4.0 H SmallGranite 4.0 H Small
$0.08

每 100 万 token 的混合价格

共 2 个模型

Replicate 提供 2 个模型,每个模型的智能、性能和价格特征各不相同。 下方对比了各模型的关键指标。

  • 智能方面,Replicate 上表现最好的模型是 Granite 4.0 H Small(6)和Granite 3.3 8B (non-reasoning)(5)。
  • 输出速度方面,最快的模型是 Granite 3.3 8B (non-reasoning)(44 t/s)和Granite 4.0 H Small(9 t/s)。
  • 延迟方面,Granite 3.3 8B (non-reasoning)(9.45 秒)和Granite 4.0 H Small(58.38 秒) 的首个答案 Token 延迟最低。
  • 价格方面,Granite 3.3 8B (non-reasoning)($0.05)和Granite 4.0 H Small($0.08) 每 100 万 token 的混合价格最低。
  • 上下文窗口方面,Granite 4.0 H Small(128k)和Granite 3.3 8B (non-reasoning)(128k) 支持 Replicate 上最大的上下文窗口。
  • Granite 3.3 8B (non-reasoning) 同时拥有最快输出和最佳价格,对吞吐量敏感且注重成本的应用很有吸引力。对于要求最高质量的任务,Granite 4.0 H Small 在智能方面领先。
Artificial Analysis Intelligence Index · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens (blended) · Lower is better

智能评测

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
查看更多

Agentic knowledge work, (Elo-500)/2000

No data available

Agentic real-world work tasks, (Elo-500)/2000

No data available

Agentic SaaS workflows

No data available

Agentic coding & terminal use

No data available
SciCodeUnder review

Coding

No data available

Reasoning & knowledge

Professional document reasoning, All-pass

No data available
CritPtUnder review

Physics reasoning

Long context reasoning

Legal agentic work, criterion pass rate

No data available

Agentic business operations

No data available

Agentic scientific research workflows in a terminal

No data available

Quantitative analysis on spreadsheets & documents

No data available

Kubernetes incident root-cause analysis

No data available

Visual reasoning

No data available

Medical long context reasoning

No data available

Intelligence Index 与价格

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

上下文窗口

上下文窗口

Context window: tokens limit · Higher is better

价格

Intelligence Index 与价格

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

性能摘要

输出速度与价格

Output speed: output tokens per second · USD per 1M tokens (blended)
Most attractive quadrant

速度

按输出速度(每秒 token 数)衡量

输出速度

Output tokens per second · Higher is better

延迟

按首 Token 延迟(秒)衡量

延迟: 首 Token 延迟

Seconds to first token received · Lower is better

端到端响应时间

Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed

端到端响应时间与价格

End-to-end response time: end-to-end seconds to output 500 tokens · USD per 1M tokens (blended)
Most attractive quadrant

进一步分析
IBM 标志
Granite 4.0 H Small
128k
开放
6*
--
9
58.38
115.40
--
Meta 标志
Llama 2 Chat 7B
4.1k
开放
6*
--
--
--
--
--
Meta 标志
Llama 3 70B
8.19k
开放
5*
--
--
--
--
--
IBM 标志
Granite 3.3 8B (non-reasoning)
128k
开放
5*
--
44
9.45
20.72
--
Meta 标志
Llama 3 8B
8.19k
开放
5*
--
--
--
--
--

关键定义

常见问题

关于 Replicate 的常见问题