Replicate: 모델 지능, 성능 및 가격

Replicate
Replicate

이 분석은 사용 사례에 가장 적합한 Replicate 제공 모델을 선택하는 데 도움을 드립니다.

가장 높은 지능

Updated
#1
Granite 4.0 H SmallGranite 4.0 H Small
6
#2
Granite 3.3 8B (non-reasoning)Granite 3.3 8B (non-reasoning)
5

Intelligence Index

총 모델 2개

가장 빠름

#1
Granite 3.3 8B (non-reasoning)Granite 3.3 8B (non-reasoning)
44 t/s
#2
Granite 4.0 H SmallGranite 4.0 H Small
14 t/s

출력 속도

총 모델 2개

가장 낮은 가격

#1
Granite 3.3 8B (non-reasoning)Granite 3.3 8B (non-reasoning)
$0.05
#2
Granite 4.0 H SmallGranite 4.0 H Small
$0.08

혼합 가격(토큰 100만 개당)

총 모델 2개

Replicate에서 지능, 성능, 가격 특성이 서로 다른 모델 2개를 제공합니다. 아래에서 모델별 주요 지표를 비교합니다.

  • Replicate에서 지능이 가장 높은 모델은 Granite 4.0 H Small(6) 및 Granite 3.3 8B (non-reasoning)(5)입니다.
  • 출력 속도가 가장 빠른 모델은 Granite 3.3 8B (non-reasoning)(44 t/s) 및 Granite 4.0 H Small(14 t/s)입니다.
  • 첫 답변 토큰까지 걸린 시간이 가장 짧은 모델은 Granite 3.3 8B (non-reasoning)(9.45초) 및 Granite 4.0 H Small(36.00초)입니다.
  • 토큰 100만 개당 혼합 가격이 가장 낮은 모델은 Granite 3.3 8B (non-reasoning)($0.05) 및 Granite 4.0 H Small($0.08)입니다.
  • Replicate에서 가장 큰 컨텍스트 창을 지원하는 모델은 Granite 4.0 H Small(128k) 및 Granite 3.3 8B (non-reasoning)(128k)입니다.
  • Granite 3.3 8B (non-reasoning)은 출력이 가장 빠르고 가격도 가장 낮아 처리량과 비용에 민감한 애플리케이션에 매력적입니다. 최고의 품질이 필요한 작업에서는 Granite 4.0 H Small의 지능이 가장 높습니다.
Updated
Artificial Analysis Intelligence Index · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens (blended) · Lower is better

지능 평가

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
더 보기

Agentic knowledge work, (Elo-500)/2000

No data available

Agentic real-world work tasks, (Elo-500)/2000

No data available

Agentic SaaS workflows

No data available

Agentic coding & terminal use

No data available
SciCodeUnder review

Coding

No data available

Reasoning & knowledge

Professional document reasoning, All-pass

No data available
CritPtUnder review

Physics reasoning

Long context reasoning

Legal agentic work, criterion pass rate

No data available

Agentic business operations

No data available

Agentic scientific research workflows in a terminal

No data available

Quantitative analysis on spreadsheets & documents

No data available

Kubernetes incident root-cause analysis

No data available

Visual reasoning

No data available

Medical long context reasoning

No data available

Intelligence Index vs. 가격

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

컨텍스트 창

컨텍스트 창

Context window: tokens limit · Higher is better

가격

Intelligence Index vs. 가격

Blended at 7:2:1 (cache-input-output) · USD per 1M tokens (blended)
Most attractive quadrant
Pareto line

성능 요약

출력 속도와 가격

Output speed: output tokens per second · USD per 1M tokens (blended)
Most attractive quadrant

속도

출력 속도(초당 토큰 수)로 측정

출력 속도

Output tokens per second · Higher is better

지연 시간

첫 토큰까지 걸린 시간(초)으로 측정

지연 시간: 첫 토큰까지 걸린 시간

Seconds to first token received · Lower is better

종단 간 응답 시간

Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed

종단 간 응답 시간과 가격

End-to-end response time: end-to-end seconds to output 500 tokens · USD per 1M tokens (blended)
Most attractive quadrant

추가 분석
IBM 로고
Granite 4.0 H Small
128k
오픈
6*
--
14
35.79
70.92
--
Meta 로고
Llama 2 Chat 7B
4.1k
오픈
6*
--
--
--
--
--
Meta 로고
Llama 3 70B
8.19k
오픈
5*
--
--
--
--
--
IBM 로고
Granite 3.3 8B (non-reasoning)
128k
오픈
5*
--
--
--
--
--
Meta 로고
Llama 3 8B
8.19k
오픈
5*
--
--
--
--
--

주요 용어 정의

자주 묻는 질문

Replicate에 관한 일반적인 질문