오픈 웨이트 모델 비교
오픈 웨이트 AI 모델을 품질, 성능, 추론 속도, 컨텍스트 창, 파라미터 수, 라이선스 세부 정보 등 주요 성능 지표로 비교하고 분석합니다.
웨이트를 다운로드할 수 있는 모델을 오픈 웨이트 모델(흔히 오픈 소스라고도 함)로 간주합니다. 자체 인프라에서 호스팅할 수 있으며 미세 조정 등을 통해 모델을 맞춤 설정할 수 있습니다.
방법론에 관한 자세한 내용은 FAQ에서 확인하세요.
주요 내용
개방성
Artificial Analysis Openness Index: Score
Openness Index assesses model openness on a 0 to 100 normalized scale (higher is more open)
Reasoning models are indicated by a lightbulb icon
오픈 웨이트 모델의 발전
Progress in Open Weights vs. Proprietary Intelligence
Artificial Analysis Intelligence Index v4.1.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR
Reasoning models are indicated by a lightbulb icon
연구소별 오픈 웨이트 언어 모델 지능 추이
Reasoning models are indicated by a lightbulb icon
크기별 오픈 웨이트 모델 지능 추이
Artificial Analysis Intelligence Index v4.1.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR
Reasoning models are indicated by a lightbulb icon
지능
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index v4.1.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR
Reasoning models are indicated by a lightbulb icon
Intelligence Evaluations
Intelligence evaluations measured independently by Artificial Analysis · Higher is better
Agentic real-world work tasks, (Elo-500)/2000
𝜏³-BankingUpdated
Agentic tool use
Agentic coding & terminal use
Coding
Humanity's Last ExamUpdated
Reasoning & knowledge
Scientific reasoning
Physics reasoning
AA-Omniscience AccuracyUpdated
Knowledge
1 - hallucination rate
AA-LCRUpdated
Long context reasoning
Agentic knowledge work, Elo
Agentic SaaS workflows
Legal agentic work, criterion pass rate
Agentic business operations
Quantitative analysis on spreadsheets & documents
Instruction following
Long-horizon agentic tasks
Kubernetes incident root-cause analysis
Visual reasoning
Reasoning models are indicated by a lightbulb icon
크기
모델 크기별 Intelligence Index
Artificial Analysis Intelligence Index v4.1.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR
Reasoning models are indicated by a lightbulb icon
Model Size: Total and Active Parameters
Comparison between total model parameters and parameters active during inference
Reasoning models are indicated by a lightbulb icon
Intelligence Index vs. Active Parameters
Artificial Analysis Intelligence Index · Active parameters at inference time
Most attractive quadrant
Pareto line
Reasoning models are indicated by a lightbulb icon
Intelligence Index vs. Total Parameters
Artificial Analysis Intelligence Index · Size in parameters (billions)
Most attractive quadrant
Pareto line
Reasoning models are indicated by a lightbulb icon
컨텍스트 창
Context Window
Context window: tokens limit · Higher is better
Reasoning models are indicated by a lightbulb icon
상세 정보
웨이트 | 제공업체 벤치마크 | ||||||||
|---|---|---|---|---|---|---|---|---|---|
Kimi K3 (max) | 60 | 2.8T 추론 시 104B 활성 | 1M | $2.3 | 38 | +12 | |||
Qwen3.8 2.4T A95B | 58 | 2.4T 추론 시 95B 활성 | 984k | $1.2 | 24 | +3 | |||
DeepSeek V4 Pro 0813 (Reasoning, Max Effort) | 53 | 1.6T 추론 시 49B 활성 | 1M | $0.7 | 69 | +5 | |||
Qwen3.8 27B (xhigh) | 52 | 27B | 256k | $0.4 | 54 | +2 | |||
Motif 3 | 47 | 314B 추론 시 13.2B 활성 | 262k | - | - | 제공되지 않음 | - | ||
MiniMax-M3 | 45 | 428B 추론 시 23B 활성 | 1M | $0.2 | 137 | +9 | |||
Inkling (xhigh) | 42 | 975B 추론 시 41B 활성 | 1M | $0.7 | 50 | +2 | |||
Nemotron 3 Ultra 550B A55B (Reasoning) | 38 | 550B 추론 시 55B 활성 | 262k | $0.5 | 188 | +6 | |||
Solar Open2 250B | 37 | 250B 추론 시 15B 활성 | 1M | - | - | - | |||
Muse Glimmer (high) | 35 | 30B | 131k | $0.2 | 109 | ||||
A.X-K2 | 35 | 692B 추론 시 33B 활성 | 262k | - | - | - | |||
K-EXAONE 2.0 0803 | 31 | 750B 추론 시 37B 활성 | 262k | - | - | 제공되지 않음 | - | ||
Mistral Medium 3.5 | 30 | 128B | 256k | $1.2 | 145 | ||||
Nemotron 3 Super 120B A12B (Reasoning) | 26 | 120.6B 추론 시 12.7B 활성 | 1M | $0.3 | 142 | ||||
gpt-oss-120b (high) | 24 | 117B 추론 시 5.1B 활성 | 131k | $0.2 | 175 | +16 | |||
Nemotron 3.5 Lightning | 24 | 31.6B 추론 시 3.6B 활성 | 1M | $0.1 | 309 | 제공되지 않음 | +3 | ||
Command A+ | 23 | 218B 추론 시 25B 활성 | 192k | - | 264 |