오픈 웨이트 모델 비교
오픈 웨이트 AI 모델을 품질, 성능, 추론 속도, 컨텍스트 창, 파라미터 수, 라이선스 세부 정보 등 주요 성능 지표로 비교하고 분석합니다.
웨이트를 다운로드할 수 있는 모델을 오픈 웨이트 모델(흔히 오픈 소스라고도 함)로 간주합니다. 자체 인프라에서 호스팅할 수 있으며 미세 조정 등을 통해 모델을 맞춤 설정할 수 있습니다.
방법론에 관한 자세한 내용은 FAQ에서 확인하세요.
개방성
Artificial Analysis Openness Index: 점수
Openness Index는 모델의 개방성을 0~100으로 정규화해 평가합니다(높을수록 더 개방적)
오픈 웨이트 모델의 발전
오픈 웨이트와 독점 모델의 지능 발전
Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
연구소별 오픈 웨이트 언어 모델 지능 추이
크기별 오픈 웨이트 모델 지능 추이
Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
지능
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Intelligence Evaluations
Intelligence evaluations measured independently by Artificial Analysis · Higher is better
AA-Briefcase v1.1Updated
Agentic knowledge work, (Elo-500)/2000
GDPval-AA v2.1Updated
Agentic real-world work tasks, (Elo-500)/2000
Agentic SaaS workflows
Agentic coding & terminal use
SciCodeUnder review
Coding
Reasoning & knowledge
Professional document reasoning, All-pass
CritPtUnder review
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Legal agentic work, criterion pass rate
Agentic business operations
Agentic scientific research workflows in a terminal
Quantitative analysis on spreadsheets & documents
Kubernetes incident root-cause analysis
Visual reasoning
Medical long context reasoning
크기
모델 크기별 Intelligence Index
Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
모델 크기: 총 매개변수와 활성 매개변수
Comparison between total model parameters and parameters active during inference (billions)
Intelligence Index와 활성 파라미터
Artificial Analysis Intelligence Index · Active parameters at inference time (billions)
Most attractive quadrant
Pareto line
Intelligence Index와 총 매개변수
Artificial Analysis Intelligence Index · Size in parameters (billions)
Most attractive quadrant
Pareto line
컨텍스트 창
컨텍스트 창
Context window: tokens limit · Higher is better
상세 정보
웨이트 | 제공업체 벤치마크 | ||||||||
|---|---|---|---|---|---|---|---|---|---|
MiMo-V2.6-Pro | 46 | 1.0T 추론 시 42B 활성 | 1M | $0.2 | 42 | ||||
GLM-5.3 (Max) | 45 | 753B 추론 시 40B 활성 | 1M | $0.9 | 77 | +22 | |||
Kimi K3 (Max) | 44 | 2.8T 추론 시 104B 활성 | 1M | $2.3 | 46 | +16 | |||
GLM 5.3 Flash | 42 | 320B 추론 시 18B 활성 | 1M | $0.1 | 52 | +19 | |||
DeepSeek V4.1 Flash (Max) | 39 | 552B 추론 시 16B 활성 | 1M | $0.2 | 214 | +20 | |||
Qwen3.8 27B (Xhigh) | 34 | 27B | 256k | $0.5 | 47 | +8 | |||
K2 Horizon 375B A23B | 31 | 375B 추론 시 23B 활성 | 524k | - | 120 | ||||
MiniMax-M3 | 29 | 428B 추론 시 23B 활성 | 1M | $0.2 | 98 | +12 | |||
Inkling (Xhigh) | 25 | 975B 추론 시 41B 활성 | 1M | $0.7 | 189 | +4 | |||
Nemotron 3 Ultra 550B A55B (Reasoning) | 23 | 550B 추론 시 55B 활성 | 262k | $0.5 | 172 | +5 | |||
Muse Glimmer (High) | 17 | 30B | 131k | $0.2 | 131 | ||||
Mistral Medium 3.5 | 14 | 128B | 256k | $1.2 | 168 |