オープンウェイトモデルの比較
品質、性能、推論速度、コンテキストウィンドウ、パラメーター数、ライセンスの詳細などの主要指標で、オープンウェイトAIモデルを比較・分析します。
ウェイトをダウンロードできるモデルをオープンウェイト(一般にオープンソースとも呼ばれます)とみなします。独自のインフラストラクチャでセルフホストでき、ファインチューニングなどによるモデルのカスタマイズも可能です。
方法論などの詳細は、よくある質問をご覧ください。
ハイライト
オープン性
Artificial Analysis Openness Index: Score
Openness Index assesses model openness on a 0 to 100 normalized scale (higher is more open)
オープンウェイトモデルの進歩
Progress in Open Weights vs. Proprietary Intelligence
Artificial Analysis Intelligence Index v4.3 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
ラボ別に見るオープンウェイト言語モデルの知能の推移
規模別に見るオープンウェイトモデルの知能の推移
Artificial Analysis Intelligence Index v4.3 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
知能
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index v4.3 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Intelligence Evaluations
Intelligence evaluations measured independently by Artificial Analysis · Higher is better
Agentic knowledge work, (Elo-500)/2000
Agentic real-world work tasks, (Elo-500)/2000
AutomationBench-AAUpdated
Agentic SaaS workflows
Agentic coding & terminal use
Coding
Reasoning & knowledge
GDP.pdfNew
Professional document reasoning, All-pass
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Legal agentic work, criterion pass rate
Agentic business operations
Quantitative analysis on spreadsheets & documents
Instruction following
Agentic tool use
Long-horizon agentic tasks
Kubernetes incident root-cause analysis
Visual reasoning
規模
モデル規模別Intelligence Index
Artificial Analysis Intelligence Index v4.3 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Model Size: Total and Active Parameters
Comparison between total model parameters and parameters active during inference
Intelligence Index vs. Active Parameters
Artificial Analysis Intelligence Index · Active parameters at inference time
Most attractive quadrant
Pareto line
Intelligence Index vs. Total Parameters
Artificial Analysis Intelligence Index · Size in parameters (billions)
Most attractive quadrant
Pareto line
コンテキストウィンドウ
Context Window
Context window: tokens limit · Higher is better
詳細
ウェイト | プロバイダーのベンチマーク | ||||||||
|---|---|---|---|---|---|---|---|---|---|
GLM-5.3 (max) | 45 | 753B 推論時に40Bが有効 | 1M | $0.9 | 73 | +14 | |||
Kimi K3 (max) | 44 | 2.8T 推論時に104Bが有効 | 1M | $2.3 | 35 | +14 | |||
GLM-5.3-Flash | 42 | 320B 推論時に18Bが有効 | 1M | $0.1 | 114 | +16 | |||
Qwen3.8 2.4T A95B | 40 | 2.4T 推論時に95Bが有効 | 984k | $1.2 | 41 | +3 | |||
DeepSeek V4.1 Flash (Reasoning, Max Effort) | 40 | 552B 推論時に16Bが有効 | 1M | $0.2 | 222 | +2 | |||
DeepSeek V4 Pro 0813 (Reasoning, Max Effort) | 36 | 1.6T 推論時に49Bが有効 | 1M | $0.7 | 91 | +7 | |||
Qwen3.8 27B (xhigh) | 34 | 27B | 256k | $0.4 | 46 | +5 | |||
K2 Horizon 375B A23B | 31 | 375B 推論時に23Bが有効 | 524k | - | - | - | |||
MiniMax-M3 | 30 | 428B 推論時に23Bが有効 | 1M | $0.2 | 103 | +12 | |||
Inkling (xhigh) | 26 | 975B 推論時に41Bが有効 | 1M | $0.7 | 88 | +3 | |||
Nemotron 3 Ultra 550B A55B (Reasoning) | 23 | 550B 推論時に55Bが有効 | 262k | $0.5 | 204 | +6 | |||
Muse Glimmer (high) | 18 | 30B | 131k | $0.2 | 91 | ||||
Mistral Medium 3.5 | 15 | 128B | 256k | $1.2 | 147 | ||||
gpt-oss-120b (high) | 12 | 117B 推論時に5.1Bが有効 | 131k | $0.2 | 226 | +17 |