オープンウェイトモデルの比較

品質、性能、推論速度、コンテキストウィンドウ、パラメーター数、ライセンスの詳細などの主要指標で、オープンウェイトAIモデルを比較・分析します。

ウェイトをダウンロードできるモデルをオープンウェイト(一般にオープンソースとも呼ばれます)とみなします。独自のインフラストラクチャでセルフホストでき、ファインチューニングなどによるモデルのカスタマイズも可能です。

方法論などの詳細は、よくある質問をご覧ください。

XiaomiのロゴMiMo-V2.6-ProとZ AIのロゴGLM-5.3 (max)はオープンウェイトモデルの中で知能が最も高く、KimiのロゴKimi K3 (max)とZ AIのロゴGLM-5.3-Flashが続きます。
Artificial Analysis Openness Index · Higher is better
Updated
Artificial Analysis Intelligence Index · Higher is better
学習可能なパラメーター数(十億単位)

オープン性

Artificial Analysis Openness Index:スコア

Openness Indexはモデルのオープン性を0~100の正規化尺度で評価します(高いほどオープン)

オープンウェイトモデルの進歩

オープンウェイトとプロプライエタリモデルの知能の進歩

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1

ラボ別に見るオープンウェイト言語モデルの知能の推移

規模別に見るオープンウェイトモデルの知能の推移

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1

知能

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
もっと見る

Agentic knowledge work, (Elo-500)/2000

Agentic real-world work tasks, (Elo-500)/2000

Agentic SaaS workflows

Agentic coding & terminal use

SciCodeUnder review

Coding

Reasoning & knowledge

Professional document reasoning, All-pass

CritPtUnder review

Physics reasoning

Long context reasoning

Legal agentic work, criterion pass rate

Agentic business operations

Agentic scientific research workflows in a terminal

Quantitative analysis on spreadsheets & documents

Kubernetes incident root-cause analysis

Visual reasoning

Medical long context reasoning

規模

モデル規模別Intelligence Index

Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1

モデルサイズ:総パラメータ数とアクティブパラメータ数

Comparison between total model parameters and parameters active during inference (billions)

Intelligence Indexと有効パラメーター数

Artificial Analysis Intelligence Index · Active parameters at inference time (billions)
Most attractive quadrant
Pareto line

Intelligence Index と総パラメータ数

Artificial Analysis Intelligence Index · Size in parameters (billions)
Most attractive quadrant
Pareto line

コンテキストウィンドウ

コンテキストウィンドウ

Context window: tokens limit · Higher is better

詳細

ウェイト
プロバイダーのベンチマーク
MiMo-V2.6-Pro
XiaomiのロゴXiaomi
46
1.0T
推論時に42Bが有効
1M
$0.2
42
XiaomiDeepInfraPrimaLabsNovita
GLM-5.3 (Max)
Z AIのロゴZ AI
45
753B
推論時に40Bが有効
1M
$0.9
77
ModalZaiSelf-hosted
+22
Kimi K3 (Max)
KimiのロゴKimi
44
2.8T
推論時に104Bが有効
1M
$2.3
46
DatabricksFireworksFireworks
+16
GLM 5.3 Flash
Z AIのロゴZ AI
42
320B
推論時に18Bが有効
1M
$0.1
52
ModalZaiBaseten
+19
DeepSeek V4.1 Flash (Max)
DeepSeekのロゴDeepSeek
39
552B
推論時に16Bが有効
1M
$0.2
214
NebiusSelf-hostedSelf-hosted
+20
Qwen3.8 27B (Xhigh)
AlibabaのロゴAlibaba
34
27B
256k
$0.5
47
DeepInfraSelf-hostedCoreWeave
+8
K2 Horizon 375B A23B
Institute of Foundation ModelsのロゴInstitute of Foundation Models
31
375B
推論時に23Bが有効
524k
-
120
Institute of Foundation Models
MiniMax-M3
MiniMaxのロゴMiniMax
29
428B
推論時に23Bが有効
1M
$0.2
98
ParasailCoreWeaveTogether AI
+12
Inkling (Xhigh)
Thinking MachinesのロゴThinking Machines
25
975B
推論時に41Bが有効
1M
$0.7
189
Self-hostedDeepInfraBaseten
+4
Nemotron 3 Ultra 550B A55B (Reasoning)
NVIDIAのロゴNVIDIA
23
550B
推論時に55Bが有効
262k
$0.5
172
CoreWeaveGMIDeepInfra
+5
Muse Glimmer (High)
MetaのロゴMeta
17
30B
131k
$0.2
131
Together AIDeepInfraSystalyzeFireworks
Mistral Medium 3.5
MistralのロゴMistral
14
128B
256k
$1.2
168
MistralSelf-hosted