大規模オープンウェイトAIモデルの比較(>150B)
パラメーター数が150Bを超えるオープンウェイトAIモデルです。
ウェイトをダウンロードできるモデルをオープンウェイト(一般にオープンソースとも呼ばれます)とみなします。独自のインフラストラクチャでセルフホストでき、ファインチューニングなどによるモデルのカスタマイズも可能です。
方法論などの詳細は、よくある質問をご覧ください。
オープン性
Artificial Analysis Openness Index:スコア
Openness Indexはモデルのオープン性を0~100の正規化尺度で評価します(高いほどオープン)
知能
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)
Intelligence Evaluations
Intelligence evaluations measured independently by Artificial Analysis · Higher is better
AA-Briefcase v1.1Updated
Agentic knowledge work, (Elo-500)/2000
GDPval-AA v2.1Updated
Agentic real-world work tasks, (Elo-500)/2000
Agentic SaaS workflows
Agentic coding & terminal use
SciCodeUnder review
Coding
Reasoning & knowledge
Professional document reasoning, All-pass
CritPtUnder review
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Legal agentic work, criterion pass rate
Agentic business operations
Agentic scientific research workflows in a terminal
Quantitative analysis on spreadsheets & documents
Kubernetes incident root-cause analysis
Visual reasoning
Medical long context reasoning
規模
モデルサイズ:総パラメータ数とアクティブパラメータ数
Comparison between total model parameters and parameters active during inference (billions)
Intelligence Indexと有効パラメーター数
Artificial Analysis Intelligence Index · Active parameters at inference time (billions)
Most attractive quadrant
Pareto line
Intelligence Index と総パラメータ数
Artificial Analysis Intelligence Index · Size in parameters (billions)
Most attractive quadrant
Pareto line
コンテキストウィンドウ
コンテキストウィンドウ
Context window: tokens limit · Higher is better
詳細
ウェイト | プロバイダーのベンチマーク | ||||||||
|---|---|---|---|---|---|---|---|---|---|
MiMo-V2.6-Pro | 46 | 1.0T 推論時に42Bが有効 | 1M | $0.2 | 47 | ||||
GLM-5.3 (Max) | 45 | 753B 推論時に40Bが有効 | 1M | $0.9 | 73 | +23 | |||
Kimi K3 (Max) | 44 | 2.8T 推論時に104Bが有効 | 1M | $2.3 | 45 | +17 | |||
GLM 5.3 Flash | 42 | 320B 推論時に18Bが有効 | 1M | $0.1 | 53 | +20 | |||
Qwen3.8 2.4T A95B | 40 | 2.4T 推論時に95Bが有効 | 262k | $1.2 | 40 | +3 | |||
Qwen3.8-Flash-Next | 40 | 180B 推論時に6Bが有効 | 256k | $0.1 | 55 | ||||
DeepSeek V4.1 Flash (Max) | 39 | 552B 推論時に16Bが有効 | 1M | $0.2 | 227 | +20 | |||
MiMo-V2.6-Flash | 38 | 309B 推論時に15Bが有効 | 1M | $0.1 | 62 | ||||
DeepSeek V4 Pro 0813 (Max) | 36 | 1.6T 推論時に49Bが有効 | 1M | $0.7 | 99 | +8 | |||
GLM-5.3 (Low) | 34 | 753B 推論時に40Bが有効 | 1M | $0.9 | 82 | ||||
Motif 3 | 34 | 314B 推論時に13.2Bが有効 | 262k | - | - | 利用不可 | - | ||
K2 Horizon 375B A23B | 31 | 375B 推論時に23Bが有効 | 524k | - | 119 |