大規模オープンウェイトAIモデルの比較(>150B)
パラメーター数が150Bを超えるオープンウェイトAIモデルです。
ウェイトをダウンロードできるモデルをオープンウェイト(一般にオープンソースとも呼ばれます)とみなします。独自のインフラストラクチャでセルフホストでき、ファインチューニングなどによるモデルのカスタマイズも可能です。
方法論などの詳細は、よくある質問をご覧ください。
ハイライト
オープン性
Artificial Analysis Openness Index: Score
Openness Index assesses model openness on a 0 to 100 normalized scale (higher is more open)
Reasoning models are indicated by a lightbulb icon
知能
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR
Reasoning models are indicated by a lightbulb icon
Intelligence Evaluations
Intelligence evaluations measured independently by Artificial Analysis · Higher is better
Agentic real-world work tasks, (Elo-500)/2000
Agentic tool use
Agentic coding & terminal use
Coding
Reasoning & knowledge
Scientific reasoning
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Agentic knowledge work, Elo
Agentic SaaS workflows
Legal agentic work, criterion pass rate
Agentic business operations
Instruction following
Long-horizon agentic tasks
Kubernetes incident root-cause analysis
Visual reasoning
Reasoning models are indicated by a lightbulb icon
規模
Model Size: Total and Active Parameters
Comparison between total model parameters and parameters active during inference
Reasoning models are indicated by a lightbulb icon
Intelligence Index vs. Active Parameters
Artificial Analysis Intelligence Index · Active parameters at inference time
Most attractive quadrant
Pareto line
Reasoning models are indicated by a lightbulb icon
Intelligence Index vs. Total Parameters
Artificial Analysis Intelligence Index · Size in parameters (billions)
Most attractive quadrant
Pareto line
Reasoning models are indicated by a lightbulb icon
コンテキストウィンドウ
Context Window
Context window: tokens limit · Higher is better
Reasoning models are indicated by a lightbulb icon
詳細
ウェイト | プロバイダーのベンチマーク | ||||||||
|---|---|---|---|---|---|---|---|---|---|
Kimi K3 (max) | 57 | 2.8T 推論時に104Bが有効 | 1M | $2.3 | 37 | +8 | |||
GLM-5.2 (max) | 51 | 753B 推論時に40Bが有効 | 1M | $0.9 | 150 | +13 | |||
DeepSeek V4 Flash 0731 (Reasoning, Max Effort) | 50 | 284B 推論時に13Bが有効 | 1M | $0.1 | 104 | +3 | |||
Kimi K3 (low) | 47 | 2.8T 推論時に104Bが有効 | 1M | $2.3 | 37 | ||||
MiniMax-M3 | 44 | 428B 推論時に23Bが有効 | 1M | $0.2 | 73 | +7 | |||
DeepSeek V4 Pro (Reasoning, Max Effort) | 44 | 1.6T 推論時に49Bが有効 | 1M | $0.2 | 64 | +8 | |||
DeepSeek V4 Pro (Reasoning, High Effort) | 43 | 1.6T 推論時に49Bが有効 | 1M | $0.2 | 66 | +6 | |||
MiMo-V2.5-Pro | 42 | 1.0T 推論時に42Bが有効 | 1M | $0.2 | 68 | +2 | |||
Kimi K2.7 Code | 42 | 1T 推論時に32Bが有効 | 256k | $0.7 | 43 | +6 | |||
Hy3 | 41 | 299B 推論時に21Bが有効 | 256k | $0.1 | 65 | ||||
Nex-N2-Pro | 41 | 397B 推論時に17Bが有効 | 262k | $0.5 | 129 | ||||
Inkling (xhigh) | 41 | 975B 推論時に41Bが有効 | 1M | $0.7 | 80 |