中規模オープンウェイトAIモデルの比較(40B~150B)
パラメーター数が40B~150BのオープンウェイトAIモデルです。
ウェイトをダウンロードできるモデルをオープンウェイト(一般にオープンソースとも呼ばれます)とみなします。独自のインフラストラクチャでセルフホストでき、ファインチューニングなどによるモデルのカスタマイズも可能です。
方法論などの詳細は、よくある質問をご覧ください。
ハイライト
オープン性
Artificial Analysis Openness Index: Score
Openness Index assesses model openness on a 0 to 100 normalized scale (higher is more open)
知能
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index v4.3 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)
Intelligence Evaluations
Intelligence evaluations measured independently by Artificial Analysis · Higher is better
Agentic knowledge work, (Elo-500)/2000
Agentic real-world work tasks, (Elo-500)/2000
AutomationBench-AAUpdated
Agentic SaaS workflows
Agentic coding & terminal use
Coding
Reasoning & knowledge
GDP.pdfNew
Professional document reasoning, All-pass
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Legal agentic work, criterion pass rate
Agentic business operations
Quantitative analysis on spreadsheets & documents
Instruction following
Agentic tool use
Long-horizon agentic tasks
Kubernetes incident root-cause analysis
Visual reasoning
規模
Model Size: Total and Active Parameters
Comparison between total model parameters and parameters active during inference
Intelligence Index vs. Active Parameters
Artificial Analysis Intelligence Index · Active parameters at inference time
Most attractive quadrant
Pareto line
Intelligence Index vs. Total Parameters
Artificial Analysis Intelligence Index · Size in parameters (billions)
Most attractive quadrant
Pareto line
コンテキストウィンドウ
Context Window
Context window: tokens limit · Higher is better
詳細
ウェイト | プロバイダーのベンチマーク | ||||||||
|---|---|---|---|---|---|---|---|---|---|
Ling-3.0-flash-VL | 25 | 124B 推論時に5.5Bが有効 | 262k | - | 145 | ||||
Ling 3.0 Flash | 21 | 124B 推論時に5.1Bが有効 | 262k | $0.0 | 329 | 利用不可 | |||
Qwen3.5 122B A10B (Non-reasoning) | 18 | 125B 推論時に10Bが有効 | 262k | $0.7 | 142 | ||||
Qwen3.5 122B A10B (Reasoning) | 16 | 125B 推論時に10Bが有効 | 262k | $0.7 | 128 | +2 | |||
Mistral Medium 3.5 | 15 | 128B | 256k | $1.2 | 147 | ||||
Nemotron 3 Super 120B A12B (Reasoning) | 14 | 120.6B 推論時に12.7Bが有効 | 1M | $0.2 | 172 | ||||
gpt-oss-120b (high) | 12 | 117B 推論時に5.1Bが有効 | 131k | $0.2 | 226 | +17 | |||
HyperNova 60B 2605 (high, based on gpt-oss-120b) | 12 | 58.7B 推論時に4.8Bが有効 | 131k | $0.1 | 349 | ||||
K2 Think V2 | 11 | 70B | 262k | - | - | - | |||
LongCat Flash Lite | 11 | 68.5B 推論時に3Bが有効 | 256k | - | - | - | |||
Mistral Small 4 (Reasoning) | 11 | 119B 推論時に6.5Bが有効 | 256k | $0.2 | 168 | ||||
Qwen3 Next 80B A3B (Reasoning) | 11 | 80B 推論時に3Bが有効 | 262k | $0.3 | 197 |