开放权重模型比较
从质量、性能、推理速度、上下文窗口、参数量和许可证详情等关键指标比较和分析开放权重 AI 模型。
如果模型权重可供下载,我们便将其视为开放权重模型。这样用户就可以在自己的基础设施上自行托管,并通过微调等方式定制模型。
如需了解我们方法论的更多详情,请参阅常见问题。
开放性
Artificial Analysis 开放性指数:得分
开放性指数以 0 至 100 的标准化评分衡量模型的开放程度(分数越高,越开放)
开放权重模型进展
开放权重与专有模型的智能水平进展
Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
各实验室开放权重语言模型智能随时间的变化
各规模开放权重模型智能随时间的变化
Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
智能
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Intelligence Evaluations
Intelligence evaluations measured independently by Artificial Analysis · Higher is better
AA-Briefcase v1.1Updated
Agentic knowledge work, (Elo-500)/2000
GDPval-AA v2.1Updated
Agentic real-world work tasks, (Elo-500)/2000
Agentic SaaS workflows
Agentic coding & terminal use
SciCodeUnder review
Coding
Reasoning & knowledge
Professional document reasoning, All-pass
CritPtUnder review
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Legal agentic work, criterion pass rate
Agentic business operations
Agentic scientific research workflows in a terminal
Quantitative analysis on spreadsheets & documents
Kubernetes incident root-cause analysis
Visual reasoning
Medical long context reasoning
规模
按模型规模划分的 Intelligence Index
Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
模型规模:总参数量与活跃参数量
Comparison between total model parameters and parameters active during inference (billions)
Intelligence Index 与活跃参数
Artificial Analysis Intelligence Index · Active parameters at inference time (billions)
Most attractive quadrant
Pareto line
Intelligence Index 与总参数量
Artificial Analysis Intelligence Index · Size in parameters (billions)
Most attractive quadrant
Pareto line
上下文窗口
上下文窗口
Context window: tokens limit · Higher is better
更多详情
权重 | 服务商基准测试 | ||||||||
|---|---|---|---|---|---|---|---|---|---|
MiMo-V2.6-Pro | 46 | 1.0T 推理时启用 42B 个参数 | 1M | $0.2 | 46 | +2 | |||
GLM-5.3 (Max) | 45 | 753B 推理时启用 40B 个参数 | 1M | $0.9 | 77 | +23 | |||
Kimi K3 (Max) | 44 | 2.8T 推理时启用 104B 个参数 | 1M | $2.3 | 44 | +18 | |||
GLM 5.3 Flash | 42 | 320B 推理时启用 18B 个参数 | 1M | $0.1 | 51 | +20 | |||
DeepSeek V4.1 Flash (Max) | 39 | 552B 推理时启用 16B 个参数 | 1M | $0.2 | 228 | +20 | |||
Qwen3.8 27B (Xhigh) | 34 | 27B | 256k | $0.5 | 47 | +8 | |||
K2 Horizon 375B A23B | 31 | 375B 推理时启用 23B 个参数 | 524k | - | 119 | ||||
MiniMax-M3 | 29 | 428B 推理时启用 23B 个参数 | 1M | $0.2 | 99 | +12 | |||
Inkling (Xhigh) | 25 | 975B 推理时启用 41B 个参数 | 1M | $0.7 | 189 | +4 | |||
Nemotron 3 Ultra 550B A55B (Reasoning) | 23 | 550B 推理时启用 55B 个参数 | 262k | $0.5 | 165 | +5 | |||
Muse Glimmer (High) | 17 | 30B | 131k | $0.2 | 107 |