开放权重模型比较
从质量、性能、推理速度、上下文窗口、参数量和许可证详情等关键指标比较和分析开放权重 AI 模型。
如果模型权重可供下载,我们便将其视为开放权重模型。这样用户就可以在自己的基础设施上自行托管,并通过微调等方式定制模型。
如需了解我们方法论的更多详情,请参阅常见问题。
亮点
开放性
Artificial Analysis Openness Index: Score
Openness Index assesses model openness on a 0 to 100 normalized scale (higher is more open)
Reasoning models are indicated by a lightbulb icon
开放权重模型进展
Progress in Open Weights vs. Proprietary Intelligence
Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR
Reasoning models are indicated by a lightbulb icon
各实验室开放权重语言模型智能随时间的变化
Reasoning models are indicated by a lightbulb icon
各规模开放权重模型智能随时间的变化
Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR
Reasoning models are indicated by a lightbulb icon
智能
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR
Reasoning models are indicated by a lightbulb icon
Intelligence Evaluations
Intelligence evaluations measured independently by Artificial Analysis · Higher is better
Agentic real-world work tasks, (Elo-500)/2000
Agentic tool use
Agentic coding & terminal use
Coding
Reasoning & knowledge
Scientific reasoning
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Agentic knowledge work, Elo
Agentic SaaS workflows
Legal agentic work, criterion pass rate
Agentic business operations
Instruction following
Long-horizon agentic tasks
Kubernetes incident root-cause analysis
Visual reasoning
Reasoning models are indicated by a lightbulb icon
规模
按模型规模划分的 Intelligence Index
Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR
Reasoning models are indicated by a lightbulb icon
Model Size: Total and Active Parameters
Comparison between total model parameters and parameters active during inference
Reasoning models are indicated by a lightbulb icon
Intelligence Index vs. Active Parameters
Artificial Analysis Intelligence Index · Active parameters at inference time
Most attractive quadrant
Pareto line
Reasoning models are indicated by a lightbulb icon
Intelligence Index vs. Total Parameters
Artificial Analysis Intelligence Index · Size in parameters (billions)
Most attractive quadrant
Pareto line
Reasoning models are indicated by a lightbulb icon
上下文窗口
Context Window
Context window: tokens limit · Higher is better
Reasoning models are indicated by a lightbulb icon
更多详情
权重 | 服务商基准测试 | ||||||||
|---|---|---|---|---|---|---|---|---|---|
Kimi K3 (max) | 57 | 2.8T 推理时启用 104B 个参数 | 1M | $2.3 | 37 | +8 | |||
GLM-5.2 (max) | 51 | 753B 推理时启用 40B 个参数 | 1M | $0.9 | 160 | +13 | |||
DeepSeek V4 Flash 0731 (Reasoning, Max Effort) | 50 | 284B 推理时启用 13B 个参数 | 1M | $0.1 | 110 | +3 | |||
MiniMax-M3 | 44 | 428B 推理时启用 23B 个参数 | 1M | $0.2 | 73 | +7 | |||
MiMo-V2.5-Pro | 42 | 1.0T 推理时启用 42B 个参数 | 1M | $0.2 | 67 | +2 | |||
Inkling (xhigh) | 41 | 975B 推理时启用 41B 个参数 | 1M | $0.7 | 80 | ||||
Nemotron 3 Ultra 550B A55B (Reasoning) | 38 | 550B 推理时启用 55B 个参数 | 262k | $0.5 | 145 | 不可用 | +6 | ||
Mistral Medium 3.5 | 30 | 128B | 256k | $1.2 | 134 | ||||
Gemma 4 31B (Reasoning) | 29 | 30.7B | 256k | - | 35 | +10 | |||
gpt-oss-120b (high) | 24 | 117B 推理时启用 5.1B 个参数 | 131k | $0.2 | 185 | +18 | |||
Command A+ | 23 | 218B 推理时启用 25B 个参数 | 192k | - | 209 |