开放权重模型比较

从质量、性能、推理速度、上下文窗口、参数量和许可证详情等关键指标比较和分析开放权重 AI 模型。

如果模型权重可供下载,我们便将其视为开放权重模型。这样用户就可以在自己的基础设施上自行托管,并通过微调等方式定制模型。

如需了解我们方法论的更多详情,请参阅常见问题。

Z AI 标志GLM-5.3 (max)Kimi 标志Kimi K3 (max) 是智能得分最高的开放权重模型,其次是 Z AI 标志GLM-5.3-FlashAlibaba 标志Qwen3.8 2.4T A95B

亮点

Artificial Analysis Openness Index · Higher is better
Updated
Artificial Analysis Intelligence Index · Higher is better
可训练参数(十亿)

开放性

Artificial Analysis Openness Index: Score

Openness Index assesses model openness on a 0 to 100 normalized scale (higher is more open)

开放权重模型进展

Progress in Open Weights vs. Proprietary Intelligence

Artificial Analysis Intelligence Index v4.3 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1

各实验室开放权重语言模型智能随时间的变化

各规模开放权重模型智能随时间的变化

Artificial Analysis Intelligence Index v4.3 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1

智能

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
See more

Agentic knowledge work, (Elo-500)/2000

Agentic real-world work tasks, (Elo-500)/2000

Agentic SaaS workflows

Agentic coding & terminal use

Coding

Reasoning & knowledge

Professional document reasoning, All-pass

Physics reasoning

Long context reasoning

Legal agentic work, criterion pass rate

Agentic business operations

Quantitative analysis on spreadsheets & documents

Instruction following

Agentic tool use

Long-horizon agentic tasks

Kubernetes incident root-cause analysis

Visual reasoning

规模

按模型规模划分的 Intelligence Index

Artificial Analysis Intelligence Index v4.3 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1

Model Size: Total and Active Parameters

Comparison between total model parameters and parameters active during inference

Intelligence Index vs. Active Parameters

Artificial Analysis Intelligence Index · Active parameters at inference time
Most attractive quadrant
Pareto line

Intelligence Index vs. Total Parameters

Artificial Analysis Intelligence Index · Size in parameters (billions)
Most attractive quadrant
Pareto line

上下文窗口

Context Window

Context window: tokens limit · Higher is better

更多详情

权重
服务商基准测试
GLM-5.3 (max)
Z AI 标志Z AI
45
753B
推理时启用 40B 个参数
1M
$0.9
72
ZaiSelf-hostedDeepInfra
+14
Kimi K3 (max)
Kimi 标志Kimi
44
2.8T
推理时启用 104B 个参数
1M
$2.3
36
ModalTogether AIDigitalOcean
+12
GLM-5.3-Flash
Z AI 标志Z AI
42
320B
推理时启用 18B 个参数
1M
$0.1
107
ZaiBasetenNovita
+16
Qwen3.8 2.4T A95B
Alibaba 标志Alibaba
40
2.4T
推理时启用 95B 个参数
984k
$1.2
40
FireworksTogether AIDeepInfra
+3
DeepSeek V4.1 Flash (Reasoning, Max Effort)
DeepSeek 标志DeepSeek
40
552B
推理时启用 16B 个参数
1M
$0.2
214
DeepSeekBaseten
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
DeepSeek 标志DeepSeek
36
1.6T
推理时启用 49B 个参数
1M
$0.7
82
DeepSeekGMISiliconFlow
+7
Qwen3.8 27B (xhigh)
Alibaba 标志Alibaba
34
27B
256k
$0.4
45
DeepInfraSelf-hostedCoreWeave
+5
K2 Horizon 375B A23B
MBZUAI Institute of Foundation Models 标志MBZUAI Institute of Foundation Models
31
375B
推理时启用 23B 个参数
524k
-
-
-
MiniMax-M3
MiniMax 标志MiniMax
30
428B
推理时启用 23B 个参数
1M
$0.2
103
ParasailCoreWeaveTogether AI
+12
Inkling (xhigh)
Thinking Machines 标志Thinking Machines
26
975B
推理时启用 41B 个参数
1M
$0.7
91
Self-hostedDeepInfraBaseten
+3
Nemotron 3 Ultra 550B A55B (Reasoning)
NVIDIA 标志NVIDIA
23
550B
推理时启用 55B 个参数
262k
$0.5
206
CoreWeaveGMIDeepInfra
+6
Muse Glimmer (high)
Meta 标志Meta
18
30B
131k
$0.2
90
Together AIDeepInfraFireworks
Mistral Medium 3.5
Mistral 标志Mistral
15
128B
256k
$1.2
148
MistralSelf-hosted
gpt-oss-120b (high)
OpenAI 标志OpenAI
12
117B
推理时启用 5.1B 个参数
131k
$0.2
228
GroqCloudflareScaleway
+17