中型开放权重 AI 模型比较(40B-150B)
参数量介于 40B 和 150B 之间的开放权重 AI 模型。
如果模型权重可供下载,我们便将其视为开放权重模型。这样用户就可以在自己的基础设施上自行托管,并通过微调等方式定制模型。
如需了解包括方法论在内的更多详情,请参阅常见问题。
开放性
Artificial Analysis 开放性指数:得分
开放性指数以 0 至 100 的标准化评分衡量模型的开放程度(分数越高,越开放)
智能
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
Estimate (independent evaluation forthcoming)
Intelligence Evaluations
Intelligence evaluations measured independently by Artificial Analysis · Higher is better
AA-Briefcase v1.1Updated
Agentic knowledge work, (Elo-500)/2000
GDPval-AA v2.1Updated
Agentic real-world work tasks, (Elo-500)/2000
Agentic SaaS workflows
Agentic coding & terminal use
SciCodeUnder review
Coding
Reasoning & knowledge
Professional document reasoning, All-pass
CritPtUnder review
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Legal agentic work, Criterion Pass Rate
Agentic business operations
Agentic scientific research workflows in a terminal
No data available
Quantitative analysis on spreadsheets & documents
Kubernetes incident root-cause analysis
Visual reasoning
Medical long context reasoning
规模
模型规模:总参数量与活跃参数量
Comparison between total model parameters and parameters active during inference (billions)
Intelligence Index 与活跃参数
Artificial Analysis Intelligence Index · Active parameters at inference time (billions)
Most attractive quadrant
Pareto line
Intelligence Index 与总参数量
Artificial Analysis Intelligence Index · Size in parameters (billions)
Most attractive quadrant
Pareto line
上下文窗口
上下文窗口
Context window: tokens limit · Higher is better
更多详情
权重 | 服务商基准测试 | ||||||||
|---|---|---|---|---|---|---|---|---|---|
Ling-3.0-flash-VL | 25 | 124B 推理时启用 5.5B 个参数 | 262k | $0.0 | 145 | ||||
Ling-3.0-flash-Fin | 23 | 124B 推理时启用 5.1B 个参数 | 262k | $0.0 | 322 | ||||
Qwen3.5 122B A10B (Non-reasoning) | 18 | 125B 推理时启用 10B 个参数 | 262k | $0.7 | 143 | ||||
Qwen3.5 122B A10B (Reasoning) | 16 | 125B 推理时启用 10B 个参数 | 262k | $0.7 | 130 | +2 | |||
Mistral Medium 3.5 | 14 | 128B | 256k | $1.2 | 165 | ||||
Nemotron 3 Super 120B A12B (Reasoning) | 13 | 120.6B 推理时启用 12.7B 个参数 | 1M | $0.3 | 151 | ||||
HyperNova 60B 2605 (High, Based on gpt-oss-120b) | 12 | 58.7B 推理时启用 4.8B 个参数 | 131k | - | - | - | |||
gpt-oss-120b (High) | 12 | 117B 推理时启用 5.1B 个参数 | 131k | $0.2 | 179 | +17 | |||
K2 Think V2 | 11 | 70B | 262k | - | - | - | |||
LongCat Flash Lite | 11 | 68.5B 推理时启用 3B 个参数 | 256k | - | - | - | |||
Mistral Small 4 (Reasoning) | 11 | 119B 推理时启用 6.5B 个参数 | 256k | $0.1 | 162 | ||||
Qwen3 Next 80B A3B (Reasoning) | 11 | 80B 推理时启用 3B 个参数 | 262k | $0.3 | 177 |