| 智能 | |||
|---|---|---|---|
| Intelligence Index | 35 | 17 | |
| AA-Briefcase v1.1 | 1425 | 472 | |
| GDPval-AA v2.1 | 1548 | 791 | |
| AutomationBench-AA | 47% | 7% | |
| Terminal-Bench 4.0 | 12% | 1% | |
| SciCode | 50% | 45% | |
| Humanity's Last Exam | 34% | 22% | |
| GDP.pdf | 12% | 10% | |
| CritPt | 11% | 3% | |
| AA-Omniscience | −18 | −33 | |
| AA-LCR v1.1 | 81% | 83% | |
| 成本 | |||
| 每 100 万 Token 价格 | $0.2298 | $0.228 | |
| 每 100 万输入 Token 价格 | $0.44 | $0.325 | |
| 每 100 万输出 Token 价格 | $1.32 | $1.35 | |
| 每 100 万缓存命中 Token 价格 | $0.014 | $0.04 | |
| 每任务成本 | $0.31 | $0.06 | |
| 运行 Intelligence Index 的成本 | US$445 | US$149 | |
| Token 使用量 | |||
| 每任务输出 Token 数 | 69k | 14k | |
| 每任务推理 Token 数 | 46k | 10k | |
| 运行 Intelligence Index 的输出 Token 数 | 170M | 59M | |
| 性能 | |||
| 输出速度 | 206 token/秒 | 124 token/秒 | |
| 首 Token 延迟 | 0.95 秒 | 1.30 秒 | |
| 首个回答 Token 时间 | 10.67 秒 | 17.47 秒 | |
| 端到端响应时间 | 13.10 秒 | 21.52 秒 | |
| 每任务耗时 | 243.26 秒 | 117.48 秒 | |
| 技术规格 | |||
| 上下文窗口 | 1000k token约 1,500 页 A4 纸(12 号 Arial 字体) | 131k token约 197 页 A4 纸(12 号 Arial 字体) | |
| 发布日期 | 2026年8月 | 2026年8月 | |
| 知识截止日期 | 2026年1月 | ||
| 总参数量 | 284B | 30B | |
| 活跃参数量 | 13B | ||
| 推理 | 是 | 是 | |
| 输入模态 | 支持:文本和图像 | 支持:文本和图像 | |
| 输出模态 | 支持:文本 | 支持:文本 | |
| 开放权重 | 否 | ||
| 许可证 | |||
| 许可证允许不受限制地商用 | 是 | ||
亮点
智能Updated
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index:开放权重与专有模型
衡量模型在特定能力和行业中的表现
Artificial Analysis 财务与会计指数
Intelligence Evaluations
Agentic knowledge work, (Elo-500)/2000
Agentic real-world work tasks, (Elo-500)/2000
Agentic SaaS workflows
Agentic coding & terminal use
Coding
Reasoning & knowledge
Professional document reasoning, All-pass
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Legal agentic work, criterion pass rate
Agentic business operations
Agentic scientific research workflows in a terminal
Quantitative analysis on spreadsheets & documents
Kubernetes incident root-cause analysis
Visual reasoning
Medical long context reasoning
AA-Briefcase v1.1Updated
AA-Briefcase Elo
AA-Omniscience
AA-Omniscience Index
开放性指数
Artificial Analysis 开放性指数:得分
Intelligence Index 比较
Intelligence Index 与每项任务的成本
单任务成本(美元,对数刻度)
Token 使用量
智能指数每项任务的输出 Token 数
成本
每项 Intelligence Index 任务的成本
运行 Artificial Analysis Intelligence Index 的成本
价格:缓存命中、输入和输出
上下文窗口
上下文窗口
速度
按输出速度(每秒 token 数)衡量
输出速度
智能指数每项任务耗时
延迟
按首 Token 延迟(秒)衡量
延迟: 首个答案 Token 延迟
端到端响应时间
Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed
端到端响应时间
模型规模(仅开放权重模型)
模型规模:总参数量与活跃参数量
常见问题
DeepSeek V4 Flash Vision (Max) 更智能。在 Artificial Analysis Intelligence Index 上,DeepSeek V4 Flash Vision (Max) 的得分为 35,而 Muse Glimmer (High) 的得分为 17。
DeepSeek V4 Flash Vision (Max) 更快。DeepSeek V4 Flash Vision (Max) 每秒生成 205.8 个 token,而 Muse Glimmer (High) 每秒生成 123.7 个 token。
两个模型每 100 万 token 的成本均为 $0.23(缓存命中/输入/输出比例为 7:2:1)。
DeepSeek V4 Flash Vision (Max) 的延迟更低。DeepSeek V4 Flash Vision (Max) 的首 Token 延迟为 0.95 秒,而 Muse Glimmer (High) 为 1.30 秒。
DeepSeek V4 Flash Vision (Max) 的上下文窗口更大。DeepSeek V4 Flash Vision (Max) 支持 1.0M 个 token,而 Muse Glimmer (High) 支持 130k 个 token。