| 智能 | |||
|---|---|---|---|
| Intelligence Index | 51 | 31 | |
| AA-Briefcase v1.1 | 1642 | 1295 | |
| GDPval-AA v2.1 | 1576 | 1349 | |
| AutomationBench-AA | 61% | 37% | |
| Terminal-Bench 4.0 | 53% | 2% | |
| SciCode | 59% | 43% | |
| Humanity's Last Exam | 55% | 32% | |
| GDP.pdf | 26% | 7% | |
| CritPt | 28% | 5% | |
| AA-Omniscience | 40 | −3 | |
| AA-LCR v1.1 | 84% | 80% | |
| 成本 | |||
| 每 100 万 Token 价格 | $2.94 | $0.00 | |
| 每 100 万输入 Token 价格 | $4.00 | $0.00 | |
| 每 100 万输出 Token 价格 | $20.00 | $0.00 | |
| 每 100 万缓存命中 Token 价格 | $0.20 | ||
| 每任务成本 | $1.34 | $0.00 | |
| 运行 Intelligence Index 的成本 | US$1,627 | US$0 | |
| Token 使用量 | |||
| 每任务输出 Token 数 | 26k | 52k | |
| 每任务推理 Token 数 | 12k | 36k | |
| 运行 Intelligence Index 的输出 Token 数 | 38M | 126M | |
| 性能 | |||
| 输出速度 | 73 token/秒 | 123 token/秒 | |
| 首 Token 延迟 | 25.18 秒 | 21.89 秒 | |
| 首个回答 Token 时间 | 25.18 秒 | 38.17 秒 | |
| 端到端响应时间 | 32.01 秒 | 42.24 秒 | |
| 每任务耗时 | 219.45 秒 | 445.76 秒 | |
| 技术规格 | |||
| 上下文窗口 | 1000k token约 1,500 页 A4 纸(12 号 Arial 字体) | 524k token约 786 页 A4 纸(12 号 Arial 字体) | |
| 发布日期 | 2026年9月 | 2026年9月 | |
| 总参数量 | 375B | ||
| 活跃参数量 | 23B | ||
| 推理 | 是 | 是 | |
| 输入模态 | 支持:文本和图像 | 支持:文本 | |
| 输出模态 | 支持:文本 | 支持:文本 | |
| 开放权重 | 否 | ||
| 许可证 | |||
| 许可证允许不受限制地商用 | 是 | ||
亮点
智能Updated
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index:开放权重与专有模型
衡量模型在特定能力和行业中的表现
Artificial Analysis 财务与会计指数
Intelligence Evaluations
Agentic knowledge work, (Elo-500)/2000
Agentic real-world work tasks, (Elo-500)/2000
Agentic SaaS workflows
Agentic coding & terminal use
Coding
Reasoning & knowledge
Professional document reasoning, All-pass
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Legal agentic work, criterion pass rate
Agentic business operations
Agentic scientific research workflows in a terminal
Quantitative analysis on spreadsheets & documents
Kubernetes incident root-cause analysis
Visual reasoning
Medical long context reasoning
AA-Briefcase v1.1Updated
AA-Briefcase Elo
AA-Omniscience
AA-Omniscience Index
开放性指数
Artificial Analysis 开放性指数:得分
Intelligence Index 比较
Intelligence Index 与每项任务的成本
单任务成本(美元,对数刻度)
Token 使用量
Output Tokens per Intelligence Index Task
成本
每项 Intelligence Index 任务的成本
运行 Artificial Analysis Intelligence Index 的成本
Pricing: Cache Hit, Input, and Output
上下文窗口
上下文窗口
速度
按输出速度(每秒 token 数)衡量
Output Speed
Time per Intelligence Index Task
延迟
按首 Token 延迟(秒)衡量
Latency: Time To First Answer Token
端到端响应时间
Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed
End-to-End Response Time
模型规模(仅开放权重模型)
模型规模:总参数量与活跃参数量
常见问题
Claude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) 更智能。在 Artificial Analysis Intelligence Index 上,Claude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) 的得分为 51,而 K2 Horizon 375B A23B 的得分为 31。
K2 Horizon 375B A23B 更快。K2 Horizon 375B A23B 每秒生成 122.9 个 token,而 Claude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) 每秒生成 73.2 个 token。
K2 Horizon 375B A23B 的延迟更低。K2 Horizon 375B A23B 的首 Token 延迟为 21.89 秒,而 Claude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) 为 25.18 秒。
Claude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) 的上下文窗口更大。Claude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) 支持 1.0M 个 token,而 K2 Horizon 375B A23B 支持 520k 个 token。