| 智能 | |||
|---|---|---|---|
| Intelligence Index | 41 | 42 | |
| AA-Briefcase v1.1 | 1461 | 1505 | |
| GDPval-AA v2.1 | 1292 | 1574 | |
| AutomationBench-AA | 53% | 57% | |
| Terminal-Bench 4.0 | 30% | 16% | |
| SciCode | 53% | 55% | |
| Humanity's Last Exam | 40% | 39% | |
| GDP.pdf | 20% | 21% | |
| CritPt | 17% | 13% | |
| AA-Omniscience | 20 | 30 | |
| AA-LCR v1.1 | 76% | 80% | |
| 成本 | |||
| 每 100 万 Token 价格 | $1.54 | $1.35 | |
| 每 100 万输入 Token 价格 | $2.00 | $2.00 | |
| 每 100 万输出 Token 价格 | $10.00 | $6.00 | |
| 每 100 万缓存命中 Token 价格 | $0.20 | $0.50 | |
| 每任务成本 | $0.59 | $1.25 | |
| 运行 Intelligence Index 的成本 | US$701 | US$1,631 | |
| Token 使用量 | |||
| 每任务输出 Token 数 | 19k | 28k | |
| 每任务推理 Token 数 | 8k | 14k | |
| 运行 Intelligence Index 的输出 Token 数 | 29M | 75M | |
| 性能 | |||
| 输出速度 | 92 token/秒 | 76 token/秒 | |
| 首 Token 延迟 | 1.34 秒 | 8.92 秒 | |
| 首个回答 Token 时间 | 1.34 秒 | 8.92 秒 | |
| 端到端响应时间 | 6.78 秒 | 15.47 秒 | |
| 每任务耗时 | 129.74 秒 | 376.77 秒 | |
| 技术规格 | |||
| 上下文窗口 | 1000k token约 1,500 页 A4 纸(12 号 Arial 字体) | 500k token约 750 页 A4 纸(12 号 Arial 字体) | |
| 发布日期 | 2026年9月 | 2026年9月 | |
| 推理 | 是 | 是 | |
| 输入模态 | 支持:文本和图像 | 支持:文本和图像 | |
| 输出模态 | 支持:文本 | 支持:文本 | |
| 开放权重 | 否 | 否 | |
亮点
智能Updated
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index:开放权重与专有模型
衡量模型在特定能力和行业中的表现
Artificial Analysis 财务与会计指数
Intelligence Evaluations
Agentic knowledge work, (Elo-500)/2000
Agentic real-world work tasks, (Elo-500)/2000
Agentic SaaS workflows
Agentic coding & terminal use
Coding
Reasoning & knowledge
Professional document reasoning, All-pass
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Legal agentic work, criterion pass rate
Agentic business operations
Agentic scientific research workflows in a terminal
Quantitative analysis on spreadsheets & documents
Kubernetes incident root-cause analysis
Visual reasoning
Medical long context reasoning
AA-Briefcase v1.1Updated
AA-Briefcase Elo
AA-Omniscience
AA-Omniscience Index
Intelligence Index 比较
Intelligence Index 与每项任务的成本
单任务成本(美元,对数刻度)
Token 使用量
Output Tokens per Intelligence Index Task
成本
每项 Intelligence Index 任务的成本
运行 Artificial Analysis Intelligence Index 的成本
Pricing: Cache Hit, Input, and Output
上下文窗口
上下文窗口
速度
按输出速度(每秒 token 数)衡量
Output Speed
Time per Intelligence Index Task
延迟
按首 Token 延迟(秒)衡量
Latency: Time To First Answer Token
端到端响应时间
Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed
End-to-End Response Time
常见问题
Grok 4.7 (Low) 更智能。在 Artificial Analysis Intelligence Index 上,Grok 4.7 (Low) 的得分为 42,而 Claude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) 的得分为 41。
Claude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) 更快。Claude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) 每秒生成 92.0 个 token,而 Grok 4.7 (Low) 每秒生成 76.3 个 token。
Grok 4.7 (Low) 更便宜。Grok 4.7 (Low) 每 100 万 token 的成本为 $1.35,而 Claude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) 为 $1.54(缓存命中/输入/输出比例为 7:2:1)。
Claude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) 的延迟更低。Claude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) 的首 Token 延迟为 1.34 秒,而 Grok 4.7 (Low) 为 8.92 秒。
Claude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) 的上下文窗口更大。Claude Sonnet 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) 支持 1.0M 个 token,而 Grok 4.7 (Low) 支持 500k 个 token。