模型比较:智能、性能与价格分析
Microevals 体验区智能
输出速度(token/秒)
延迟(秒)
价格(美元/100 万 token)
上下文窗口
亮点
智能
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index by Open Weights / Proprietary
Intelligence Evaluations
Agentic real-world work tasks, (Elo-500)/2000
Agentic tool use
Agentic coding & terminal use
Coding
Reasoning & knowledge
Scientific reasoning
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Agentic knowledge work, Elo
Agentic SaaS workflows
Legal agentic work, criterion pass rate
Agentic business operations
Instruction following
Long-horizon agentic tasks
Kubernetes incident root-cause analysis
Visual reasoning
AA-Briefcase
AA-Briefcase Elo
AA-Omniscience
AA-Omniscience Index
开放性指数
Artificial Analysis Openness Index: Score
Intelligence Index 比较
Intelligence Index vs. Cost per Intelligence Index Task
Token 使用量
Output Tokens per Intelligence Index Task
成本
Cost per Intelligence Index Task
Cost to Run Artificial Analysis Intelligence Index
Pricing: Cache Hit, Input, and Output
上下文窗口
Context Window
速度
按输出速度(每秒 token 数)衡量
Output Speed
Time per Intelligence Index Task
延迟
按首 Token 延迟(秒)衡量
Latency: Time To First Answer Token
端到端响应时间
Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed
End-to-End Response Time
模型规模(仅开放权重模型)
Model Size: Total and Active Parameters
常见问题
在已评测的 175 个模型中,Claude Opus 5 (Adaptive Reasoning, Max Effort) 目前以 61 分领跑 Artificial Analysis Intelligence Index。
按 Intelligence Index 排名,顶尖 AI 模型是:1. Claude Opus 5 (Adaptive Reasoning, Max Effort)(61)、2. Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)(60)、3. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)(60)、4. GPT-5.6 Sol (max)(59)和5. Claude Opus 5 (Adaptive Reasoning, High Effort)(59)。
Celeris-1 最快,输出速度为每秒 2,033.7 个 token,其次是 Mercury 2(741.8 t/s)和 LFM2.5-VL-1.6B(474.5 t/s)。
Nova Micro 最实惠,混合价格为每 100 万 token $0.03,其次是 Sarvam 30B (high)($0.03)和 Gemma 4 E4B (Non-reasoning)($0.03)。
Gemini 2.5 Flash-Lite (Non-reasoning) 的首 Token 延迟最低,为 0.33 秒,其次是 Command A+(0.44 秒)和 Gemini 2.5 Flash (Non-reasoning)(0.47 秒)。
Kimi K3 (max) 是排名最高的开放权重模型,Intelligence Index 得分为 57。总计评测的 175 个模型中,有 99 个开放权重模型。
按 Intelligence Index 排名,顶尖开放权重 AI 模型是:1. Kimi K3 (max)(57)、2. GLM-5.2 (max)(51)和3. DeepSeek V4 Flash 0731 (Reasoning, Max Effort)(50)。
在 130 个推理模型中,Claude Opus 5 (Adaptive Reasoning, Max Effort) 以 61 的 Intelligence Index 得分领先。推理模型会先通过扩展思考来解决复杂问题,再给出回答。
模型会在多个维度上进行比较,包括智能(质量)、价格、输出速度(每秒 token 数)、延迟(首 Token 延迟)、端到端响应时间和上下文窗口大小。性能指标通过标准化提示词,在 591 个模型上直接测量。
点击图表中的任意模型名称或行,即可打开该模型的专属页面,查看详细指标并与类似模型直接比较。您还可以使用模型选择器,自定义每张图表中显示的模型。 查看排行榜