亮点
智能
根据我们的独立评测衡量领先 AI 模型的智能
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index by Open Weights / Proprietary
Cost per Intelligence Index Task
Intelligence Index vs. Cost per Intelligence Index Task
前沿语言模型的智能变化
领先编程智能体在端到端软件工程任务中的表现、成本和执行时间
Artificial Analysis Coding Agent Index
图像与视频
图像竞技场和视频竞技场排行榜中的顶尖模型,包含 95% 置信区间
文生图排行榜
语音
文本转语音竞技场、语音转文本和语音到语音评测中的顶尖模型
Text to Speech Arena Leaderboard
Artificial Analysis Agentic Index
Intelligence Evaluations
Agentic real-world work tasks, (Elo-500)/2000
Agentic tool use
Agentic coding & terminal use
Coding
Reasoning & knowledge
Scientific reasoning
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Agentic knowledge work, Elo
Agentic SaaS workflows
Legal agentic work, criterion pass rate
Agentic business operations
Instruction following
Long-horizon agentic tasks
Kubernetes incident root-cause analysis
Visual reasoning
AA-Briefcase
AA-Briefcase 是一项面向长周期知识工作的前沿智能体评测,通过要求交付电子表格、演示文稿和备忘录等成果的真实商业工作流来测试智能体
AA-Briefcase Elo
AA-Omniscience
AA-Omniscience 是一项知识与幻觉基准测试,奖励准确回答、惩罚不当猜测,并全面展示不同模型在各领域生成事实可靠内容的能力
AA-Omniscience Index
GDPval-AA v2
GDPval-AA v2 在广泛职业中使用具有真实经济价值的任务评测 AI 模型
GDPval-AA v2 Leaderboard
Artificial Analysis 开放性指数根据模型各个组成部分的可获取性和透明度,评估模型的“开放”程度。
Artificial Analysis Openness Index: Components
Artificial Analysis Openness Index vs. Artificial Analysis Intelligence Index
输出 Token
根据我们的独立评测统计领先 AI 模型的输出 token 数
Output Tokens per Intelligence Index Task
成本
根据我们的独立评测分析领先 AI 模型的价格和实际成本
Cost per Intelligence Index Task
Cost to Run Artificial Analysis Intelligence Index
Pricing: Cache Hit, Input, and Output
速度与延迟
比较第一方 API 的性能