亮点
智能Updated
根据我们的独立评测衡量领先 AI 模型的智能
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index by Open Weights / Proprietary
Cost per Intelligence Index Task
Intelligence Index vs. Cost per Intelligence Index Task
Intelligence Index vs. Cost per Intelligence Index Task, by Model Release
前沿语言模型的智能变化
编程智能体指数Updated
领先编程智能体在端到端软件工程任务中的表现、成本和执行时间
Artificial Analysis Coding Agent Index
Artificial Analysis Coding Agent Index vs. 每个任务的成本
图像与视频
图像竞技场和视频竞技场排行榜中的顶尖模型,包含 95% 置信区间
文生图排行榜
语音
文本转语音竞技场、语音转文本和语音到语音评测中的顶尖模型
Provider Voice Arena Quality Elo
Artificial Analysis Finance & Accounting Index
Intelligence Evaluations
Agentic knowledge work, (Elo-500)/2000
Agentic real-world work tasks, (Elo-500)/2000
Agentic SaaS workflows
Agentic coding & terminal use
Coding
Reasoning & knowledge
Professional document reasoning, All-pass
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Legal agentic work, criterion pass rate
Agentic business operations
Quantitative analysis on spreadsheets & documents
Instruction following
Agentic tool use
Long-horizon agentic tasks
Kubernetes incident root-cause analysis
Visual reasoning
AA-Briefcase
AA-Briefcase 是一项面向长周期知识工作的前沿智能体评测,通过要求交付电子表格、演示文稿和备忘录等成果的真实商业工作流来测试智能体
AA-Briefcase Elo
AA-AnalystAgent
AA-AnalystAgent 是一项针对真实电子表格和文档进行端到端定量分析的基准测试,考察业务分析师和数据分析师日常所做的工作
AA-AnalystAgent pass^5
AA-Omniscience
AA-Omniscience 是一项知识与幻觉基准测试,奖励准确回答、惩罚不当猜测,并全面展示不同模型在各领域生成事实可靠内容的能力
AA-Omniscience Index
GDPval-AA v2
GDPval-AA v2 在广泛职业中使用具有真实经济价值的任务评测 AI 模型
GDPval-AA v2 Leaderboard
Artificial Analysis 开放性指数根据模型各个组成部分的可获取性和透明度,评估模型的“开放”程度。
Artificial Analysis Openness Index: Components
Artificial Analysis Openness Index vs. Artificial Analysis Intelligence Index
输出 Token
根据我们的独立评测统计领先 AI 模型的输出 token 数
Output Tokens per Intelligence Index Task
成本
根据我们的独立评测分析领先 AI 模型的价格和实际成本
Cost per Intelligence Index Task
Cost to Run Artificial Analysis Intelligence Index
Pricing: Cache Hit, Input, and Output
速度与延迟
比较第一方 API 的性能