亮点

Updated
Artificial Analysis Intelligence Index · Higher is better
Output tokens per second · Higher is better
Weighted average cost (USD) per Intelligence Index task · Lower is better
新增语言模型评测 · 9月15日
Qwen3.8 Max (0902)Qwen3.8 Max (0902)
新增语言模型评测 · 9月13日
K2 Horizon 0.9BK2 Horizon 0.9B
新增语言模型评测 · 9月13日
K2 Horizon 3.7BK2 Horizon 3.7B
新增语言模型评测 · 9月13日
K2 Horizon 7BK2 Horizon 7B
新增语言模型评测 · 9月13日
K2 Horizon MoVA 36B A4BK2 Horizon MoVA 36B A4B
新增语言模型评测 · 9月10日
Agnes 3.0 FlashAgnes 3.0 Flash
新增语言模型评测 · 9月10日
DeepSeek V4.1 Flash (Reasoning, Max Effort)DeepSeek V4.1 Flash (Reasoning, Max Effort)
新增语言模型评测 · 9月10日
Ling-3.0-flash-VLLing-3.0-flash-VL
新文章发布 · 9月9日
Benchmarking GPT-6 Astra
新文章发布 · 9月7日
OpenBMB releases MiniCPM5-2B
新文章发布 · 9月7日
Announcing the Artificial Analysis Intelligence Index v4.3
方法论更新 · 9月7日
Artificial AnalysisArtificial Analysis Intelligence Index v4.3
新增语言模型评测 · 9月7日
MiniCPM5-2BMiniCPM5-2B
新文章发布 · 9月4日
Announcing Artificial Analysis Intelligence Index v4.2
新增语言模型评测 · 9月3日
GPT-6 Astra (low)GPT-6 Astra (low)
新增语言模型评测 · 9月3日
GPT-6 Astra (medium)GPT-6 Astra (medium)
新增语言模型评测 · 9月3日
GPT-6 Astra (high)GPT-6 Astra (high)
新增语言模型评测 · 9月3日
GPT-6 Astra (xhigh)GPT-6 Astra (xhigh)
新增语言模型评测 · 9月3日
GPT-6 Astra (max)GPT-6 Astra (max)
新增语言模型评测 · 9月3日
K2 Horizon 375B A23BK2 Horizon 375B A23B查看更多

智能Updated

根据我们的独立评测衡量领先 AI 模型的智能

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.3 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1

Artificial Analysis Intelligence Index by Open Weights / Proprietary

Artificial Analysis Intelligence Index v4.3 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1

Cost per Intelligence Index Task

Weighted average cost (USD) per Artificial Analysis Intelligence Index task, segmented by token type. Lower is better

Intelligence Index vs. Cost per Intelligence Index Task

Artificial Analysis Intelligence Index · Weighted average cost (USD) per Artificial Analysis Intelligence Index task
Most attractive quadrant
Pareto line

Intelligence Index vs. Cost per Intelligence Index Task, by Model Release

All reasoning and effort variants of each selected release · Weighted average cost (USD) per Artificial Analysis Intelligence Index task
Most attractive quadrant
Pareto line

前沿语言模型的智能变化

Artificial Analysis Intelligence Index v4.3 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1

领先编程智能体在端到端软件工程任务中的表现、成本和执行时间

Artificial Analysis Coding Agent Index

Artificial Analysis Coding Agent Index v1.5 incorporates 3 benchmarks: DeepSWE v1.1, Terminal-Bench 4.0, and SWE-Atlas-QnA · Higher is better

Artificial Analysis Coding Agent Index vs. 每个任务的成本

Artificial Analysis Coding Agent Index vs. 每个任务的平均按 token 计费 API 成本(USD)
Most attractive quadrant
Pareto line

图像与视频

图像竞技场和视频竞技场排行榜中的顶尖模型,包含 95% 置信区间

文生图排行榜

来自图像竞技场盲测偏好投票的 Elo 分数。在此查看完整排行榜。

语音

文本转语音竞技场、语音转文本和语音到语音评测中的顶尖模型

Provider Voice Arena Quality Elo

Arena Elo: average Elo rating of the model · Higher is better

衡量模型在特定能力和行业中的表现

Artificial Analysis Finance & Accounting Index

Incorporates 7 evaluations: AA-Omniscience, GDPval-AA v2, AA-Briefcase, Humanity's Last Exam, AutomationBench-AA, AA-LCR v1.1, GDP.pdf · Higher is better

Intelligence Evaluations

Intelligence evaluations measured independently by Artificial Analysis · Higher is better
See more

Agentic knowledge work, (Elo-500)/2000

Agentic real-world work tasks, (Elo-500)/2000

Agentic SaaS workflows

Agentic coding & terminal use

Coding

Reasoning & knowledge

Professional document reasoning, All-pass

Physics reasoning

Long context reasoning

Legal agentic work, criterion pass rate

Agentic business operations

Quantitative analysis on spreadsheets & documents

Instruction following

Agentic tool use

Long-horizon agentic tasks

Kubernetes incident root-cause analysis

Visual reasoning

AA-Briefcase

AA-Briefcase 是一项面向长周期知识工作的前沿智能体评测,通过要求交付电子表格、演示文稿和备忘录等成果的真实商业工作流来测试智能体

AA-Briefcase Elo

AA-Briefcase is an agentic knowledge work benchmark developed by Artificial Analysis. AA-Briefcase Elo is a combined metric that aggregates rubric pass rate, analytical quality Elo and presentation Elo · Higher is better

AA-AnalystAgent

AA-AnalystAgent 是一项针对真实电子表格和文档进行端到端定量分析的基准测试,考察业务分析师和数据分析师日常所做的工作

AA-AnalystAgent pass^5

Share of end-to-end quantitative analysis tasks solved on all five attempts · Higher is better

AA-Omniscience

AA-Omniscience 是一项知识与幻觉基准测试,奖励准确回答、惩罚不当猜测,并全面展示不同模型在各领域生成事实可靠内容的能力

AA-Omniscience Index

AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct.

GDPval-AA v2

GDPval-AA v2 在广泛职业中使用具有真实经济价值的任务评测 AI 模型

GDPval-AA v2 Leaderboard

Elo rating for performance on real-world work tasks · Anchored to a human baseline of 1,000 · Higher is better
Human Baseline (1,000)

Artificial Analysis 开放性指数根据模型各个组成部分的可获取性和透明度,评估模型的“开放”程度。

Artificial Analysis Openness Index: Components

Openness Index underlying score contribution by components, up to a maximum of 18 (higher is more open)

Artificial Analysis Openness Index vs. Artificial Analysis Intelligence Index

Most attractive quadrant
Pareto line

输出 Token

根据我们的独立评测统计领先 AI 模型的输出 token 数

Output Tokens per Intelligence Index Task

Weighted average number of output tokens used to run one task in the Artificial Analysis Intelligence Index

成本

根据我们的独立评测分析领先 AI 模型的价格和实际成本

Cost per Intelligence Index Task

Weighted average cost (USD) per Artificial Analysis Intelligence Index task, segmented by token type. Lower is better

Cost to Run Artificial Analysis Intelligence Index

Cost (USD) to run all evaluations in the Artificial Analysis Intelligence Index

Pricing: Cache Hit, Input, and Output

Price (USD per M Tokens)

速度与延迟

比较第一方 API 的性能

Output Speed

Output tokens per second · Higher is better

Time per Intelligence Index Task

Weighted average decode time (minutes) per task; excludes TTFT and overhead time · Lower is better

服务商

Endpoint Accuracy Index: gpt-oss-120b (high)

v1.0 · Composite of BFCL v4-500, HLE-250 and AA-LCR-25 run against each provider endpoint · Percentage of the reference endpoint, with 95% confidence interval · Higher is better

Output Speed vs. Price: gpt-oss-120b (high)

Output tokens per second · USD per 1M tokens (blended) · 10,000 input tokens
Most attractive quadrant
Pareto line

价格(缓存命中、输入与输出):gpt-oss-120b (high)

Price (USD per M Tokens) · Lower is better · 10,000 input tokens

Output Speed: gpt-oss-120b (high)

Output speed: output tokens per second · 10,000 input tokens