Artificial Analysis 能力指数方法论
概述
Artificial Analysis 能力指数衡量模型在特定使用场景下的表现,这些场景既可以是编程这类宽泛的能力,也可以是法律工作这类专业领域。
每一项组成指数的基准测试都由 Artificial Analysis 独立运行,之后才将其得分合并计入指数。与所有评测指标一样,能力指数存在局限性,未必能直接适用于每一个使用场景,但在比较模型于某一领域真正重要的工作上的表现时,它们提供了有用的综合参考。
底层的各项基准测试,包括每一项如何运行和评分,均记录在智能基准测试方法论中。
指数构成明细
指数分为技能类(Skill)和行业类(Industry)两种。技能类指数衡量跨领域适用的某项能力,例如编程(Coding)或智能体(Agentic)类工作,其分数按各组成基准测试的等权平均计算。行业类指数针对单一职业或领域,例如法律(Legal)、医疗与医学(Healthcare & Medical)或金融与会计(Finance & Accounting),并根据各项能力在该领域真实任务中出现的频率对其加权。
我们的任务权重参考了 O*NET 风格的工作活动分类体系。与所有评测指标一样,这些指数存在局限性,未必能直接对应到每一个使用场景。下表列出了各指数的组成部分及其权重。
| 指数 | 类型 | 权重 | 能力 | 评测 | 说明 |
|---|---|---|---|---|---|
| Agentic Index | Skill | 50% | Agentic Knowledge Work | GDPval-AA v2 | Tool use, planning, and multi-step execution of real-world knowledge work |
| 50% | Agentic Customer Interaction | 𝜏³-Banking | Multi-turn, tool-using customer-support workflows over a knowledge base | ||
| Finance & Accounting Index | Industry | 30% | Business Knowledge | AA-Omniscience | Domain recall in accounting, corporate finance, economics, and investments |
| 30% | Agentic Knowledge Work | GDPval-AA v2 | Tool use, planning, and multi-step task execution, such as running spreadsheets, querying ERPs, or coordinating a workflow | ||
| 20% | Reasoning | HLE | Multi-step quantitative and analytic reasoning, used for sensitivity analysis, valuation, and structured problem solving | ||
| 10% | Agentic Customer Interaction | 𝜏³-Banking | Multi-turn, tool-using customer workflows over an unstructured knowledge base | ||
| 5% | Long-Context | LCR | Reading and reasoning across long financial filings, deal documents, and research notes | ||
| 5% | Non-Hallucination | AA-Omniscience | Avoiding fabricated figures or citations | ||
| Strategy & Ops Index | Industry | 30% | Business Knowledge | AA-Omniscience | Working knowledge of business processes, accounting basics, and operational concepts |
| 35% | Agentic Knowledge Work | GDPval-AA v2 | Tool use, planning, and orchestrating multi-step office workflows end-to-end | ||
| 30% | Agentic Customer Interaction | 𝜏³-Banking | Multi-turn dialogue and tool use to resolve customer and stakeholder requests | ||
| 5% | Long-Context | LCR | Holding context across long threads, policies, and records | ||
| Legal Index | Industry | 35% | Legal Knowledge | AA-Omniscience | Recall of statutes, doctrines, and procedure across jurisdictions |
| 25% | Agentic Knowledge Work | GDPval-AA v2 | Running matter-management workflows, drafting pipelines, and tool-augmented research | ||
| 10% | Long-Context | LCR | Reading and synthesizing across contracts, discovery productions, and case-law packets | ||
| 10% | Non-Hallucination | AA-Omniscience | Avoiding fabricated case cites or invented statutes | ||
| 5% | Agentic Customer Interaction | 𝜏³-Banking | Multi-turn, tool-using client intake and support workflows | ||
| 15% | Reasoning | HLE | Multi-step argumentation, statutory interpretation, and weighing conflicting authorities | ||
| Healthcare & Medical Index | Industry | 35% | Medical & Health Knowledge | AA-Omniscience | Clinical knowledge across diagnosis, pharmacology, and care pathways |
| 25% | Agentic Knowledge Work | GDPval-AA v2 | Tool use, planning, and orchestrating EHR/pharmacy workflows end-to-end | ||
| 15% | Non-Hallucination | AA-Omniscience | Avoiding fabricated drug interactions, doses, or guidelines | ||
| 15% | Reasoning | HLE | Multi-step clinical reasoning across biology and medicine | ||
| 10% | Agentic Customer Interaction | 𝜏³-Banking | Multi-turn, tool-using patient and member support workflows | ||
| Engineering Index | Industry | 35% | Engineering Knowledge | AA-Omniscience | Domain recall across civil, electrical, mechanical, and other engineering disciplines |
| 35% | Reasoning | HLE, GPQA Diamond, Crit-Pt | Multi-step quantitative reasoning for derivations, sizing calculations, and design trade-offs | ||
| 25% | Agentic Knowledge Work | GDPval-AA v2 | Tool use, planning, and multi-step execution of engineering deliverables | ||
| 5% | Agentic Terminal Use | Terminal-Bench v2.1 | Operating real terminal environments for builds, scripts, system administration, and debugging | ||
| Economics Index | Industry | 35% | Economics Knowledge | AA-Omniscience | Recall across micro and macroeconomics, public finance, and markets |
| 35% | Reasoning | HLE | Multi-step quantitative and analytic reasoning for modeling, estimation, and inference | ||
| 15% | Agentic Knowledge Work | GDPval-AA v2 | Tool use, planning, and multi-step execution of analytical deliverables | ||
| 15% | Long-Context | LCR | Reading and reasoning across long reports, datasets, and research notes |