Artificial Analysis 能力指数方法论

概述

Artificial Analysis 能力指数衡量模型在特定使用场景下的表现,这些场景既可以是编程这类宽泛的能力,也可以是法律工作这类专业领域。

每一项组成指数的基准测试都由 Artificial Analysis 独立运行,之后才将其得分合并计入指数。与所有评测指标一样,能力指数存在局限性,未必能直接适用于每一个使用场景,但在比较模型于某一领域真正重要的工作上的表现时,它们提供了有用的综合参考。

底层的各项基准测试,包括每一项如何运行和评分,均记录在智能基准测试方法论中。

指数构成明细

指数分为技能类(Skill)和行业类(Industry)两种。技能类指数衡量跨领域适用的某项能力,例如编程(Coding)或智能体(Agentic)类工作,其分数按各组成基准测试的等权平均计算。行业类指数针对单一职业或领域,例如法律(Legal)、医疗与医学(Healthcare & Medical)或金融与会计(Finance & Accounting),并根据各项能力在该领域真实任务中出现的频率对其加权。

我们的任务权重参考了 O*NET 风格的工作活动分类体系。与所有评测指标一样,这些指数存在局限性,未必能直接对应到每一个使用场景。下表列出了各指数的组成部分及其权重。

指数类型权重能力评测说明
Agentic IndexSkill50%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and multi-step execution of real-world knowledge work
50%Agentic Customer Interaction𝜏³-BankingMulti-turn, tool-using customer-support workflows over a knowledge base
Finance & Accounting IndexIndustry30%Business KnowledgeAA-OmniscienceDomain recall in accounting, corporate finance, economics, and investments
30%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and multi-step task execution, such as running spreadsheets, querying ERPs, or coordinating a workflow
20%ReasoningHLEMulti-step quantitative and analytic reasoning, used for sensitivity analysis, valuation, and structured problem solving
10%Agentic Customer Interaction𝜏³-BankingMulti-turn, tool-using customer workflows over an unstructured knowledge base
5%Long-ContextLCRReading and reasoning across long financial filings, deal documents, and research notes
5%Non-HallucinationAA-OmniscienceAvoiding fabricated figures or citations
Strategy & Ops IndexIndustry30%Business KnowledgeAA-OmniscienceWorking knowledge of business processes, accounting basics, and operational concepts
35%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and orchestrating multi-step office workflows end-to-end
30%Agentic Customer Interaction𝜏³-BankingMulti-turn dialogue and tool use to resolve customer and stakeholder requests
5%Long-ContextLCRHolding context across long threads, policies, and records
Legal IndexIndustry35%Legal KnowledgeAA-OmniscienceRecall of statutes, doctrines, and procedure across jurisdictions
25%Agentic Knowledge WorkGDPval-AA v2Running matter-management workflows, drafting pipelines, and tool-augmented research
10%Long-ContextLCRReading and synthesizing across contracts, discovery productions, and case-law packets
10%Non-HallucinationAA-OmniscienceAvoiding fabricated case cites or invented statutes
5%Agentic Customer Interaction𝜏³-BankingMulti-turn, tool-using client intake and support workflows
15%ReasoningHLEMulti-step argumentation, statutory interpretation, and weighing conflicting authorities
Healthcare & Medical IndexIndustry35%Medical & Health KnowledgeAA-OmniscienceClinical knowledge across diagnosis, pharmacology, and care pathways
25%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and orchestrating EHR/pharmacy workflows end-to-end
15%Non-HallucinationAA-OmniscienceAvoiding fabricated drug interactions, doses, or guidelines
15%ReasoningHLEMulti-step clinical reasoning across biology and medicine
10%Agentic Customer Interaction𝜏³-BankingMulti-turn, tool-using patient and member support workflows
Engineering IndexIndustry35%Engineering KnowledgeAA-OmniscienceDomain recall across civil, electrical, mechanical, and other engineering disciplines
35%ReasoningHLE, GPQA Diamond, Crit-PtMulti-step quantitative reasoning for derivations, sizing calculations, and design trade-offs
25%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and multi-step execution of engineering deliverables
5%Agentic Terminal UseTerminal-Bench v2.1Operating real terminal environments for builds, scripts, system administration, and debugging
Economics IndexIndustry35%Economics KnowledgeAA-OmniscienceRecall across micro and macroeconomics, public finance, and markets
35%ReasoningHLEMulti-step quantitative and analytic reasoning for modeling, estimation, and inference
15%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and multi-step execution of analytical deliverables
15%Long-ContextLCRReading and reasoning across long reports, datasets, and research notes