Artificial Analysisの能力指数の方法論

概要

Artificial Analysisの能力指数は、コーディングのような幅広い能力から法務のような専門分野まで、特定のユースケースでモデルがどの程度優れた性能を発揮するかを測定します。

各構成ベンチマークは、スコアを指数に統合する前にArtificial Analysisが独立して実行します。ほかの評価指標と同様に、能力指数には制約があり、すべてのユースケースに直接適用できるとは限りません。それでも、特定の分野で重要な業務についてモデルを比較するうえで、有用な総合評価となります。

基礎となる各ベンチマークの実行方法と採点方法は、知能ベンチマークの方法論に記載しています。

指数の内訳

指数は、スキル別または業界別に分類されます。スキル指数は、コーディングやエージェント型タスクなど、分野を横断して適用できる能力を測定し、構成ベンチマークの等加重平均として算出します。業界指数は、法務、ヘルスケア・医療、財務・会計など、単一の職種や分野を対象とし、その分野の実際の業務で各能力が現れる頻度に応じて一連の能力を重み付けします。

タスクの重み付けには、O*NET形式の業務活動分類を使用しています。ほかの評価指標と同様に、これらの指数には制約があり、すべてのユースケースに直接対応するとは限りません。以下の表に、各指数の構成要素と重みを示します。

指数種類重み能力評価説明
Agentic IndexSkill50%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and multi-step execution of real-world knowledge work
50%Agentic Customer Interaction𝜏³-BankingMulti-turn, tool-using customer-support workflows over a knowledge base
Finance & Accounting IndexIndustry30%Business KnowledgeAA-OmniscienceDomain recall in accounting, corporate finance, economics, and investments
30%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and multi-step task execution, such as running spreadsheets, querying ERPs, or coordinating a workflow
20%ReasoningHLEMulti-step quantitative and analytic reasoning, used for sensitivity analysis, valuation, and structured problem solving
10%Agentic Customer Interaction𝜏³-BankingMulti-turn, tool-using customer workflows over an unstructured knowledge base
5%Long-ContextLCRReading and reasoning across long financial filings, deal documents, and research notes
5%Non-HallucinationAA-OmniscienceAvoiding fabricated figures or citations
Strategy & Ops IndexIndustry30%Business KnowledgeAA-OmniscienceWorking knowledge of business processes, accounting basics, and operational concepts
35%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and orchestrating multi-step office workflows end-to-end
30%Agentic Customer Interaction𝜏³-BankingMulti-turn dialogue and tool use to resolve customer and stakeholder requests
5%Long-ContextLCRHolding context across long threads, policies, and records
Legal IndexIndustry35%Legal KnowledgeAA-OmniscienceRecall of statutes, doctrines, and procedure across jurisdictions
25%Agentic Knowledge WorkGDPval-AA v2Running matter-management workflows, drafting pipelines, and tool-augmented research
10%Long-ContextLCRReading and synthesizing across contracts, discovery productions, and case-law packets
10%Non-HallucinationAA-OmniscienceAvoiding fabricated case cites or invented statutes
5%Agentic Customer Interaction𝜏³-BankingMulti-turn, tool-using client intake and support workflows
15%ReasoningHLEMulti-step argumentation, statutory interpretation, and weighing conflicting authorities
Healthcare & Medical IndexIndustry35%Medical & Health KnowledgeAA-OmniscienceClinical knowledge across diagnosis, pharmacology, and care pathways
25%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and orchestrating EHR/pharmacy workflows end-to-end
15%Non-HallucinationAA-OmniscienceAvoiding fabricated drug interactions, doses, or guidelines
15%ReasoningHLEMulti-step clinical reasoning across biology and medicine
10%Agentic Customer Interaction𝜏³-BankingMulti-turn, tool-using patient and member support workflows
Engineering IndexIndustry35%Engineering KnowledgeAA-OmniscienceDomain recall across civil, electrical, mechanical, and other engineering disciplines
35%ReasoningHLE, GPQA Diamond, Crit-PtMulti-step quantitative reasoning for derivations, sizing calculations, and design trade-offs
25%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and multi-step execution of engineering deliverables
5%Agentic Terminal UseTerminal-Bench v2.1Operating real terminal environments for builds, scripts, system administration, and debugging
Economics IndexIndustry35%Economics KnowledgeAA-OmniscienceRecall across micro and macroeconomics, public finance, and markets
35%ReasoningHLEMulti-step quantitative and analytic reasoning for modeling, estimation, and inference
15%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and multi-step execution of analytical deliverables
15%Long-ContextLCRReading and reasoning across long reports, datasets, and research notes