Artificial Analysisの能力指数の方法論

概要

Artificial Analysisの能力指数は、法務、医療、金融といった特定の専門的なユースケースでモデルがどの程度優れた性能を発揮するかを測定します。

各構成ベンチマークを独立して実行したうえで、そのスコアを指数に統合します。ほかの評価指標と同様に、能力指数には制約があり、すべてのユースケースに直接適用できるとは限りません。それでも、特定の分野で重要な業務についてモデルを比較するうえで、有用な総合評価となります。

基礎となる各ベンチマークの実行方法と採点方法は、知能ベンチマークの方法論に記載しています。

指数の内訳

各指数は、法務、ヘルスケア・医療、財務・会計など、単一の職種や分野を対象とし、その分野の実際の業務で各能力が現れる頻度に応じて一連の能力を重み付けします。

タスクの重み付けには、O*NET形式の業務活動分類を使用しています。以下の表に、各指数の構成要素と重みを示します。

指数重み能力評価説明
Finance & Accounting Index30%Business KnowledgeAA-OmniscienceDomain recall in accounting, corporate finance, economics, and investments
30%Agentic Knowledge WorkGDPval-AA v2, AA-BriefcaseTool use, planning, and multi-step task execution, such as building spreadsheets, running an analysis, or coordinating a workflow
20%ReasoningHLEMulti-step quantitative and analytic reasoning, used for sensitivity analysis, valuation, and structured problem solving
10%Agentic Tool UseAutomationBench-AACompleting finance workflows across business apps such as spreadsheets, email, and accounting tools without breaking guardrails
5%Long-ContextLCR, GDP.pdfReading and reasoning across long financial filings, deal documents, and PDF reports
5%Non-HallucinationAA-OmniscienceAvoiding fabricated figures or citations
Strategy & Ops Index30%Business KnowledgeAA-OmniscienceWorking knowledge of business processes, accounting basics, and operational concepts
35%Agentic Knowledge WorkGDPval-AA v2, AA-BriefcaseTool use, planning, and orchestrating multi-step office workflows end-to-end
30%Agentic Tool UseAutomationBench-AACompleting operations, HR, marketing, sales, and support workflows across business apps without breaking guardrails
5%Long-ContextLCR, GDP.pdfHolding context across long threads, policies, records, and PDF documents
Legal Index35%Legal KnowledgeAA-OmniscienceRecall of statutes, doctrines, and procedure across jurisdictions
25%Agentic Knowledge WorkGDPval-AA v2, AA-BriefcaseRunning matter-management workflows, drafting pipelines, and tool-augmented research
15%ReasoningHLEMulti-step argumentation, statutory interpretation, and weighing conflicting authorities
10%Long-ContextLCR, GDP.pdfReading and synthesizing across contracts, discovery productions, case-law packets, and PDF filings
10%Non-HallucinationAA-OmniscienceAvoiding fabricated case cites or invented statutes
5%Agentic Tool UseAutomationBench-AACompleting client support and operations workflows across business apps without breaking guardrails
Healthcare & Medical Index30%Medical & Health KnowledgeAA-OmniscienceClinical knowledge across diagnosis, pharmacology, and care pathways
25%Agentic Knowledge WorkGDPval-AA v2, AA-BriefcaseTool use, planning, and orchestrating EHR/pharmacy workflows end-to-end
15%Long-Context ReasoningMLCR-AASynthesising findings across long, fragmented patient records and claims files
10%Non-HallucinationAA-OmniscienceAvoiding fabricated drug interactions, doses, or guidelines
10%ReasoningHLEMulti-step clinical reasoning across biology and medicine
10%Agentic Tool UseAutomationBench-AACompleting patient support and operations workflows across business apps without breaking guardrails
Engineering Index35%Engineering KnowledgeAA-OmniscienceDomain recall across civil, electrical, mechanical, and other engineering disciplines
30%ReasoningHLE, CritPtMulti-step quantitative reasoning for derivations, sizing calculations, and design trade-offs
20%Agentic Knowledge WorkGDPval-AA v2, AA-BriefcaseTool use, planning, and multi-step execution of engineering deliverables
15%Agentic Terminal UseTerminal-Bench v4.0Operating real terminal environments for builds, scripts, system administration, and debugging
Economics Index35%Economics KnowledgeAA-OmniscienceRecall across micro and macroeconomics, public finance, and markets
35%ReasoningHLEMulti-step quantitative and analytic reasoning for modeling, estimation, and inference
25%Agentic Knowledge WorkGDPval-AA v2, AA-BriefcaseTool use, planning, and multi-step execution of analytical deliverables
5%Long-Context ReasoningLCRReading and reasoning across long reports, datasets, and research notes