Artificial Analysisの能力指数の方法論
概要
Artificial Analysisの能力指数は、法務、医療、金融といった特定の専門的なユースケースでモデルがどの程度優れた性能を発揮するかを測定します。
各構成ベンチマークを独立して実行したうえで、そのスコアを指数に統合します。ほかの評価指標と同様に、能力指数には制約があり、すべてのユースケースに直接適用できるとは限りません。それでも、特定の分野で重要な業務についてモデルを比較するうえで、有用な総合評価となります。
基礎となる各ベンチマークの実行方法と採点方法は、知能ベンチマークの方法論に記載しています。
指数の内訳
各指数は、法務、ヘルスケア・医療、財務・会計など、単一の職種や分野を対象とし、その分野の実際の業務で各能力が現れる頻度に応じて一連の能力を重み付けします。
タスクの重み付けには、O*NET形式の業務活動分類を使用しています。以下の表に、各指数の構成要素と重みを示します。
| 指数 | 重み | 能力 | 評価 | 説明 |
|---|---|---|---|---|
| Finance & Accounting Index | 30% | Business Knowledge | AA-Omniscience | Domain recall in accounting, corporate finance, economics, and investments |
| 30% | Agentic Knowledge Work | GDPval-AA v2, AA-Briefcase | Tool use, planning, and multi-step task execution, such as building spreadsheets, running an analysis, or coordinating a workflow | |
| 20% | Reasoning | HLE | Multi-step quantitative and analytic reasoning, used for sensitivity analysis, valuation, and structured problem solving | |
| 10% | Agentic Tool Use | AutomationBench-AA | Completing finance workflows across business apps such as spreadsheets, email, and accounting tools without breaking guardrails | |
| 5% | Long-Context | LCR, GDP.pdf | Reading and reasoning across long financial filings, deal documents, and PDF reports | |
| 5% | Non-Hallucination | AA-Omniscience | Avoiding fabricated figures or citations | |
| Strategy & Ops Index | 30% | Business Knowledge | AA-Omniscience | Working knowledge of business processes, accounting basics, and operational concepts |
| 35% | Agentic Knowledge Work | GDPval-AA v2, AA-Briefcase | Tool use, planning, and orchestrating multi-step office workflows end-to-end | |
| 30% | Agentic Tool Use | AutomationBench-AA | Completing operations, HR, marketing, sales, and support workflows across business apps without breaking guardrails | |
| 5% | Long-Context | LCR, GDP.pdf | Holding context across long threads, policies, records, and PDF documents | |
| Legal Index | 35% | Legal Knowledge | AA-Omniscience | Recall of statutes, doctrines, and procedure across jurisdictions |
| 25% | Agentic Knowledge Work | GDPval-AA v2, AA-Briefcase | Running matter-management workflows, drafting pipelines, and tool-augmented research | |
| 15% | Reasoning | HLE | Multi-step argumentation, statutory interpretation, and weighing conflicting authorities | |
| 10% | Long-Context | LCR, GDP.pdf | Reading and synthesizing across contracts, discovery productions, case-law packets, and PDF filings | |
| 10% | Non-Hallucination | AA-Omniscience | Avoiding fabricated case cites or invented statutes | |
| 5% | Agentic Tool Use | AutomationBench-AA | Completing client support and operations workflows across business apps without breaking guardrails | |
| Healthcare & Medical Index | 30% | Medical & Health Knowledge | AA-Omniscience | Clinical knowledge across diagnosis, pharmacology, and care pathways |
| 25% | Agentic Knowledge Work | GDPval-AA v2, AA-Briefcase | Tool use, planning, and orchestrating EHR/pharmacy workflows end-to-end | |
| 15% | Long-Context Reasoning | MLCR-AA | Synthesising findings across long, fragmented patient records and claims files | |
| 10% | Non-Hallucination | AA-Omniscience | Avoiding fabricated drug interactions, doses, or guidelines | |
| 10% | Reasoning | HLE | Multi-step clinical reasoning across biology and medicine | |
| 10% | Agentic Tool Use | AutomationBench-AA | Completing patient support and operations workflows across business apps without breaking guardrails | |
| Engineering Index | 35% | Engineering Knowledge | AA-Omniscience | Domain recall across civil, electrical, mechanical, and other engineering disciplines |
| 30% | Reasoning | HLE, CritPt | Multi-step quantitative reasoning for derivations, sizing calculations, and design trade-offs | |
| 20% | Agentic Knowledge Work | GDPval-AA v2, AA-Briefcase | Tool use, planning, and multi-step execution of engineering deliverables | |
| 15% | Agentic Terminal Use | Terminal-Bench v4.0 | Operating real terminal environments for builds, scripts, system administration, and debugging | |
| Economics Index | 35% | Economics Knowledge | AA-Omniscience | Recall across micro and macroeconomics, public finance, and markets |
| 35% | Reasoning | HLE | Multi-step quantitative and analytic reasoning for modeling, estimation, and inference | |
| 25% | Agentic Knowledge Work | GDPval-AA v2, AA-Briefcase | Tool use, planning, and multi-step execution of analytical deliverables | |
| 5% | Long-Context Reasoning | LCR | Reading and reasoning across long reports, datasets, and research notes |