Artificial Analysisの能力指数の方法論
概要
Artificial Analysisの能力指数は、コーディングのような幅広い能力から法務のような専門分野まで、特定のユースケースでモデルがどの程度優れた性能を発揮するかを測定します。
各構成ベンチマークは、スコアを指数に統合する前にArtificial Analysisが独立して実行します。ほかの評価指標と同様に、能力指数には制約があり、すべてのユースケースに直接適用できるとは限りません。それでも、特定の分野で重要な業務についてモデルを比較するうえで、有用な総合評価となります。
基礎となる各ベンチマークの実行方法と採点方法は、知能ベンチマークの方法論に記載しています。
指数の内訳
指数は、スキル別または業界別に分類されます。スキル指数は、コーディングやエージェント型タスクなど、分野を横断して適用できる能力を測定し、構成ベンチマークの等加重平均として算出します。業界指数は、法務、ヘルスケア・医療、財務・会計など、単一の職種や分野を対象とし、その分野の実際の業務で各能力が現れる頻度に応じて一連の能力を重み付けします。
タスクの重み付けには、O*NET形式の業務活動分類を使用しています。ほかの評価指標と同様に、これらの指数には制約があり、すべてのユースケースに直接対応するとは限りません。以下の表に、各指数の構成要素と重みを示します。
| 指数 | 種類 | 重み | 能力 | 評価 | 説明 |
|---|---|---|---|---|---|
| Agentic Index | Skill | 50% | Agentic Knowledge Work | GDPval-AA v2 | Tool use, planning, and multi-step execution of real-world knowledge work |
| 50% | Agentic Customer Interaction | 𝜏³-Banking | Multi-turn, tool-using customer-support workflows over a knowledge base | ||
| Finance & Accounting Index | Industry | 30% | Business Knowledge | AA-Omniscience | Domain recall in accounting, corporate finance, economics, and investments |
| 30% | Agentic Knowledge Work | GDPval-AA v2 | Tool use, planning, and multi-step task execution, such as running spreadsheets, querying ERPs, or coordinating a workflow | ||
| 20% | Reasoning | HLE | Multi-step quantitative and analytic reasoning, used for sensitivity analysis, valuation, and structured problem solving | ||
| 10% | Agentic Customer Interaction | 𝜏³-Banking | Multi-turn, tool-using customer workflows over an unstructured knowledge base | ||
| 5% | Long-Context | LCR | Reading and reasoning across long financial filings, deal documents, and research notes | ||
| 5% | Non-Hallucination | AA-Omniscience | Avoiding fabricated figures or citations | ||
| Strategy & Ops Index | Industry | 30% | Business Knowledge | AA-Omniscience | Working knowledge of business processes, accounting basics, and operational concepts |
| 35% | Agentic Knowledge Work | GDPval-AA v2 | Tool use, planning, and orchestrating multi-step office workflows end-to-end | ||
| 30% | Agentic Customer Interaction | 𝜏³-Banking | Multi-turn dialogue and tool use to resolve customer and stakeholder requests | ||
| 5% | Long-Context | LCR | Holding context across long threads, policies, and records | ||
| Legal Index | Industry | 35% | Legal Knowledge | AA-Omniscience | Recall of statutes, doctrines, and procedure across jurisdictions |
| 25% | Agentic Knowledge Work | GDPval-AA v2 | Running matter-management workflows, drafting pipelines, and tool-augmented research | ||
| 10% | Long-Context | LCR | Reading and synthesizing across contracts, discovery productions, and case-law packets | ||
| 10% | Non-Hallucination | AA-Omniscience | Avoiding fabricated case cites or invented statutes | ||
| 5% | Agentic Customer Interaction | 𝜏³-Banking | Multi-turn, tool-using client intake and support workflows | ||
| 15% | Reasoning | HLE | Multi-step argumentation, statutory interpretation, and weighing conflicting authorities | ||
| Healthcare & Medical Index | Industry | 35% | Medical & Health Knowledge | AA-Omniscience | Clinical knowledge across diagnosis, pharmacology, and care pathways |
| 25% | Agentic Knowledge Work | GDPval-AA v2 | Tool use, planning, and orchestrating EHR/pharmacy workflows end-to-end | ||
| 15% | Non-Hallucination | AA-Omniscience | Avoiding fabricated drug interactions, doses, or guidelines | ||
| 15% | Reasoning | HLE | Multi-step clinical reasoning across biology and medicine | ||
| 10% | Agentic Customer Interaction | 𝜏³-Banking | Multi-turn, tool-using patient and member support workflows | ||
| Engineering Index | Industry | 35% | Engineering Knowledge | AA-Omniscience | Domain recall across civil, electrical, mechanical, and other engineering disciplines |
| 35% | Reasoning | HLE, GPQA Diamond, Crit-Pt | Multi-step quantitative reasoning for derivations, sizing calculations, and design trade-offs | ||
| 25% | Agentic Knowledge Work | GDPval-AA v2 | Tool use, planning, and multi-step execution of engineering deliverables | ||
| 5% | Agentic Terminal Use | Terminal-Bench v2.1 | Operating real terminal environments for builds, scripts, system administration, and debugging | ||
| Economics Index | Industry | 35% | Economics Knowledge | AA-Omniscience | Recall across micro and macroeconomics, public finance, and markets |
| 35% | Reasoning | HLE | Multi-step quantitative and analytic reasoning for modeling, estimation, and inference | ||
| 15% | Agentic Knowledge Work | GDPval-AA v2 | Tool use, planning, and multi-step execution of analytical deliverables | ||
| 15% | Long-Context | LCR | Reading and reasoning across long reports, datasets, and research notes |