Artificial Analysis Capability Indices Methodology

Overview

Artificial Analysis Capability Indices measure how well models perform on specific professional use cases, such as legal, healthcare, or finance work.

We run every component benchmark independently before combining its scores into the index. Like all evaluation metrics, capability indices have limitations and may not apply directly to every use case, but they provide a useful synthesis for comparing models on the work that matters to a given domain.

The underlying benchmarks, including how each one is run and scored, are documented in the Intelligence Benchmarking methodology.

Index Breakdowns

Each index targets a single profession or field, such as Legal, Healthcare & Medical, or Finance & Accounting, and weights a set of capabilities by how often each appears in real-world tasks for that field.

Our task weightings draw on an O*NET-style taxonomy of work activities. The table below shows the components and weights for each index.

IndexWeightCapabilityEvaluationsDescription
Finance & Accounting Index30%Business KnowledgeAA-OmniscienceDomain recall in accounting, corporate finance, economics, and investments
30%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and multi-step task execution, such as running spreadsheets, querying ERPs, or coordinating a workflow
20%ReasoningHLEMulti-step quantitative and analytic reasoning, used for sensitivity analysis, valuation, and structured problem solving
10%Agentic Customer Interaction𝜏³-BankingMulti-turn, tool-using customer workflows over an unstructured knowledge base
5%Long-ContextLCRReading and reasoning across long financial filings, deal documents, and research notes
5%Non-HallucinationAA-OmniscienceAvoiding fabricated figures or citations
Strategy & Ops Index30%Business KnowledgeAA-OmniscienceWorking knowledge of business processes, accounting basics, and operational concepts
35%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and orchestrating multi-step office workflows end-to-end
30%Agentic Customer Interaction𝜏³-BankingMulti-turn dialogue and tool use to resolve customer and stakeholder requests
5%Long-ContextLCRHolding context across long threads, policies, and records
Legal Index35%Legal KnowledgeAA-OmniscienceRecall of statutes, doctrines, and procedure across jurisdictions
25%Agentic Knowledge WorkGDPval-AA v2Running matter-management workflows, drafting pipelines, and tool-augmented research
10%Long-ContextLCRReading and synthesizing across contracts, discovery productions, and case-law packets
10%Non-HallucinationAA-OmniscienceAvoiding fabricated case cites or invented statutes
5%Agentic Customer Interaction𝜏³-BankingMulti-turn, tool-using client intake and support workflows
15%ReasoningHLEMulti-step argumentation, statutory interpretation, and weighing conflicting authorities
Healthcare & Medical Index30%Medical & Health KnowledgeAA-OmniscienceClinical knowledge across diagnosis, pharmacology, and care pathways
25%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and orchestrating EHR/pharmacy workflows end-to-end
15%Long-Context ReasoningMLCR-AASynthesising findings across long, fragmented patient records and claims files
10%Non-HallucinationAA-OmniscienceAvoiding fabricated drug interactions, doses, or guidelines
10%ReasoningHLEMulti-step clinical reasoning across biology and medicine
10%Agentic Customer Interaction𝜏³-BankingMulti-turn, tool-using patient and member support workflows
Engineering Index35%Engineering KnowledgeAA-OmniscienceDomain recall across civil, electrical, mechanical, and other engineering disciplines
35%ReasoningHLE, GPQA Diamond, Crit-PtMulti-step quantitative reasoning for derivations, sizing calculations, and design trade-offs
25%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and multi-step execution of engineering deliverables
5%Agentic Terminal UseTerminal-Bench v2.1Operating real terminal environments for builds, scripts, system administration, and debugging
Economics Index35%Economics KnowledgeAA-OmniscienceRecall across micro and macroeconomics, public finance, and markets
35%ReasoningHLEMulti-step quantitative and analytic reasoning for modeling, estimation, and inference
15%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and multi-step execution of analytical deliverables
15%Long-ContextLCRReading and reasoning across long reports, datasets, and research notes