All articles

September 14, 2026

Announcing Artificial Analysis Capability Indices v1.1

We are updating the Capability Indices v1.1 with stronger domain tuning, combining slices of core Intelligence Index v4.3 evaluations alongside specialized evaluations.

The Capability Indices map tasks from O*NET occupations to the benchmarks that represent them, weighting each benchmark by how often its capability appears across tasks. Capability Indices v1.1 incorporates updates to Intelligence Index v4.2 and v4.3, slicing by domain when available.

Changelog

IndexAddedDeleted
Finance & AccountingAgentic Tool Use (AutomationBench-AA, Finance)
Agentic Knowledge Work (AA-Briefcase)
Long-Context (GDP.pdf)
Agentic Customer Interaction (𝜏³-Banking)
Strategy & OpsAgentic Tool Use (AutomationBench-AA, Operations)
Agentic Knowledge Work (AA-Briefcase)
Long-Context (GDP.pdf)
Agentic Customer Interaction (𝜏³-Banking)
LegalAgentic Tool Use (AutomationBench-AA, Operations and Support)
Agentic Knowledge Work (AA-Briefcase)
Long-Context (GDP.pdf)
Agentic Customer Interaction (𝜏³-Banking)
Healthcare & MedicalAgentic Tool Use (AutomationBench-AA, Operations and Support)
Long-Context Reasoning (MLCR-AA)
Agentic Knowledge Work (AA-Briefcase)
Agentic Customer Interaction (𝜏³-Banking)
EngineeringAgentic Terminal Use (Terminal-Bench v4.0)
Agentic Knowledge Work (AA-Briefcase)
Agentic Terminal Use (Terminal-Bench v2.1)
Reasoning (GPQA Diamond)
EconomicsAgentic Knowledge Work (AA-Briefcase)-

Changes by Index

Finance & Accounting

Agentic Tool Use added, sourced from the Finance slice of AutomationBench-AA. Agentic Customer Interaction removed. AA-Briefcase added alongside GDPval-AA v2 in Agentic Knowledge Work, and GDP.pdf added alongside LCR in Long-Context.

CapabilityEvaluationsv1.0v1.1
Business KnowledgeAA-Omniscience30%30%
Agentic Knowledge WorkGDPval-AA v2, AA-Briefcase30%30%
ReasoningHLE20%20%
Agentic Tool UseAutomationBench-AA0%10%
Long-ContextLCR, GDP.pdf5%5%
Non-HallucinationAA-Omniscience5%5%

Strategy & Ops

Agentic Tool Use added, sourced from the Operations slice of AutomationBench-AA to reflect the importance of tool use in day-to-day operational tasks. Agentic Customer Interaction removed. AA-Briefcase added alongside GDPval-AA v2 in Agentic Knowledge Work, and GDP.pdf added alongside LCR in Long-Context.

CapabilityEvaluationsv1.0v1.1
Agentic Knowledge WorkGDPval-AA v2, AA-Briefcase35%35%
Business KnowledgeAA-Omniscience30%30%
Agentic Tool UseAutomationBench-AA0%30%
Long-ContextLCR, GDP.pdf5%5%

Legal

Agentic Tool Use added, sourced from the Operations and Support slices of AutomationBench-AA to reflect the role of tool use in legal workflows. Agentic Customer Interaction removed. AA-Briefcase added alongside GDPval-AA v2 in Agentic Knowledge Work, and GDP.pdf added alongside LCR in Long-Context.

CapabilityEvaluationsv1.0v1.1
Legal KnowledgeAA-Omniscience35%35%
Agentic Knowledge WorkGDPval-AA v2, AA-Briefcase25%25%
ReasoningHLE15%15%
Long-ContextLCR, GDP.pdf10%10%
Non-HallucinationAA-Omniscience10%10%
Agentic Tool UseAutomationBench-AA0%5%

Healthcare & Medical

Long-Context Reasoning added, sourced from MLCR-AA. This evaluates reasoning across lengthy clinical records. Agentic Tool Use added, sourced from the Operations and Support slices of AutomationBench-AA to cover operational workflows relevant to healthcare. Agentic Customer Interaction removed. AA-Briefcase added alongside GDPval-AA v2 in Agentic Knowledge Work.

CapabilityEvaluationsv1.0v1.1
Medical & Health KnowledgeAA-Omniscience35%30%
Agentic Knowledge WorkGDPval-AA v2, AA-Briefcase25%25%
Long-Context ReasoningMLCR-AA0%15%
Non-HallucinationAA-Omniscience15%10%
ReasoningHLE15%10%
Agentic Tool UseAutomationBench-AA0%10%

Engineering

Agentic Terminal Use updated to Terminal-Bench v4.0. GPQA Diamond removed from Reasoning. AA-Briefcase added alongside GDPval-AA v2 in Agentic Knowledge Work.

CapabilityEvaluationsv1.0v1.1
Engineering KnowledgeAA-Omniscience35%35%
ReasoningHLE, CritPt35%30%
Agentic Knowledge WorkGDPval-AA v2, AA-Briefcase25%20%
Agentic Terminal UseTerminal-Bench 4.05%15%

Economics

AA-Briefcase added alongside GDPval-AA v2 in Agentic Knowledge Work.

CapabilityEvaluationsv1.0v1.1
Economics KnowledgeAA-Omniscience35%35%
ReasoningHLE35%35%
Agentic Knowledge WorkGDPval-AA v2, AA-Briefcase15%25%
Long-Context ReasoningLCR15%5%

Full results and methodology

Explore all Capability Indices: https://artificialanalysis.ai/models/capabilities

Read the methodology: https://artificialanalysis.ai/methodology/capability-indices