All capability indexes

Economics Index

Assesses model performance across the economics domain. Capabilities evaluated include domain-specific knowledge (microeconomics, macroeconomics, public finance), analysis and forecasting, research synthesis, quantitative modeling, and more.

See representative workflows

The Artificial Analysis Economics Index combines performance across benchmarks chosen for economics work, spanning economics knowledge, agentic execution, reasoning, and long-context reading. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.

This composite metric provides a single score for tracking model performance across economics tasks. All underlying benchmarks are run independently by Artificial Analysis. See our Intelligence Benchmarking Methodology for how evaluations are conducted.

CapabilityWeightEvaluations
Economics Knowledge35%AA-Omniscience Business Accuracy
Reasoning35%HLE
Agentic Knowledge Work15%GDPval-AA v2
Long-Context15%LCR

Score

Artificial Analysis Economics Index

Incorporates 4 evaluations: AA-Omniscience, Humanity's Last Exam, GDPval-AA v2, AA-LCR · Higher is better
Reasoning models are indicated by a lightbulb icon

Artificial Analysis Economics Index: Capability Breakdown

Incorporates 4 evaluations: AA-Omniscience, Humanity's Last Exam, GDPval-AA v2, AA-LCR · Segmented by contribution
Reasoning models are indicated by a lightbulb icon

Capability Breakdown

Artificial Analysis Economics Index: Economics Knowledge

Incorporates 1 evaluation: AA-Omniscience · Higher is better
Reasoning models are indicated by a lightbulb icon

Representative Workflows

Real-world workflows that exercise the capabilities the Economics Index weights most heavily.

Analysis & forecastingEconomics KnowledgeReasoning

Example: Estimate the impact of a proposed tariff change by applying incidence and elasticity theory to trade data, then quantify the resulting welfare trade-offs.

Research & synthesisEconomics KnowledgeReasoningLong-Context

Example: Reconcile two studies reaching opposite conclusions on the same minimum-wage question to read both in full, compare their identification strategies, and recommend the more defensible interpretation.

Modelling & toolingReasoningAgentic Knowledge Work

Example: Build a forecasting workbook in Python that ingests several FRED data series to run a baseline ARIMA model, chart the projections, and produce a short written interpretation of the outputs.

Release Date

Artificial Analysis Economics Index vs. Release Date

Most attractive region

Cost

Artificial Analysis Economics Index: Cost per Task

Average cost per task (USD), broken down by input, cache hit, cache write, reasoning, and answer tokens

Average cost per task in the index. Costs are split by input, cache hit, cache write, reasoning, and answer token pricing where canonical token counts are available.

Artificial Analysis Economics Index: Total Cost

Total cost (USD) to run the index

The cost to run the index, calculated using the model's input and output token pricing and the number of tokens used.

Speed

Artificial Analysis Economics Index: Time per Task

Weighted average decode time (minutes) per task; excludes TTFT and overhead time · Lower is better

The weighted average time (minutes) per index task. This is calculated by dividing output tokens per task by output speed, weighted by the relative weights of each benchmark in the index.

Output Tokens

Artificial Analysis Economics Index: Output Tokens per Task

Output tokens used to run one task, broken down by reasoning and answer tokens

The average number of answer and reasoning tokens produced per benchmark task in this index.

Frequently Asked Questions

Based on the Artificial Analysis Economics Index, the top-performing AI models for economics work are currently Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (62), Claude Opus 4.8 (Adaptive Reasoning, Max Effort) (56), and GPT-5.6 Sol (max) (56). Rankings are updated as new models are released.

Yes. The Economics Index from Artificial Analysis is an independent benchmark of how AI models perform on economics work. It measures performance on the economics domain, including economics knowledge, quantitative reasoning, agentic execution, and long-context analysis.

The Economics Index is a composite benchmark from Artificial Analysis that assesses model performance across the economics domain. Capabilities evaluated include domain-specific knowledge (microeconomics, macroeconomics, public finance), analysis and forecasting, research synthesis, quantitative modeling, and more.

The Economics Index is calculated as a weighted average of its capability sub-scores. The sub-scores and their weights are: Economics Knowledge (35%), Reasoning (35%), Agentic Knowledge Work (15%), and Long-Context (15%).

The Economics Index includes AA-Omniscience Business Accuracy, HLE, GDPval-AA v2, and LCR.

Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) currently has the highest Economics Index score, with a score of 62 among models with published results. View model

A higher Economics Index score indicates stronger overall performance across the benchmarks that make up the index. For a specific use case, individual benchmark results may be more informative than the composite score.