Economics Index
Assesses model performance across the economics domain. Capabilities evaluated include domain-specific knowledge (microeconomics, macroeconomics, public finance), analysis and forecasting, research synthesis, quantitative modeling, and more.
代表的なワークフローを見るThe Artificial Analysis Economics Index combines performance across benchmarks chosen for economics work, spanning economics knowledge, agentic execution, reasoning, and long-context reading. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across economics tasks. 基礎となるすべてのベンチマークは、Artificial Analysisが独立して実施しています。評価方法の詳細は、知能ベンチマークの方法論をご覧ください。
| 能力 | 重み | 評価 |
|---|---|---|
| Economics Knowledge | 35% | AA-Omniscience Business Accuracy |
| Reasoning | 35% | HLE |
| Agentic Knowledge Work | 25% | GDPval-AA v2、AA-Briefcase |
| Long-Context Reasoning | 5% | LCR |
スコア
Artificial Analysis Economics Index
Artificial Analysis Economics Index:能力の内訳
能力の内訳
Artificial Analysis Economics Index:Economics Knowledge
代表的なワークフロー
Economics Indexで特に重視される能力を測る、実際の業務を想定したワークフローです。
例:Estimate the impact of a proposed tariff change by applying incidence and elasticity theory to trade data, then quantify the resulting welfare trade-offs.
例:Reconcile two studies reaching opposite conclusions on the same minimum-wage question to read both in full, compare their identification strategies, and recommend the more defensible interpretation.
例:Build a forecasting workbook in Python that ingests several FRED data series to run a baseline ARIMA model, chart the projections, and produce a short written interpretation of the outputs.
コスト
Artificial Analysis Economics Index:タスク当たりのコスト
Artificial Analysis Economics Index vs. タスク当たりのコスト
速度
Artificial Analysis Economics Index:タスク当たりの時間
出力トークン
Artificial Analysis Economics Index:タスク当たりの出力トークン
リリース日
Artificial Analysis Economics Index vs. リリース日
よくある質問
Artificial AnalysisのEconomics Indexによると、現在経済学の業務で最も高い性能を示すAIモデルはClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (63)、Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) (62)、Claude Opus 5 (Adaptive Reasoning, Max Effort) (61)です。ランキングは新しいモデルのリリースに合わせて更新されます。
はい。Artificial AnalysisのEconomics Indexは、経済学の業務におけるAIモデルの性能を測る独立したベンチマークです。経済学の知識、定量的推論、エージェント型の実行、長文コンテキスト分析を評価します。
Economics Indexは、経済学分野におけるモデルの性能を評価するArtificial Analysisの複合ベンチマークです。ミクロ経済学、マクロ経済学、公共財政などの専門知識に加え、分析と予測、研究の統合、定量モデリングなどを評価します。
Economics Indexは、各能力のサブスコアを加重平均して算出します。サブスコアとその重みは、Economics Knowledge (35%)、Reasoning (35%)、Agentic Knowledge Work (25%)、Long-Context Reasoning (5%)です。
Economics IndexにはAA-Omniscience Business Accuracy、HLE、GDPval-AA v2、AA-Briefcase、LCRが含まれます。
結果が公開されているモデルの中では、現在Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)がEconomics Indexで最高の63を記録しています。 モデルを見る
Economics Indexのスコアが高いほど、指数を構成するベンチマーク全体での性能が優れていることを示します。特定の用途では、複合スコアよりも個別のベンチマーク結果のほうが参考になる場合があります。