経済学指数
経済学分野におけるモデルの性能を評価します。ミクロ経済学、マクロ経済学、財政学に関する専門知識、分析・予測、研究の統合、定量モデリングなどを対象とします。
代表的なワークフローを見るThe Artificial Analysis Economics Index combines performance across benchmarks chosen for economics work, spanning economics knowledge, agentic execution, reasoning, and long-context reading. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across economics tasks. 基礎となるすべてのベンチマークは、Artificial Analysisが独立して実施しています。評価方法の詳細は、知能ベンチマークの方法論をご覧ください。
| 能力 | 重み | 評価 |
|---|---|---|
| 経済学知識 | 35% | AA-Omniscience ビジネスの正確性 |
| 推論 | 35% | HLE |
| エージェント型知識作業 | 25% | GDPval-AA v2.1、AA-Briefcase v1.1 |
| 長文脈推論 | 5% | LCR |
スコア
Artificial Analysis 経済学指数
Artificial Analysis 経済学指数:能力の内訳
能力の内訳
Artificial Analysis 経済学指数:経済学知識
代表的なワークフロー
経済学指数で特に重視される能力を測る、実際の業務を想定したワークフローです。
例:Estimate the impact of a proposed tariff change by applying incidence and elasticity theory to trade data, then quantify the resulting welfare trade-offs.
例:Reconcile two studies reaching opposite conclusions on the same minimum-wage question to read both in full, compare their identification strategies, and recommend the more defensible interpretation.
例:Build a forecasting workbook in Python that ingests several FRED data series to run a baseline ARIMA model, chart the projections, and produce a short written interpretation of the outputs.
コスト
Artificial Analysis 経済学指数:タスク当たりのコスト
Artificial Analysis 経済学指数 vs. タスク当たりのコスト
速度
Artificial Analysis 経済学指数:タスク当たりの時間
出力トークン
Artificial Analysis 経済学指数:タスク当たりの出力トークン
リリース日
Artificial Analysis 経済学指数 vs. リリース日
よくある質問
Artificial Analysisの経済学指数によると、現在経済学の業務で最も高い性能を示すAIモデルはClaude Opus 5.5 (Max, Default Fallback) (66)、Claude Opus 5.5 (Xhigh, Default Fallback) (63)、Claude Fable 5.1 (Max, Default Fallback) (63)です。ランキングは新しいモデルのリリースに合わせて更新されます。
はい。Artificial Analysisの経済学指数は、経済学の業務におけるAIモデルの性能を測る独立したベンチマークです。経済学の知識、定量的推論、エージェント型の実行、長文コンテキスト分析を評価します。
経済学指数は、経済学分野におけるモデルの性能を評価するArtificial Analysisの複合ベンチマークです。ミクロ経済学、マクロ経済学、公共財政などの専門知識に加え、分析と予測、研究の統合、定量モデリングなどを評価します。
経済学指数は、各能力のサブスコアを加重平均して算出します。サブスコアとその重みは、経済学知識 (35%)、推論 (35%)、エージェント型知識作業 (25%)、長文脈推論 (5%)です。
経済学指数にはAA-Omniscience ビジネスの正確性、HLE、GDPval-AA v2.1、AA-Briefcase v1.1、LCRが含まれます。
結果が公開されているモデルの中では、現在Claude Opus 5.5 (Max, Default Fallback)が経済学指数で最高の66を記録しています。 モデルを見る
経済学指数のスコアが高いほど、指数を構成するベンチマーク全体での性能が優れていることを示します。特定の用途では、複合スコアよりも個別のベンチマーク結果のほうが参考になる場合があります。