Economics Index
Assesses model performance across the economics domain. Capabilities evaluated include domain-specific knowledge (microeconomics, macroeconomics, public finance), analysis and forecasting, research synthesis, quantitative modeling, and more.
查看代表性工作流The Artificial Analysis Economics Index combines performance across benchmarks chosen for economics work, spanning economics knowledge, agentic execution, reasoning, and long-context reading. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across economics tasks. 所有底层基准测试均由 Artificial Analysis 独立运行。有关评测的实施方式,请参阅我们的智能基准测试方法论。
| 能力 | 权重 | 评测 |
|---|---|---|
| Economics Knowledge | 35% | AA-Omniscience Business Accuracy |
| Reasoning | 35% | HLE |
| Agentic Knowledge Work | 25% | GDPval-AA v2和AA-Briefcase |
| Long-Context Reasoning | 5% | LCR |
得分
Artificial Analysis Economics Index
Artificial Analysis Economics Index:能力明细
能力明细
Artificial Analysis Economics Index:Economics Knowledge
代表性工作流
这些真实工作流重点检验 Economics Index 中权重最高的能力。
示例:Estimate the impact of a proposed tariff change by applying incidence and elasticity theory to trade data, then quantify the resulting welfare trade-offs.
示例:Reconcile two studies reaching opposite conclusions on the same minimum-wage question to read both in full, compare their identification strategies, and recommend the more defensible interpretation.
示例:Build a forecasting workbook in Python that ingests several FRED data series to run a baseline ARIMA model, chart the projections, and produce a short written interpretation of the outputs.
成本
Artificial Analysis Economics Index:单任务成本
Artificial Analysis Economics Index 与单任务成本
速度
Artificial Analysis Economics Index:单任务耗时
输出 token
Artificial Analysis Economics Index:单任务输出 token
发布日期
Artificial Analysis Economics Index 与发布日期
常见问题
根据 Artificial Analysis Economics Index,目前在经济学工作上表现最佳的 AI 模型是 Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (63)、Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) (62)和Claude Opus 5 (Adaptive Reasoning, Max Effort) (61)。新模型发布后,排行榜会随之更新。
有。Artificial Analysis Economics Index 是一项独立基准测试,用于衡量 AI 模型在经济学工作上的表现。它评估经济学知识、定量推理、智能体执行和长上下文分析等能力。
Economics Index 是 Artificial Analysis 推出的综合基准测试,用于评估模型在经济学领域的表现。评估能力包括微观经济学、宏观经济学、公共财政等专业知识,以及分析与预测、研究综合、定量建模等。
Economics Index 按各项能力子分数的加权平均值计算。各项子分数及其权重为:Economics Knowledge (35%)、Reasoning (35%)、Agentic Knowledge Work (25%)和Long-Context Reasoning (5%)。
Economics Index 包含 AA-Omniscience Business Accuracy、HLE、GDPval-AA v2、AA-Briefcase和LCR。
在已公布结果的模型中,Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) 目前以 63 分位居 Economics Index 榜首。 查看模型
Economics Index 得分越高,表示模型在构成该指数的各项基准测试中整体表现越强。对于特定用例,单项基准测试结果可能比综合得分更具参考价值。