Economics Index
Assesses model performance across the economics domain. Capabilities evaluated include domain-specific knowledge (microeconomics, macroeconomics, public finance), analysis and forecasting, research synthesis, quantitative modeling, and more.
대표 워크플로 보기The Artificial Analysis Economics Index combines performance across benchmarks chosen for economics work, spanning economics knowledge, agentic execution, reasoning, and long-context reading. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across economics tasks. 모든 기반 벤치마크는 Artificial Analysis가 독립적으로 실행합니다. 평가 수행 방식은 지능 벤치마킹 방법론을 참조하세요.
| 역량 | 가중치 | 평가 |
|---|---|---|
| Economics Knowledge | 35% | AA-Omniscience Business Accuracy |
| Reasoning | 35% | HLE |
| Agentic Knowledge Work | 25% | GDPval-AA v2 및 AA-Briefcase |
| Long-Context Reasoning | 5% | LCR |
점수
Artificial Analysis Economics Index
Artificial Analysis Economics Index: 역량 세부 분석
역량 세부 분석
Artificial Analysis Economics Index: Economics Knowledge
대표 워크플로
Economics Index에서 가장 높은 가중치를 둔 역량을 평가하는 실제 워크플로입니다.
예시: Estimate the impact of a proposed tariff change by applying incidence and elasticity theory to trade data, then quantify the resulting welfare trade-offs.
예시: Reconcile two studies reaching opposite conclusions on the same minimum-wage question to read both in full, compare their identification strategies, and recommend the more defensible interpretation.
예시: Build a forecasting workbook in Python that ingests several FRED data series to run a baseline ARIMA model, chart the projections, and produce a short written interpretation of the outputs.
비용
Artificial Analysis Economics Index: 작업당 비용
Artificial Analysis Economics Index vs. 작업당 비용
속도
Artificial Analysis Economics Index: 작업당 시간
출력 토큰
Artificial Analysis Economics Index: 작업당 출력 토큰
출시일
Artificial Analysis Economics Index vs. 출시일
자주 묻는 질문
Artificial Analysis Economics Index에 따르면 현재 경제학 업무에서 가장 뛰어난 AI 모델은 Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (63), Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) (62) 및 Claude Opus 5 (Adaptive Reasoning, Max Effort) (61)입니다. 순위는 새 모델 출시에 맞춰 업데이트됩니다.
네. Artificial Analysis의 Economics Index는 경제학 업무에서 AI 모델의 성능을 측정하는 독립적인 벤치마크입니다. 경제학 지식, 정량적 추론, 에이전트 실행, 긴 컨텍스트 분석을 평가합니다.
Economics Index는 경제학 분야에서 모델의 성능을 평가하는 Artificial Analysis의 종합 벤치마크입니다. 미시경제학, 거시경제학, 공공재정 관련 전문 지식과 분석 및 예측, 연구 종합, 정량 모델링 등을 평가합니다.
Economics Index는 각 역량 하위 점수의 가중 평균으로 계산합니다. 하위 점수와 가중치는 Economics Knowledge (35%), Reasoning (35%), Agentic Knowledge Work (25%) 및 Long-Context Reasoning (5%)입니다.
Economics Index에는 AA-Omniscience Business Accuracy, HLE, GDPval-AA v2, AA-Briefcase 및 LCR이 포함됩니다.
결과가 공개된 모델 중 현재 Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)의 Economics Index 점수가 63로 가장 높습니다. 모델 보기
Economics Index 점수가 높을수록 지수를 구성하는 벤치마크 전반에서 더 뛰어난 성능을 보인다는 뜻입니다. 특정 사용 사례에서는 종합 점수보다 개별 벤치마크 결과가 더 유용할 수 있습니다.