경제학 지수
경제학 전반에서 모델 성능을 평가합니다. 미시경제학, 거시경제학, 공공재정에 관한 전문 지식, 분석 및 예측, 연구 종합, 정량 모델링 등을 평가합니다.
대표 워크플로 보기The Artificial Analysis Economics Index combines performance across benchmarks chosen for economics work, spanning economics knowledge, agentic execution, reasoning, and long-context reading. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across economics tasks. 모든 기반 벤치마크는 Artificial Analysis가 독립적으로 실행합니다. 평가 수행 방식은 지능 벤치마킹 방법론을 참조하세요.
| 역량 | 가중치 | 평가 |
|---|---|---|
| 경제학 지식 | 35% | AA-Omniscience 비즈니스 정확도 |
| 추론 | 35% | HLE |
| 에이전트형 지식 작업 | 25% | GDPval-AA v2.1 및 AA-Briefcase v1.1 |
| 장문 맥락 추론 | 5% | LCR |
점수
Artificial Analysis 경제학 지수
Artificial Analysis 경제학 지수: 역량 세부 분석
역량 세부 분석
Artificial Analysis 경제학 지수: 경제학 지식
대표 워크플로
경제학 지수에서 가장 높은 가중치를 둔 역량을 평가하는 실제 워크플로입니다.
예시: Estimate the impact of a proposed tariff change by applying incidence and elasticity theory to trade data, then quantify the resulting welfare trade-offs.
예시: Reconcile two studies reaching opposite conclusions on the same minimum-wage question to read both in full, compare their identification strategies, and recommend the more defensible interpretation.
예시: Build a forecasting workbook in Python that ingests several FRED data series to run a baseline ARIMA model, chart the projections, and produce a short written interpretation of the outputs.
비용
Artificial Analysis 경제학 지수: 작업당 비용
Artificial Analysis 경제학 지수 vs. 작업당 비용
속도
Artificial Analysis 경제학 지수: 작업당 시간
출력 토큰
Artificial Analysis 경제학 지수: 작업당 출력 토큰
출시일
Artificial Analysis 경제학 지수 vs. 출시일
자주 묻는 질문
Artificial Analysis 경제학 지수에 따르면 현재 경제학 업무에서 가장 뛰어난 AI 모델은 Claude Opus 5.5 (Max, Default Fallback) (66), Claude Opus 5.5 (Xhigh, Default Fallback) (63) 및 Claude Fable 5.1 (Max, Default Fallback) (63)입니다. 순위는 새 모델 출시에 맞춰 업데이트됩니다.
네. Artificial Analysis의 경제학 지수는 경제학 업무에서 AI 모델의 성능을 측정하는 독립적인 벤치마크입니다. 경제학 지식, 정량적 추론, 에이전트 실행, 긴 컨텍스트 분석을 평가합니다.
경제학 지수는 경제학 분야에서 모델의 성능을 평가하는 Artificial Analysis의 종합 벤치마크입니다. 미시경제학, 거시경제학, 공공재정 관련 전문 지식과 분석 및 예측, 연구 종합, 정량 모델링 등을 평가합니다.
경제학 지수는 각 역량 하위 점수의 가중 평균으로 계산합니다. 하위 점수와 가중치는 경제학 지식 (35%), 추론 (35%), 에이전트형 지식 작업 (25%) 및 장문 맥락 추론 (5%)입니다.
경제학 지수에는 AA-Omniscience 비즈니스 정확도, HLE, GDPval-AA v2.1, AA-Briefcase v1.1 및 LCR이 포함됩니다.
결과가 공개된 모델 중 현재 Claude Opus 5.5 (Max, Default Fallback)의 경제학 지수 점수가 66로 가장 높습니다. 모델 보기
경제학 지수 점수가 높을수록 지수를 구성하는 벤치마크 전반에서 더 뛰어난 성능을 보인다는 뜻입니다. 특정 사용 사례에서는 종합 점수보다 개별 벤치마크 결과가 더 유용할 수 있습니다.