Economics Index
Assesses model performance across the economics domain. Capabilities evaluated include domain-specific knowledge (microeconomics, macroeconomics, public finance), analysis and forecasting, research synthesis, quantitative modeling, and more.
Repräsentative Workflows ansehenThe Artificial Analysis Economics Index combines performance across benchmarks chosen for economics work, spanning economics knowledge, agentic execution, reasoning, and long-context reading. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across economics tasks. Alle zugrunde liegenden Benchmarks werden unabhängig von Artificial Analysis durchgeführt. Wie die Evaluierungen ablaufen, erläutert unsere Methodik für das Intelligenz-Benchmarking.
| Fähigkeit | Gewichtung | Evaluierungen |
|---|---|---|
| Economics Knowledge | 35 % | AA-Omniscience Business Accuracy |
| Reasoning | 35 % | HLE |
| Agentic Knowledge Work | 25 % | GDPval-AA v2 und AA-Briefcase |
| Long-Context Reasoning | 5 % | LCR |
Punktzahl
Artificial Analysis Economics Index
Artificial Analysis Economics Index: Aufschlüsselung der Fähigkeiten
Aufschlüsselung der Fähigkeiten
Artificial Analysis Economics Index: Economics Knowledge
Repräsentative Workflows
Praxisnahe Workflows, die besonders die von Economics Index am stärksten gewichteten Fähigkeiten prüfen.
Beispiel: Estimate the impact of a proposed tariff change by applying incidence and elasticity theory to trade data, then quantify the resulting welfare trade-offs.
Beispiel: Reconcile two studies reaching opposite conclusions on the same minimum-wage question to read both in full, compare their identification strategies, and recommend the more defensible interpretation.
Beispiel: Build a forecasting workbook in Python that ingests several FRED data series to run a baseline ARIMA model, chart the projections, and produce a short written interpretation of the outputs.
Kosten
Artificial Analysis Economics Index: Kosten pro Aufgabe
Artificial Analysis Economics Index vs. Kosten pro Aufgabe
Geschwindigkeit
Artificial Analysis Economics Index: Zeit pro Aufgabe
Ausgabe-Token
Artificial Analysis Economics Index: Ausgabe-Token pro Aufgabe
Veröffentlichungsdatum
Artificial Analysis Economics Index vs. Veröffentlichungsdatum
Häufig gestellte Fragen
Laut dem Artificial Analysis Economics Index sind Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (63), Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) (62) und Claude Opus 5 (Adaptive Reasoning, Max Effort) (61) derzeit die leistungsstärksten KI-Modelle für wirtschaftswissenschaftliche Aufgaben. Die Rangliste wird bei der Veröffentlichung neuer Modelle aktualisiert.
Ja. Der Economics Index von Artificial Analysis ist ein unabhängiger Benchmark für die Leistung von KI-Modellen bei wirtschaftswissenschaftlichen Aufgaben. Er misst ökonomisches Wissen, quantitatives Schlussfolgern, agentische Ausführung und die Analyse langer Kontexte.
Der Economics Index ist ein zusammengesetzter Benchmark von Artificial Analysis, der die Modellleistung in den Wirtschaftswissenschaften bewertet. Geprüft werden unter anderem Fachwissen zu Mikroökonomie, Makroökonomie und öffentlichen Finanzen, Analyse und Prognose, Forschungssynthese sowie quantitative Modellierung.
Der Economics Index wird als gewichteter Durchschnitt seiner Teilpunktzahlen berechnet. Die Teilpunktzahlen und ihre Gewichtungen sind: Economics Knowledge (35 %), Reasoning (35 %), Agentic Knowledge Work (25 %) und Long-Context Reasoning (5 %).
Der Economics Index enthält AA-Omniscience Business Accuracy, HLE, GDPval-AA v2, AA-Briefcase und LCR.
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) erzielt derzeit mit 63 die höchste Punktzahl im Economics Index unter den Modellen mit veröffentlichten Ergebnissen. Modell ansehen
Eine höhere Punktzahl im Economics Index steht für eine insgesamt stärkere Leistung in den Benchmarks des Index. Für einen bestimmten Anwendungsfall können einzelne Benchmark-Ergebnisse aussagekräftiger sein als die zusammengesetzte Punktzahl.