Methodik der Fähigkeitsindizes von Artificial Analysis
Überblick
Die Fähigkeitsindizes von Artificial Analysis messen, wie gut Modelle in bestimmten Anwendungsfällen abschneiden – von breit einsetzbaren Fähigkeiten wie Programmierung bis zu Berufsfeldern wie dem Rechtswesen.
Jeder Komponenten-Benchmark wird von Artificial Analysis unabhängig ausgeführt, bevor seine Werte im Index zusammengeführt werden. Wie alle Evaluationsmetriken haben auch Fähigkeitsindizes ihre Grenzen und lassen sich möglicherweise nicht direkt auf jeden Anwendungsfall übertragen. Sie bieten jedoch eine nützliche Zusammenfassung für den Vergleich von Modellen bei den Aufgaben, die für einen bestimmten Bereich relevant sind.
Die zugrunde liegenden Benchmarks sowie ihre jeweilige Durchführung und Bewertung sind in der Methodik für das Intelligenz-Benchmarking dokumentiert.
Aufschlüsselung der Indizes
Indizes sind entweder fähigkeits- oder branchenbezogen. Fähigkeitsindizes messen eine bereichsübergreifende Fähigkeit, etwa Programmierung oder agentische Aufgaben, und werden als gleich gewichteter Durchschnitt ihrer Komponenten-Benchmarks berechnet. Branchenindizes sind auf einen einzelnen Beruf oder Fachbereich wie Rechtswesen, Gesundheit und Medizin oder Finanzen und Rechnungswesen ausgerichtet. Sie gewichten verschiedene Fähigkeiten danach, wie häufig diese in realen Aufgaben des jeweiligen Bereichs vorkommen.
Unsere Aufgabengewichtungen basieren auf einer an O*NET angelehnten Taxonomie von Arbeitsaktivitäten. Wie alle Evaluationsmetriken haben diese Indizes ihre Grenzen und lassen sich möglicherweise nicht direkt auf jeden Anwendungsfall übertragen. Die folgende Tabelle zeigt die Komponenten und Gewichtungen jedes Index.
| Index | Typ | Gewichtung | Fähigkeit | Evaluationen | Beschreibung |
|---|---|---|---|---|---|
| Agentic Index | Skill | 50% | Agentic Knowledge Work | GDPval-AA v2 | Tool use, planning, and multi-step execution of real-world knowledge work |
| 50% | Agentic Customer Interaction | 𝜏³-Banking | Multi-turn, tool-using customer-support workflows over a knowledge base | ||
| Finance & Accounting Index | Industry | 30% | Business Knowledge | AA-Omniscience | Domain recall in accounting, corporate finance, economics, and investments |
| 30% | Agentic Knowledge Work | GDPval-AA v2 | Tool use, planning, and multi-step task execution, such as running spreadsheets, querying ERPs, or coordinating a workflow | ||
| 20% | Reasoning | HLE | Multi-step quantitative and analytic reasoning, used for sensitivity analysis, valuation, and structured problem solving | ||
| 10% | Agentic Customer Interaction | 𝜏³-Banking | Multi-turn, tool-using customer workflows over an unstructured knowledge base | ||
| 5% | Long-Context | LCR | Reading and reasoning across long financial filings, deal documents, and research notes | ||
| 5% | Non-Hallucination | AA-Omniscience | Avoiding fabricated figures or citations | ||
| Strategy & Ops Index | Industry | 30% | Business Knowledge | AA-Omniscience | Working knowledge of business processes, accounting basics, and operational concepts |
| 35% | Agentic Knowledge Work | GDPval-AA v2 | Tool use, planning, and orchestrating multi-step office workflows end-to-end | ||
| 30% | Agentic Customer Interaction | 𝜏³-Banking | Multi-turn dialogue and tool use to resolve customer and stakeholder requests | ||
| 5% | Long-Context | LCR | Holding context across long threads, policies, and records | ||
| Legal Index | Industry | 35% | Legal Knowledge | AA-Omniscience | Recall of statutes, doctrines, and procedure across jurisdictions |
| 25% | Agentic Knowledge Work | GDPval-AA v2 | Running matter-management workflows, drafting pipelines, and tool-augmented research | ||
| 10% | Long-Context | LCR | Reading and synthesizing across contracts, discovery productions, and case-law packets | ||
| 10% | Non-Hallucination | AA-Omniscience | Avoiding fabricated case cites or invented statutes | ||
| 5% | Agentic Customer Interaction | 𝜏³-Banking | Multi-turn, tool-using client intake and support workflows | ||
| 15% | Reasoning | HLE | Multi-step argumentation, statutory interpretation, and weighing conflicting authorities | ||
| Healthcare & Medical Index | Industry | 35% | Medical & Health Knowledge | AA-Omniscience | Clinical knowledge across diagnosis, pharmacology, and care pathways |
| 25% | Agentic Knowledge Work | GDPval-AA v2 | Tool use, planning, and orchestrating EHR/pharmacy workflows end-to-end | ||
| 15% | Non-Hallucination | AA-Omniscience | Avoiding fabricated drug interactions, doses, or guidelines | ||
| 15% | Reasoning | HLE | Multi-step clinical reasoning across biology and medicine | ||
| 10% | Agentic Customer Interaction | 𝜏³-Banking | Multi-turn, tool-using patient and member support workflows | ||
| Engineering Index | Industry | 35% | Engineering Knowledge | AA-Omniscience | Domain recall across civil, electrical, mechanical, and other engineering disciplines |
| 35% | Reasoning | HLE, GPQA Diamond, Crit-Pt | Multi-step quantitative reasoning for derivations, sizing calculations, and design trade-offs | ||
| 25% | Agentic Knowledge Work | GDPval-AA v2 | Tool use, planning, and multi-step execution of engineering deliverables | ||
| 5% | Agentic Terminal Use | Terminal-Bench v2.1 | Operating real terminal environments for builds, scripts, system administration, and debugging | ||
| Economics Index | Industry | 35% | Economics Knowledge | AA-Omniscience | Recall across micro and macroeconomics, public finance, and markets |
| 35% | Reasoning | HLE | Multi-step quantitative and analytic reasoning for modeling, estimation, and inference | ||
| 15% | Agentic Knowledge Work | GDPval-AA v2 | Tool use, planning, and multi-step execution of analytical deliverables | ||
| 15% | Long-Context | LCR | Reading and reasoning across long reports, datasets, and research notes |