Methodik der Fähigkeitsindizes von Artificial Analysis

Überblick

Die Fähigkeitsindizes von Artificial Analysis messen, wie gut Modelle in bestimmten Anwendungsfällen abschneiden – von breit einsetzbaren Fähigkeiten wie Programmierung bis zu Berufsfeldern wie dem Rechtswesen.

Jeder Komponenten-Benchmark wird von Artificial Analysis unabhängig ausgeführt, bevor seine Werte im Index zusammengeführt werden. Wie alle Evaluationsmetriken haben auch Fähigkeitsindizes ihre Grenzen und lassen sich möglicherweise nicht direkt auf jeden Anwendungsfall übertragen. Sie bieten jedoch eine nützliche Zusammenfassung für den Vergleich von Modellen bei den Aufgaben, die für einen bestimmten Bereich relevant sind.

Die zugrunde liegenden Benchmarks sowie ihre jeweilige Durchführung und Bewertung sind in der Methodik für das Intelligenz-Benchmarking dokumentiert.

Aufschlüsselung der Indizes

Indizes sind entweder fähigkeits- oder branchenbezogen. Fähigkeitsindizes messen eine bereichsübergreifende Fähigkeit, etwa Programmierung oder agentische Aufgaben, und werden als gleich gewichteter Durchschnitt ihrer Komponenten-Benchmarks berechnet. Branchenindizes sind auf einen einzelnen Beruf oder Fachbereich wie Rechtswesen, Gesundheit und Medizin oder Finanzen und Rechnungswesen ausgerichtet. Sie gewichten verschiedene Fähigkeiten danach, wie häufig diese in realen Aufgaben des jeweiligen Bereichs vorkommen.

Unsere Aufgabengewichtungen basieren auf einer an O*NET angelehnten Taxonomie von Arbeitsaktivitäten. Wie alle Evaluationsmetriken haben diese Indizes ihre Grenzen und lassen sich möglicherweise nicht direkt auf jeden Anwendungsfall übertragen. Die folgende Tabelle zeigt die Komponenten und Gewichtungen jedes Index.

IndexTypGewichtungFähigkeitEvaluationenBeschreibung
Agentic IndexSkill50%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and multi-step execution of real-world knowledge work
50%Agentic Customer Interaction𝜏³-BankingMulti-turn, tool-using customer-support workflows over a knowledge base
Finance & Accounting IndexIndustry30%Business KnowledgeAA-OmniscienceDomain recall in accounting, corporate finance, economics, and investments
30%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and multi-step task execution, such as running spreadsheets, querying ERPs, or coordinating a workflow
20%ReasoningHLEMulti-step quantitative and analytic reasoning, used for sensitivity analysis, valuation, and structured problem solving
10%Agentic Customer Interaction𝜏³-BankingMulti-turn, tool-using customer workflows over an unstructured knowledge base
5%Long-ContextLCRReading and reasoning across long financial filings, deal documents, and research notes
5%Non-HallucinationAA-OmniscienceAvoiding fabricated figures or citations
Strategy & Ops IndexIndustry30%Business KnowledgeAA-OmniscienceWorking knowledge of business processes, accounting basics, and operational concepts
35%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and orchestrating multi-step office workflows end-to-end
30%Agentic Customer Interaction𝜏³-BankingMulti-turn dialogue and tool use to resolve customer and stakeholder requests
5%Long-ContextLCRHolding context across long threads, policies, and records
Legal IndexIndustry35%Legal KnowledgeAA-OmniscienceRecall of statutes, doctrines, and procedure across jurisdictions
25%Agentic Knowledge WorkGDPval-AA v2Running matter-management workflows, drafting pipelines, and tool-augmented research
10%Long-ContextLCRReading and synthesizing across contracts, discovery productions, and case-law packets
10%Non-HallucinationAA-OmniscienceAvoiding fabricated case cites or invented statutes
5%Agentic Customer Interaction𝜏³-BankingMulti-turn, tool-using client intake and support workflows
15%ReasoningHLEMulti-step argumentation, statutory interpretation, and weighing conflicting authorities
Healthcare & Medical IndexIndustry35%Medical & Health KnowledgeAA-OmniscienceClinical knowledge across diagnosis, pharmacology, and care pathways
25%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and orchestrating EHR/pharmacy workflows end-to-end
15%Non-HallucinationAA-OmniscienceAvoiding fabricated drug interactions, doses, or guidelines
15%ReasoningHLEMulti-step clinical reasoning across biology and medicine
10%Agentic Customer Interaction𝜏³-BankingMulti-turn, tool-using patient and member support workflows
Engineering IndexIndustry35%Engineering KnowledgeAA-OmniscienceDomain recall across civil, electrical, mechanical, and other engineering disciplines
35%ReasoningHLE, GPQA Diamond, Crit-PtMulti-step quantitative reasoning for derivations, sizing calculations, and design trade-offs
25%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and multi-step execution of engineering deliverables
5%Agentic Terminal UseTerminal-Bench v2.1Operating real terminal environments for builds, scripts, system administration, and debugging
Economics IndexIndustry35%Economics KnowledgeAA-OmniscienceRecall across micro and macroeconomics, public finance, and markets
35%ReasoningHLEMulti-step quantitative and analytic reasoning for modeling, estimation, and inference
15%Agentic Knowledge WorkGDPval-AA v2Tool use, planning, and multi-step execution of analytical deliverables
15%Long-ContextLCRReading and reasoning across long reports, datasets, and research notes