Methodik der Fähigkeitsindizes von Artificial Analysis

Überblick

Die Fähigkeitsindizes von Artificial Analysis messen, wie gut Modelle in bestimmten beruflichen Anwendungsfällen abschneiden, etwa im Rechtswesen, im Gesundheitswesen oder im Finanzbereich.

Wir führen jeden Komponenten-Benchmark unabhängig aus, bevor seine Werte im Index zusammengeführt werden. Wie alle Evaluationsmetriken haben auch Fähigkeitsindizes ihre Grenzen und lassen sich möglicherweise nicht direkt auf jeden Anwendungsfall übertragen. Sie bieten jedoch eine nützliche Zusammenfassung für den Vergleich von Modellen bei den Aufgaben, die für einen bestimmten Bereich relevant sind.

Die zugrunde liegenden Benchmarks sowie ihre jeweilige Durchführung und Bewertung sind in der Methodik für das Intelligenz-Benchmarking dokumentiert.

Aufschlüsselung der Indizes

Jeder Index ist auf einen einzelnen Beruf oder Fachbereich wie Rechtswesen, Gesundheit und Medizin oder Finanzen und Rechnungswesen ausgerichtet. Er gewichtet verschiedene Fähigkeiten danach, wie häufig diese in realen Aufgaben des jeweiligen Bereichs vorkommen.

Unsere Aufgabengewichtungen basieren auf einer an O*NET angelehnten Taxonomie von Arbeitsaktivitäten. Die folgende Tabelle zeigt die Komponenten und Gewichtungen jedes Index.

IndexGewichtungFähigkeitEvaluationenBeschreibung
Finance & Accounting Index30%Business KnowledgeAA-OmniscienceDomain recall in accounting, corporate finance, economics, and investments
30%Agentic Knowledge WorkGDPval-AA v2, AA-BriefcaseTool use, planning, and multi-step task execution, such as building spreadsheets, running an analysis, or coordinating a workflow
20%ReasoningHLEMulti-step quantitative and analytic reasoning, used for sensitivity analysis, valuation, and structured problem solving
10%Agentic Tool UseAutomationBench-AACompleting finance workflows across business apps such as spreadsheets, email, and accounting tools without breaking guardrails
5%Long-ContextLCR, GDP.pdfReading and reasoning across long financial filings, deal documents, and PDF reports
5%Non-HallucinationAA-OmniscienceAvoiding fabricated figures or citations
Strategy & Ops Index30%Business KnowledgeAA-OmniscienceWorking knowledge of business processes, accounting basics, and operational concepts
35%Agentic Knowledge WorkGDPval-AA v2, AA-BriefcaseTool use, planning, and orchestrating multi-step office workflows end-to-end
30%Agentic Tool UseAutomationBench-AACompleting operations, HR, marketing, sales, and support workflows across business apps without breaking guardrails
5%Long-ContextLCR, GDP.pdfHolding context across long threads, policies, records, and PDF documents
Legal Index35%Legal KnowledgeAA-OmniscienceRecall of statutes, doctrines, and procedure across jurisdictions
25%Agentic Knowledge WorkGDPval-AA v2, AA-BriefcaseRunning matter-management workflows, drafting pipelines, and tool-augmented research
15%ReasoningHLEMulti-step argumentation, statutory interpretation, and weighing conflicting authorities
10%Long-ContextLCR, GDP.pdfReading and synthesizing across contracts, discovery productions, case-law packets, and PDF filings
10%Non-HallucinationAA-OmniscienceAvoiding fabricated case cites or invented statutes
5%Agentic Tool UseAutomationBench-AACompleting client support and operations workflows across business apps without breaking guardrails
Healthcare & Medical Index30%Medical & Health KnowledgeAA-OmniscienceClinical knowledge across diagnosis, pharmacology, and care pathways
25%Agentic Knowledge WorkGDPval-AA v2, AA-BriefcaseTool use, planning, and orchestrating EHR/pharmacy workflows end-to-end
15%Long-Context ReasoningMLCR-AASynthesising findings across long, fragmented patient records and claims files
10%Non-HallucinationAA-OmniscienceAvoiding fabricated drug interactions, doses, or guidelines
10%ReasoningHLEMulti-step clinical reasoning across biology and medicine
10%Agentic Tool UseAutomationBench-AACompleting patient support and operations workflows across business apps without breaking guardrails
Engineering Index35%Engineering KnowledgeAA-OmniscienceDomain recall across civil, electrical, mechanical, and other engineering disciplines
30%ReasoningHLE, CritPtMulti-step quantitative reasoning for derivations, sizing calculations, and design trade-offs
20%Agentic Knowledge WorkGDPval-AA v2, AA-BriefcaseTool use, planning, and multi-step execution of engineering deliverables
15%Agentic Terminal UseTerminal-Bench v4.0Operating real terminal environments for builds, scripts, system administration, and debugging
Economics Index35%Economics KnowledgeAA-OmniscienceRecall across micro and macroeconomics, public finance, and markets
35%ReasoningHLEMulti-step quantitative and analytic reasoning for modeling, estimation, and inference
25%Agentic Knowledge WorkGDPval-AA v2, AA-BriefcaseTool use, planning, and multi-step execution of analytical deliverables
5%Long-Context ReasoningLCRReading and reasoning across long reports, datasets, and research notes