Artificial Analysis 역량 지수 방법론

개요

Artificial Analysis 역량 지수는 법률, 의료, 금융 업무처럼 특정 전문 사용 사례에서 모델이 얼마나 뛰어난 성능을 보이는지 측정합니다.

지수를 구성하는 모든 벤치마크를 독립적으로 실행한 후 그 점수를 지수로 결합합니다. 모든 평가 지표와 마찬가지로 역량 지수에는 한계가 있으며 모든 사용 사례에 직접 적용되지는 않을 수 있지만, 특정 분야에서 중요한 업무를 기준으로 모델을 비교하는 데 유용한 종합 정보를 제공합니다.

각 기반 벤치마크의 실행 및 채점 방법은 지능 벤치마킹 방법론에 설명되어 있습니다.

지수 구성 분석

각 지수는 법률, 의료 및 보건, 금융 및 회계 같은 하나의 직업 또는 분야를 대상으로 하며, 해당 분야의 실제 업무에서 각 역량이 나타나는 빈도에 따라 일련의 역량에 가중치를 부여합니다.

업무 가중치는 O*NET 방식의 업무 활동 분류 체계를 바탕으로 합니다. 아래 표는 각 지수의 구성 요소와 가중치를 보여 줍니다.

지수가중치역량평가설명
Finance & Accounting Index30%Business KnowledgeAA-OmniscienceDomain recall in accounting, corporate finance, economics, and investments
30%Agentic Knowledge WorkGDPval-AA v2, AA-BriefcaseTool use, planning, and multi-step task execution, such as building spreadsheets, running an analysis, or coordinating a workflow
20%ReasoningHLEMulti-step quantitative and analytic reasoning, used for sensitivity analysis, valuation, and structured problem solving
10%Agentic Tool UseAutomationBench-AACompleting finance workflows across business apps such as spreadsheets, email, and accounting tools without breaking guardrails
5%Long-ContextLCR, GDP.pdfReading and reasoning across long financial filings, deal documents, and PDF reports
5%Non-HallucinationAA-OmniscienceAvoiding fabricated figures or citations
Strategy & Ops Index30%Business KnowledgeAA-OmniscienceWorking knowledge of business processes, accounting basics, and operational concepts
35%Agentic Knowledge WorkGDPval-AA v2, AA-BriefcaseTool use, planning, and orchestrating multi-step office workflows end-to-end
30%Agentic Tool UseAutomationBench-AACompleting operations, HR, marketing, sales, and support workflows across business apps without breaking guardrails
5%Long-ContextLCR, GDP.pdfHolding context across long threads, policies, records, and PDF documents
Legal Index35%Legal KnowledgeAA-OmniscienceRecall of statutes, doctrines, and procedure across jurisdictions
25%Agentic Knowledge WorkGDPval-AA v2, AA-BriefcaseRunning matter-management workflows, drafting pipelines, and tool-augmented research
15%ReasoningHLEMulti-step argumentation, statutory interpretation, and weighing conflicting authorities
10%Long-ContextLCR, GDP.pdfReading and synthesizing across contracts, discovery productions, case-law packets, and PDF filings
10%Non-HallucinationAA-OmniscienceAvoiding fabricated case cites or invented statutes
5%Agentic Tool UseAutomationBench-AACompleting client support and operations workflows across business apps without breaking guardrails
Healthcare & Medical Index30%Medical & Health KnowledgeAA-OmniscienceClinical knowledge across diagnosis, pharmacology, and care pathways
25%Agentic Knowledge WorkGDPval-AA v2, AA-BriefcaseTool use, planning, and orchestrating EHR/pharmacy workflows end-to-end
15%Long-Context ReasoningMLCR-AASynthesising findings across long, fragmented patient records and claims files
10%Non-HallucinationAA-OmniscienceAvoiding fabricated drug interactions, doses, or guidelines
10%ReasoningHLEMulti-step clinical reasoning across biology and medicine
10%Agentic Tool UseAutomationBench-AACompleting patient support and operations workflows across business apps without breaking guardrails
Engineering Index35%Engineering KnowledgeAA-OmniscienceDomain recall across civil, electrical, mechanical, and other engineering disciplines
30%ReasoningHLE, CritPtMulti-step quantitative reasoning for derivations, sizing calculations, and design trade-offs
20%Agentic Knowledge WorkGDPval-AA v2, AA-BriefcaseTool use, planning, and multi-step execution of engineering deliverables
15%Agentic Terminal UseTerminal-Bench v4.0Operating real terminal environments for builds, scripts, system administration, and debugging
Economics Index35%Economics KnowledgeAA-OmniscienceRecall across micro and macroeconomics, public finance, and markets
35%ReasoningHLEMulti-step quantitative and analytic reasoning for modeling, estimation, and inference
25%Agentic Knowledge WorkGDPval-AA v2, AA-BriefcaseTool use, planning, and multi-step execution of analytical deliverables
5%Long-Context ReasoningLCRReading and reasoning across long reports, datasets, and research notes