AI Model Evaluations
Independent evaluations of AI models across reasoning, knowledge, coding and agentic capabilities. Filter by the skills and knowledge domains.
Artificial Analysis Intelligence Index v4.3.2
A composite benchmark aggregating ten challenging evaluations to provide a holistic measure of AI capabilities across mathematics, science, coding, and reasoning.
Artificial Analysis Openness Index
A composite measure providing an industry standard to communicate model openness for users and developers.
AA-Briefcase v1.1: Agentic Knowledge Work Benchmark
A private evaluation developed by Artificial Analysis for frontier agentic capability in long-horizon knowledge work, testing agents on realistic business workflows that require deliverables such as spreadsheets, presentations, and memos.
GDPval-AA v2.1 Leaderboard
GDPval-AA v2.1 is Artificial Analysis' evaluation framework for OpenAI's GDPval dataset. It tests AI models on real-world tasks across 44 occupations and 9 major industries. Models are given shell access and web browsing capabilities in an agentic loop via Stirrup to solve tasks, with Elo ratings derived from blind pairwise comparisons.
APEX-Agents-AA Benchmark Leaderboard
Artificial Analysis' implementation of the APEX-Agents benchmark, testing AI agents on long-horizon, cross-application tasks in professional-services environments with realistic application tooling.
AA-AnalystAgent Benchmark Leaderboard
Artificial Analysis' data analysis benchmark, testing AI agents on their ability to work with spreadsheets and documents to answer quantitative questions a Business Analyst or Data Analyst would face day-to-day.
AutomationBench-AA: Agentic SaaS Workflow Benchmark
A benchmark measuring agentic task completion across simulated SaaS application environments, scoring the share of each task's objectives completed without guardrail violations.
Harvey LAB-AA Benchmark Leaderboard
Artificial Analysis' implementation of Harvey's Legal Agent Benchmark (LAB), testing AI agents on real-world legal work from Harvey's dataset of 120 private tasks spanning 24 legal practice areas. The agent reads case documents in a sandbox and produces legal deliverables (e.g., memos, disclosure schedules, deposition summaries), graded criterion-by-criterion by a single LLM rubric judge.
GDP.pdf Benchmark Leaderboard
Artificial Analysis' implementation of Surge AI's GDP.pdf benchmark, testing whether language models can reason over long, real-world professional documents and satisfy detailed task-specific criteria.
EnterpriseOps-Gym-AA Benchmark Leaderboard
Artificial Analysis' independent implementation of ServiceNow's EnterpriseOps-Gym, an agentic benchmark testing whether LLM agents can complete stateful, multi-step enterprise workflows across eight business domains via live tool use, graded on the final state of the underlying databases.
Terminal-Bench 4.0 Benchmark Leaderboard
A harder 66-task benchmark of complex terminal work across software, machine learning, science, operations, security, hardware, and media, with recalibrated compute and time allowances and improved instructions, environments, and verifiers.
Terminal-Bench-Science 0.1 Benchmark Leaderboard
A 70-task benchmark of research workflows authored and reviewed by domain experts across the life, physical, earth, mathematical, and engineering sciences, each completed in a terminal and checked by its own set of tests.
Artificial Analysis Long Context Reasoning Benchmark Leaderboard
A challenging benchmark measuring language models' ability to extract, reason about, and synthesize information from long-form documents ranging from 10k to 100k tokens (measured using the cl100k_base tokenizer).
AA-Omniscience: Knowledge and Hallucination Benchmark
A benchmark measuring factual recall and hallucination across various economically relevant domains.
SciCode Benchmark Leaderboard
A scientist-curated coding benchmark featuring 288 test set subproblems from 80 laboratory problems across 16 scientific disciplines.
Humanity's Last Exam Benchmark Leaderboard
A frontier-level benchmark with 2,500 expert-vetted questions across mathematics, sciences, and humanities, designed to be the final closed-ended academic evaluation.
CritPt Benchmark Leaderboard
A benchmark designed to test LLMs on research-level physics reasoning tasks, featuring 71 composite research challenges.
GPQA Diamond Benchmark Leaderboard
The most challenging 198 questions from GPQA, where PhD experts achieve 65% accuracy but skilled non-experts only reach 34% despite web access.
ITBench-AA Benchmark Leaderboard
Artificial Analysis' implementation of IBM's ITBench benchmark, testing AI agents on Kubernetes incident root-cause analysis from offline incident snapshots. The agent inspects alerts, events, traces, and topology and identifies the contributing-factor entities (deployments, pods, namespaces, network policies, etc.) responsible for the failure.
MMMU-Pro Benchmark Leaderboard
An enhanced MMMU benchmark that eliminates shortcuts and guessing strategies to more rigorously test multimodal models across 30 academic disciplines.
IFBench Benchmark Leaderboard
A benchmark evaluating precise instruction-following generalization on 58 diverse, verifiable out-of-domain constraints that test models' ability to follow specific output requirements.
Medical Long Context Reasoning (MLCR-AA)
An open benchmark from Wisedocs measuring how well models reason over long, fragmented medical records, performing the multi-document synthesis claims professionals rely on when reviewing insurance and healthcare cases.
𝜏³-Banking Benchmark Leaderboard
A fintech customer-support benchmark from the 𝜏-Knowledge framework that tests whether agents can navigate a large unstructured knowledge base and execute multi-step tool calls to resolve realistic banking workflows.
Terminal-Bench 2.1 Benchmark Leaderboard
A verified refresh of Terminal-Bench 2.0 — 89 curated tasks across software engineering, system administration, data processing, model training, and security, with environment and instruction fixes so scores reflect agent capability rather than environment gaps.
Terminal-Bench Hard Benchmark Leaderboard
An agentic benchmark evaluating AI capabilities in terminal environments through software engineering, system administration, and data processing tasks.
𝜏²-Bench Telecom Benchmark Leaderboard
A dual-control conversational AI benchmark simulating technical support scenarios where both agent and user must coordinate actions to resolve telecom service issues.
MMLU-Pro Benchmark Leaderboard
An enhanced version of MMLU with 12,000 graduate-level questions across 14 subject areas, featuring ten answer options and deeper reasoning requirements.
LiveCodeBench Benchmark Leaderboard
A contamination-free coding benchmark that continuously harvests fresh competitive programming problems from LeetCode, AtCoder, and CodeForces, evaluating code generation, self-repair, and execution.
MATH-500 Benchmark Leaderboard
A 500-problem subset from the MATH dataset, featuring competition-level mathematics across six domains including algebra, geometry, and number theory.
AIME 2025 Benchmark Leaderboard
All 30 problems from the 2025 American Invitational Mathematics Examination, testing olympiad-level mathematical reasoning with integer answers from 000-999.
Global-MMLU-Lite Benchmark Leaderboard
A lightweight, multilingual version of MMLU, designed to evaluate knowledge and reasoning skills across a diverse range of languages and cultural contexts.