All evaluations

Harvey LAB-AA v1.1 Benchmark Leaderboard

Artificial Analysis' implementation of Harvey's Legal Agent Benchmark (LAB), testing AI agents on real-world legal work from Harvey's dataset of 120 private tasks spanning 24 legal practice areas. v1.1, built in collaboration with Harvey, checks every deliverable against the task's source documents for hallucinations, so work that passes every rubric criterion but contains a material hallucination receives no credit. A panel of three judges grades each deliverable criterion-by-criterion.
See example tasks

What's new in v1.1

v1.1 supersedes v1.0, and scores are not comparable across versions. Harvey's studies with legal professionals found hallucinations to be a primary factor in which outputs lawyers prefer, so we collaborated with Harvey to add a hallucination check to scoring.

  • Harvey's latest private dataset, with improvements to the tasks and criteria
  • A new hallucination check of every deliverable against the task's source documents, with material hallucinations penalized in scoring
  • Judges updated to GPT-6 Sol, Grok 4.7 and Claude Opus 5.5. We found minimal self-preference and average all three to mitigate any bias

Harvey LAB-AA v1.1

Grok 4.7 (Xhigh) scores the highest on Harvey LAB-AA v1.1 with a score of 9.4%, followed by Muse Spark 1.3 (Max) with a score of 8.9% and GPT-6 Astra (Max) with a score of 8.6%

Score

Harvey LAB-AA v1.1: Hallucination-Gated All-Pass Rate

Average over tasks of the share of the judge panel finding every rubric criterion passed, with a task zeroed on any material hallucination · Independently benchmarked by Artificial Analysis

Hallucinations

Harvey LAB-AA v1.1: Hallucinations per Task

Upheld hallucination flags per task, split by severity · Material flags would mislead a reader on a substantive point, minor flags are real errors unlikely to affect the legal interpretation, and only material flags affect the headline score · Lower is better

Hallucination-Gated Near-Pass Rate

Harvey LAB-AA v1.1: Hallucination-Gated Near-Pass Rate

Average over tasks of the share of the judge panel that missed 0, <=1 and <=2 criteria · Tasks with at least 1 material hallucination count at 0% · Higher is better
Sort by criteria missed

Score Comparisons

Harvey LAB-AA v1.1: Hallucination-Gated All-Pass Rate vs. Artificial Analysis Intelligence Index

Hallucination-Gated All-Pass Rate is the average share of the judge panel passing every rubric criterion, with a task zeroed on any material hallucination · Artificial Analysis Intelligence Index measures general capability across the full index eval suite · Higher is better on both axes
Most attractive quadrant

Cost

Harvey LAB-AA v1.1: Cost per Task

Average cost per task (USD), broken down by input, cache hit, cache write, reasoning, and answer tokens

Token Usage

Harvey LAB-AA v1.1: Output Tokens per Task

Output tokens used to run one task, broken down by reasoning and answer tokens

Speed

Harvey LAB-AA v1.1: Time per Task

Weighted average decode time (minutes) per task; excludes TTFT and overhead time · Lower is better

Turns

Harvey LAB-AA v1.1: Average Turns per Task

Average number of model turns per Harvey LAB-AA v1.1 task · Lower is better

Score by Practice Area

Harvey LAB-AA v1.1: Hallucination-Gated All-Pass Rate by Practice Area (Normalized)

Harvey LAB-AA v1.1 Hallucination-Gated All-Pass Rate by legal practice area · Five tasks per practice area · Scores are normalized per practice area across the models shown, where green represents the highest score for that practice area and red represents the lowest score for that practice area

Model Size (Open Weights Models Only)

Harvey LAB-AA v1.1: Hallucination-Gated All-Pass Rate vs. Total Parameters

Harvey LAB-AA v1.1 Hallucination-Gated All-Pass Rate · Size in parameters (billions) · Open weights models only
Most attractive quadrant
Pareto line

Score vs. Release Date

Harvey LAB-AA v1.1: Hallucination-Gated All-Pass Rate vs. Release Date

Most attractive region

Example Tasks & Submissions

Browse representative Harvey LAB tasks from the public task set, the reference files each model was given, and the deliverables it produced.

Mergers & Acquisitions

Instructions

Review the attached acquisition data room contracts and internal memo for change of control and assignment provisions, and prepare a comprehensive deal team report.

Output: coc-analysis-report.docx

Deliverables

Expected outputs the model must produce

  • coc-analysis-report.docxA comprehensive deal team report analyzing change of control and assignment provisions across the target’s material contracts.

Reference files

Provided to the model

Open
Open
Open
Open
Open
Open
Open
Open
Open
Open
Open
Open
Open
Open
Open
Open
Open
Open
Open

Model submissions

Deliverables produced by each model

Claude Opus 5.5 (max with fallback) - coc-analysis-report.docx
Open

Leaderboard

Creator
Name
Hallucination-Gated All-Pass Rate
Criterion Pass Rate
Material Hallucinations per Task
Cost per Task
Release Date
1
SpaceXAI logoSpaceXAI
Grok 4.7 (Xhigh)9.4%92.8%0.74$9.49Sep 2026
2
Meta logoMeta
Muse Spark 1.3 (Max)8.9%96.0%1.68$4.21Sep 2026
3
OpenAI logoOpenAI
GPT-6 Astra (Max)8.6%90.6%0.03$13.67Sep 2026
4
OpenAI logoOpenAI
GPT-6.1 Sol (Max)6.9%90.6%0.09$2.48Sep 2026
5
Anthropic logoAnthropic
Claude Fable 5.1 (Max, Default Fallback)6.4%92.4%1.32$21.73Sep 2026
6
Kimi logoKimi
Kimi K3 (Max)5.3%93.0%2.09$6.19Jul 2026
7
Anthropic logoAnthropic
Claude Opus 5.5 (Max, Default Fallback)4.2%90.0%0.59$17.93Sep 2026
8
OpenAI logoOpenAI
GPT-6 Sol (Max)3.6%86.9%0.07$3.33Sep 2026
9
OpenAI logoOpenAI
GPT-6 Luna (Max)3.3%86.0%0.38$0.22Sep 2026
10
Mistral logoMistral
Mistral Large 4 Preview3.1%93.0%2.11$2.54Oct 2026
11
Anthropic logoAnthropic
Claude Sonnet 5.5 (Max, Default Fallback)2.8%92.1%1.03$14.03Sep 2026
12
Anthropic logoAnthropic
Claude Haiku 5.5 (Xhigh)2.8%88.0%1.61$0.43Oct 2026
13
DeepSeek logoDeepSeek
DeepSeek V4.1 Flash (Max)1.7%92.5%2.58$0.59Sep 2026
14
Anthropic logoAnthropic
Claude Haiku 5.5 (Low)1.7%87.5%1.91$0.17Oct 2026
15
Anthropic logoAnthropic
Claude Haiku 5.5 (High)1.1%87.9%1.68$0.32Oct 2026

Explore Evaluations

Artificial Analysis Intelligence Index v4.3.2Artificial Analysis Intelligence Index v4.3.2

A composite benchmark aggregating ten challenging evaluations to provide a holistic measure of AI capabilities across mathematics, science, coding, and reasoning.

Artificial Analysis Openness IndexArtificial Analysis Openness Index

A composite measure providing an industry standard to communicate model openness for users and developers.

AA-Briefcase v1.1: Agentic Knowledge Work BenchmarkAA-Briefcase v1.1: Agentic Knowledge Work Benchmark

A private evaluation developed by Artificial Analysis for frontier agentic capability in long-horizon knowledge work, testing agents on realistic business workflows that require deliverables such as spreadsheets, presentations, and memos.

GDPval-AA v2.1 LeaderboardGDPval-AA v2.1 Leaderboard

GDPval-AA v2.1 is Artificial Analysis' evaluation framework for OpenAI's GDPval dataset. It tests AI models on real-world tasks across 44 occupations and 9 major industries. Models are given shell access and web browsing capabilities in an agentic loop via Stirrup to solve tasks, with Elo ratings derived from blind pairwise comparisons.

APEX-Agents-AA Benchmark LeaderboardAPEX-Agents-AA Benchmark Leaderboard

Artificial Analysis' implementation of the APEX-Agents benchmark, testing AI agents on long-horizon, cross-application tasks in professional-services environments with realistic application tooling.

AA-AnalystAgent Benchmark LeaderboardAA-AnalystAgent Benchmark Leaderboard

Artificial Analysis' data analysis benchmark, testing AI agents on their ability to work with spreadsheets and documents to answer quantitative questions a Business Analyst or Data Analyst would face day-to-day.

AutomationBench-AA: Agentic SaaS Workflow BenchmarkAutomationBench-AA: Agentic SaaS Workflow Benchmark

A benchmark measuring agentic task completion across simulated SaaS application environments, scoring the share of each task's objectives completed without guardrail violations.

Harvey LAB-AA v1.1 Benchmark LeaderboardHarvey LAB-AA v1.1 Benchmark Leaderboard

Artificial Analysis' implementation of Harvey's Legal Agent Benchmark (LAB), testing AI agents on real-world legal work from Harvey's dataset of 120 private tasks spanning 24 legal practice areas. v1.1, built in collaboration with Harvey, checks every deliverable against the task's source documents for hallucinations, so work that passes every rubric criterion but contains a material hallucination receives no credit. A panel of three judges grades each deliverable criterion-by-criterion.

GDP.pdf Benchmark LeaderboardGDP.pdf Benchmark Leaderboard

Artificial Analysis' implementation of Surge AI's GDP.pdf benchmark, testing whether language models can reason over long, real-world professional documents and satisfy detailed task-specific criteria.

EnterpriseOps-Gym-AA Benchmark LeaderboardEnterpriseOps-Gym-AA Benchmark Leaderboard

Artificial Analysis' independent implementation of ServiceNow's EnterpriseOps-Gym, an agentic benchmark testing whether LLM agents can complete stateful, multi-step enterprise workflows across eight business domains via live tool use, graded on the final state of the underlying databases.

Terminal-Bench 4.0 Benchmark LeaderboardTerminal-Bench 4.0 Benchmark Leaderboard

A harder 66-task benchmark of complex terminal work across software, machine learning, science, operations, security, hardware, and media, with recalibrated compute and time allowances and improved instructions, environments, and verifiers.

Terminal-Bench-Science 0.1 Benchmark LeaderboardTerminal-Bench-Science 0.1 Benchmark Leaderboard

A 70-task benchmark of research workflows authored and reviewed by domain experts across the life, physical, earth, mathematical, and engineering sciences, each completed in a terminal and checked by its own set of tests.

Artificial Analysis Long Context Reasoning Benchmark LeaderboardArtificial Analysis Long Context Reasoning Benchmark Leaderboard

A challenging benchmark measuring language models' ability to extract, reason about, and synthesize information from long-form documents ranging from 10k to 100k tokens (measured using the cl100k_base tokenizer).

AA-Omniscience: Knowledge and Hallucination BenchmarkAA-Omniscience: Knowledge and Hallucination Benchmark

A benchmark measuring factual recall and hallucination across various economically relevant domains.

SciCode Benchmark LeaderboardSciCode Benchmark Leaderboard

A scientist-curated coding benchmark featuring 288 test set subproblems from 80 laboratory problems across 16 scientific disciplines.

Humanity's Last Exam Benchmark LeaderboardHumanity's Last Exam Benchmark Leaderboard

A frontier-level benchmark with 2,500 expert-vetted questions across mathematics, sciences, and humanities, designed to be the final closed-ended academic evaluation.

CritPt Benchmark LeaderboardCritPt Benchmark Leaderboard

A benchmark designed to test LLMs on research-level physics reasoning tasks, featuring 71 composite research challenges.

GPQA Diamond Benchmark Leaderboard

The most challenging 198 questions from GPQA, where PhD experts achieve 65% accuracy but skilled non-experts only reach 34% despite web access.

ITBench-AA Benchmark LeaderboardITBench-AA Benchmark Leaderboard

Artificial Analysis' implementation of IBM's ITBench benchmark, testing AI agents on Kubernetes incident root-cause analysis from offline incident snapshots. The agent inspects alerts, events, traces, and topology and identifies the contributing-factor entities (deployments, pods, namespaces, network policies, etc.) responsible for the failure.

MMMU-Pro Benchmark LeaderboardMMMU-Pro Benchmark Leaderboard

An enhanced MMMU benchmark that eliminates shortcuts and guessing strategies to more rigorously test multimodal models across 30 academic disciplines.

IFBench Benchmark LeaderboardIFBench Benchmark Leaderboard

A benchmark evaluating precise instruction-following generalization on 58 diverse, verifiable out-of-domain constraints that test models' ability to follow specific output requirements.

Medical Long Context Reasoning (MLCR-AA)Medical Long Context Reasoning (MLCR-AA)

An open benchmark from Wisedocs measuring how well models reason over long, fragmented medical records, performing the multi-document synthesis claims professionals rely on when reviewing insurance and healthcare cases.

𝜏³-Banking Benchmark Leaderboard𝜏³-Banking Benchmark Leaderboard

A fintech customer-support benchmark from the 𝜏-Knowledge framework that tests whether agents can navigate a large unstructured knowledge base and execute multi-step tool calls to resolve realistic banking workflows.

Terminal-Bench 2.1 Benchmark LeaderboardTerminal-Bench 2.1 Benchmark Leaderboard

A verified refresh of Terminal-Bench 2.0 — 89 curated tasks across software engineering, system administration, data processing, model training, and security, with environment and instruction fixes so scores reflect agent capability rather than environment gaps.

Terminal-Bench Hard Benchmark LeaderboardTerminal-Bench Hard Benchmark Leaderboard

An agentic benchmark evaluating AI capabilities in terminal environments through software engineering, system administration, and data processing tasks.

𝜏²-Bench Telecom Benchmark Leaderboard𝜏²-Bench Telecom Benchmark Leaderboard

A dual-control conversational AI benchmark simulating technical support scenarios where both agent and user must coordinate actions to resolve telecom service issues.

MMLU-Pro Benchmark LeaderboardMMLU-Pro Benchmark Leaderboard

An enhanced version of MMLU with 12,000 graduate-level questions across 14 subject areas, featuring ten answer options and deeper reasoning requirements.

LiveCodeBench Benchmark LeaderboardLiveCodeBench Benchmark Leaderboard

A contamination-free coding benchmark that continuously harvests fresh competitive programming problems from LeetCode, AtCoder, and CodeForces, evaluating code generation, self-repair, and execution.

MATH-500 Benchmark LeaderboardMATH-500 Benchmark Leaderboard

A 500-problem subset from the MATH dataset, featuring competition-level mathematics across six domains including algebra, geometry, and number theory.

AIME 2025 Benchmark LeaderboardAIME 2025 Benchmark Leaderboard

All 30 problems from the 2025 American Invitational Mathematics Examination, testing olympiad-level mathematical reasoning with integer answers from 000-999.

Global-MMLU-Lite Benchmark LeaderboardGlobal-MMLU-Lite Benchmark Leaderboard

A lightweight, multilingual version of MMLU, designed to evaluate knowledge and reasoning skills across a diverse range of languages and cultural contexts.