Announcing Harvey LAB-AA: evaluating AI agents on real-world legal work
Harvey LAB-AA (Legal Agent Benchmark) is our implementation of Harvey's new agentic legal benchmark, evaluating language models on real-world legal work across 24 practice areas.
Models are tested on a private set of 120 legal tasks built by the team at Harvey, spanning practice areas from corporate M&A and capital markets to tax, litigation, and bankruptcy. Models work to create the legal outputs specified in each task, and each task is graded against a rubric of binary criteria. The primary metric we present is the all-pass rate: the share of tasks where all criteria in the rubric are satisfied, reflecting the high standard of real-world professional legal deliverables.
Score
Harvey LAB-AA: All-pass Rate
Claude Fable 5 (max, with Opus 4.8 fallback) leads Harvey LAB-AA with a 14.2% all-pass rate, after falling back to Claude Opus 4.8 on only one task. This is almost double the next best models, Claude Opus 4.8 (max) and GLM-5.2 (max), which tie at 7.5%, followed by MiniMax-M3 at 6.7% and Claude Sonnet 5 at 5.0%.
Frontier legal work is far from solved: most models pass a majority of individual rubric criteria but fully satisfy the requirements of very few tasks. The best model still leaves ~86% of professional legal deliverables incomplete, and 13 of the 28 models evaluated at launch fully pass zero tasks. Only four models score above 90% on criterion pass rate: Claude Fable 5 (93.6%), Claude Opus 4.8 (91.1%), GLM-5.2 (91.0%), and Claude Sonnet 5 (90.1%).
For up-to-date results see the Harvey LAB-AA evaluation page. Charts show data as at 7 July 2026.
Cost
Harvey LAB-AA: Cost per Task
Token Usage
Harvey LAB-AA: Output Tokens per Task
Speed
Harvey LAB-AA: Time per Task
Turns
Harvey LAB-AA Benchmark Leaderboard: Average Turns per Task
Score vs. Release Date
Harvey LAB-AA: All-pass Rate vs. Release Date
Example Tasks & Submissions
Browse representative Harvey LAB tasks from the public task set, the reference files each model was given, and the deliverables it produced.
Instructions
Review the attached acquisition data room contracts and internal memo for change of control and assignment provisions, and prepare a comprehensive deal team report.
Output: coc-analysis-report.docx
Deliverables
Expected outputs the model must produce
- coc-analysis-report.docxA comprehensive deal team report analyzing change of control and assignment provisions across the target’s material contracts.
Reference files
Provided to the model
Model submissions
Deliverables produced by each model
How Harvey LAB-AA differs from Harvey's LAB
Harvey LAB-AA is our independent reimplementation of Harvey's evaluation, and there are several key differences to the original version:
- Models are run on our Stirrup agent harness, enabling features such as context compaction rather than failure when reaching context limits, with simplified Artificial Analysis-authored agent and judge prompts
- We do not include Harvey's custom tools and document-generation skill scripts (e.g. pptx, docx), instead providing a simple code execution tool to reflect raw model ability
- Deliverables must match the exact filename specified, rather than fuzzy matching when models produce incorrect filenames
- Grading uses a single Gemini 3.1 Pro judge, tested to be well-calibrated against a frontier panel
Harvey LAB-AA resources
- The leaderboard and full results live on the Harvey LAB-AA evaluation page, updated as new models are released
- The methodology page documents the full implementation, including the agent and grading prompts
- Harvey's original LAB announcement introduces the benchmark and its design
- A public set of representative tasks is available on GitHub
- Harvey LAB-AA runs on Stirrup, our open-source agent framework
Read the latest

Announcing the Artificial Analysis Cyber Index Alliance
The Artificial Analysis Cyber Index Alliance brings together industry partners to set a new standard for evaluating how AI models perform on enterprise cyber defense tasks. The Alliance launches alongside the Artificial Analysis Cyber Index, which combines three partner-contributed and open benchmarks to evaluate how well agents find and fix vulnerabilities.
September 28, 2026

Claude Sonnet 5.5 reaches #2 on the Artificial Analysis Intelligence Index
Anthropic's new Sonnet model scores 56, just 2 points behind Opus 5.5 (max), but at the highest Output Tokens per Task we have measured
September 28, 2026

GPT-6 Sol and Luna push the cost efficiency frontier
GPT-6 Sol and Luna push the cost efficiency frontier by halving cost relative to GPT-5.6 Sol and Luna. Intelligence Index and Coding Agent Index scores remain level with GPT-5.6, with progress in some evaluations and regressions in others
September 22, 2026