Announcing Harvey LAB-AA v1.1: adding hallucination checks to raise the bar for agentic legal work
We collaborated with Harvey on Harvey LAB-AA v1.1, which updates our scoring methodology for the Legal Agent Benchmark (LAB) for AI agents doing real-world, agentic, legal work. Every deliverable is now checked for hallucinations against its source documents, a three-judge panel grades every rubric criterion, and the new headline metric, Hallucination-Gated All-Pass Rate, only credits a task when the deliverables satisfy every rubric criterion and contain no material hallucinations.
Every model works through 120 private legal tasks built by the team at Harvey, from corporate M&A and capital markets to tax, litigation and bankruptcy. This is the first step in enhancing the methodology for Harvey LAB-AA. In future updates, we're working with Harvey to better account for the full set of factors lawyers value, including usability features like style and tone.
Score
Harvey LAB-AA v1.1: Hallucination-Gated All-Pass Rate
Grok 4.7 (xhigh) leads Harvey LAB-AA v1.1 with a 9.4% Hallucination-Gated All-Pass Rate, narrowly ahead of Muse Spark 1.3 (max) at 8.9% and GPT-6 Astra (max) at 8.6%. GPT-6.1 Sol (max) follows at 6.9%, then Claude Fable 5.1 (max, with fallback) at 6.4%, Kimi K3 (max) at 5.3% and Claude Opus 5.5 (max, with fallback) at 4.2%. The Claude models ran with Anthropic's fallback enabled, but never used it.
Accounting for hallucinations changes the rankings on Harvey LAB-AA. Without the hallucination gate, Muse Spark 1.3 would lead clearly with a 26.7% All-Pass Rate, but two thirds of those passes contain a material hallucination. GPT-6 Astra keeps almost all of its passes (8.9% to 8.6%) and moves from joint 10th on all-pass to 3rd on Hallucination-Gated All-Pass Rate. GPT-6.1 Sol falls relatively little, from 7.5% to 6.9%. Across the models tested, more than 60% of otherwise passing results contain a material hallucination, and several models score 0% after accounting for hallucinations.
Most models complete the large majority of rubric criteria: 16 models pass 85.6-96.0% of criteria. The rubrics assess each deliverable holistically, rather than only through criteria selected for difficulty. We report the Criterion Pass Rate before the hallucination gate alongside material hallucinations per task to show rubric coverage and grounding separately.
For up-to-date results see the Harvey LAB-AA evaluation page. This article shows data as at 8 October 2026.
Changes in Harvey LAB-AA v1.1
We collaborated with Harvey to update Harvey LAB-AA to v1.1. The update changes how the benchmark is graded and scored, so v1.1 results are not directly comparable to previously published v1.0 numbers:
- Hallucination auditing. Every deliverable a model submits is now audited against the task's source documents in a two-stage check; a task with no usable submission scores zero and is not audited. A first judge flags candidate hallucinations across three categories - contradictions of the source materials, fabricated source content, and specific assertions with no support in the sources - and a second pass by the same judge re-checks each flag against that task's source documents, dismissing flags that do not hold up and classifying the rest as material or minor.
- Hallucination-Gated All-Pass Rate headline. The headline metric is now the Hallucination-Gated All-Pass Rate: a task counts only when the deliverables satisfy every rubric criterion and contain no material hallucination.
- Criterion Pass Rate and material hallucinations. We report the share of rubric criteria passed before the hallucination gate, averaged across the three judges and pooled over all criteria, alongside the mean number of material hallucinations per checked task. The scatter compares these two metrics directly.
- Three-judge rubric panel. Rubric criteria are now graded by a panel of three LLM judges - GPT-6 Sol, Grok 4.7, and Claude Opus 5.5 - with verdicts averaged across the panel, replacing the single judge used in v1.0. We reviewed these judges and found minimal self-preference for their own model families, and averaging across all three mitigates any biases that emerge.
- Dataset update. v1.1 uses the latest private dataset from Harvey (v1.1.0), which includes improvements to the tasks and criteria.
How Harvey LAB-AA differs from Harvey's LAB
Harvey LAB-AA is our independent reimplementation of Harvey's evaluation, and there are several key differences to the original version:
- Models are run on our Stirrup agent harness, enabling features such as context compaction rather than failure when reaching context limits, with simplified Artificial Analysis-authored agent and judge prompts
- We do not include Harvey's custom tools and document-generation skill scripts (e.g. pptx, docx), instead providing a simple code execution tool to reflect raw model ability
- Deliverables must match the exact filename specified, rather than fuzzy matching when models produce incorrect filenames
Hallucinations
In human preference studies run by Harvey, human experts flagged hallucinations as a primary determining factor in which response they prefer out of two otherwise comprehensive answers. Our hallucination check focuses on errors that could materially affect real legal work, and takes a conservative approach to flagging them. It distinguishes unsupported task-specific claims from general legal knowledge and places the latter out of scope, so models are not penalized for drawing on case law or statutes beyond the source files.
Every deliverable a model submits is audited for hallucinations against the task's source documents. A task with no usable submission scores zero and is not audited nor used in the hallucinations per task calculation. A hallucination is a claim in the work product that the sources do not support, and each falls into one of three categories:
- A contradiction of the source materials.
- Fabricated source content.
- A specific assertion with no support in the record.
Each hallucination is graded material or minor. A material hallucination would mislead a reader on a substantive point, such as a wrong contractually required date. A minor hallucination is a real error that is unlikely to meaningfully affect the legal interpretation of a deliverable. Only material hallucinations affect scoring: having one or more zeroes a task's score in the Hallucination-Gated All-Pass Rate. Criterion Pass Rate measures rubric coverage before the hallucination gate. Material and minor hallucination counts are reported per task the check ran on; minor hallucinations do not affect scores.
How the hallucination check works
Every deliverable is checked against the task's source documents in three steps.
Find possible hallucinations
The judge compares the submission with the task sources and flags statements for review.
Task source · draft partnership return excerpt
Scheduled principal amortization of $6,000,000 was made during 2023, reducing the outstanding balance from $220,000,000 (beginning of year, inclusive of $6,000,000 classified as current) to $214,000,000 (end of year, inclusive of $6,000,000 classified as current in Line 15).
The submission · discussion list
“debt schedule $220M term loan + $214M notes”
The same submission · exposure schedule
“$6,400,000 *293/365 $5,136,986”
Separate submission, same task · issues memo
“The $960,000 overpayment was credited to 2024”
The rubric and hallucination check assess different aspects of the submission.
Illustrative scoring example
2 of 3 judges passed every criterion
27 of 30 judge verdicts were passes; unaffected by hallucinations
The GPT-6 models hallucinate least: GPT-6 Astra averages 0.03 material hallucinations per task (4 hallucinations across all 120 tasks) and GPT-6 Sol 0.07 (8 hallucinations across all 120 tasks), while Gemini 3.8 Flash averages the most in the launch set, at 13.96 per task. Completing the criteria and not hallucinating are different skills: Muse Spark 1.3 passes the most criteria (96.0%) but averages 1.68 material hallucinations per task, against 0.03 for GPT-6 Astra. Every open-weights model averages at least 2.09 material hallucinations per task. A model's hallucination rate depends far more on the model than on the practice area.
Harvey LAB-AA v1.1: Hallucinations per Task
How we chose the hallucination checker
We use GPT-6 Sol (high) for both passes of the hallucination check, separately from the three-judge rubric panel. We compared six candidates for the hallucination judge: GPT-6 Sol, GPT-6 Luna, Grok 4.7, Claude Opus 5.5, Claude Sonnet 5.5 and Gemini 3.8 Flash, each at high reasoning effort, on the same 20 tasks and deliverables from eight evaluated models.
Across this subset of tasks, GPT-6 Sol and GPT-6 Luna generally identified more material hallucinations, while Claude Sonnet 5.5 and Gemini 3.8 Flash identified far fewer. Claude Opus 5.5 fell between Grok 4.7 and Claude Sonnet 5.5 in every model row. In total, Opus upheld 99 material hallucinations, compared with 219 for Grok, 57 for Sonnet and 470 for GPT-6 Sol. GPT-6 Sol identified more material hallucinations despite the check's conservative approach to flagging errors, which excludes general legal knowledge not contained in the source documents.
All six checkers found no material hallucinations in GPT-6 Astra's outputs, while GPT-6 Sol ranged from 0 to 0.20 per task. Both were among the models with the fewest material hallucinations under every checker.
Harvey LAB-AA v1.1: Hallucination Judge Model Comparison
| MODEL BEING CHECKED | HALLUCINATION JUDGE | |||||
|---|---|---|---|---|---|---|
| GPT-6 Sol(high) | GPT-6 Luna(high) | Grok 4.7(high) | Claude Opus 5.5(high) | Claude Sonnet 5.5(high) | Gemini 3.8 Flash(high) | |
| GPT-6 Astra (max) | ||||||
| GPT-6 Sol (max) | ||||||
| Grok 4.7 (xhigh) | ||||||
| Claude Opus 5.5 (max with fallback) | ||||||
| Claude Fable 5.1 (max with fallback) | ||||||
| Muse Spark 1.3 (max) | ||||||
| Kimi K3 (max) | ||||||
| Gemini 3.8 Flash (high) | ||||||
| Total material hallucinations per task | 23.50 | 19.35 | 10.95 | 4.95 | 2.85 | 1.40 |
Hallucination-Gated Near-Pass Rate
Harvey LAB-AA v1.1: Hallucination-Gated Near-Pass Rate
Cost
Harvey LAB-AA v1.1: Cost per Task
Token Usage
Harvey LAB-AA v1.1: Output Tokens per Task
Speed
Harvey LAB-AA v1.1: Time per Task
Turns
Harvey LAB-AA v1.1: Average Turns per Task
Model Size (Open Weights Models Only)
Harvey LAB-AA v1.1: Hallucination-Gated All-Pass Rate vs. Compute Proxy
Score vs. Release Date
Harvey LAB-AA v1.1: Hallucination-Gated All-Pass Rate vs. Release Date
Example Tasks & Submissions
Browse representative Harvey LAB tasks from the public task set, the reference files each model was given, and the deliverables it produced.
Instructions
Review the attached acquisition data room contracts and internal memo for change of control and assignment provisions, and prepare a comprehensive deal team report.
Output: coc-analysis-report.docx
Deliverables
Expected outputs the model must produce
- coc-analysis-report.docxA comprehensive deal team report analyzing change of control and assignment provisions across the target’s material contracts.
Reference files
Provided to the model
Model submissions
Deliverables produced by each model
Harvey LAB-AA resources
- The leaderboard and full results live on the Harvey LAB-AA evaluation page, updated as new models are released
- The methodology page documents the full implementation, including the agent and rubric grading prompts
- Harvey's original LAB announcement introduces the benchmark and its design
- A public set of representative tasks is available on GitHub
- Harvey LAB-AA runs on Stirrup, our open-source agent framework
Read the latest

Introducing trusted-access models to the Artificial Analysis Cyber Index
GPT-6 Sol (Daybreak Blue) now leads the Cyber Index
October 8, 2026

Anthropic has released Claude Haiku 5.5
Claude Haiku 5.5 scores 43 on the Artificial Analysis Intelligence Index, up 26 points one year after the last Haiku release
October 7, 2026

Mistral has released Mistral Large 4, making France home to the most intelligent model outside the US and China
Mistral has released Mistral Large 4, scoring 38 on the Artificial Analysis Intelligence Index; France is back to having the most intelligent model from outside the US and China
October 6, 2026