All articles

October 8, 2026

Announcing Harvey LAB-AA v1.1: adding hallucination checks to raise the bar for agentic legal work

We collaborated with Harvey on Harvey LAB-AA v1.1, which updates our scoring methodology for the Legal Agent Benchmark (LAB) for AI agents doing real-world, agentic, legal work. Every deliverable is now checked for hallucinations against its source documents, a three-judge panel grades every rubric criterion, and the new headline metric, Hallucination-Gated All-Pass Rate, only credits a task when the deliverables satisfy every rubric criterion and contain no material hallucinations.
Every model works through 120 private legal tasks built by the team at Harvey, from corporate M&A and capital markets to tax, litigation and bankruptcy. This is the first step in enhancing the methodology for Harvey LAB-AA. In future updates, we're working with Harvey to better account for the full set of factors lawyers value, including usability features like style and tone.

Score

Harvey LAB-AA v1.1: Hallucination-Gated All-Pass Rate

Average over tasks of the share of the judge panel finding every rubric criterion passed, with a task zeroed on any material hallucination · Independently benchmarked by Artificial Analysis

Grok 4.7 (xhigh) leads Harvey LAB-AA v1.1 with a 9.4% Hallucination-Gated All-Pass Rate, narrowly ahead of Muse Spark 1.3 (max) at 8.9% and GPT-6 Astra (max) at 8.6%. GPT-6.1 Sol (max) follows at 6.9%, then Claude Fable 5.1 (max, with fallback) at 6.4%, Kimi K3 (max) at 5.3% and Claude Opus 5.5 (max, with fallback) at 4.2%. The Claude models ran with Anthropic's fallback enabled, but never used it.

Accounting for hallucinations changes the rankings on Harvey LAB-AA. Without the hallucination gate, Muse Spark 1.3 would lead clearly with a 26.7% All-Pass Rate, but two thirds of those passes contain a material hallucination. GPT-6 Astra keeps almost all of its passes (8.9% to 8.6%) and moves from joint 10th on all-pass to 3rd on Hallucination-Gated All-Pass Rate. GPT-6.1 Sol falls relatively little, from 7.5% to 6.9%. Across the models tested, more than 60% of otherwise passing results contain a material hallucination, and several models score 0% after accounting for hallucinations.

Most models complete the large majority of rubric criteria: 16 models pass 85.6-96.0% of criteria. The rubrics assess each deliverable holistically, rather than only through criteria selected for difficulty. We report the Criterion Pass Rate before the hallucination gate alongside material hallucinations per task to show rubric coverage and grounding separately.

For up-to-date results see the Harvey LAB-AA evaluation page. This article shows data as at 8 October 2026.

Changes in Harvey LAB-AA v1.1

We collaborated with Harvey to update Harvey LAB-AA to v1.1. The update changes how the benchmark is graded and scored, so v1.1 results are not directly comparable to previously published v1.0 numbers:

  • Hallucination auditing. Every deliverable a model submits is now audited against the task's source documents in a two-stage check; a task with no usable submission scores zero and is not audited. A first judge flags candidate hallucinations across three categories - contradictions of the source materials, fabricated source content, and specific assertions with no support in the sources - and a second pass by the same judge re-checks each flag against that task's source documents, dismissing flags that do not hold up and classifying the rest as material or minor.
  • Hallucination-Gated All-Pass Rate headline. The headline metric is now the Hallucination-Gated All-Pass Rate: a task counts only when the deliverables satisfy every rubric criterion and contain no material hallucination.
  • Criterion Pass Rate and material hallucinations. We report the share of rubric criteria passed before the hallucination gate, averaged across the three judges and pooled over all criteria, alongside the mean number of material hallucinations per checked task. The scatter compares these two metrics directly.
  • Three-judge rubric panel. Rubric criteria are now graded by a panel of three LLM judges - GPT-6 Sol, Grok 4.7, and Claude Opus 5.5 - with verdicts averaged across the panel, replacing the single judge used in v1.0. We reviewed these judges and found minimal self-preference for their own model families, and averaging across all three mitigates any biases that emerge.
  • Dataset update. v1.1 uses the latest private dataset from Harvey (v1.1.0), which includes improvements to the tasks and criteria.

How Harvey LAB-AA differs from Harvey's LAB

Harvey LAB-AA is our independent reimplementation of Harvey's evaluation, and there are several key differences to the original version:

  • Models are run on our Stirrup agent harness, enabling features such as context compaction rather than failure when reaching context limits, with simplified Artificial Analysis-authored agent and judge prompts
  • We do not include Harvey's custom tools and document-generation skill scripts (e.g. pptx, docx), instead providing a simple code execution tool to reflect raw model ability
  • Deliverables must match the exact filename specified, rather than fuzzy matching when models produce incorrect filenames

Hallucinations

In human preference studies run by Harvey, human experts flagged hallucinations as a primary determining factor in which response they prefer out of two otherwise comprehensive answers. Our hallucination check focuses on errors that could materially affect real legal work, and takes a conservative approach to flagging them. It distinguishes unsupported task-specific claims from general legal knowledge and places the latter out of scope, so models are not penalized for drawing on case law or statutes beyond the source files.

Every deliverable a model submits is audited for hallucinations against the task's source documents. A task with no usable submission scores zero and is not audited nor used in the hallucinations per task calculation. A hallucination is a claim in the work product that the sources do not support, and each falls into one of three categories:

  • A contradiction of the source materials.
  • Fabricated source content.
  • A specific assertion with no support in the record.

Each hallucination is graded material or minor. A material hallucination would mislead a reader on a substantive point, such as a wrong contractually required date. A minor hallucination is a real error that is unlikely to meaningfully affect the legal interpretation of a deliverable. Only material hallucinations affect scoring: having one or more zeroes a task's score in the Hallucination-Gated All-Pass Rate. Criterion Pass Rate measures rubric coverage before the hallucination gate. Material and minor hallucination counts are reported per task the check ran on; minor hallucinations do not affect scores.

How the hallucination check works

Every deliverable is checked against the task's source documents in three steps.

Find possible hallucinations

The judge compares the submission with the task sources and flags statements for review.

Tax compliance review · submission excerpts3 examples

Task source · draft partnership return excerpt

Scheduled principal amortization of $6,000,000 was made during 2023, reducing the outstanding balance from $220,000,000 (beginning of year, inclusive of $6,000,000 classified as current) to $214,000,000 (end of year, inclusive of $6,000,000 classified as current in Line 15).

The submission · discussion list

“debt schedule $220M term loan + $214M notes”

Possible double counting of debt

The same submission · exposure schedule

“$6,400,000 *293/365 $5,136,986”

Possible error in the prorated calculation

Separate submission, same task · issues memo

“The $960,000 overpayment was credited to 2024”

Says the $960,000 was already credited, but the source asks for confirmation before filing.

The rubric and hallucination check assess different aspects of the submission.

1 material example1 minor · 1 not upheld

Illustrative scoring example

All-Pass RateShare of judges that pass every criterion
67%67%

2 of 3 judges passed every criterion

Criterion Pass RateShare of judge verdicts that are passes
90%90%

27 of 30 judge verdicts were passes; unaffected by hallucinations

Three verified examples from the same tax task. The material and minor examples share a submission; the dismissed example comes from another submission. These excerpts are not a full task audit.

The GPT-6 models hallucinate least: GPT-6 Astra averages 0.03 material hallucinations per task (4 hallucinations across all 120 tasks) and GPT-6 Sol 0.07 (8 hallucinations across all 120 tasks), while Gemini 3.8 Flash averages the most in the launch set, at 13.96 per task. Completing the criteria and not hallucinating are different skills: Muse Spark 1.3 passes the most criteria (96.0%) but averages 1.68 material hallucinations per task, against 0.03 for GPT-6 Astra. Every open-weights model averages at least 2.09 material hallucinations per task. A model's hallucination rate depends far more on the model than on the practice area.

Harvey LAB-AA v1.1: Hallucinations per Task

Upheld hallucination flags per task, split by severity · Material flags would mislead a reader on a substantive point, minor flags are real errors unlikely to affect the legal interpretation, and only material flags affect the headline score · Lower is better

How we chose the hallucination checker

We use GPT-6 Sol (high) for both passes of the hallucination check, separately from the three-judge rubric panel. We compared six candidates for the hallucination judge: GPT-6 Sol, GPT-6 Luna, Grok 4.7, Claude Opus 5.5, Claude Sonnet 5.5 and Gemini 3.8 Flash, each at high reasoning effort, on the same 20 tasks and deliverables from eight evaluated models.

Across this subset of tasks, GPT-6 Sol and GPT-6 Luna generally identified more material hallucinations, while Claude Sonnet 5.5 and Gemini 3.8 Flash identified far fewer. Claude Opus 5.5 fell between Grok 4.7 and Claude Sonnet 5.5 in every model row. In total, Opus upheld 99 material hallucinations, compared with 219 for Grok, 57 for Sonnet and 470 for GPT-6 Sol. GPT-6 Sol identified more material hallucinations despite the check's conservative approach to flagging errors, which excludes general legal knowledge not contained in the source documents.

All six checkers found no material hallucinations in GPT-6 Astra's outputs, while GPT-6 Sol ranged from 0 to 0.20 per task. Both were among the models with the fewest material hallucinations under every checker.

Harvey LAB-AA v1.1: Hallucination Judge Model Comparison

Mean upheld material hallucinations per task · Subset of 20 tasks and 8 models · Lower is better
Mean material hallucinations per task
0481217
Harvey LAB-AA v1.1: Hallucination Judge Model Comparison. Mean upheld material hallucinations per task · Subset of 20 tasks and 8 models · Lower is better.
MODEL BEING CHECKEDHALLUCINATION JUDGE
GPT-6 Sol(high)GPT-6 Luna(high)Grok 4.7(high)Claude Opus 5.5(high)Claude Sonnet 5.5(high)Gemini 3.8 Flash(high)
GPT-6 Astra (max)
GPT-6 Sol (max)
Grok 4.7 (xhigh)
Claude Opus 5.5 (max with fallback)
Claude Fable 5.1 (max with fallback)
Muse Spark 1.3 (max)
Kimi K3 (max)
Gemini 3.8 Flash (high)
Total material hallucinations per task23.5019.3510.954.952.851.40
Six hallucination judges compared on a fixed subset of 20 tasks and eight models. Cells show mean upheld material hallucinations per task. This subset comparison does not change the full 120-task leaderboard, which uses GPT-6 Sol (high) as the hallucination checker.

Hallucination-Gated Near-Pass Rate

Allowing near misses in terms of criteria puts GPT-6 Astra and GPT-6.1 Sol ahead. GPT-6 Astra rises from an 8.6% Hallucination-Gated All-Pass Rate to a 20.3% Hallucination-Gated Near-Pass Rate with one missed criterion and 31.7% with two missed criteria, and GPT-6.1 Sol from 6.9% to 20.3% and 29.6%, against 15.3% and 23.1% for Grok 4.7. The next biggest gains go to GPT-6 Sol (3.6% to 18.1%) and Claude Sonnet 5.5 (2.8% to 16.9%). Models whose passes are wiped out by hallucinations gain little, since a material hallucination still zeroes the task in every band: GLM-5.3 reaches only 2.2% even with two misses and Gemini 3.8 Flash stays at 0%.

Harvey LAB-AA v1.1: Hallucination-Gated Near-Pass Rate

Average over tasks of the share of the judge panel that missed 0, <=1 and <=2 criteria · Tasks with at least 1 material hallucination count at 0% · Higher is better
Sort by criteria missed

Cost

Top-scoring models aren't the most expensive. Grok 4.7 leads at ~$9.50 per task, under half the cost of Claude Fable 5.1 (~$21.70), the most expensive model. Muse Spark 1.3 represents strong performance for its cost, landing in second for ~$4.20 per task.

Harvey LAB-AA v1.1: Cost per Task

Average cost per task (USD), broken down by input, cache hit, cache write, reasoning, and answer tokens

Token Usage

Generating more output tokens does not necessarily translate to a higher score. GPT-6 Astra scores 8.6% on ~81k output tokens per task, under half of Grok 4.7's ~180k, while the three Claude models generate the most output tokens (~202k to ~562k per task) and score 2.8% to 6.4%.

Harvey LAB-AA v1.1: Output Tokens per Task

Output tokens used to run one task, broken down by reasoning and answer tokens

Speed

Stronger models tend to take longer. Estimated decode time is ~33 minutes per task for both Grok 4.7 and Claude Fable 5.1. Muse Spark 1.3 scores second at ~12 minutes per task. These estimates exclude time to first token and other overhead.

Harvey LAB-AA v1.1: Time per Task

Weighted average decode time (minutes) per task; excludes TTFT and overhead time · Lower is better

Turns

The top scorers don't run the longest loops. Grok 4.7 averages ~63 turns per task and GPT-6 Astra ~54. Claude Sonnet 5.5 runs the longest, ~179 turns, and scores 2.8%.

Harvey LAB-AA v1.1: Average Turns per Task

Average number of model turns per Harvey LAB-AA v1.1 task · Lower is better

Model Size (Open Weights Models Only)

Harvey LAB-AA v1.1: Hallucination-Gated All-Pass Rate vs. Compute Proxy

Harvey LAB-AA v1.1 Hallucination-Gated All-Pass Rate · Compute proxy
Most attractive quadrant
Pareto line

Score vs. Release Date

Harvey LAB-AA v1.1: Hallucination-Gated All-Pass Rate vs. Release Date

Most attractive region

Example Tasks & Submissions

Browse representative Harvey LAB tasks from the public task set, the reference files each model was given, and the deliverables it produced.

Mergers & Acquisitions

Instructions

Review the attached acquisition data room contracts and internal memo for change of control and assignment provisions, and prepare a comprehensive deal team report.

Output: coc-analysis-report.docx

Deliverables

Expected outputs the model must produce

  • coc-analysis-report.docxA comprehensive deal team report analyzing change of control and assignment provisions across the target’s material contracts.

Reference files

Provided to the model

Open
Open
Open
Open
Open
Open
Open
Open
Open
Open
Open
Open
Open
Open
Open
Open
Open
Open
Open

Model submissions

Deliverables produced by each model

Claude Opus 5.5 (max with fallback) - coc-analysis-report.docx
Open

Harvey LAB-AA resources