所有评测

AA-Briefcase v1.1: Agentic Knowledge Work Benchmark

A private evaluation developed by Artificial Analysis for frontier agentic capability in long-horizon knowledge work, testing agents on realistic business workflows that require deliverables such as spreadsheets, presentations, and memos.

AA-Briefcase 通过四个持续数周的知识工作项目评测模型,共包含数千个输入文件和 91 项任务。在不同场景中,模型必须完成数据科学、产品管理和企业战略等领域的真实专业工作流程。每个场景都是智能体依次推进的多周工作流程,每周包含多项任务。每项任务都需要生成成果,并依据检查项量表进行评分。虽然同一场景中的任务会跨周共享文件和上下文,但模型目前会独立运行每项任务,不会沿用自己先前的提交结果。

AA-Briefcase数据科学产品管理银行运营重工业战略第 1 周第 2 周第 3 周第 4 周第 5 周第 6 周任务 1任务 2任务 3任务 4评测场景AA-Briefcase 围绕 四个主要知识工作领域 构建,每个领域 都是一个独立的项目场景周次每个场景 由多个周次组成,用于模拟 专业工作流程任务每周最多包含 五项不同任务,智能体需在 场景和周次上下文中 完成检查每项任务 从客观量表检查、分析质量和呈现质量 三个维度评分,每次比较 都采用任务特定标准量表各项标准按通过或未通过进行二元评判分析质量与另一模型进行成对比较呈现质量与另一模型进行成对比较

每项任务按三类检查进行评分:

量表

每项检查按通过或未通过二元评分

模型是否遵循任务指示、识别散落在源文件中的要求、使用正确证据并得出正确结论?

分析质量

两两比较

与另一模型的提交结果相比,哪份成果更全面、分析更严谨、论据更充分?

呈现质量

两两比较

与另一模型的提交结果相比,哪一份呈现得更专业?

我们已通过 Hugging Face 发布第五个公开场景,用于展示场景结构、提交和评分方式。该场景仅供演示,不计入 AA-Briefcase 官方结果。

结果

AA-Briefcase Elo

AA-Briefcase v1.1 is an agentic knowledge work benchmark developed by Artificial Analysis. AA-Briefcase Elo is a combined metric that aggregates rubric pass rate, analytical quality Elo and presentation Elo · Higher is better
Not publicly available

AA-Briefcase Elo 与每项任务成本

AA-Briefcase Elo · 每项任务成本(美元)
Most attractive quadrant
Pareto line

成本

AA-Briefcase 每项任务成本

运行 AA-Briefcase 时每项任务的平均成本(美元),根据 token 使用量和模型定价(含代表性缓存命中率)计算
Not publicly available

示例任务、提交内容与评分

探索来自公开尽职调查场景的一个代表性 AA-Briefcase 周次,该场景可通过 Hugging Face 获取。此处展示的输出和评分说明了 AA-Briefcase 的评估内容。得分基于一组 代表性模型。此代表性场景中的提交内容和评审结果不会计入模型的 AA-Briefcase Elo 或其他基准得分。
模型
market_overview.pdf
打开
market_overview.tex
打开

得分对比

AA-Briefcase Elo 与 Artificial Analysis Intelligence Index

AA-Briefcase Elo · Artificial Analysis Intelligence Index
Most attractive quadrant

按文件类型划分的结果

按交付物文件类型(Excel、PowerPoint、PDF、Word、其他)划分的 AA-Briefcase 表现。

按文件类型划分的 AA-Briefcase 量表通过率(标准化)

按交付物文件类型划分的量表通过率 · 得分按文件类型在所有测试模型中进行标准化,绿色表示该文件类型的最高得分,红色表示最低得分

Token 使用量

AA-Briefcase 每项任务的输出 Token

每项 AA-Briefcase 任务平均消耗的推理和回答 token
Not publicly available

速度

每项任务耗时

Wall-clock time (minutes) per task: answer and reasoning generation plus tool execution time · Lower is better

轮次

每项任务平均轮次

Average number of model turns per AA-Briefcase task · Lower is better
Not publicly available

工具使用情况

各智能体在 AA-Briefcase 中的工具调用情况:按工具类别统计的调用次数、每轮平均工具调用次数以及来源库探索覆盖率。

AA-Briefcase 工具调用构成(每项任务平均)

每项 AA-Briefcase 任务的平均工具调用次数,按意图分组
Not publicly available

模型规模(仅开放权重模型)

AA-Briefcase Elo 与总参数量

AA-Briefcase Elo · 参数规模(十亿) · 仅限开放权重模型
Most attractive quadrant

得分 vs. 发布日期

AA-Briefcase Elo 与发布日期

AA-Briefcase Elo · 模型发布日期
Most attractive region

排行榜

开发者
名称
Elo
置信区间
发布日期
1
Anthropic logoAnthropic
Claude Sonnet 5.5 (Max, Default Fallback)1824-12 / +132026年9月
2
Anthropic logoAnthropic
Claude Opus 5.5 (Max, Default Fallback)1808-13 / +122026年9月
3
Anthropic logoAnthropic
Claude Opus 5.5 (Xhigh, Default Fallback)1768-11 / +112026年9月
4
Anthropic logoAnthropic
Claude Sonnet 5.5 (Xhigh, Default Fallback)1751-14 / +132026年9月
5
Anthropic logoAnthropic
Claude Opus 5.5 (High, Default Fallback)1690-10 / +102026年9月
6
Anthropic logoAnthropic
Claude Fable 5.1 (Max, Default Fallback)1676-10 / +102026年9月
7
Anthropic logoAnthropic
Claude Opus 5 (Max)1662-8 / +92026年7月
8
Anthropic logoAnthropic
Claude Fable 5.1 (Xhigh, Default Fallback)1657-8 / +92026年9月
9
SpaceXAI logoSpaceXAI
Grok 4.7 (Xhigh)1645-10 / +112026年9月
10
Anthropic logoAnthropic
Claude Sonnet 5.5 (High, Default Fallback)1640-9 / +112026年9月
11
Anthropic logoAnthropic
Claude Opus 5 (Xhigh)1636-8 / +92026年7月
12
SpaceXAI logoSpaceXAI
Grok 4.7 (High)1633-10 / +112026年9月
13
Anthropic logoAnthropic
Claude Opus 5.5 (Medium, Default Fallback)1628-10 / +112026年9月
14
Alibaba logoAlibaba
Qwen3.8 Max (0902)1621-10 / +102026年9月
15
Meta logoMeta
Muse Spark 1.3 (Max)1583-9 / +102026年9月

探索评测

Artificial Analysis Intelligence Index v4.3.2Artificial Analysis Intelligence Index v4.3.2

A composite benchmark aggregating ten challenging evaluations to provide a holistic measure of AI capabilities across mathematics, science, coding, and reasoning.

Artificial Analysis Openness IndexArtificial Analysis Openness Index

A composite measure providing an industry standard to communicate model openness for users and developers.

AA-Briefcase v1.1: Agentic Knowledge Work BenchmarkAA-Briefcase v1.1: Agentic Knowledge Work Benchmark

A private evaluation developed by Artificial Analysis for frontier agentic capability in long-horizon knowledge work, testing agents on realistic business workflows that require deliverables such as spreadsheets, presentations, and memos.

GDPval-AA v2.1 LeaderboardGDPval-AA v2.1 Leaderboard

GDPval-AA v2.1 is Artificial Analysis' evaluation framework for OpenAI's GDPval dataset. It tests AI models on real-world tasks across 44 occupations and 9 major industries. Models are given shell access and web browsing capabilities in an agentic loop via Stirrup to solve tasks, with Elo ratings derived from blind pairwise comparisons.

APEX-Agents-AA Benchmark LeaderboardAPEX-Agents-AA Benchmark Leaderboard

Artificial Analysis' implementation of the APEX-Agents benchmark, testing AI agents on long-horizon, cross-application tasks in professional-services environments with realistic application tooling.

AA-AnalystAgent Benchmark LeaderboardAA-AnalystAgent Benchmark Leaderboard

Artificial Analysis' data analysis benchmark, testing AI agents on their ability to work with spreadsheets and documents to answer quantitative questions a Business Analyst or Data Analyst would face day-to-day.

AutomationBench-AA: Agentic SaaS Workflow BenchmarkAutomationBench-AA: Agentic SaaS Workflow Benchmark

A benchmark measuring agentic task completion across simulated SaaS application environments, scoring the share of each task's objectives completed without guardrail violations.

Harvey LAB-AA Benchmark LeaderboardHarvey LAB-AA Benchmark Leaderboard

Artificial Analysis' implementation of Harvey's Legal Agent Benchmark (LAB), testing AI agents on real-world legal work from Harvey's dataset of 120 private tasks spanning 24 legal practice areas. The agent reads case documents in a sandbox and produces legal deliverables (e.g., memos, disclosure schedules, deposition summaries), graded criterion-by-criterion by a single LLM rubric judge.

GDP.pdf Benchmark LeaderboardGDP.pdf Benchmark Leaderboard

Artificial Analysis' implementation of Surge AI's GDP.pdf benchmark, testing whether language models can reason over long, real-world professional documents and satisfy detailed task-specific criteria.

EnterpriseOps-Gym-AA Benchmark LeaderboardEnterpriseOps-Gym-AA Benchmark Leaderboard

Artificial Analysis' independent implementation of ServiceNow's EnterpriseOps-Gym, an agentic benchmark testing whether LLM agents can complete stateful, multi-step enterprise workflows across eight business domains via live tool use, graded on the final state of the underlying databases.

Terminal-Bench 4.0 Benchmark LeaderboardTerminal-Bench 4.0 Benchmark Leaderboard

A harder 66-task benchmark of complex terminal work across software, machine learning, science, operations, security, hardware, and media, with recalibrated compute and time allowances and improved instructions, environments, and verifiers.

Terminal-Bench-Science 0.1 Benchmark LeaderboardTerminal-Bench-Science 0.1 Benchmark Leaderboard

A 70-task benchmark of research workflows authored and reviewed by domain experts across the life, physical, earth, mathematical, and engineering sciences, each completed in a terminal and checked by its own set of tests.

Artificial Analysis Long Context Reasoning Benchmark LeaderboardArtificial Analysis Long Context Reasoning Benchmark Leaderboard

A challenging benchmark measuring language models' ability to extract, reason about, and synthesize information from long-form documents ranging from 10k to 100k tokens (measured using the cl100k_base tokenizer).

AA-Omniscience: Knowledge and Hallucination BenchmarkAA-Omniscience: Knowledge and Hallucination Benchmark

A benchmark measuring factual recall and hallucination across various economically relevant domains.

SciCode Benchmark LeaderboardSciCode Benchmark Leaderboard

A scientist-curated coding benchmark featuring 288 test set subproblems from 80 laboratory problems across 16 scientific disciplines.

Humanity's Last Exam Benchmark LeaderboardHumanity's Last Exam Benchmark Leaderboard

A frontier-level benchmark with 2,500 expert-vetted questions across mathematics, sciences, and humanities, designed to be the final closed-ended academic evaluation.

CritPt Benchmark LeaderboardCritPt Benchmark Leaderboard

A benchmark designed to test LLMs on research-level physics reasoning tasks, featuring 71 composite research challenges.

GPQA Diamond Benchmark Leaderboard

The most challenging 198 questions from GPQA, where PhD experts achieve 65% accuracy but skilled non-experts only reach 34% despite web access.

ITBench-AA Benchmark LeaderboardITBench-AA Benchmark Leaderboard

Artificial Analysis' implementation of IBM's ITBench benchmark, testing AI agents on Kubernetes incident root-cause analysis from offline incident snapshots. The agent inspects alerts, events, traces, and topology and identifies the contributing-factor entities (deployments, pods, namespaces, network policies, etc.) responsible for the failure.

MMMU-Pro Benchmark LeaderboardMMMU-Pro Benchmark Leaderboard

An enhanced MMMU benchmark that eliminates shortcuts and guessing strategies to more rigorously test multimodal models across 30 academic disciplines.

IFBench Benchmark LeaderboardIFBench Benchmark Leaderboard

A benchmark evaluating precise instruction-following generalization on 58 diverse, verifiable out-of-domain constraints that test models' ability to follow specific output requirements.

Medical Long Context Reasoning (MLCR-AA)Medical Long Context Reasoning (MLCR-AA)

An open benchmark from Wisedocs measuring how well models reason over long, fragmented medical records, performing the multi-document synthesis claims professionals rely on when reviewing insurance and healthcare cases.

𝜏³-Banking Benchmark Leaderboard𝜏³-Banking Benchmark Leaderboard

A fintech customer-support benchmark from the 𝜏-Knowledge framework that tests whether agents can navigate a large unstructured knowledge base and execute multi-step tool calls to resolve realistic banking workflows.

Terminal-Bench 2.1 Benchmark LeaderboardTerminal-Bench 2.1 Benchmark Leaderboard

A verified refresh of Terminal-Bench 2.0 — 89 curated tasks across software engineering, system administration, data processing, model training, and security, with environment and instruction fixes so scores reflect agent capability rather than environment gaps.

Terminal-Bench Hard Benchmark LeaderboardTerminal-Bench Hard Benchmark Leaderboard

An agentic benchmark evaluating AI capabilities in terminal environments through software engineering, system administration, and data processing tasks.

𝜏²-Bench Telecom Benchmark Leaderboard𝜏²-Bench Telecom Benchmark Leaderboard

A dual-control conversational AI benchmark simulating technical support scenarios where both agent and user must coordinate actions to resolve telecom service issues.

MMLU-Pro Benchmark LeaderboardMMLU-Pro Benchmark Leaderboard

An enhanced version of MMLU with 12,000 graduate-level questions across 14 subject areas, featuring ten answer options and deeper reasoning requirements.

LiveCodeBench Benchmark LeaderboardLiveCodeBench Benchmark Leaderboard

A contamination-free coding benchmark that continuously harvests fresh competitive programming problems from LeetCode, AtCoder, and CodeForces, evaluating code generation, self-repair, and execution.

MATH-500 Benchmark LeaderboardMATH-500 Benchmark Leaderboard

A 500-problem subset from the MATH dataset, featuring competition-level mathematics across six domains including algebra, geometry, and number theory.

AIME 2025 Benchmark LeaderboardAIME 2025 Benchmark Leaderboard

All 30 problems from the 2025 American Invitational Mathematics Examination, testing olympiad-level mathematical reasoning with integer answers from 000-999.

Global-MMLU-Lite Benchmark LeaderboardGlobal-MMLU-Lite Benchmark Leaderboard

A lightweight, multilingual version of MMLU, designed to evaluate knowledge and reasoning skills across a diverse range of languages and cultural contexts.