すべての評価

AA-Briefcase v1.1: Agentic Knowledge Work Benchmark

A private evaluation developed by Artificial Analysis for frontier agentic capability in long-horizon knowledge work, testing agents on realistic business workflows that require deliverables such as spreadsheets, presentations, and memos.

AA-Briefcaseは、数千の入力ファイルと合計91件のタスクからなる、数週間にわたる4つの知識労働プロジェクトでモデルを評価します。各シナリオでは、データサイエンス、プロダクト管理、企業戦略などの分野における現実的な専門業務を完了する必要があります。各シナリオはエージェントが順番に進める複数週のワークフローで、週ごとに複数のタスクがあります。各タスクの成果物はチェック項目のルーブリックで評価されます。シナリオ内のタスクは週をまたいでファイルとコンテキストを共有しますが、現在、モデルは以前の提出物を引き継がず、各タスクを独立した実行として完了します。

AA-Briefcaseデータサイエンスプロダクト管理銀行業務重工業戦略第1週第2週第3週第4週第5週第6週タスク1タスク2タスク3タスク4評価シナリオAA-Briefcaseは 4つの主要な知識労働分野で構成されそれぞれが独立したプロジェクトシナリオになっています週各シナリオは 専門的なワークフローを再現する 複数の週で構成されますタスク各週には最大5件のタスクがあり エージェントはシナリオと週のコンテキスト内で完了しますチェック各タスクは 客観的なルーブリックチェック 分析品質 プレゼンテーション品質の3観点でタスク固有の基準により採点されますルーブリック基準ごとの二値合否分析品質別モデルとのペア比較プレゼンテーション別モデルとのペア比較

各タスクは3種類のチェックで評価されます。

ルーブリック

チェックごとの二値合否

モデルはタスクの指示に従い、ソースファイルに隠れた要件を特定し、適切な根拠を使用して正しい結論に達したか?

分析品質

ペア比較

別のモデルの提出物と比べ、どちらの成果物がより網羅的で、分析的に厳密かつ十分な根拠に基づいているか?

プレゼンテーション

ペア比較

別のモデルの提出物と比べ、どちらがより専門的に仕上げられているか?

シナリオの構造、提出、採点を示す公開用の5つ目のシナリオをHugging Faceで公開しています。これはデモ専用で、AA-Briefcaseの公式結果には含まれません。

結果

AA-Briefcase Elo

AA-Briefcase v1.1 is an agentic knowledge work benchmark developed by Artificial Analysis. AA-Briefcase Elo is a combined metric that aggregates rubric pass rate, analytical quality Elo and presentation Elo · Higher is better
Not publicly available

AA-Briefcase Eloとタスク当たりコストの比較

AA-Briefcase Elo · タスク当たりコスト(USD)
Most attractive quadrant
Pareto line

コスト

AA-Briefcaseのタスク当たりコスト

AA-Briefcaseのタスク当たり平均実行コスト(USD)。トークン使用量と、代表的なキャッシュヒット率を含むモデル料金から算出
Not publicly available

タスク、提出物、採点の例

Hugging Faceで公開されているDue Diligenceシナリオから、AA-Briefcaseの代表的な1週間を確認できます。ここに示す出力と採点はAA-Briefcaseの評価内容を例示しています。スコアは代表的なモデル群について表示されます。この代表シナリオの提出物と判定は、モデルのAA-Briefcase Eloやその他のベンチマークスコアには加算されません。
モデル
market_overview.pdf
開く
market_overview.tex
開く

スコア比較

AA-Briefcase EloとArtificial Analysis Intelligence Indexの比較

AA-Briefcase Elo · Artificial Analysis Intelligence Index
Most attractive quadrant

ファイル形式別の結果

成果物のファイル形式(Excel、PowerPoint、PDF、Word、その他)別のAA-Briefcase成績。

AA-Briefcaseのファイル形式別ルーブリック合格率(正規化)

成果物のファイル形式別のルーブリック合格率 · スコアはテストした全モデルを対象にファイル形式ごとに正規化されています。緑はそのファイル形式で最も高いスコア、赤は最も低いスコアを表します

トークン使用量

AA-Briefcaseのタスク当たり出力トークン

AA-Briefcaseのタスク当たりに消費した推論トークンと回答トークンの平均
Not publicly available

速度

タスク当たり時間

Wall-clock time (minutes) per task: answer and reasoning generation plus tool execution time · Lower is better

ターン数

タスク当たり平均ターン数

Average number of model turns per AA-Briefcase task · Lower is better
Not publicly available

ツール使用量

AA-Briefcase実行中に各エージェントが行ったツール呼び出し。ツールカテゴリ別の回数、ターン当たり平均呼び出し回数、ソース探索率を示します。

AA-Briefcaseのツール呼び出しの内訳(タスク当たり平均)

AA-Briefcaseのタスク当たりの平均ツール呼び出し回数(目的別)
Not publicly available

モデルサイズ(オープンウェイトモデルのみ)

AA-Briefcase Eloと総パラメーター数の比較

AA-Briefcase Elo · パラメーター数(10億) · オープンウェイトモデルのみ
Most attractive quadrant

スコア vs. リリース日

AA-Briefcase Eloとリリース日の比較

AA-Briefcase Elo · モデルのリリース日
Most attractive region

リーダーボード

開発元
名前
Elo
CI
リリース日
1
Anthropic logoAnthropic
Claude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)1824-12 / +132026年9月
2
Anthropic logoAnthropic
Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback)1822-12 / +122026年9月
3
Anthropic logoAnthropic
Claude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)1780-11 / +112026年9月
4
Anthropic logoAnthropic
Claude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback)1751-13 / +142026年9月
5
Anthropic logoAnthropic
Claude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback)1705-10 / +102026年9月
6
Anthropic logoAnthropic
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)1678-10 / +102026年9月
7
Anthropic logoAnthropic
Claude Opus 5 (Adaptive Reasoning, Max Effort)1673-10 / +112026年7月
8
Anthropic logoAnthropic
Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)1669-9 / +102026年9月
9
SpaceXAI logoSpaceXAI
Grok 4.7 (Xhigh)1657-9 / +92026年9月
10
Anthropic logoAnthropic
Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)1649-9 / +102026年7月
11
Anthropic logoAnthropic
Claude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback)1642-10 / +102026年9月
12
Anthropic logoAnthropic
Claude Sonnet 5.5 (Adaptive Reasoning, High Effort, Default Fallback)1639-11 / +112026年9月
13
SpaceXAI logoSpaceXAI
Grok 4.7 (High)1632-10 / +102026年9月
14
Alibaba logoAlibaba
Qwen3.8 Max (0902)1621-10 / +112026年9月
15
Anthropic logoAnthropic
Claude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback)1592-9 / +102026年9月

評価を探す

Artificial Analysis Intelligence Index v4.3.2Artificial Analysis Intelligence Index v4.3.2

A composite benchmark aggregating ten challenging evaluations to provide a holistic measure of AI capabilities across mathematics, science, coding, and reasoning.

Artificial Analysis Openness IndexArtificial Analysis Openness Index

A composite measure providing an industry standard to communicate model openness for users and developers.

AA-Briefcase v1.1: Agentic Knowledge Work BenchmarkAA-Briefcase v1.1: Agentic Knowledge Work Benchmark

A private evaluation developed by Artificial Analysis for frontier agentic capability in long-horizon knowledge work, testing agents on realistic business workflows that require deliverables such as spreadsheets, presentations, and memos.

GDPval-AA v2.1 LeaderboardGDPval-AA v2.1 Leaderboard

GDPval-AA v2.1 is Artificial Analysis' evaluation framework for OpenAI's GDPval dataset. It tests AI models on real-world tasks across 44 occupations and 9 major industries. Models are given shell access and web browsing capabilities in an agentic loop via Stirrup to solve tasks, with Elo ratings derived from blind pairwise comparisons.

APEX-Agents-AA Benchmark LeaderboardAPEX-Agents-AA Benchmark Leaderboard

Artificial Analysis' implementation of the APEX-Agents benchmark, testing AI agents on long-horizon, cross-application tasks in professional-services environments with realistic application tooling.

AA-AnalystAgent Benchmark LeaderboardAA-AnalystAgent Benchmark Leaderboard

Artificial Analysis' data analysis benchmark, testing AI agents on their ability to work with spreadsheets and documents to answer quantitative questions a Business Analyst or Data Analyst would face day-to-day.

AutomationBench-AA: Agentic SaaS Workflow BenchmarkAutomationBench-AA: Agentic SaaS Workflow Benchmark

A benchmark measuring agentic task completion across simulated SaaS application environments, scoring the share of each task's objectives completed without guardrail violations.

Harvey LAB-AA Benchmark LeaderboardHarvey LAB-AA Benchmark Leaderboard

Artificial Analysis' implementation of Harvey's Legal Agent Benchmark (LAB), testing AI agents on real-world legal work from Harvey's dataset of 120 private tasks spanning 24 legal practice areas. The agent reads case documents in a sandbox and produces legal deliverables (e.g., memos, disclosure schedules, deposition summaries), graded criterion-by-criterion by a single LLM rubric judge.

GDP.pdf Benchmark LeaderboardGDP.pdf Benchmark Leaderboard

Artificial Analysis' implementation of Surge AI's GDP.pdf benchmark, testing whether language models can reason over long, real-world professional documents and satisfy detailed task-specific criteria.

EnterpriseOps-Gym-AA Benchmark LeaderboardEnterpriseOps-Gym-AA Benchmark Leaderboard

Artificial Analysis' independent implementation of ServiceNow's EnterpriseOps-Gym, an agentic benchmark testing whether LLM agents can complete stateful, multi-step enterprise workflows across eight business domains via live tool use, graded on the final state of the underlying databases.

Terminal-Bench 4.0 Benchmark LeaderboardTerminal-Bench 4.0 Benchmark Leaderboard

A harder 66-task benchmark of complex terminal work across software, machine learning, science, operations, security, hardware, and media, with recalibrated compute and time allowances and improved instructions, environments, and verifiers.

Terminal-Bench-Science 0.1 Benchmark LeaderboardTerminal-Bench-Science 0.1 Benchmark Leaderboard

A 70-task benchmark of research workflows authored and reviewed by domain experts across the life, physical, earth, mathematical, and engineering sciences, each completed in a terminal and checked by its own set of tests.

Artificial Analysis Long Context Reasoning Benchmark LeaderboardArtificial Analysis Long Context Reasoning Benchmark Leaderboard

A challenging benchmark measuring language models' ability to extract, reason about, and synthesize information from long-form documents ranging from 10k to 100k tokens (measured using the cl100k_base tokenizer).

AA-Omniscience: Knowledge and Hallucination BenchmarkAA-Omniscience: Knowledge and Hallucination Benchmark

A benchmark measuring factual recall and hallucination across various economically relevant domains.

SciCode Benchmark LeaderboardSciCode Benchmark Leaderboard

A scientist-curated coding benchmark featuring 288 test set subproblems from 80 laboratory problems across 16 scientific disciplines.

Humanity's Last Exam Benchmark LeaderboardHumanity's Last Exam Benchmark Leaderboard

A frontier-level benchmark with 2,500 expert-vetted questions across mathematics, sciences, and humanities, designed to be the final closed-ended academic evaluation.

CritPt Benchmark LeaderboardCritPt Benchmark Leaderboard

A benchmark designed to test LLMs on research-level physics reasoning tasks, featuring 71 composite research challenges.

GPQA Diamond Benchmark Leaderboard

The most challenging 198 questions from GPQA, where PhD experts achieve 65% accuracy but skilled non-experts only reach 34% despite web access.

ITBench-AA Benchmark LeaderboardITBench-AA Benchmark Leaderboard

Artificial Analysis' implementation of IBM's ITBench benchmark, testing AI agents on Kubernetes incident root-cause analysis from offline incident snapshots. The agent inspects alerts, events, traces, and topology and identifies the contributing-factor entities (deployments, pods, namespaces, network policies, etc.) responsible for the failure.

MMMU-Pro Benchmark LeaderboardMMMU-Pro Benchmark Leaderboard

An enhanced MMMU benchmark that eliminates shortcuts and guessing strategies to more rigorously test multimodal models across 30 academic disciplines.

IFBench Benchmark LeaderboardIFBench Benchmark Leaderboard

A benchmark evaluating precise instruction-following generalization on 58 diverse, verifiable out-of-domain constraints that test models' ability to follow specific output requirements.

Medical Long Context Reasoning (MLCR-AA)Medical Long Context Reasoning (MLCR-AA)

An open benchmark from Wisedocs measuring how well models reason over long, fragmented medical records, performing the multi-document synthesis claims professionals rely on when reviewing insurance and healthcare cases.

𝜏³-Banking Benchmark Leaderboard𝜏³-Banking Benchmark Leaderboard

A fintech customer-support benchmark from the 𝜏-Knowledge framework that tests whether agents can navigate a large unstructured knowledge base and execute multi-step tool calls to resolve realistic banking workflows.

Terminal-Bench 2.1 Benchmark LeaderboardTerminal-Bench 2.1 Benchmark Leaderboard

A verified refresh of Terminal-Bench 2.0 — 89 curated tasks across software engineering, system administration, data processing, model training, and security, with environment and instruction fixes so scores reflect agent capability rather than environment gaps.

Terminal-Bench Hard Benchmark LeaderboardTerminal-Bench Hard Benchmark Leaderboard

An agentic benchmark evaluating AI capabilities in terminal environments through software engineering, system administration, and data processing tasks.

𝜏²-Bench Telecom Benchmark Leaderboard𝜏²-Bench Telecom Benchmark Leaderboard

A dual-control conversational AI benchmark simulating technical support scenarios where both agent and user must coordinate actions to resolve telecom service issues.

MMLU-Pro Benchmark LeaderboardMMLU-Pro Benchmark Leaderboard

An enhanced version of MMLU with 12,000 graduate-level questions across 14 subject areas, featuring ten answer options and deeper reasoning requirements.

LiveCodeBench Benchmark LeaderboardLiveCodeBench Benchmark Leaderboard

A contamination-free coding benchmark that continuously harvests fresh competitive programming problems from LeetCode, AtCoder, and CodeForces, evaluating code generation, self-repair, and execution.

MATH-500 Benchmark LeaderboardMATH-500 Benchmark Leaderboard

A 500-problem subset from the MATH dataset, featuring competition-level mathematics across six domains including algebra, geometry, and number theory.

AIME 2025 Benchmark LeaderboardAIME 2025 Benchmark Leaderboard

All 30 problems from the 2025 American Invitational Mathematics Examination, testing olympiad-level mathematical reasoning with integer answers from 000-999.

Global-MMLU-Lite Benchmark LeaderboardGlobal-MMLU-Lite Benchmark Leaderboard

A lightweight, multilingual version of MMLU, designed to evaluate knowledge and reasoning skills across a diverse range of languages and cultural contexts.