Artificial Analysis 智能基准测试方法论

Artificial Analysis Intelligence Index v4.3.2

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index结合了一套全面的评估数据集,以评估推理、知识、数学和编程方面的语言模型能力。

它是整体语言模型智能的有用综合,可用于比较语言模型。与所有评估指标一样,它也有局限性,可能并不直接适用于每个用例。然而,我们相信,它是语言模型之间比当今存在的任何其他指标更有用的综合比较。

Artificial Analysis Intelligence Index v4.3.2 包含 10 项评估:AA-Briefcase v1.1、GDPval-AA v2.1、AutomationBench-AA、Terminal-Bench 4.0、SciCode、AA-LCR v1.1、AA-Omniscience、Humanity's Last Exam、GDP.pdf 和 CritPt。我们的方法强调公平性和实际应用价值。

我们估计Artificial Analysis Intelligence Index的95%置信区间小于±1% - 基于对Artificial Analysis Intelligence Index v4.3.2中包含的所有评估数据集的某些模型进行超过10次重复的实验。个别评估结果的置信区间可能大于±1%。我们期待将来披露统计分析的更多细节。

Artificial Analysis Intelligence Index 是一套以文本为主的英语评估。我们将图像输入、语音输入和多语言能力作为独立于 Intelligence Index 评估套件的基准测试。

Intelligence Index评估套件

Intelligence Index 是四个类别的加权平均值:智能体 (30%)、编程 (20%)、科学推理 (20%) 和通用能力 (30%)。权重设置侧重智能体任务。各评估所属类别及其权重如下所示。

类别评估题目重复次数响应类型评分Intelligence Index
权重
工具
使用
私有测试集
智能体 (30%)AA-Briefcase v1.191 项任务,涵盖 4 个场景1智能体完成任务并输出文件综合 Elo,汇总基于评分量表的任务成功度、分析质量和呈现质量的成对比较15%✓
GDPval-AA v2.1220 项任务1智能体完成任务并输出文件评审小组进行成对比较 (Elo),以 DeepSeek V4.1 Flash (max) 的 1600 分为锚点,冻结后归一化10%✓✗
AutomationBench-AA657 项任务1使用 REST API 工具实现 SaaS 工作流程自动化按目标完成情况评分,违反防护规则的任务得零分5%✓
编程 (20%)Terminal-Bench 4.0663基于终端的任务执行测试套件 pass/fail,pass@110%✗
SciCode288 个子问题 (测试集)3Python 代码(必须通过所有单元测试)代码执行,pass@1;在提示词中加入科学家标注的背景信息,按子问题评分10%✗✗
通用能力 (30%)AA-Omniscience6,0001开放式回答准确度 (10%) 和 1 - 幻觉率 (5%) 作为单独的组件15%✗
GDP.pdf100 项任务,涵盖 10 个领域5基于长 PDF 的自由格式答案主要指标为 All-pass,另报告按任务等权汇总的 Mean Pass10%✗✗
AA-LCR v1.11003开放式回答LLM 判定答案等价性,pass@15%✗✗
科学推理 (20%)HLE (Humanity's Last Exam)2,1581开放式回答LLM 判定答案等价性,pass@110%✗✗
CritPt705Python 函数、符号表达式、数值答案官方评分服务器,pass@110%✗

补充评测

除了Intelligence Index套件之外,我们还进行了一系列涵盖多语言、视觉、数学和其他能力的附加评估。这些单独报告,不包含在Intelligence Index分数中。

Artificial Analysis Multilingual Index:衡量模型的多语言能力,基于 Global-MMLU-Lite 在支持语言中的评估。我们支持以下语言:

  • 🇬🇧 English
  • 🇨🇳 Chinese
  • 🇮🇳 Hindi
  • 🇪🇸 Spanish
  • 🇫🇷 French
  • 🇸🇦 Arabic
  • 🇧🇩 Bangla
  • 🇵🇹 Portuguese
  • 🇮🇩 Indonesian
  • 🇯🇵 Japanese
  • 🇰🇪 Swahili
  • 🇩🇪 German
  • 🇰🇷 Korean
  • 🇮🇹 Italian
  • 🇳🇬 Yoruba
  • 🇲🇲 Burmese
类别评估题目重复次数响应类型评分工具
使用
私有测试集
智能体𝜏³-Banking975带知识检索的双重控制智能体-用户模拟后端数据库状态评估,pass@1✓✗
Harvey LAB-AA120 项任务1智能体生成法律工作成果并输出文件单个 LLM 评审依据评分量表评判各项标准,pass@1✓✓
APEX-Agents-AA452 项任务3智能体完成专业服务任务基于评分标准的本地文件评分,pass@1✓✗
AA-AnalystAgent80,涵盖 14 个领域5智能体执行 Python 代码,并给出自由格式的最终答案LLM 判定答案正确与否,数值预检查可覆盖该判定,pass^5✓✓
ITBench-AA59 个场景 (公开 + 私有)3根据离线 Kubernetes 事件快照进行结构化 JSON 根本原因诊断经 LLM 归一化的实体匹配,完全召回时的平均精确率✓
EnterpriseOps-Gym-AA1,117 项 oracle 模式任务 (8 个领域)3在可重置的企业评估环境服务器上进行多轮 MCP 工具调用基于结果的SQL状态验证器,严格的pass@1成功率✓✗
Terminal-Bench-Science 0.170 (5 个领域)3基于终端的任务执行测试套件 pass/fail,pass@1✗
通用能力IFBench2945开放式回答提取和规则驱动评估,pass@1✗✗
MLCR-AA60 道题目 (专家级 + 复合级)3开放式回答简洁性门槛加 LLM 评审小组(3 名评审对完整性和准确性进行多数表决),pass@1✗
其他Global-MMLU-Lite~ 6,000(每种语言 ~ 400)1多项选择(4个选项)正则表达式提取,pass@1✗✗
MMMU Pro1,7301多项选择(10个选项)正则表达式提取,pass@1✗✗

智能评测原则

我们的评估方法遵循四个核心原则:

  • 标准化:所有模型均在相同的条件下进行评估,并具有一致的提示策略、温度设置和评估标准。
  • 公正:我们采用评估技术,避免对正确遵循提示中的说明给出答案的模型进行不公平的惩罚。这包括使用清晰的提示、强大的答案提取方法和灵活的答案验证来适应模型输出中的有效变化。
  • 零样本指令提示:我们使用明确的指令进行评估,不提供示例或演示,以测试模型在没有少样本学习的情况下遵循指令的能力。这种方式适用于现代指令微调模型和对话模型。
  • 透明:我们公开我们的方法,包括提示模板、评估标准和限制。

通用测试参数

我们使用以下设置测试所有评估:

  • 温度: 0 用于非推理模型,0.6 用于推理模型(除非模型实验室推荐其他温度)
  • 最大输出 Token 数:
    • 非推理模型:16,384 个Token(当模型具有较小的上下文窗口或较低的最大输出Token上限时向下调整)
    • 推理模型:允许的最大输出Token,如模型创建者所披露(每个推理模型的自定义设置)
  • 错误处理:
    • API 失败时自动重试(最多 30 次尝试)
    • 所有 30 次重试均失败的问题均需手动审核。持续 API 故障导致问题的结果不会发布。专有模型的所有可用 API 阻止某个问题的错误可能会降低分数(这种影响并不重要)
  • 评分方法:我们通常在评估中使用 pass@1 评分,模型必须在第一次尝试时给出正确的答案。对于多次重复的评估,pass@1 是通过汇总所有重复的结果来计算的。计算如下:
    其中,如果尝试 i 正确,则 pi = 1,否则为 0,k 是所有重复中的测试实例总数。

我们维护所有评估数据集的内部副本。下面列出了我们所选数据集的来源。

在 Artificial Analysis Intelligence Index 评估中,我们优先使用各模型 API 提供商报告的 Token 数量,准确计算运行 Intelligence Index 的成本。极少数情况下,提供商未提供 Token 数量,我们会使用标准分词器作为后备方案。性能基准测试采用不同方式:使用 o200k_base 分词器在客户端计算 Token 数,使不同模型处理相同文本时的 Token 计数口径一致。报告缓存命中率和成本时,我们将这些 Token 数与模型典型缓存命中率的实时测量结合,而非依赖评估运行时的一次性测量。

我们使用 e2b 作为智能体基准测试的主要沙箱提供商。

Artificial Analysis Intelligence Index 评测

构成当前Artificial Analysis Intelligence Index的评估,按能力分组。

智能体

AA-Briefcase v1.1

  • 状态: 纳入Artificial Analysis Intelligence Index v4.3.2,权重为 15%。
  • 说明:AA-Briefcase 是一项新基准,用于在行业专家设计的复杂项目中,评估模型完成真实知识工作任务的能力。评估项目持续数周,每个项目包含许多相互关联的任务和数千份输入源文件。AA-Briefcase 结合评分量表评审和两两比较,评估可验证的任务成功度、分析质量和呈现质量,从而全面衡量智能体处理知识工作的能力。
  • 相较 AA-Briefcase v1 的变化:v1.1 仅改变 Elo 评分的拟合方式。我们使用 Crowd-BT 模型拟合评分。对于提交结果 与 的比较,拟合后的能力参数分别为 与 ,标注者质量为 :
    我们定义 ,给出:
    其中, 根据过往 AA-Briefcase 评审结果,按评分维度分别拟合。评分量表部分由确定性比较决定,不由评审模型判定,因此该维度的 。Elo 分数虽有变化,排名顺序基本保持不变。
  • 示例数据集: https://huggingface.co/datasets/ArtificialAnalysis/AA-Briefcase-Lite
  • 智能体运行框架: https://github.com/ArtificialAnalysis/Stirrup
  • 实现:
    • 每个 AA-Briefcase 场景都是一个实际的多周业务问题,被组织为一个多周工作流程,智能体按顺序完成该工作流程,每周执行 2-5 个任务。尽管场景中的任务在数周内共享文件和上下文,但模型目前以独立运行的方式完成每项任务,而无需继承其之前提交的内容。智能体接收任务描述和可访问的源文件,然后生成最终的可交付文件,而无需在执行期间进行实时交互或迭代反馈。
    • 场景源池包括共享文件和特定于周的文件,混合真实、增强和合成材料。源文件旨在包含真实的专业资料,例如 Slack 导出、电子表格、PDF、访谈记录、市场研究、标准文档、应用商店页面、董事会材料、电子邮件和其他业务记录。本周晚些时候的任务可能会收到标准化的基本案例文件(为每个模型提供相同的参考工作产品),因此每个任务保持独立运行,同时保持一周的连续性。
    • 模型提交使用 Stirrup 在一周范围的 E2B 沙箱中运行。
      • 轮次:每个任务智能体运行最多 500 轮。
      • 工具: 为智能体提供了一个代码执行工具,可在沙箱内运行 shell 命令和代码,以及下面的完成工具(当模型支持视觉时,还提供一个视图图像工具)。沙箱无法访问互联网,因此智能体只能使用提供的源文件。
      • 沙箱::每个场景/周沙箱都是根据该周的源文件构建的,并预装了用于文档处理和科学计算的标准Python软件包和系统工具。
      • 任务结束工具::一个完成工具,智能体调用该工具来提交摘要及其可交付成果的绝对路径(验证为实际文件,而不是目录或丢失的路径),以及一个abandon_task_finish(放弃)工具,仅当它断定任务确实不可能时才调用该工具。
    • 提示词:跨生成和评分使用的提示:
      • 智能体系统提示词:
        You are an AI agent working on a specific task within a multi-week simulated workplace scenario. Each task is part of a longer workflow; your job is to complete the current task using the tools provided in up to 500 steps, then submit your deliverables.
        
        When you are done you must call the `finish` tool as your final step, passing a brief summary of what you accomplished and a list of absolute paths for every deliverable file.
        
        If you have genuinely concluded that the task cannot be completed — for example because required inputs are missing, a hard dependency is unavailable, or the request itself is incoherent — call the `abandon_task_finish` tool with a brief reason instead. Do not use it to escape difficulty.
        
        You cannot interact with the user during the task. Record any clarifying assumptions you made in your finish summary.
      • 智能体任务提示词:
        <execution_context>
        ## Sandbox
        
        You operate inside an isolated Linux container through the `code_exec` tool, which runs shell commands and lets you read, create, and edit files. Commands run as the unprivileged user `user` (UID 1000), starting from `/home/user`. Passwordless `sudo` exists but is rarely needed, since your home directory is fully writable.
        
        Every command runs independently: no working directory, environment variable, or other shell state carries over from one call to the next. Prefer absolute paths for both files and commands, and do not navigate with `cd` across calls — a `cd` in one command is gone by the next, so relying on it leaves you silently operating in the wrong place. When a step genuinely needs a different directory, chain it into the same command (e.g. `cd /home/user/work && python build.py`).
        
        ## No network
        
        The container has no outbound connectivity, and there is no proxy, allowlist, or flag that can turn it on — treat the environment as permanently offline. Anything that reaches for the internet will fail, including package installs (`pip`, `npm`, `apt`), remote `git` operations, and any HTTP/HTTPS client request from any language.
        
        Identify a network block by its error signature rather than by guessing: failed name resolution (`Could not resolve host`, `Temporary failure in name resolution`), an unreachable route (`Network is unreachable`, a refused or timed-out connection to a public host), or a stalled TLS handshake. When you see these, the failure is structural — do not retry the same call and do not hunt for a workaround (mirrors, alternate hosts, cached copies). Re-plan using only what is already installed and what ships inside your workspace.
        
        ## Filesystem
        
        - Writable: everything under `/home/user/` plus `/tmp`. Use these for deliverables, intermediate files, and caches.
        - Read-only inputs:
        - `/home/user/shared/` — reference material shared across the whole scenario
        - `/home/user/week/` — documents specific to this week's tasks
        Copy these into a working folder before transforming them rather than editing them in place.
        
        ## Runtime
        
        A broad scientific-computing and document-processing stack is already installed, so confirm what is present before assuming a gap:
        - Python 3.13 with the usual data stack (numpy, pandas, polars, scipy), plotting (matplotlib, plotly), the scikit-learn ML family, and document tooling (python-docx, python-pptx, openpyxl, PyMuPDF, pdfplumber, reportlab, weasyprint, Pillow, opencv), plus Playwright.
        - System tools include LibreOffice, Pandoc, Tesseract, FFmpeg, ImageMagick, Ghostscript, TeX Live, OpenJDK, Chromium, jq, and git.
        - Check availability with `pip show <pkg>` or `which <tool>` instead of installing — installs fail offline, but almost anything you would reach for is already here.
        - matplotlib runs headless (`MPLBACKEND=Agg`): write figures to files; never call `plt.show()`.
        - Commands are terminated after 20 minutes. Keep them bounded, persist intermediate results to disk, and split long jobs into smaller steps.
        
        ## Submitting your work
        
        Finish by calling the `finish` tool — anything not submitted through it is not graded. Your call must include:
        1. A short summary of what you accomplished.
        2. Absolute paths to every deliverable (files only, not folders).
        
        Save each deliverable directly in `/home/user` under the exact filename the task asks for — not in a subdirectory.
        
        Save deliverables as ordinary, visible files. Do not leave the only copy of your work in a dot-prefixed file or directory (e.g. `.submission.txt`, `.outputs/report.md`), including inside an archive; a `.zip` is fine when the task explicitly asks for one. Assume your files will be opened and edited by others after submission, so write them to last.
        
        If the task genuinely cannot be completed, call the `abandon_task_finish` tool with a brief reason instead. Use it only when you have concluded the work is impossible — not to escape a difficult task.
        </execution_context>
        
        <scenario_overview>
        {scenario_overview}
        </scenario_overview>
        
        <week_overview>
        {week_overview}
        </week_overview>
        
        <task_description>
        {task}
        </task_description>
        
        <deliverables>
        Submit these files, by exact name, saved directly in `/home/user`:
        {expected_output_filenames}
        </deliverables>
        
        Please begin working on the task now.
      • 二进制评分标准提示:
        You are grading a submitted deliverable against one binary rubric check.
        
        The user message contains:
        - the task instructions,
        - the rubric item,
        - the submitted artifact content.
        
        Submitted artifacts may appear as text blocks, image blocks, or parser notes for unsupported content.
        
        Use only evidence from the submitted artifact content. Do not infer facts from filenames, task instructions, or rubric text unless the submitted artifact content supports them.
        
        Beyond the task instructions and rubric in the user message, you only ever receive the submitted artifact itself, never the external source files it cites. Do not fail an item merely because you cannot open or cross-check a cited source — judge citations on whether they are present, specific, and well-formed in the submission, not on whether the source's contents can be independently confirmed.
        
        Return a strict binary judgment:
        - passed=true only if the pass criteria are satisfied.
        - passed=false if any required element is missing, materially wrong, unsupported, or not evidenced.
        
        Write concise reasoning that cites submitted artifact evidence or the absence of evidence.
        Do not award partial credit.
    • 每项任务都根据两种检查方式进行评分。 评分标准检查是针对单个提交进行评分的二元通过/失败标准。 成对检查比较同一任务的两个提交,并返回首选提交或平局。有两种:分析质量(其输出具有更深入、结构更好的分析)和演示(其输出更专业地呈现)。
    • 评分量表评审、分析质量两两比较和呈现质量两两比较,各自由三名评审模型组成的小组负责,而非单一评审,以减少对同一模型或模型系列提交结果的偏好。评分量表评审使用最大 effort 的 Claude Opus 4.8、高 reasoning 的 GPT-5.5 和高 reasoning 的 Gemini 3.1 Pro Preview。两两比较使用高 effort 的 Claude Opus 5、中 reasoning 的 GPT-5.6 Sol 和高 reasoning 的 Gemini 3.8 Flash。每项评分量表判定和每次由 LLM 评判的两两比较,都由从相应小组中抽样选出的一名评审负责,抽样在各项检查和对比之间保持均衡。为保证结果可比,同一项评分量表检查始终由同一评审判定。AA-Briefcase Elo 是本评估的主要指标,汇总分析质量 Elo、呈现质量 Elo 和评分量表通过率;评分量表表现通过合成对决和最大似然 Elo 聚合转换为 Elo。每个评分维度均使用 Crowd-BT 模型拟合。
    • 纳入 Intelligence Index:AA-Briefcase v1.1 的综合 Elo 分数在模型加入时冻结,并以 clamp((Elo - 500) / 2000) 归一化后纳入 Intelligence Index。这一映射与 GDPval-AA v2.1 相同。Elo 量表以 GPT-5.5 (medium) 的 1000 分为锚点,固定的归一化范围则使模型对 Intelligence Index 的贡献随时间保持稳定。随着模型在该评估中取得进展,Artificial Analysis 可能更新参考参数,以维持 Intelligence Index 的有效区分度。

GDPval-AA v2.1

  • 说明: GDPval-AA v2.1是针对OpenAI的GDPval数据集的Artificial Analysis评估框架。它评估了语言模型在具有经济价值的任务上的能力,涵盖了对美国 GDP 做出贡献的关键部门的 44 个职业。
  • 与 GDPval-AA v2 的更改: v2.1 仅更改 Elo 比例的固定方式:
    • 我们通过将 DeepSeek V4.1 Flash (max) 固定在 1600 来锚定比例。
    • 我们使用 Crowd-BT 模型拟合评分。对于提交结果 与 的比较,拟合后的能力参数分别为 与 ,标注者质量为 :
      我们定义 ,给出:
      其中, 根据过往 GDPval-AA 评审结果拟合。
    • 虽然 Elo 分数发生变化,但排名顺序在很大程度上保持不变。
  • 论文: https://arxiv.org/abs/2510.04374
  • 智能体运行框架: https://github.com/ArtificialAnalysis/Stirrup
  • 数据集:
    • 我们的评估基于来自 https://huggingface.co/datasets/openai/gdpval 的公共黄金 OpenAI GDPval 数据集
    • 数据集中的某些 Microsoft Office 文件缺少元数据部分或格式错误的关系条目,导致 LibreOffice 无法打开它们。我们添加了最少的缺失元数据并修复了格式错误的条目以确保兼容性。文档正文、幻灯片内容和布局未更改。
  • 实现:此评估包括两个阶段:
    • 任务提交 – 模型被赋予一项任务,并需要生成一个或多个文件。
    • 配对评分 – 从三名前沿LLM评委组成的小组中抽取的评委对同一任务的两个提交内容进行盲目排名,每个提交内容均由不同的模型创建。
    • Elo计算:收集成对排名后,我们通过最大似然估计将它们拟合到Crowd-BT模型,并使用三明治估计器计算置信区间以建立最终的Elo指标。我们通过将 DeepSeek V4.1 Flash (max) 固定在 1600 来锚定 Elo 等级。所有其他评级均相对于该锚定进行拟合。
    • 纳入 Intelligence Index:GDPval-AA v2.1 的Elo 分数在模型加入时冻结,并以 clamp((Elo - 500) / 2000) 归一化后纳入 Intelligence Index。Elo 量表以 DeepSeek V4.1 Flash (max) 的 1600 分为锚点,固定的归一化范围则使模型对 Intelligence Index 的贡献随时间保持稳定。随着模型在该评估中取得进展,Artificial Analysis 可能更新参考参数,以维持 Intelligence Index 的有效区分度。
  • 任务提交详细信息:
    • 所有模型均使用我们的开源智能体工具Stirrup运行。在该工具中,为模型提供了一个代码执行环境(E2B 沙箱),以及可自行调用的以下六个工具:
      • Web Fetch – 从网页中获取并提取主要内容作为 markdown。
      • 网页搜索 – 使用Brave Search API搜索网页;返回前 5 个结果,包括主要、URL 和说明。
      • 查看图像 – 从沙箱中读取并显示图像文件(.png、.jpg、.jpeg)作为本机图像Token以供LLM使用。该工具仅适用于具有视觉支持的模型。在发送到模型之前,图像会缩小至最大 1 兆像素。
      • Code Exec – 通过code_exec工具在沙箱中执行bash命令;返回退出代码、stdout 和 stderr。
      • 完成 – 发出任务完成信号并指定要提交的文件。
      • 放弃任务 – 表示模型不相信它可以完成任务,并给出一个简短的原因,而不是提交文件。
    • 对于每个任务,都会使用与给定任务关联的参考文件初始化一个新的 E2B 沙箱,并预安装一系列与任务集相关的包。我们的包集合基于原始 GDPval 论文中公开的环境,并在 v2 中扩展了附加依赖项(包括完整的 TeX Live LaTeX 工具链和构建工具)。

      系统包是在 Debian trixie 基础镜像中解析的完整固定传递闭包,因此大多数条目都是我们直接安装的包的依赖项

    • 我们通过插入相关任务提示、参考文件和完成工具详细信息的指令来提示智能体。
  • 执行限制:
    • LLM有 250 轮时间来完成任务。单轮被定义为辅助消息及其工具调用(如果有)。当模型接近极限时,它会收到剩余转弯预算的通知。
    • 如果模型不相信自己可以完成任务,则可以通过放弃任务工具提前结束运行,并提供简短的原因而不是提交文件。
    • 如果模型在完成给定回合后超过其上下文窗口的 70%,智能体会要求其总结任务状态、已完成的工作、当前文件、剩余步骤和重要上下文,然后清除较早的回合历史记录,同时保留任务提示和摘要以供继续。

任务提交系统提示:

You are an AI agent completing a standalone professional task. Your job is to use the provided tools to produce the requested deliverables within 250 steps, then submit your work.

When you are done, call the `finish` tool as your final step with:
1. A brief summary of what you accomplished.
2. Absolute paths to every deliverable file.

If you have genuinely concluded that the task cannot be completed because required inputs are missing, a hard dependency is unavailable, or the request is incoherent, call the `abandon_task_finish` tool with a brief reason instead. Do not use it to escape difficulty.

You cannot interact with the user during the task. Make reasonable assumptions when needed and record them in your finish summary.

任务提交提示:

## Runtime

You are running in an isolated Linux sandbox. Use the `code_exec` tool to read, create, and modify files. Commands run as the non-root user `user` (UID 1000). Default working directory is `/home/user`.

Every command runs independently: no working directory, environment variable, or other shell state carries over from one call to the next. Prefer absolute paths for both files and commands, and do not navigate with `cd` across calls — a `cd` in one command is gone by the next, so relying on it leaves you silently operating in the wrong place. When a step genuinely needs a different directory, chain it into the same command (e.g. `cd /home/user/work && python build.py`).

A broad scientific-computing and document-processing stack is already installed, so confirm what is present before assuming a gap:
- Python 3.13 with the usual data stack (numpy, pandas, polars, scipy), plotting (matplotlib, plotly), the scikit-learn ML family, and document tooling (python-docx, python-pptx, openpyxl, PyMuPDF, pdfplumber, reportlab, weasyprint, Pillow, opencv), plus Playwright.
- System tools include LibreOffice, Pandoc, Tesseract, FFmpeg, ImageMagick, Ghostscript, TeX Live, OpenJDK, Chromium, jq, and git.
- Commands are terminated after 10 minutes. Keep them bounded, persist intermediate results to disk, and split long jobs into smaller steps.

## Reference Files Location

(This section appears only when the task includes reference files.)

The reference files for the task are available in your environment's file system.

Here are their paths:

- [absolute path to each reference file]

## Completing Your Work

In order to complete the task you must use the `finish` tool to submit your work. If you do not use the `finish` tool you will fail this task!

As a last resort if you really cannot make any meaningful progress, use `abandon_task_finish` with a brief reason instead of submitting files.

**Required in your finish call:**
1. A brief summary of what you accomplished
2. A list of **ABSOLUTE file paths** for the required output files (Do not submit folders).

## Task

Here is the task you need to complete:

[task description]

Please begin working on the task now.
  • 上下文溢出:如果下一个模型调用(或汇总请求本身)超出了上下文窗口,智能体将继续展开较早的回合,直到汇总成功。
  • 任务完成:要完成任务,LLM必须调用完成工具,提供已完成工作的摘要以及打算提交的文件的路径。这个工具可以随时使用。
  • 评分:我们分两个阶段对模型提交之间的配对匹配进行采样:
    • 平衡采样:我们首先对每个模型进行不同的采样,平衡任务、评委和对手之间的暴露,以形成初始评级。
    • 主动采样:在初始阶段之后,我们过渡到 Elo 知情采样,优先考虑具有相似评级的模型之间的配对,以便在每次比较时获取最多的信息。我们在整个过程中保持每个模型内任务的平衡暴露。
    • 提交内容被随机匿名为提交 A 和 B,以减轻评分者模型的任何模型或位置偏差。
    • 比赛由来自领先实验室的三名前沿 LLM 评委进行评分,每个评委均以其默认推理设置运行:GPT-5.6 Sol (medium reasoning)、Gemini 3.8 Flash (high reasoning)和 Claude Opus 5 (high effort)。我们在每次比较时都会在评委之间进行抽样。初始任务、所有参考文件和所有提交文件都会被解析并作为上下文提供给法官。
    • 基于文档的文件(.pdf、.docx、.pptx、.xlsx 等)被解析为文本和图像。我们提取 .zip 文件并分别解析每个单独的文件。对于包含音频或视频文件的任务,比较将路由到Gemini 3.8 Flash,它可以本地处理这些模式。此上下文嵌入在评分提示中,要求法官确定提交内容 A 和 B 中哪一个对任务的响应更好。
    • 最终评分:我们的最终 Elo 分数是通过所有成对比较的最大似然估计计算得出的 Bradley-Terry 评分(平局算作双方各半胜),锚定于 1600 处的 DeepSeek V4.1 Flash (max)。95% 置信区间的计算公式为用于量化评级不确定性的三明治估计器。

AutomationBench-AA

  • 说明:AutomationBench-AA 是 Artificial Analysis 对 Zapier 的 AutomationBench 的评估实现。它使用 REST API 作为工具接口,测试模型能否完成跨多个模拟业务应用的真实 SaaS 工作流。
  • 论文: https://arxiv.org/abs/2604.18934
  • 排行榜: https://zapier.com/benchmarks
  • 代码仓库: https://github.com/zapier/AutomationBench
  • 数据集:
    • 我们评估来自 AutomationBench 数据集版本 1.0.6 的私有 657 任务保留分割
    • 这些任务涵盖六个业务领域:财务、HR、营销、运营、销售和支持
    • 它们在模拟应用程序环境中运行,其中包括 Gmail、Google Sheets、Slack、Salesforce、Zendesk、Jira 和 HubSpot 等产品
  • 实现:
    • 我们在 AutomationBench 多轮环境中运行每个任务一次,上限为 50 轮。模型使用 API 工具集,通过结构化工具调用发现和调用所需的 REST 端点
    • 我们将每个 AutomationBench 断言分类为目标(必须由智能体实现)或护栏(最初通过且不得被智能体破坏)
    • 使用对最终环境状态的编程检查来对目标和护栏进行评分。 AutomationBench-AA不使用单独的LLM评委进行评分
    • 对于主要得分,如果模型违反任何护栏,则任务将收到 0 分。如果没有违反护栏,任务将收到模型完成的目标的百分比。出错的任务也得 0 分
    • 每项任务都属于一个业务域,因此域细分是任务集的互斥子集。应用程序故障并不相互排斥:一项任务可能涉及多个应用程序,因此其目标和护栏断言可能会影响多个应用程序

编程

Terminal-Bench 4.0

  • 说明: Terminal-Bench 4.0 版本,由Stanford University研究人员、Laude Institute和开源社区开发。涵盖软件工程、系统管理、数据处理、模型训练和安全性,每项任务都由自己的验证套件进行评分。
  • 排行榜: https://www.tbench.ai/?version=4
  • 数据集: https://github.com/harbor-framework/terminal-bench
  • 实现:
    • 我们使用 mini-swe-agent 工具评估完整的 Terminal-Bench 4.0 数据集(66 个任务),每个任务的 pass@1 平均得分超过 3 次重复
    • 每个任务都有自己的一组测试。我们遵循 Terminal-Bench 方法:仅当每个测试都通过时,任务才通过,评分在每个任务自己独立的验证器容器中运行,与智能体环境隔离,超过超时的验证器将被视为失败
    • 我们对智能体的评估应用以下约束:
      • 最大智能体步数限制为 500
      • 任务超时和沙箱资源遵循上游任务定义
    • 所有其他智能体配置均遵循 mini-swe-agent 默认值,包括交互式迷你配置和提示、本机 bash 工具,并且没有上下文压缩或摘要:智能体始终会看到其完整记录

SciCode

  • 说明: Python 编程来解决科学计算任务。
  • 论文: https://arxiv.org/abs/2407.13168
  • 数据集: https://scicode-bench.github.io/
  • 实现:
    • 我们使用提示中包含的科学家注释的背景信息进行测试
    • 我们报告子问题级别评分
    • Pass@1 评估标准
    • SciCode 步骤脚本在隔离执行器上进行评分,执行超时为 300 秒(数据集 v1.0.1)

通用

AA-Omniscience

  • 说明: AA-Omniscience 是一种知识和幻觉基准,用于衡量事实可靠性、奖励精确知识并惩罚不正确的猜测或幻觉。它提供了对模型在不同知识领域区分已知和未知的能力的详细评估。
  • 论文: https://arxiv.org/abs/2511.13029
  • 数据集: https://huggingface.co/datasets/ArtificialAnalysis/AA-Omniscience-Public
  • 实现:
    • 该基准测试由涵盖 42 个主题的 6,000 个问题组成,包括商业、人文与社会科学、健康、法律、软件工程以及科学、工程和数学。
    • 使用AA-Omniscience Index对模型进行评分,该指数为正确答案评分,为幻觉反应扣分,并保持弃权中立,奖励对错误猜测的弃权
    • 每个答案均分为CORRECT、INCORRECT、PARTIAL_ANSWER或NOT_ATTEMPTED 基于模型的响应和参考答案。使用GPT-5.6 Luna (medium)作为评审模型
    • 纳入 Intelligence Index: AA-Omniscience为Intelligence Index贡献了两个组成部分:(1) 准确性 - 正确答案的比例,权重为总体指数的 10%,以及 (2) 非幻觉率 - 计算方法为 1 减去幻觉率,加权为整体指数的 5%(AA-Omniscience 的 15% 份额)。

GDP.pdf

  • 说明:Artificial Analysis 对 Surge AI 的 GDP.pdf 的实现。该基准测试评估模型能否对真实工作中的长篇专业文档进行推理,并满足任务特定的标准。
  • 论文: https://arxiv.org/abs/2607.11192
  • 数据集: surgeai/GDP.pdf
  • 评估集::涵盖 10 个专业领域的 100 项任务,基于 4,592 个 PDF 页面,并根据 1,275 条标准进行评分。我们会尝试每项任务五次,因此每个模型的固定分母为 500 次尝试。
  • 文档准备和交付:我们使用LiteParse 2.5.0准备每个源PDF,并启用英文OCR。每个模型都会收到每个页面的完整提取文本。当端点支持时,我们首先按页面顺序发送页面图像,然后发送任务并在单个用户消息中完成提取的文本。没有图像输入的模型仅接收文本。我们以 150 DPI 渲染页面图像,并在模型上下文或提供商有效负载限制需要时将其降低到最低 72 DPI。我们还可以将不透明页面从 PNG 转换为 JPEG。当端点限制一个请求可以携带的图像数量时,我们将页面组合成复合图像,每个图像首先两页,上限更严格时最多可达四页,每个单元格都标有其页码。每张图像过去四页,图像仅覆盖前几页;其余页面保留在提取的文本中。当任一应用时,提示会告诉模型页面图像是合成的,以及图像覆盖停止的页码。如果图像在适应后无法适应上下文或请求大小限制,我们将单独发送完整的提取文本。我们不会截断或总结提取的文本。模型无需浏览或工具即可一次性回答。

    与 Surge AI 的实现不同,我们不使用 API 的文档输入功能。这些功能对用户不透明,位于模型层之上,因此可能引入由 API 产品决策和提供商造成的差异,无法在同等条件下比较模型。

  • 评审: GPT-5.6 Luna Medium 独立评审每个标准。每次调用都会收到任务提示、参赛者答案和一项标准,但不会收到源 PDF 或参赛者身份。仅当每个标准都有结论时,我们才会接受任务评分。我们将错误、缺失尝试和终端输入失败评分为零。
  • 报告的指标:
    • All-pass是主要指标:所有 500 次尝试中每项标准都通过的比例。
    • Mean Pass是次要指标:我们计算每次尝试的标准通过率,然后在任务和重复之间以相等的权重对其进行平均。
    • 各领域的细分结果采用同样的按任务等权汇总的 Mean Pass 计算方式。我们不报告各领域的 All-pass。
  • 成本与速度的统计范围:公布的成本仅包含被评估模型的调用,不包含评审模型调用、PDF 预处理和 OCR。我们根据输出 Token 用量和模型输出速度估算每项任务的耗时。该估算不包含评审调用、PDF 预处理和 OCR,因此并非端到端评估耗时。

与 Surge AI 实施的差异

Artificial AnalysisSurge AI
文件输入使用 LiteParse 和 OCR 提取的文本,以及支持图像模型的页面图像原始 PDF 发送到提供商的文档输入
评审模型GPT-5.6 Luna MediumGemini 3.5 Flash

任务集是共享的,但文档输入和判断不同,所以Artificial Analysis和Surge分数没有直接可比性。

AA-LCR v1.1

  • 说明:通过测试模型对多个长文档的推理能力来评估长上下文性能(使用 cl100k_base 分词器测量,约 100k Token)。
  • 相较 AA-LCR 的变化:新增系统提示词以明确评分说明,修正了 16 个参考答案,并使用 GPT-5.6 Luna (medium) 评分。分数无法与 v1.0 直接比较。
  • 数据集: https://huggingface.co/datasets/ArtificialAnalysis/AA-LCR
  • 实现:
    • 100 道高难度文本题,涵盖 7 类文档:公司报告、行业报告、政府咨询文件、学术文献、法律文件、营销材料和调查报告。
    • 每道题约有 100k 输入 Token(使用 cl100k_base 分词器测量),因此模型须支持至少 128K 的上下文窗口,才能获得此基准测试的分数。运行整个基准测试需要约 3M 不重复的输入 Token,涉及约 230 份文档(输出 Token 数通常因模型而异)。
    • 使用 GPT-5.6 Luna (medium) 作为答案等价性判定器,以 pass@1 方式评估模型响应。

科学推理

HLE (Humanity's Last Exam)

  • 说明:Centre for AI Safety(由Dan Hendrycks领导)的最新前沿学术基准。
  • 论文: https://arxiv.org/abs/2501.14249v2
  • 数据集: https://huggingface.co/datasets/cais/hle
  • 实现:
    • 2,158 数学、人文和自然科学领域的纯文本问题(自 2025 年 5 月修订版起,总共包含 2,500 个问题 - 我们使用纯文本子集来实现模型之间的最大可比性)
    • 我们注意到,HLE 作者透露,他们的数据集管理过程涉及基于 GPT-4o、Gemini 1.5 Pro、Claude 3.5 Sonnet 测试的对抗性选择问题, o1、o1-mini 和 o1-preview(后两个仅适用于纯文本问题)。因此,我们不鼓励将这些模型与 HLE 管理过程中未使用的模型进行直接比较,因为数据集可能对管理过程中使用的模型存在偏见。
    • 使用改编自原始 HLE 论文的等式检查器 LLM 提示进行评估,使用 GPT-5.6 Luna (medium),使用 pass@1 评分(在下面找到提示)

CritPt

  • 说明: 研究级物理推理基准,涉及多个子领域的未发表的前沿物理问题。
  • 论文: https://arxiv.org/abs/2509.26574
  • 网站: https://critpt.com/
  • 代码仓库: https://github.com/CritPt-Benchmark/CritPt
  • 数据集: https://huggingface.co/datasets/CritPt-Benchmark/CritPt
  • 实现:
    • 我们与 CritPt 团队合作,为所有 70 个测试集挑战(不包括示例挑战)实施“挑战”级组件
    • 我们对每个问题重复 5 次,评分为 pass@1
    • 使用两步解析方法调用模型,其中第一步要求模型通过推理完成挑战,第二步将响应格式化为预期的代码格式以进行评分(请参阅 CritPt 评估页面上的示例提示进行解析)
    • Token使用和成本估计反映了这两个步骤(推理和答案解析)
    • 答案格式包括数值、SymPy中的符号表达式和Python函数(使用测试用例评估)
    • 官方的 CritPt 评分服务器用于评估所有挑战响应的正确性。评分 API 访问权限根据具体情况授予经批准的实验室和研究人员 - 发送电子邮件critpt@artificialanalysis.ai提出请求,并参阅Artificial AnalysisAPI文档了解详细信息

补充评测详情

智能体

Harvey LAB-AA

  • 说明:Harvey LAB-AA 是 Artificial Analysis 对 Harvey 的 Legal Agent Benchmark (LAB) 的实现,使用 Harvey 的数据集,其中包含涵盖 24 个法律执业领域的 120 项非公开任务。每项任务中,智能体在沙箱内阅读案件文件,并生成法律工作成果,包括备忘录、披露安排、证言摘要、修订标记文档等。一个 LLM 评审根据任务专用评分量表逐项评判成果。量表中的每个标准都是独立的二元通过/不通过标准,以此全面衡量智能体处理真实法律工作的能力。
  • 示例数据集::资源管理器中显示的五个公共示例任务取自 Harvey 的公共示例,网址为 https://github.com/harveyai/harvey-labs。主要数据是在 Harvey 的私人 120 任务数据集上生成的,该数据集未公开发布。
  • 智能体运行框架: https://github.com/ArtificialAnalysis/Stirrup
  • 实现:
    • 每项任务都是 24 个实践领域之一的独立法律工作任务。智能体接收任务指令、一组只读输入文档以及它必须生成的可交付成果的确切文件名,然后在一次运行中完成任务,而无需在执行期间进行实时交互或迭代反馈。
    • 输入文档是任务的案例材料 - 合同、协议、备忘录、成绩单和其他法律记录 - 在沙箱中以只读方式呈现。智能体将它们复制到工作文件夹中,使用沙箱中可用的文档处理工具读​​取它们,并将其可交付成果(通常为 .docx、.xlsx 或 .md)直接写入任务指定的确切文件名下的主目录中。
    • 所有模型均使用我们的开源智能体工具Stirrup运行。
      • 轮次:每个任务智能体最多运行 200 轮。
      • 工具:在该工具中,为模型提供了一个沙盒代码执行环境,并且还为具有视觉功能的模型提供了一个图像查看器工具,该工具可以从沙箱中读取图像文件作为模型的本机图像Token。沙箱无法访问互联网,因此智能体只能使用提供的输入文档和映像中预装的软件。
      • 沙箱::每个任务都在一个独立的 Linux 沙箱中运行,该沙箱由共享智能体评估基础映像 (Debian + Python 3.13) 构建,并配有文档处理工具,例如 Pandoc、poppler/pdftotext、LibreOffice、python-docx、预安装了 python-pptx、openpyxl、pdfplumber、PyMuPDF 和 markitdown。任务的输入文档在运行时处于只读状态;各个 shell 命令将在 20 分钟后终止。
      • 任务结束工具::一个完成工具,智能体调用该工具来提交摘要及其可交付成果的绝对路径(验证为实际文件,而不是目录或丢失的路径),以及一个abandon_task_finish(放弃)工具,仅当它断定任务确实不可能时才调用该工具。
    • 与 Harvey 基准的差异: Harvey LAB-AA 是Artificial Analysis的独立重新实现,因此我们的数据不能直接与 Harvey 自己发布的结果进行比较。主要区别:
      • 提交的内容必须与任务说明中指定的文件名完全匹配。几乎未命中的文件名被视为未生成,这比 Harvey 的尽力匹配更严格,并且可能会降低我们相对于他们的分数。
      • 仅当没有产生任何可交付成果时,一项标准才会彻底失败,而无需向法官展示。部分提交(其中存在一些标准声明的文件)仍会被判断,任何缺失的文件都会被标记为不存在。
      • 使用Gemini 3.1 Pro作为评审模型。
      • 我们在 E2B 沙箱中运行 Stirrup 的本机 shell 工具,并使用 Artificial Analysis 编写的智能体和判断提示,而不是 Harvey 的沙箱和自定义工具。
      • Harvey 的原始实现为智能体配备了自定义工具和文档生成技能脚本(例如,用于生成 .docx、.xlsx 和 .pptx 文件)。我们不提供这些,因此智能体使用沙箱中的通用工具生成这些文件。
    • 每项任务都有一份评分量表,包含权重相同、彼此独立的二元通过/不通过标准。每个标准均根据从该标准指定的成果文件中提取的文本来评判。评审只使用文本:评审模型查看成果的提取文本、任务说明和该标准的匹配条件,给出严格的通过或不通过判定,不给部分分数。
    • 报告了两个主要指标:标准通过率,即可交付成果满足的原子通过/失败评分标准的比例(高于标准的平均值),以及全部通过率,即每个标准都通过且没有部分计分的任务比例。标准通过率是整个网站显示的默认指标。
    • 提示词:跨生成和评分使用的提示:
      • 智能体系统提示词:
        You are an AI agent completing a professional legal-work task. Use the tools provided to read the input documents, produce the requested deliverable files, and submit them within {max_turns} steps.
        
        When you are done you must call the `{finish_tool_name}` tool as your final step, passing a brief summary of what you accomplished and a list of absolute paths for every deliverable file.
        
        If you have genuinely concluded that the task cannot be completed - for example because required inputs are missing or a hard dependency is unavailable - call the `{abandon_task_finish}` tool with a brief reason instead. Do not use it to escape difficulty.
        
        You cannot interact with the user during the task. Make reasonable assumptions when needed and record them in your finish summary.
      • 智能体任务提示词:
        <execution_context>
        ## Sandbox
        
        You operate inside an isolated Linux sandbox through the `code_exec` tool, which runs shell commands and lets you read, create, and edit files. Commands run as the unprivileged user `user` (UID 1000).
        
        Files you write persist on disk across calls, but **shell state does not**: each command runs in a fresh shell, so no working directory, environment variable, or other shell state carries from one call to the next. Always use absolute paths for files, and do not navigate with `cd` across calls - a `cd` in one command is gone by the next. When a step genuinely needs a different directory, chain it into the same command (e.g. `cd /home/user && python build.py`).
        
        ## No network
        
        The sandbox has no outbound connectivity, and there is no proxy, allowlist, or flag that turns it on - treat it as permanently offline. Anything that reaches the internet will fail: package installs (`pip`, `npm`, `apt`), remote `git`, and any HTTP/HTTPS request.
        
        Recognise a network block by its error signature - failed name resolution (`Could not resolve host`, `Temporary failure in name resolution`), an unreachable route (`Network is unreachable`), or a stalled connection - rather than guessing. When you see these the failure is structural: do not retry the same call or hunt for a workaround (mirrors, alternate hosts, cached copies). Re-plan using only what is already installed and the files in your workspace.
        
        ## Filesystem
        
        - Writable: everything under `/home/user/` plus `/tmp`. Use these for deliverables, intermediate files, and caches.
        - Read-only inputs: `/home/user/documents` - the task's input documents. Copy these into a working folder before transforming them rather than editing them in place.
        
        ## Runtime
        
        A document-processing stack is already installed - check what is present before assuming a gap:
        
        - **Reading inputs**: `pandoc` or `python3 -c "import docx; ..."` for Word; `pdftotext` or `python3 -c "import pdfplumber; ..."` for PDFs; `python3 -c "import openpyxl; ..."` for Excel; `markitdown <path>` as a general-purpose extractor for .docx, .xlsx, .pptx, and .pdf. `libreoffice` (the `soffice` binary) is also installed - use `soffice --headless --convert-to pdf <path>` to convert any Office format (.docx/.xlsx/.pptx, including legacy .doc/.xls) when the python parsers fall short.
        - **Producing deliverables**:
          - `.docx`: `python3 -c "from docx import Document; ..."` or `pandoc -o out.docx`.
          - `.xlsx`: `python3 -c "import openpyxl; ..."`.
          - `.md` and other plain text: write directly with `cat`/`tee`/your script.
        - Check availability with `pip show <pkg>` or `which <tool>` rather than installing - installs fail offline, but the document stack above is already present.
        - Commands are terminated after {command_timeout_minutes} minutes. Keep them bounded, persist intermediate results to disk, and split long jobs into smaller steps.
        
        ## Submitting your work
        
        Finish by calling the `{finish_tool_name}` tool - anything not submitted through it is not graded. Your call must include:
        1. A short summary of what you accomplished.
        2. Absolute paths to every deliverable (files only, not folders).
        
        Save each deliverable directly in `/home/user` under the exact filename the task asks for - not in a subdirectory. Save deliverables as ordinary, visible files - do not leave the only copy of your work in a dot-prefixed file or directory (e.g. `.report.docx`, `.output/report.docx`). Assume your files will be opened and edited by others after submission.
        
        If the task genuinely cannot be completed, call the `{abandon_task_finish}` tool with a brief reason instead. Use it only when you have concluded the work is impossible - not to escape a difficult task.
        </execution_context>
        
        <task>
        ### {title}
        
        {instructions}
        </task>
        
        <deliverables>
        Submit these files, by exact name, saved directly in `/home/user`:
        {expected_deliverables}
        </deliverables>
        
        Please begin working on the task now.
      • 评判系统提示(任务上下文和工作产品):
        You are evaluating a legal AI agent's work product against one binary quality criterion.
        
        <task_context_for_work_product>
        The work product below was produced for this legal task. Use the task only as context for what the deliverables were meant to address - judge the work product, not the task.
        
        {task_title}
        
        {task_instructions}
        </task_context_for_work_product>
        
        <work_product>
        {agent_output}
        </work_product>
      • 评判标准提示:
        <criterion>
        <title>
        {criterion_title}
        </title>
        <match_criteria>
        {match_criteria}
        </match_criteria>
        </criterion>
        
        Return `pass` only if the work product satisfies the criterion as described; otherwise `fail`.

APEX-Agents-AA

  • 说明: APEX-Agents-AA 是 Artificial Analysis 对 Mercor 的 APEX-Agents 基准测试的独立实现。它评估跨投资银行、管理咨询和法律等专业服务环境中的长期、跨应用智能体工作。
  • 论文: https://arxiv.org/abs/2601.14242
  • 数据集:
    • 我们的评估基于来自 https://huggingface.co/datasets/mercor/apex-agents 的公共 APEX-Agents 数据集
    • 我们评估了公开的 480 个任务版本中的 452 个任务(不包括具有外部运行时依赖性的 Investment Banking Worlds 244 和 246)
  • 实现:
    • 每个任务重复运行 3 次,并使用 pass@1 进行评分 - 仅当所有评分项都满足时,重复才会通过,排行榜分数是重复的平均通过率
    • 所有模型均使用我们的开源智能体运行框架 Stirrup 运行,每个任务的轮次上限为 200 次
    • 智能体在Archipelago环境内运行,并通过其网关公开的MCP服务器访问工作场所工具
    • 智能体从一个小型元工具工具带开始,并且必须使用以下方式显式管理 MCP 支持的工具:
      • 列出工具 – 显示当前可用的工具
      • 检查工具 – 在添加工具之前检查它
      • 添加工具 – 为智能体提供由 MCP 支持的工具
      • 删除工具 – 删除不再需要的工具
    • 智能体人还收到:
      • 待办事项写入 - 创建或更新客服人员的待办事项列表。它可以替换完整列表或通过待办事项 ID 合并更新,并且所有待办事项必须在最终提交被接受之前完成或取消
      • 完成 - 提交客服人员的最终答案以及完成状态。这是提交最终答案的唯一方式,并且只有完成提交后才能进行评分
    • MCP 工具调用有 60 秒的超时时间。当需要使用 20k 字符头部和 5k 字符尾部摘录将工具输出截断为 24k Token预算时。图像输入在返回到模型之前被压缩到大约 1 MP
    • 评分是通过Archipelago本地文件评分器在本地运行的。每次重复都会根据任务评分标准,使用通过“完成”提交的最终答案以及初始和最终世界快照之间的文件系统差异进行评分。仅当每个主要项都得到满足时,重复才会通过。以“低”推理的Gemini 3 Flash作为LLM评判

AA-AnalystAgent

  • 说明: AA-AnalystAgent 是Artificial Analysis的端到端数据分析基准。智能体回答跨业务和科学领域的定量问题,使用提供的源电子表格和文档作为主要输入,并在沙盒代码执行环境中执行 Python。 AA-AnalystAgent 作为独立排行榜进行报告,不是Artificial Analysis Intelligence Index的组成部分。
  • 智能体运行框架: https://github.com/ArtificialAnalysis/Stirrup
  • 数据集:
    • AA-AnalystAgent 是一个私人持有的基准测试;问题集、参考答案和源文件不公开发布,以限制污染风险
    • 涵盖 14 个商业和科学领域的 80 个定量问题,包括环境报告、贸易和商品统计、医疗保健支出报告、水文和天气数据、政府拨款、能源成本模型、财务模型和项目时间表
    • 问题涵盖五个功能工作流程原型,涵盖实际分析师工作的范围:来源查找和诊断、过滤器和总计、比率、趋势和敏感性、损益建模以及现金、资产负债表和估值建模
    • 每个问题都与一个包含参考电子表格和文档(xlsx、docx)的文件夹配对,这些文件夹已上传到智能体的工作区中。智能体提供人工编写的参考答案,并由评分者在评分时使用
    • 参考答案由Artificial Analysis独立验证
  • 实现:
    • 每个问题都有 5 次独立的重复。排行榜分数为 pass^5 — 5 次尝试中每一次正确回答问题的比例:
      其中,如果对问题 j 的尝试 i 正确,则 pji = 1,否则为 0,n 是问题数。这与我们在其他评估中使用的 pass@1 评分不同:分析师智能体只有在其答案无需重新检查的情况下才有效,因此主要指标会奖励重现正确答案,而不是偶尔达到正确答案
    • 除了 pass^5 之外,我们还计算 pass@1(每次尝试的平均通过率,在所有重复中汇总)和 pass@5(至少一次尝试解决问题的比例),这将模型的可靠性与其上限分开
    • 所有模型均使用我们的开源智能体工具 Stirrup 作为智能体运行,每个任务的旋转上限为 100 次
    • 该智能体提供了一个小型工具集,涵盖在隔离的 Linux 沙箱(Python 3.12,安装了问题的参考文件和预安装的一组固定的标准 Python 数据分析库)中执行代码、URL 获取、支持视觉模型的图像查看以及最终答案提交。模型被指示仅提交答案值(例如仅数字或标签),无需解释
    • 每个答案都会根据保留的参考答案进行二进制正确或错误评级。每个单元格都会发送给LLM法官,以便对工件进行评分和法官成本核算保持统一。然后,确定性数字等价预检查会覆盖明确情况下的判断:在相同的单位约定中,在问题要求的精度下,等于参考值的答案是有保证的通过。预检查是单方面的——它永远不会失败——所以它无法解决的所有问题都会保留法官的裁决。如果法官在单元格上返回格式错误的响应,则预检查可以解决,预检查仍会记录通过。 Gemini 3 Flash (Reasoning)担任LLM评委

智能体提示:系统会提示智能体使用以下模板,插入问题的参考文件、任务和完成工具名称:

You are tasked with answering a data analysis question.

## Environment

The `code_exec` tool provides access to a Linux-based execution environment with a full file system where you can create, read, and modify files.

Python 3.12 is the default runtime. Use `python script.py` to run scripts.
The following Python packages are preinstalled (pinned versions):

- numpy 2.4.4, numpy-financial 1.0.0, pandas 3.0.2, scipy 1.17.1, polars 1.40.0
- matplotlib 3.10.8, seaborn 0.13.2
- scikit-learn 1.7.2, statsmodels 0.14.4
- openpyxl 3.1.5, xlrd 2.0.2, python-docx 1.2.0, formulas 1.3.4
- PyMuPDF 1.27.2.2, pdfplumber 0.11.9
- Pillow 12.2.0, requests 2.33.1, beautifulsoup4 4.13.4
- tqdm 4.67.3, tabulate 0.10.0, sympy 1.14.0

## Reference Files

The following reference files are available in your workspace:

<reference_files>
{reference_files}
</reference_files>

## Task

<task>
{task}
</task>

## Submitting Your Answer

When you have determined the answer, use the `{finish_tool_name}` tool to submit it.
Your answer should be a concise, direct response to the question.
If the question asks for a number, provide just the number.
If the question asks for a name or label, provide just that.
Do NOT include explanations in your answer — only the final answer value.

评分者提示:每个(模型、问题)响应都会发送给 LLM 法官,并显示以下提示,插入原始问题、保留的参考答案和智能体提交的答案。然后,数字预检查可能会推翻该判决,如上所述:

You are an expert evaluator grading a data analyst's response to a question.

Decide whether the response is correct or incorrect, judged against the reference answer and the standard a professional data analyst working in this question's domain would be held to. Focus on the substance of the answer, not prose style. Be objective and consistent, and give a brief explanation for your verdict.

First identify exactly what the reference requires — the specific value(s), item(s), or label(s) — and what the response actually commits to, then compare them directly before deciding.

Apply these conventions:

- Format directives are binding. If the question specifies a form or precision — a number of decimal places, "as a percentage", "to the nearest cent", a cell reference, particular units — the response must satisfy it. A right value in the wrong requested form is incorrect.
- Equivalent representations of the same value are correct. Thousands separators, currency symbols, surrounding whitespace, and trailing zeros are immaterial; a percentage and its decimal fraction (e.g. 12.84% and 0.1284) are the same value; adding or omitting a "%" sign never changes correctness when the digits already match the value the reference states; a quantity stated in the dataset's native units (e.g. thousands) matches the same amount written in full.
- Judge precision by value, to a sensible number of significant figures. When the question states a precision, require exactly that. When it does not, accept any answer that is a correct rounding of the reference value — reference answers often carry more decimal places than are meaningful (e.g. a dollar figure written as 64792.44714), and a competent analyst rounds sensibly, so do NOT reject an answer merely for having fewer decimals than the reference. Reject an answer only when its value genuinely differs from the reference (a wrong figure, not a coarser rounding of the same value) or when it discards so much precision that it misstates the quantity.
- Match every required item. If the question asks for more than one item (e.g. "which two tasks"), the response is correct only if it identifies exactly the reference's items. Judge the single set the response commits to and ignore hedged alternatives phrased as "(or ...)"; a response naming different items than the reference — however plausible — is incorrect.
- Honor explicit acceptance and rejection clauses in the reference answer. If the reference names specific values as acceptable or as not acceptable, follow it exactly.

## Question

{question_prompt}

## Reference Answer

{reference_answer}

## Response to Evaluate

{model_response}

EnterpriseOps-Gym-AA

  • 状态:独立评估(不属于Artificial Analysis Intelligence Index v4.3.2的一部分)
  • 说明: EnterpriseOps-Gym-AA 是 Artificial Analysis 对 ServiceNow 的 EnterpriseOps-Gym 基准的独立实现,该基准评估 AI 智能体在现实的企业工作流程中进行有状态的多步骤规划和工具使用。智能体通过工具操作实时企业系统,并根据底层数据库的最终状态而不是其确切的操作顺序进行评分。
  • 论文: https://arxiv.org/abs/2603.13594
  • 数据集: https://huggingface.co/datasets/ServiceNow-AI/EnterpriseOps-Gym
  • 智能体运行框架: https://github.com/ArtificialAnalysis/Stirrup
  • 领域:我们评估基准中全部八个企业领域的 oracle 模式任务:客户服务管理 (CSM)、人力资源 (HR)、IT 服务管理 (ITSM)、电子邮件、日历、Teams、Drive,以及在同一工作流中协调多个上述系统的混合任务。
  • 实现:
    • 每项任务都在隔离且可重置的沙箱中运行:相关企业系统作为独立评估环境服务器启动,各自通过实时 Model Context Protocol (MCP) 服务器提供工具,并由注入合成数据的任务专用 SQLite 数据库支持。每项任务克隆自己的数据库,使运行相互隔离且可复现。
    • 我们仅在其oracle工具模式中运行基准测试:为智能体提供任务所需的工具集,将计划和执行与工具检索隔离开来。源数据集的干扰工具模式未运行。
    • 所有模型均使用我们的开源智能体工具 Stirrup 在标准原因与行动工具使用循环中运行,每个任务的轮次上限为 100 轮。每个任务重复运行 3 次,主要得分是重复次数的平均值。
    • 评分是基于结果的。智能体完成后,每个任务数据库的最终状态都会被快照并使用基准的 SQL 验证器进行检查,这些验证器测试目标完成情况、状态和完整性约束、权限和流程合规性以及是否存在意外副作用。
    • 报告了两个指标。主要成功率是严格的pass@1:只有当任务通过了每一个验证者时才算成功。我们还报告验证者通过率,即通过的单个验证者检查的比例,作为更细粒度的辅助指标。
    • 与 ServiceNow 基准的差异: EnterpriseOps-Gym-AA 是我们的独立实现,在我们自己的 Stirrup 运行框架和智能体提示上运行,因此我们的数字不能直接与论文中报告的结果进行比较。

Terminal-Bench-Science 0.1

  • 状态:独立评估(不属于Artificial Analysis Intelligence Index v4.3.2的一部分)
  • 说明: Terminal-Bench-Science 是由Stanford University研究人员与Terminal-Bench和Harbor团队共同开发的开放式学术合作项目,其中包括全球各机构科学家的贡献。其研究工作流程任务取自真实的科学实践,并由专家对每一项进行策划和审查。每个任务都有自己的一组测试,智能体必须通过在终端中工作来通过这些测试。 0.1.0版本涵盖了生命、物理、数学、工程和地球科学。
  • 引用: https://doi.org/10.5281/zenodo.22110254
  • 排行榜: https://terminal-bench-science.ai/
  • 数据集: https://github.com/harbor-framework/terminal-bench-science
  • 实现:
    • 我们使用 mini-swe-agent 工具评估完整的 Terminal-Bench-Science 0.1.0 版本(70 项任务:19 项生命科学、17 项物理科学、17 项数学科学、9 项工程科学、8 项地球科学),并计算 pass@1 平均分每个任务重复 3 次以上
    • 每个任务都有自己的一组测试。仅当每个测试都通过时,任务才通过,并且评分在每个任务自己独立的验证器容器中运行,与智能体的环境隔离
    • 我们对智能体的评估应用以下约束:
      • 我们将智能体限制为 1,000 步
      • 任务超时和沙箱资源遵循上游任务定义
    • 所有其他智能体配置均遵循 mini-swe-agent 默认值

ITBench-AA

  • 状态:独立评估(不属于Artificial Analysis Intelligence Index v4.3.2的一部分)
  • 说明: ITBench-AA 是 Artificial Analysis 对 IBM 的 ITBench 基准测试的独立实现,用于评估站点可靠性工程 (SRE) 上的 AI 智能体: Kubernetes事件根本原因分析。
  • 论文: https://arxiv.org/abs/2502.05352
  • 代码仓库: https://huggingface.co/datasets/ArtificialAnalysis/ITBench-AA
  • 智能体运行框架: https://github.com/ArtificialAnalysis/Stirrup
  • 数据集:
    • 我们评估了 59 个 Kubernetes 事件任务:其中 40 个来自 IBM 的公共 ITBench SRE 版本,以及 ITBench 团队与我们共享的 19 个私有任务。主要得分是两个部分的平均分
    • 每个任务都是一个离线 Kubernetes 事件快照,其中包含警报、事件、跟踪、指标、日志和应用程序拓扑,烘焙到特定于场景的沙箱中并安装在 /home/user 下
  • 实现:
    • 每项任务重复运行 3 次。主要指标是完全召回时的精确率:只要遗漏任何参考答案中的根本原因实体,该次运行就得 0.0 分;否则,按提交的实体计算精确率。
    • 所有模型均使用我们的开源智能体运行框架 Stirrup 运行,每个任务的轮次上限为 100 轮。智能体循环通知语言模型在最后 20 个回合中其回合限制已接近。
    • 为智能体提供一个 run_shell 工具来检查快照,以及一个 finish 工具来提交其最终答案。它必须将结构化 JSON 诊断写入 /home/user/agent_output.json,其中包含负责事件的独立根本原因 Kubernetes 实体的最小集合,并为每个实体提供推理和证据,同时排除下游症状
    • 评分仅使用 LLM 判断将提交的 contributing_factors 标准化为真实规范实体和别名组。
    • 标准化后,真实别名组将合并到评分组中,因此 pod 等等效实体及其相应的部署/服务算作相同的预测。如果别名组的任何成员被标记为根本原因,则合并组将被评分为根本原因目标;预测同一别名组中的多个实体仅计算一次。
    • 如果遗漏任何根本原因评分组,则完全召回的精确度计算为 0.0。如果没有遗漏任何根本原因组,则为 true_positives / (true_positives + false_positives),其中不匹配的预测和映射到非根本原因组的预测将计为误报。
    • 使用具有中等推理能力的 GPT-5.5 作为评分器模型,用于将模型的输出与每个任务的基本事实进行比较
  • 生成提示:
    **Task**:
    
    You are an expert SRE (Site Reliability Engineer) and Kubernetes SRE Support Agent investigating a production incident from OFFLINE snapshot data.
    
    ====================================================================
    # INCIDENT SNAPSHOT DATA LOCATION
    ====================================================================
    Your incident data and working directory is located in
    - /home/user
    
    The final output must be written to /home/user/agent_output.json
    
    Available Python packages:
    - `drain3==0.9.11`
    - `numpy==2.4.5`
    - `pandas==3.0.3`
    
    Both `python` and `python3` are available and use the same environment.
    
    Your objective is to generate a **JSON diagnosis** identifying the root causes of the incident — the minimal set of independent Kubernetes entities whose failures directly explain the incident.
    
    Requirements:
    - Provide reasoning and evidence for every listed entity.
    - When the JSON file is ready, call the provided finish tool and submit `/home/user/agent_output.json`.
    
    All entities MUST use the format: `namespace/Kind/name`
    
    Examples:
    - `otel-demo/Deployment/ad` (Deployment named "ad" in namespace "otel-demo")
    - `otel-demo/Service/frontend` (Service named "frontend")
    - `cluster/Node/worker-node-1` (cluster-scoped resource)
    
    DO NOT include UIDs in the entity name.
    
    ====================================================================
    ## Output Format
    ====================================================================
    Output must consist solely of the final diagnosis in the specified JSON format below — do **not** include any additional text, markdown, or comments:
    
    ```json
    {
    "contributing_factors": [
      {
        "name": "namespace/Kind/name",
        "reasoning": "A short, clear, human-readable explanation for why this entity is a root cause. Reference evidence where possible.",
        "evidence": "Concise summary of supporting facts — relevant alerts, events, logs, traces, or metrics. Plain string."
      }
    ]
    }
    ```
    
    ====================================================================
    # RULES FOR INCLUSION
    ====================================================================
    
    **Only include an entity if both of the following are true:**
    
    1. **There is qualifying evidence** — it appears in at least one of: a firing alert, a Kubernetes event, an error/warning log line, a metric anomaly, or trace evidence directly tied to the incident window. A passing mention in an unrelated log is not sufficient.
    
    2. **It passes the irreducibility test** — you cannot fully explain its failure by pointing to another entity already in the list. Ask: *"If I remove this entity, does my explanation of the incident become incomplete?"* If yes, include it. If another entity already accounts for it, leave it out.
    
    **Do not include** downstream effects, symptoms, or intermediates — only the independent upstream causes.
    
    **Example (exhausted ResourceQuota blocking pod scheduling):**
    
    Causal chain: ResourceQuota exhausted → ReplicaSet cannot schedule pods → Deployment degraded
    
    - ✅ `otel-demo/ResourceQuota/otel-demo-mem-quota` — memory limit exhausted; directly blocks pod creation. Include.
    - ❌ `otel-demo/ReplicaSet/ad-7f9d4b` — failed only because the quota above was exhausted. Exclude.
    - ❌ `otel-demo/Deployment/ad` — degraded as a downstream consequence. Exclude.
    
    **Multiple entries are allowed only if they are truly independent** — two separate upstream causes that do not explain each other.
    
    When in doubt, prefer the most specific Kubernetes object that independently introduced the failure.
    
    ====================================================================
    # INVESTIGATION WORKFLOW
    ====================================================================
    
    ### Phase 1 — Context Discovery
    List available files (alerts, logs, events, topology).
    
    ### Phase 2 — Symptom Analysis
    Read all alert files. Compute:
    - Start time, End time, Duration, Frequency
    
    ### Phase 3 — Hypothesis Generation
    - Create initial hypotheses (e.g. "checkout pods OOMKilled", "redis latency spike").
    - Create a validation plan for each hypothesis.
    
    ### Phase 4 — Evidence Collection Loop
    - Use tools (and generated python code) to gather log, event, metrics, trace evidence.
    - Validate or refute each hypothesis using real data.
    - Explain firing alerts as soon as you find supporting evidence.
    
    ### Phase 5 — Causal Chain Construction
    Build a causal chain like
    `[Config Error] → [CrashLoop] → [Service Down] → [Frontend 5xx]`
    
    ### Phase 6 — Conclusion
    Ensure:
    - All alerts are explained in the reasoning/evidence for the root causes, but do not add downstream entities only to account for alerts
    - All included entities pass the irreducibility test
    - JSON is written to `/home/user/agent_output.json`
    - Call the finish tool and submit the file
  • 评分提示:
    You are an expert AI evaluator specializing in Root Cause Analysis (RCA) for complex software systems.
    
    You will be provided with:
    
    1. A **Ground Truth (GT)** JSON object containing entity definitions.
    2. A **Generated Response** JSON object containing predicted entities.
    
    Your job is only to normalize generated entities to ground-truth entities.
    
    Ground Truth fields such as `groups`, `aliases`, `filter`, and `kind` may appear either at the top level of `GT` or under `GT.spec`. Treat `GT.spec` as the ground-truth payload when present.
    
    -----
    
    ### Normalization Rules
    
    Before any downstream scoring can occur, you must accurately normalize entities from the `Generated Response` to the `Ground Truth`.
    
    This process must be based on **explicit evidence** from the entity's metadata.
    You must not infer or guess mappings based on an entity's position in a causal chain.
    
    Only normalize entities from `Generated Response.contributing_factors`.
    
    An entity from the `Generated Response` can only be mapped to a `Ground Truth` entity if a **Confident Match** can be established.
    
    **Definition of a Confident Match:**
    A generated entity is a confident match to a ground-truth entity only if its `name` field, or other explicit identifying metadata, clearly corresponds to the `filter` and `kind` of a ground-truth entity.
    
    **Alias Handling:**
    The `GT.aliases` field contains arrays of equivalent entity IDs.
    If a generated entity clearly matches an entity in an alias group, you may normalize it to the matching GT entity ID from that alias group.
    
    **Workload Kind Equivalence:**
    Treat `Deployment` and `Pod` as equivalent for normalization when the namespace and workload name correspond. For example, `otel-demo/Deployment/checkout` is a confident match for a GT `Pod` entity whose filter matches checkout pods in the `otel-demo` namespace.
    
    **Entity Name Format:**
    Generated entities use the format `namespace/Kind/name`.
    
    Examples:
    - `otel-demo/Deployment/flagd`
    - `otel-demo/Service/frontend`
    - `otel-demo/Pod/checkout-8546fdc74d-d68cn`
    
    Confident match examples:
    - A generated entity with `name: "otel-demo/Service/adservice"` is a confident match for the GT entity with `id: "ad-service-1"` and `filter: [".*adservice\\\\b"]`.
    - A generated entity with `name: "otel-demo/Service/adservice"` can match `ad-pod-1` only if the GT alias set makes that link explicit, for example `["ad-pod-1", "ad-service-1"]`.
    - If `GT.aliases` contains `["load-generator-pod-1", "load-generator-service-1"]`, then normalizing a generated `load-generator-service-1` match to that alias group is valid.
    - A generated `chaos-mesh/Schedule/...` entity whose name matches a GT filter is a confident match for the spawned chaos resource of any kind, provided name and namespace correspond.
    - A generated entity with `name: "67cbd7fe98a0776a"` and no other identifying evidence is not a confident match.
    
    If a generated entity does not have a confident match, leave it unmatched and set its normalized GT entity ID to `null`.
    
    Preserve the original order of the generated `contributing_factors`.
    
    -----
    
    ### Output Format
    
    Return only a single JSON object with this shape:
    
    ```json
    {
    "contributing_factor_entities": [
      {
        "submitted_entity_name": "namespace/Kind/name",
        "normalized_gt_entity_id": "ground-truth-entity-id-or-null",
        "reasoning": "brief explanation of why this is a confident match or why it is unmatched"
      }
    ]
    }
    ```
    
    Rules:
    - Include one item for every generated entity in `contributing_factors`.
    - Preserve input order.
    - Use `normalized_gt_entity_id: null` when there is no confident match.
    - Return only valid JSON.
    
    Given the following Ground Truth (GT) and Generated Response, normalize the generated contributing-factor entities to the Ground Truth.
    
    ## Ground Truth (GT):
    ```json
    {ground_truth}
    ```
    
    ## Generated Response:
    ```json
    {generated_response}
    ```
    
    ## Task:
    1. Look only at `Generated Response.contributing_factors`.
    2. For each such entity, determine whether there is a confident match in the Ground Truth.
    3. If there is a confident match, return the matched ground-truth entity ID.
    4. If there is not a confident match, return `normalized_gt_entity_id: null`.
    5. Do not score anything. Return only the normalization result JSON.

通用

IFBench

  • 状态:独立评估(不属于Artificial Analysis Intelligence Index v4.3.2的一部分)。 IFBench 在 v4.1 中已从Intelligence Index中删除,但我们继续在新模型版本上运行它。
  • 说明: 评估模型在单轮中遵循精确指令的能力的基准。它测试广泛的技能,包括计数、格式化和句子处理。
  • 论文: https://arxiv.org/abs/2507.02833
  • 数据集: https://huggingface.co/datasets/allenai/IFBench_test
  • 实现:
    • 使用单轮 IFBench 数据集,包含 294 个问题
    • 我们对每个问题重复 5 次,评分为 pass@1
    • 我们使用 allenai/IFBench 的官方源代码评估响应
    • 我们采用松散评估模式来稳健地评估指令遵循,通过检查模型输出的几种变化(例如,有或没有第一行和最后一行,以及删除星号)来解释无关的文本或格式。
    • 我们的分数代表提示级别的准确性(所有问题和重复的平均值)
    • 我们不使用 IFBench 的多轮版本,它使用不同的数据集

MLCR-AA(医疗长上下文推理)

  • 状态:独立评估(不属于 Artificial Analysis Intelligence Index v4.3.2);是 Artificial Analysis Healthcare & Medical Index 的组成部分
  • 说明:MLCR-AA 是 Artificial Analysis 对 Wisedocs 开放基准 MLCR (Medical Long Context Reasoning) 的评估,衡量模型能否对冗长且分散的医疗记录进行推理,完成理赔专业人员在审查保险和医疗案例时所做的多文档综合分析,例如重建时间线、因果关系、治疗模式和理赔相关性。
  • 代码: Wisedocs-AI/medical-long-context-reasoning
  • 公共数据集: Wisedocs/mlcr-dataset
  • 关键细节:
    • 大约 25,000-64,000 个Token的真实合成医疗案例
    • 问题分为六个难度等级,从定位单个事实到专家级临床综合和复合、多部分推理
    • 通过简洁性门并包含答案的回答由三名 LLM 评委组成的小组进行评分;准确性和完整性均由多数票决定
    • 长度超过参考答案五倍的回答无法通过简洁性门,并且在不进行判断的情况下得分为零;仅当响应通过该门并被判定为完整且准确时,总体通过率才算作响应
    • Artificial Analysis评估一组私人保留的两种最难的问题类型(专家级临床综合和复合、多部分推理):60 个问题,每个运行 3 次重复。该私有数据集与公开发布的数据集是分开的
    • Artificial Analysis将总体通过率(仅当回答简洁且被认为完整且准确时才计分)报告为主要分数。判断准确性和完整性是判断答案中的条件率;简洁涵盖所有答复。这些细分与主要分数一起显示; pass@1

其他

Global-MMLU-Lite

  • 状态:独立评估(不属于Artificial Analysis Intelligence Index v4.3.2的一部分);为Artificial Analysis Multilingual Index提供支持
  • 说明: MMLU 的轻量级多语言版本,旨在评估各种语言和文化背景下的知识和推理技能。
  • 数据集: CohereLabs/Global-MMLU-Lite
  • 关键细节:
    • 约 6,000 个问题(每种支持的语言约 400 个问题)
    • 多项选择(4个选项)
    • 正则表达式提取,pass@1

MMMU Pro

  • 状态:独立评估(不属于Artificial Analysis Intelligence Index v4.3.2的一部分);多模态(视觉)推理基准
  • 说明:增强的 MMMU 基准,消除了捷径和猜测策略,可以更严格地测试跨 30 个学科的多模态模型。
  • 数据集: MMMU/MMMU_Pro
  • 关键细节:
    • 1,730 个问题
    • 多项选择(10个选项)
    • 正则表达式提取,pass@1

历史评测

我们已经废弃或取代了评估。我们将他们的方法保留在这里,以供参考和历史可比性;它们不再是Artificial Analysis Intelligence Index或我们主动报告的一部分。

GPQA Diamond(研究生水平抗谷歌搜索问答基准测试)

  • 状态::在 v4.2 中从Artificial Analysis Intelligence Index中删除,一直是 v4.1.1 及之前版本的组成部分。我们仍然在新模型版本上运行它,并将其作为独立评估进行报告。
  • 说明:科学知识和推理基准。
  • 子集::选择钻石子集(198 个问题)以实现最大准确度和判别力
  • 论文: https://arxiv.org/abs/2311.12022
  • 数据集: https://github.com/openai/simple-evals/blob/main/gpqa_eval.py
  • 关键细节:
    • 涵盖生物学、物理和化学的 198 个问题 - 我们测试了完整 GPQA 数据集(总共 448 个问题)的 GPQA Diamond 子集,该子集被原作者定义为最高质量的子集,其中专家都回答正确,大多数非专家回答错误
    • 4选项多项选择格式
    • 基于正则表达式的答案提取,带有 pass@1 评分(下面的提示和正则表达式)

𝜏³-Banking

  • 状态::在 v4.3 中从Artificial Analysis Intelligence Index中删除,一直是 v4.2 及之前版本的组成部分。我们仍然在新模型版本上运行它,并将其作为独立评估进行报告。
  • 说明:Sierra 开发的 𝜏-Knowledge 框架中的金融科技客户支持领域,评估智能体能否将大型非结构化知识库的检索与通过工具执行的多步骤账户变更协调起来。
  • 论文: https://arxiv.org/abs/2603.04370
  • 博客: sierra.ai/blog/bench-advancing-agent-benchmarking-to-knowledge-and-voice
  • 数据集: https://github.com/sierra-research/tau2-bench
  • 实现:
    • 智能体处理约 700 个相互关联的策略文档(约 195K Token,21 个产品类别),并且必须找到相关策略、对其进行推理,并执行多步骤的工具调用序列 - 包括仅在文档中引用而不是明确列出的工具
    • 我们评估完整的 𝜏³-Banking 任务套件(97 个任务),每个任务重复 5 次,并报告重复次数的平均值 pass@1,运行上游 tau2-bench v1.0.1 数据集和评分器
    • 结果是根据实际后端数据库状态(例如,是否提出争议或是否发放临时信用)而不是对话质量进行评分
    • 我们使用GPT-5.4 Mini (medium reasoning)作为用户模拟器和自然语言断言判断
    • 为了对银行语料库进行知识检索,我们在原始 𝜏-Bench 工具中启用 BM25 词法搜索和 grep (bm25_grep 模式)
    • 我们将每次任务运行限制为最多 200 步,这是 𝜏-Knowledge 文本模式运行的默认参考设置。这里的“步”采用 𝜏-Bench 运行框架的定义,指模拟中传递的每条消息,包括用户模拟器的轮次,而不仅是被评估模型的轮次。

Terminal-Bench 2.1

  • 状态:在Intelligence Index v4.3 中被Terminal-Bench 4.0 取代,一直是v4.2 及之前版本的组成部分。它仍然是Coding Index的一部分。
  • 说明: 经过验证的 Terminal-Bench 更新,由Stanford University研究人员、Laude Institute和开源社区开发。在软件工程、系统管理、数据处理、模型训练和安全方面保留相同的 89 个策划任务,并通过环境和指令修复来使分数反映智能体能力而不是环境差距。
  • 论文: https://arxiv.org/abs/2601.11868
  • 排行榜: tbench.ai/leaderboard/terminal-bench/2.1
  • 实现:
    • 我们在 E2B 沙箱环境中使用 Terminus 2 智能体工具评估完整的 Terminal-Bench 2.1 数据集(89 个任务),每个任务的 pass@1 评分平均超过 3 次重复
    • 每个任务都附带一个验证套件,智能体必须通过与终端交互来满足该套件的要求 - 仅当每个测试都通过时,任务才被视为成功
    • 我们对智能体的评估应用以下约束:
      • 最大“情节”(模型在终端检查当前状态并计划一系列下一步行动)限制为 250
      • 每个任务智能体超时设置为两小时(7,200 秒),或者任务自己指定的超时(较长,远高于典型任务持续时间)
    • 在我们的测试中,这些约束主要限制模型陷入不成功循环的情况,并且由于这些约束,我们没有看到一致的性能差异

Terminal-Bench Hard

  • 注: 已被 Terminal-Bench 2.1 取代,我们将继续使用它。 Terminal-Bench Hard 是 v4.1 之前的Artificial Analysis Intelligence Index的组成部分
  • 说明:由Stanford University研究人员、Laude Institute和开源社区开发的智能体基准,于 2025 年发布。Terminal-Bench 评估智能体和模型解决各种任务(包括软件工程、系统管理和游戏)的能力场景)使用终端界面。
  • 页面: https://www.tbench.ai/
  • 数据集注册表: https://www.tbench.ai/registry
  • 实现:
    • 我们实现了terminal-bench-core数据集的“hard”子集,使用截至2025年8月14日的最新数据集版本(提交74221fb);我们评估该子集中的 44 个任务(由于原始数据集中的外部依赖性问题,少数任务被排除在外)
    • 我们使用 Terminus 2 智能体工具来评估这个“hard”子集,以确保模型之间的一致性,并根据 pass@1 评分对模型进行评分,每个任务的总体平均重复次数超过 3 次
    • 在 Terminal-Bench 框架中,每个任务都应用了一组特定的测试,如果所有测试都通过,则视为成功,否则视为不成功
    • 我们对智能体的评估应用以下约束:
      • 最大“情节”(模型在终端检查当前状态并计划一系列下一步行动)限制为 100
      • 我们将每个任务的全局超时设置为两小时(7,200 秒);实际上,100 集的限制是约束条件
      • 模型每次重复每个任务最多限制 100 万个累积输入标记
    • 在我们的测试中,这些约束主要限制模型陷入不成功循环的情况,并且由于这些约束,我们没有看到一致的性能差异

𝜏²-Bench Telecom

  • 注: 已被 𝜏³-Banking 取代,我们将继续使用它。 𝜏²-Bench 在 v4.1 之前,电信是Artificial Analysis Intelligence Index的组成部分
  • 说明:由 Sierra 为“双重控制”场景中的对话式 AI 智能体开发的基准,使用语言模型模拟智能体和用户角色来测试规划、工具使用和指导/通信。
  • 论文: https://arxiv.org/abs/2506.07982
  • 博客: sierra.ai/resources/research/tau-squared-bench
  • 数据集: https://github.com/sierra-research/tau2-bench
  • 实现:
    • 𝜏²-Bench 中引入的“电信”域包含 114 个任务(从总共 2,285 个以编程方式生成的任务中进行子采样),具有不同的“意图”来描述任务是否与服务、移动数据或 MMS 问题相关。我们对电信领域进行全面评估,每个任务重复 3 次,并使用 pass@1 评分作为 3 次尝试的平均值来报告分数
    • 在此基准测试中,结果“世界状态”决定智能体是否成功 - 例如,智能体完成任务后用户的手机数据是否正常工作
    • 完整的 𝜏²-Bench 套件包括 3 种执行模式,在消融研究中具有不同的规划和沟通水平;我们实现了“默认”双重控制模式,完全模拟且独立的用户和助理智能体
    • 我们使用 Qwen3 235B A22B 2507 (Non-reasoning)作为用户智能体模拟器,以确保一致的检查点可用性和对推理设置的完全控制以及强大的基础智能
    • 我们对执行施加限制,将每个任务重复的步骤限制为最多 100 个

MATH-500

  • 注:从Artificial Analysis Intelligence Index和我们的主动报告中退出。
  • 说明: MATH 基准的 500 个问题子集,涵盖高中数学竞赛,涵盖一系列科目和难度级别。
  • 数据集: huggingface.co/datasets/HuggingFaceH4/MATH-500

AIME 2025(美国数学邀请赛)

  • 注:不再参与我们的活跃报道;不再属于Artificial Analysis Intelligence Index v4.3.2。
  • 说明:来自 2025 年American Invitational Mathematics Examination的高级数学解题数据集。
  • 数据集: 2025 AIME I & 2025 AIME II
  • 关键细节:
    • 严格的数字答案格式(整数 1–999)
    • 通过@1 得分,每个问题重复 10 次
    • 基于脚本的评分,以 SymPy 标准化 + 答案等价性判定器 LLM 作为备份

MMLU-Pro(多任务语言理解基准测试,Pro 版本)

LiveCodeBench

提示词模板、答案提取与评测

多项选择题(GPQA、MMLU-Pro)

我们通过以下指令提示提示多项选择评估。该提示由Artificial Analysis独立开发,并通过各种消融研究仔细验证。我们评估认为,与传统的完成式多项选择评估方法或我们测试的其他指令提示相比,此提示是一种更清晰、因此更公平的方法。

GPQA 使用四个选项 (A–D)。 MMLU-Pro 使用十个选项 (A–J);我们使用相同的结构和额外的选择。

Answer the following multiple choice question. The last line of your response should be in the following format: 'Answer: A/B/C/D' (e.g. 'Answer: A').

{Question}

A) {A}
B) {B}
C) {C}
D) {D}
Answer the following multiple choice question. The last line of your response should be in the following format: 'Answer: A/B/C/D/E/F/G/H/I/J' (e.g. 'Answer: A').

{Question}

A) {A}
B) {B}
C) {C}
D) {D}
E) {E}
F) {F}
G) {G}
H) {H}
I) {I}
J) {J}

多项选择题提取正则表达式

我们使用多阶段方法提取多项选择答案来处理各种答案格式。对于单字母回复,我们直接使用该字母。否则,我们首先尝试匹配寻找正式 "Answer: X" 格式的主要模式(考虑可选的 markdown 格式):

主要模式:

(?i)[\*\_]{0,2}Answer[\*\_]{0,2}\s*:[\s\*\_]{0,2}\s*([A-Z])(?![a-zA-Z0-9])

如果主要模式失败,我们会依次尝试以下后备模式来捕获各种答案格式:

  • LaTeX 方框表示法(例如 \boxed{A} 或 \boxed{The answer is A})
    \boxed\{[^}]*([A-Z])[^}]*\}
  • 自然语言(例如,"answer is B")
    answer is ([a-zA-Z])
  • 带括号(例如,"answer is (C")
    answer is \\(([a-zA-Z])
  • 选择格式(例如,"D) some answer text")
    ([A-Z])\)\s*[^A-Z]*
  • 明确的陈述(例如,"E is the correct answer")
    ([A-Z])\s+is\s+the\s+correct\s+answer
  • 回复末尾的独立信件
    ([A-Z])\s*$
  • 字母后跟句点(例如,"F.")
    ([A-Z])\s*\.
  • 字母后跟非单词字符
    ([A-Z])\s*[^\w]

我们总是将找到的最后一个匹配项用于响应中的自我更正。

等价性判定 LLM

对于开放式答案(HLE、AA-LCR)的评估,我们使用答案等价性判定器LLM来确定模型的响应在语义上是否与正确答案等效。这种方法使用语言模型来评估两个答案是否具有相同的含义,即使措辞不同。答案等价性判定器评估语义等价性,而不是要求精确的字符串匹配,这对于存在多个有效短语的问题尤其重要。

HLE 和 AA-LCR 共享一个答案等价性判定器,GPT-5.6 Luna (medium),根据人类判断的验证进行选择。以下各节详细介绍了每次评估的具体答案等价性判定器提示。

HLE (Humanity's Last Exam)

我们使用以下文本作为系统提示来提示 HLE,精确答案和多项选择问题之间存在差异 - 这些提示与 Hendrycks 等人的原始实现相比没有任何更改。对于不支持系统提示的型号,我们将系统提示添加到用户提示之前。

System Prompt for Exact Answer Questions:

"Your response should be in the following format:
Explanation: {your explanation for your final answer}
Exact Answer: {your succinct, final answer}
Confidence: {your confidence score between 0% and 100% for your answer}"

System Prompt for Multiple Choice Questions:

"Your response should be in the following format:
Explanation: {your explanation for your answer choice}
Answer: {your chosen answer}
Confidence: {your confidence score between 0% and 100% for your answer}"

HLE 等价性判定提示词

我们在 GPT-5.6 Luna (medium) 上使用以下提示来检查响应答案是否正确 - 该提示再次使用,与 Hendrycks 等人的原始实现没有任何更改。

Judge whether the following [response] to [question] is correct or not based on the precise and unambiguous [correct_answer] below.

[question]: {question}

[response]: {response}

Your judgement must be in the format and criteria specified below:

extracted_final_answer: The final exact answer extracted from the [response]. Put the extracted answer as 'None' if there is no exact, final answer to extract from the response.

[correct_answer]: {correct_answer}

reasoning: Explain why the extracted_final_answer is correct or incorrect based on [correct_answer], focusing only on if there are meaningful differences between [correct_answer] and the extracted_final_answer. Do not comment on any background to the problem, do not attempt to solve the problem, do not argue for any answer different than [correct_answer], focus only on whether the answers match.

correct: Answer 'yes' if extracted_final_answer matches the [correct_answer] given above, or is within a small margin of error for numerical problems. Answer 'no' otherwise, i.e. if there if there is any inconsistency, ambiguity, non-equivalency, or if the extracted answer is incorrect.


confidence: The extracted confidence score between 0|\%| and 100|\%| from [response]. Put 100 if there is no confidence score available.

LCR 等价性判定提示词

Assess whether the following CANDIDATE ANSWER is CORRECT or INCORRECT.
For the CANDIDATE ANSWER to be correct, it must be consistent with the OFFICIAL ANSWER.

The question, for reference only: {question}
The OFFICIAL ANSWER: {official_answer}
CANDIDATE ANSWER TO ASSESS: {candidate_answer}

Reply only with CORRECT or INCORRECT.

数学题(AIME 2025)

我们使用以下指令提示提示 AIME:

Solve the following math problem step by step. Put your answer inside \\boxed{{}}.

{Question}

Remember to put your answer inside \\boxed{{}}.

数学等价性判定提示词

如上所述,我们使用语言模型相等性检查器来补充基于脚本的评分。我们使用 Llama 3.3 70B 的以下提示来检查两个答案是否相等。该提示由 OpenAI 开发并在其 simple-evals 存储库中发布。

Look at the following two expressions (answers to a math problem) and judge whether they are equivalent. Only perform trivial simplifications

Examples:

  Expression 1: $2x+3$
  Expression 2: $3+2x$

Yes

  Expression 1: 3/2
  Expression 2: 1.5

Yes

  Expression 1: $x^2+2x+1$
  Expression 2: $y^2+2y+1$

No

  Expression 1: $x^2+2x+1$
  Expression 2: $(x+1)^2$

Yes

  Expression 1: 3245/5
  Expression 2: 649

No
(these are actually equal, don't mark them equivalent if you need to do nontrivial simplifications)

  Expression 1: 2/(-3)
  Expression 2: -2/3

Yes
(trivial simplifications are allowed)

  Expression 1: 72 degrees
  Expression 2: 72

Yes
(give benefit of the doubt to units)

  Expression 1: 64
  Expression 2: 64 square feet

Yes
(give benefit of the doubt to units)

---

YOUR TASK


Respond with only "Yes" or "No" (without quotes). Do not include a rationale.

  Expression 1: %(expression1)s
  Expression 2: %(expression2)s

代码生成任务

SciCode

我们使用以下提示来提示 SciCode,该提示与 Tian 等人的科学家注释背景提示的原始实现没有任何变化。

PROBLEM DESCRIPTION:
You will be provided with problem steps along with background knowledge necessary for solving the problem. Your task will be to develop a Python solution focused on the next step of the problem-solving process.

PROBLEM STEPS AND FUNCTION CODE:
Here, you'll find the Python code for the initial steps of the problem-solving process. This code is integral to building the solution.

{problem_steps_str}

NEXT STEP - PROBLEM STEP AND FUNCTION HEADER:
This part will describe the next step in the problem-solving process. A function header will be provided, and your task is to develop the Python code for this next step based on the provided description and function header.

{next_step_str}

DEPENDENCIES:
Use only the following dependencies in your solution. Do not include these dependencies at the beginning of your code.

{dependencies}

RESPONSE GUIDELINES:
Now, based on the instructions and information provided above, write the complete and executable Python program for the next step in a single block.
Your response should focus exclusively on implementing the solution for the next step, adhering closely to the specified function header and the context provided by the initial steps.
Your response should NOT include the dependencies and functions of all previous steps. If your next step function calls functions from previous steps, please make sure it uses the headers provided without modification.
DO NOT generate EXAMPLE USAGE OR TEST CODE in your response. Please make sure your response python code in format of ```python```.

LiveCodeBench

我们使用以下提示来提示 LiveCodeBench,与原始团队对 LiveCodeBench 提示的原始实现没有任何更改。但我们注意到,我们没有应用 LiveCodeBench 团队使用的自定义系统提示 - 我们没有使用他们的通用系统提示,也没有对某些模型使用他们的自定义系统提示。

Questions with starter code:

### Question:
{question.question_content}

### Format: You will use the following starter code to write the solution to the problem and enclose your code within delimiters.
```python
{question.starter_code}
```

### Answer: (use the provided format with backticks)


Questions without starter code:

### Question:
{question.question_content}

### Format: Read the inputs from stdin solve the problem and write the answer to stdout (do not directly test on the sample inputs). Enclose your code within delimiters as follows. Ensure that when the python program runs, it reads the inputs, runs the algorithm and writes output to STDOUT.
```python
# YOUR CODE HERE
```

### Answer: (use the provided format with backticks

代码提取正则表达式

我们使用以下正则表达式从响应中提取代码:

(?<=```python\n)((?:\n|.)+?)(?=\n```)

版本历史

版本4.3.2

2026 年 9 月至今

  • GDPval-AA v2.1:Elo 量表现在以 DeepSeek V4.1 Flash (max) 的 1600 分为锚点,并使用 Crowd-BT 模型拟合评分。
  • AA-Briefcase v1.1:锚点仍为 GPT-5.5 (medium) 的 1000 分,但每个两两比较的评分维度现在均使用 Crowd-BT 模型拟合。

版本4.3.1

2026 年 9 月

  • 将成对评审小组更新为当前模型版本:AA-Briefcase 和 GDPval-AA v2 成对比较现在使用 Claude Opus 5、GPT-5.6 Sol 和 Gemini 3.8 Flash。评分细则面板未更改。

版本4.3

2026 年 9 月

  • 在智能体类别中将 𝜏³-Banking 替换为 AutomationBench-AA (5%)
  • 将 Terminal-Bench 2.1 替换为 Terminal-Bench 4.0(66 个任务,mini-swe-agent 工具)
  • GDP.pdf 图像处理升级:改进了对超出 API 限制的图像的处理,现在可以在更广泛的情况下调整大小,以确保图像可以传递到模型
  • 权重:GDPval-AA v2 (10%)、AA-Briefcase (15%)、AutomationBench-AA (5%)、Terminal-Bench 4.0 (10%)、 SciCode (10%)、AA-LCR (5%)、AA-Omniscience 准确度 (10%) 和非幻觉 (5%)、HLE (10%)、 GDP.pdf (10%)、CritPt (10%)

版本4.2

2026 年 9 月

  • 将 AA-Briefcase (15%) 添加到智能体类别
  • 将 GDP.pdf (10%) 添加到常规类别
  • 从Intelligence Index中删除了GPQA Diamond
  • 升级AA-LCR至v1.1
  • 将 SciCode 评分超时从 60 秒提高到 300 秒,并隔离脚本执行,使运行较慢但正确的代码不再被判为失败;按 v1.0.1 重新评分。
  • 重新平衡权重:GDPval-AA v2 (10%)、𝜏³-Banking (5%)、AA-Briefcase (15%)、Terminal-Bench 2.1 (10%)、 SciCode (10%)、AA-LCR (5%)、AA-Omniscience 准确度 (10%) 和非幻觉 (5%)、HLE (10%)、 GDP.pdf (10%)、CritPt (10%)

版本4.1.1

2026年8月—2026年9月

  • 将 𝜏³-Banking 移至上游 tau2-bench v1.0.1 数据集和评分器
  • 将 HLE、AA-LCR 和 AA-Omniscience 的评审模型升级为 GPT-5.6 Luna (medium),分别替代 GPT-4o (Aug '24)、Qwen3 235B A22B 2507 Non-Reasoning 和 Gemini 3 Flash Preview (Reasoning)

版本4.1

2026年6月—2026年8月

  • 将 GDPval-AA 升级到 GDPval-AA v2:升级的沙箱具有新的和扩展的依赖项,Elo 分数重新基线为 1000 分的人类专家表现,由三位前沿 LLM 评委组成的小组,回合限制扩大到 250 回合有能力提前退出
  • 将 Terminal-Bench Hard 替换为 Terminal-Bench 2.1(更高的回合限制,无Token限制)
  • 将 𝜏²-Bench 电信替换为 𝜏³-Banking
  • 从Intelligence Index中删除了IFBench(我们继续在新型号版本上运行它)
  • 调整类别权重以进一步强调智能体任务:智能体(34%)、编程(24%)、科学推理(24%)、一般(18%),其中AA-Omniscience分为准确性(8%)和非幻觉(4%)部分
  • 升级的Token和成本指标以更好地反映实际成本,包括缓存命中率和缓存Token定价

版本4.0.4

2026年3月—2026年6月

  • 在弃用之前的评分机模型 Gemini 3 Pro Preview 后,将 GDPval-AA 的评分机模型更新为 Gemini 3.1 Pro Preview

版本4.0.3

2026年2月—2026年3月

  • 在弃用之前的评分器模型Gemini 2.5 Flash (09-2025) (Reasoning)后,将 Omniscience 的评分器模型更新为Gemini 3 Flash Preview (Reasoning)

版本4.0.2

2026年1月—2026年2月

  • 在修订后,将Intelligence Index中的GDPval-AA Elo分数重新锚定为最新值,以提高对罕见代码沙箱故障的稳健性

版本4.0.1

2026 年 1 月

  • 改进了 Terminal-Bench 对 44 个任务的硬评估,删除了由于固定提交时原始数据集中的外部依赖问题而导致的一小部分任务

版本4.0

2026 年 1 月

  • 添加了GDPval-AA(现实世界的知识工作)
  • 添加了AA-Omniscience(知识和幻觉)
  • 添加了CritPt(物理推理)
  • 从Intelligence Index中删除了MMLU-Pro、LiveCodeBench、AIME 2025
  • 新的基于类别的权重结构:智能体(25%)、编程(25%)、一般(25%)、科学推理(25%)

版本3.0

2025年9月2日—2025年12月

  • 添加了 Terminal-Bench Hard(智能体工作流程)
  • 添加了 𝜏²-Bench 电信(智能体工作流程)
  • 在Intelligence Index中包含MMLU-Pro和LiveCodeBench
  • 更新权重

版本2.2

2025年8月6日—2025年9月1日

  • 新增 Artificial Analysis Long Context Reasoning
  • 更新权重

版本2.1

2025年8月5日—2025年8月6日

  • 添加了IFBench
  • 添加了AIME 2025
  • 删除了MATH-500
  • 删除了AIME 2024
  • 更新权重

版本2.0

2025年2月11日—2025年8月4日

版本1.0—1.3

2024年1月—2025年2月10日