Coding Agent Index 方法论
概述
Artificial Analysis 对编程智能体进行端到端软件工程任务的基准测试。目标是衡量智能体完成真实编程工作的能力,以及其在结果、可靠性、token 用量、成本和执行时间等方面的表现差异。
Coding Agent Index 页面上的公开结果基于任务级别的基准测试尝试,并汇总为单项评测得分、汇总效率指标以及 Artificial Analysis Coding Agent Index。
本页面重点介绍公开的 Artificial Analysis Coding Agent Index 的构建方式、当前包含的基准测试组件,以及公开的 pass@1、成本、token 用量和执行时间指标的计算方法。
Artificial Analysis Coding Agent Index
当前公开的 Artificial Analysis Coding Agent Index 是一项综合基准测试分数,由公开编程智能体套件中已配置的基准测试组件构建而成。
该指数的目的并非将所有编程工作归结为一种基准测试任务类型。不同的编程智能体在仓库问答、实现与修复缺陷任务以及重度终端操作工作流上的表现可能差异很大。该指数的作用是将这些不同的基准测试族汇总为一个顶层性能视图,同时在其下保留各项基准测试的详细拆分。
指数组件
当前公开指数包含以下基准测试组件:
| 评测 | 领域 | 任务 | 每任务尝试次数 | 响应类型 | 评分 |
|---|---|---|---|---|---|
| DeepSWE | 长周期软件工程 | 113 | 3 | 代码补丁 / 仓库变更 | 程序验证器 pass/fail,pass@1 |
| Terminal-Bench v2 | 智能体终端操作 | 84* | 3 | 基于终端的任务执行 | 测试套件 pass/fail,pass@1 |
| SWE-Atlas-QnA | 仓库问答 | 124 | 3 | 开放式回答 | Scale AI Task Resolve Rate(二元 pass/fail),pass@1 |
* Terminal-Bench v2 原始包含 89 个任务;我们因环境兼容性问题排除了其中五个任务。
评测任务
当前公开指数涵盖 3 个基准测试组件中的 321 个评测任务。
- Add iterable collection combinators to true-myth
- Abort pending body reads on shutdown
- Add rolling min, max, median, and quantile methods
- Format BigQuery pipe syntax queries correctly
- Add unified manifest stream output across Helm commands
- Add a deterministic CookieStore with modern Set-Cookie parsing
- Preserve restored query state in persisted snapshots
- Add an error-accumulating Validated container
- Add a persistent analysis cache to Vulture
- Add multi-module memory snapshots to wazero
- Add JSONPath query APIs to orderedmap and Starlark modules
- Add a per-origin circuit breaker to ofetch
- Add stepped slices for arrays and strings
- Add `matchEach` to ts-pattern
- Harden module loading, cache introspection, and script flags
- Add action pinning linting for actions and reusable workflows
- Add typed window function builders with OVER clauses
- Add typed blend range access and blend-if compositing
- Add deterministic map conflict detection to Y.Map writes
- Add dependency-aware async initialization to the container
- Add multipart response parsing to HTTPX
- Add scoped per-rule ignore markers to Obsidian Linter
- Partition report files by launcher and expand report templates
- Add keyset cursor pagination to `$find`
- Add ShapeIndex encoding and decoding
- Add entity snapshot and rollback APIs to Koota
- Restore RichLog follow-state parity and expand reflow behavior
- Add duration-aware sharding to Vitest
- Add grouped test phases with synchronized barriers
- Add bidirectional TOML table converters
- Add go:embed directive support for interpreted packages
- Add task snapshots, inspection, and diffing to aiomonitor
- Add drift detection and compliance baselines
- Expose accumulated streamed function-call args in SDK surfaces
- Add grouping-set and window-frame SQL helpers
- Add rule evaluation profiling to Rego
- Add typed variable bindings to Anko
- Add multiplexed ordered streams over KCP
- Add single-active-consumer priority and cancel tracking to virtual transports
- Reconstruct template strings in partial evaluation output
- Add bounded-memory spilling to SCC aggregation
- Fix isolated Go-side calls for Tengo callables and closures
- Add interprocedural taint checks for Bandit injection sinks
- Add deterministic multi-key sorting to fd
- Add task graph export with JSON, DOT, and text output
- Add default arguments to Anko function parameters
- Add deprecation, sunset, and successor headers to FastAPI routes
- Add a checker for broken doc comment links
- Add boundary modes to `@stencil`
- Add retry-aware publishing audit logs
- Add method declarations and interface dispatch to Scriggo
- Add partial structuring with error recovery to cattrs
- Add conditional required attributes to schemas
- Add worktree merge conflict handling
- Add session bundle recording and replay to IPython
- Add flattened dataclass fields to Mashumaro field options
- Add build-time grammar conflict analysis to participle
- Add XML diff, patch, and merge operations to etree
- Add streaming JSON iteration to HTTPX responses
- Validate daemon watch, status, and log lifecycle
- Add recursive schema composition to Valibot
- Add input key aliases to name mapping
- Add durability callbacks and wait APIs for sync writes
- Add duration encoding to TableVectorizer
- Add safe import checkpoints and invariant validation
- Add shorthand expansion and compression to the lexer
- Add RFC 5545 timezone interoperability to dateutil recurrence parsing
- Add atomic signal selectors to Kea
- Add consistent hash policy support to TrafficPolicy
- Fix PromQL label sorting across typed and untyped values
- Format CREATE TABLE DDL and add DDL parsing helpers
- Add trap coredump generation to wasmi
- Add hierarchical evaluation cancellation to Boa
- Add value-based query predicates to Koota
- Add request coalescing to `Runnable`
- Add explicit resource management declarations to the parser
- Add lazy recursive schemas with DTO and JSON Schema export
- Persist the fitted feature schema across evaluate, predict, serve, and export
- Add composite trait aspects to Koota
- Add tube multiplexing to pwntools
- Add scoped state data to state machine callbacks and history
- Add destructuring bindings to Tengo
- Add JSON Schema refs and dependency keywords
- Add link format conversion between wiki and markdown syntax
- Add implicit HEAD and automatic OPTIONS responses to FastAPI routes
- Add transparent encryption to dump uploads
- Add conditional option dependencies to Optique
- Preserve structure needed by stylesheet selectors
- Add policy-based alerting for failures, latency, and SSL expiry
- Add incremental cache controls to Bandit
- Coalesce qualifying choices into character classes
- Preserve ANSI resets during truncation and styling
- Add async autocomplete options and fetch lifecycle handling
- Add config file parsing to Cliffy commands
- Add HTML document format handling to Dasel
- Add CSS Grid layout to the Box component
- Add `\multicolumn` column spans to array-like environments
- Add a deferred mutation buffer to batch entity changes
- Add automatic table of contents generation for Obsidian linter
- Implement a deterministic IntersectionObserver in Happy DOM
- Add pair-level relation tracking modifiers
- Add error stack serialization to SuperJSON
- Complete Kitty keyboard phases and stable fallback key metadata
- Add structured nosec directives for regions and next line
- Add bail-on-test-failure handling to Testem
- Add try/catch error recovery to expr
- Add configurable array merge strategies to Helm value coalescing
- Implement recursive agent delegation through delegate_task tool calls
- Add SSE streaming endpoints to HttpApi
- Add transactional reload status and rollback tracking to Prometheus
- Add GraphQL incremental delivery with @defer and @stream
- Add dead-lettering, TTL, and overflow handling to virtual queues
- Reuse one toolbar across multiple Quill editors
- adaptive-rejection-sampler
- bn-fit-modify
- break-filter-js-from-html
- build-cython-ext
- build-pmars
- build-pov-ray
- caffe-cifar-10
- cancel-async-tasks
- chess-best-move
- circuit-fibsqrt
- cobol-modernization
- code-from-image
- compile-compcert
- configure-git-webserver
- constraints-scheduling
- count-dataset-tokens
- crack-7z-hash
- custom-memory-heap-crash
- db-wal-recovery
- distribution-search
- dna-assembly
- dna-insert
- extract-elf
- extract-moves-from-video
- feal-differential-cryptanalysis
- financial-document-processor
- fix-code-vulnerability
- fix-git
- fix-ocaml-gc
- gcode-to-text
- git-leak-recovery
- git-multibranch
- headless-terminal
- hf-model-inference
- install-windows-3.11
- kv-store-grpc
- large-scale-text-editing
- largest-eigenval
- llm-inference-batching-scheduler
- log-summary-date-ranges
- mailman
- make-mips-interpreter
- mcmc-sampling-stan
- merge-diff-arc-agi-task
- model-extraction-relu-logits
- modernize-scientific-stack
- mteb-leaderboard
- mteb-retrieve
- multi-source-data-merger
- nginx-request-logging
- openssl-selfsigned-cert
- overfull-hbox
- password-recovery
- path-tracing
- path-tracing-reverse
- polyglot-c-py
- polyglot-rust-c
- portfolio-optimization
- protein-assembly
- prove-plus-comm
- pypi-server
- pytorch-model-cli
- pytorch-model-recovery
- qemu-alpine-ssh
- qemu-startup
- query-optimize
- raman-fitting
- regex-chess
- regex-log
- reshard-c4-data
- rstan-to-pystan
- sam-cell-seg
- sanitize-git-repo
- schemelike-metacircular-eval
- sparql-university
- sqlite-db-truncate
- sqlite-with-gcov
- torch-pipeline-parallelism
- torch-tensor-parallelism
- train-fasttext
- tune-mjcf
- video-processing
- vulnerable-secret
- winning-avg-corewars
- 6905333b74f22949d97ba998
- 6905333b74f22949d97ba999
- 6905333b74f22949d97ba99a
- 6905333b74f22949d97ba99b
- 6905333b74f22949d97ba99d
- 6905333b74f22949d97ba99f
- 6905333b74f22949d97ba9a2
- 6905333b74f22949d97ba9a3
- 6905333b74f22949d97ba9a4
- 6905333b74f22949d97ba9a5
- 6905333b74f22949d97ba9a6
- 6905333b74f22949d97ba9a7
- 6905333b74f22949d97ba9a8
- 6905333b74f22949d97ba9a9
- 6905333b74f22949d97ba9aa
- 6905333b74f22949d97ba9ab
- 6905333b74f22949d97ba9ac
- 6905333b74f22949d97ba9ad
- 6905333b74f22949d97ba9ae
- 6905333b74f22949d97ba9af
- 6905333b74f22949d97ba9b1
- 6905333b74f22949d97ba9b2
- 6905333b74f22949d97ba9b3
- 6905333b74f22949d97ba9b5
- 6905333b74f22949d97ba9b6
- 6905333b74f22949d97ba9b7
- 6905333b74f22949d97ba9b8
- 6905333b74f22949d97ba9ba
- 6905333b74f22949d97ba9bb
- 6905333b74f22949d97ba9bc
- 6905333b74f22949d97ba9bd
- 6905333b74f22949d97ba9be
- 6905333b74f22949d97ba9bf
- 6905333b74f22949d97ba9c0
- 6905333b74f22949d97ba9c1
- 6905333b74f22949d97ba9c2
- 6905333b74f22949d97ba9c3
- 6905333b74f22949d97ba9c4
- 6905333b74f22949d97ba9c5
- 6905333b74f22949d97ba9c6
- 6905333b74f22949d97ba9c8
- 6905333b74f22949d97ba9c9
- 6905333b74f22949d97ba9ca
- 6905333b74f22949d97ba9cb
- 6905333b74f22949d97ba9cc
- 6905333b74f22949d97ba9cd
- 6905333b74f22949d97ba9ce
- 6905333b74f22949d97ba9cf
- 6905333b74f22949d97ba9d0
- 6905333b74f22949d97ba9d1
- 6905333b74f22949d97ba9d2
- 6905333b74f22949d97ba9d3
- 6905333b74f22949d97ba9d4
- 6905333b74f22949d97ba9d5
- 6905333b74f22949d97ba9d6
- 6905333b74f22949d97ba9d7
- 6905333b74f22949d97ba9d8
- 6905333b74f22949d97ba9d9
- 6905333b74f22949d97ba9db
- 6905333b74f22949d97ba9dc
- 6905333b74f22949d97ba9dd
- 6905333b74f22949d97ba9de
- 6905333b74f22949d97ba9e0
- 6905333b74f22949d97ba9e1
- 6905333b74f22949d97ba9e3
- 6905333b74f22949d97ba9e4
- 6905333b74f22949d97ba9e5
- 6905333b74f22949d97ba9e7
- 6905333b74f22949d97ba9e8
- 6905333b74f22949d97ba9e9
- 6905333b74f22949d97ba9eb
- 6905333b74f22949d97ba9ee
- 6905333b74f22949d97ba9f0
- 6905333b74f22949d97ba9f1
- 6905333b74f22949d97ba9f2
- 6905333b74f22949d97ba9f4
- 6905333b74f22949d97ba9f5
- 6905333b74f22949d97ba9f7
- 6905333b74f22949d97ba9f8
- 6905333b74f22949d97ba9f9
- 6905333b74f22949d97ba9fa
- 6905333b74f22949d97ba9fb
- 6905333b74f22949d97ba9fc
- 6905333b74f22949d97ba9fd
- 6905333b74f22949d97ba9ff
- 6905333b74f22949d97baa01
- 6905333b74f22949d97baa02
- 6905333b74f22949d97baa03
- 6905333b74f22949d97baa04
- 6905333b74f22949d97baa05
- 6905333b74f22949d97baa06
- 6905333b74f22949d97baa07
- 6905333b74f22949d97baa09
- 6905333b74f22949d97baa0b
- 6905333b74f22949d97baa0c
- 6905333b74f22949d97baa0d
- 6905333b74f22949d97baa0f
- 6905333b74f22949d97baa10
- 6905333b74f22949d97baa11
- 6905333b74f22949d97baa12
- 6905333b74f22949d97baa14
- 6905333b74f22949d97baa15
- 6905333b74f22949d97baa16
- 6905333b74f22949d97baa17
- 6905333b74f22949d97baa19
- 6905333b74f22949d97baa1a
- 6905333b74f22949d97baa1b
- 6905333b74f22949d97baa1c
- 6905333b74f22949d97baa1d
- 6905333b74f22949d97baa1e
- 6905333b74f22949d97baa1f
- 6905333b74f22949d97baa20
- 6905333b74f22949d97baa21
- 6905333b74f22949d97baa22
- 6905333b74f22949d97baa23
- 6905333b74f22949d97baa24
- 6905333b74f22949d97baa25
- 6905333b74f22949d97baa26
- 6905333b74f22949d97baa27
- 6905333b74f22949d97baa28
- 6905333b74f22949d97baa2a
- 6905333b74f22949d97baa2b
- 6905333b74f22949d97baa2c
- 6905333b74f22949d97baa2d
指数汇总的内容
对于每个智能体变体,Artificial Analysis 为每个包含的基准测试组件计算 pass@1 得分,然后将这些组件得分汇总为公开指数。
同一基准测试套件还支撑着基准测试页面上的公开汇总效率指标,包括运行成本、token 用量和执行时间。这意味着性能视图和效率视图对齐的是同一组底层基准测试覆盖范围,而非来自无关的运行。
评分与结果
pass@1 结果
每个评测尝试都从基准测试评估器获得一个 pass@1 结果。SWE-Atlas-QnA 采用与 Scale AI 的 Task Resolve Rate 方法论对齐的二元 pass/fail 评分:通过的尝试得 1 分,失败或出错的尝试得 0 分。
| 术语 | 定义 |
|---|---|
| 二元 pass@1 | 一种测试套件评估结果,任务通过得 1 分,失败得 0 分。 |
单项评测得分
对于每项评测,Artificial Analysis 首先对每个任务的三次评测尝试取平均值,然后对这些任务级 pass@1 得分再取平均值,以确保每个任务具有相同的权重。
效率指标
成本、token 用量和执行时间以当前公开编程智能体基准测试套件中每任务尝试的汇总平均值报告。
- 运行成本:基于服务商 token 定价而非消费者套餐的每任务平均按 token 付费 API 成本。
- Token 用量:每任务平均输入 token、缓存 token、缓存写入 token、推理 token 和输出 token。
- 执行时间:每任务平均实际运行时间,包括完整任务耗时以及可用时的智能体耗时子集。
如果某项指标缺少遥测数据,这些缺失值将从相应平均值中排除,而非视为零。
在成本指标中,当服务商定价支持区分时,缓存输入与非缓存输入分开计算;当服务商对创建提示词缓存状态收费时,缓存写入费用也会被纳入。此举旨在比按固定单价估算更准确地反映按 token 付费的 API 定价。
智能体设置
公开基准测试行代表的是智能体变体,而不仅仅是模型名称。可能影响行为的设置在报告中分别列出。
除非另有说明,我们使用每个智能体的默认推理设置,以使基准测试反映默认的用户体验。
随着新评测和智能体变体的加入,基准测试方法论可能会随时间演变,但公开比较旨在反映已发布基准测试套件中同类智能体变体之间的对等对比。
版本历史
版本 1.3
2026 年 7 月至今
- 优化了 SWE-Atlas-QnA 的二元 pass/fail 评分,使其与 Scale AI 的 Task Resolve Rate 方法论对齐
版本 1.2
2026 年 7 月
- 将 SWE-Atlas-QnA 的评分方式从评分标准奖励改为二元 pass/fail,要求所有评分标准项均通过才能将任务标记为正确
版本 1.1
2026 年 6 月至 2026 年 7 月
- 新增 DeepSWE(长周期软件工程)
- 从 Coding Agent Index 中移除 SWE-Bench-Pro-Hard-AA
版本 1.0
2026 年 5 月至 2026 年 6 月
- 初始版本,包含 SWE-Bench-Pro-Hard-AA(代码生成)、Terminal-Bench v2(智能体终端操作)和 SWE-Atlas-QnA(仓库问答)