Coding Agent Indexの方法論
概要
Artificial Analysisでは、エンドツーエンドのソフトウェアエンジニアリングタスクでコーディングエージェントをベンチマークしています。現実的なコーディング作業をエージェントがどの程度完遂できるか、また成果、信頼性、トークン使用量、コスト、実行時間によって性能がどう異なるかを測定することが目的です。
Coding Agent Indexページの公開結果は、タスク単位のベンチマーク試行から構築され、評価ごとのスコア、統合した効率性指標、Artificial Analysis Coding Agent Indexに集約されます。
このページでは、公開されているArtificial Analysis Coding Agent Indexの構築方法、現在含まれているベンチマーク構成要素、公開されるpass@1、コスト、トークン使用量、実行時間の各指標の算出方法を説明します。
Artificial Analysis Coding Agent Index
現在公開しているArtificial Analysis Coding Agent Indexは、公開コーディングエージェントスイートに設定されたベンチマーク構成要素から算出する複合ベンチマークスコアです。
この指数の目的は、すべてのコーディング作業を1種類のベンチマークタスクに集約することではありません。コーディングエージェントの性能は、リポジトリに関するQ&A、実装やバグ修正のタスク、ターミナル操作が中心のワークフローで大きく異なる場合があります。指数は、ベンチマークごとの内訳を維持しながら、こうした異なるベンチマーク群を1つの上位の性能ビューにまとめるためのものです。
指数の構成要素
現在の公開指数には、次のベンチマーク構成要素が含まれています。
| 評価 | 分野 | タスク | タスクあたりの試行回数 | 回答形式 | 採点 |
|---|---|---|---|---|---|
| DeepSWE | 長期的なソフトウェアエンジニアリング | 113 | 3 | コードパッチ / リポジトリの変更 | プログラム検証のpass/fail、pass@1 |
| Terminal-Bench v2 | エージェントによるターミナル操作 | 84* | 3 | ターミナルベースのタスク実行 | テストスイートのpass/fail、pass@1 |
| SWE-Atlas-QnA | リポジトリに関するQ&A | 124 | 3 | 自由回答 | Scale AI Task Resolve Rate(バイナリpass/fail)、pass@1 |
* Terminal-Bench v2には本来89件のタスクがありますが、環境互換性の問題により5件を除外しています。
評価対象タスク
現在の公開指数は、3件のベンチマーク構成要素にわたる321件の評価対象タスクを網羅しています。
- Add iterable collection combinators to true-myth
- Abort pending body reads on shutdown
- Add rolling min, max, median, and quantile methods
- Format BigQuery pipe syntax queries correctly
- Add unified manifest stream output across Helm commands
- Add a deterministic CookieStore with modern Set-Cookie parsing
- Preserve restored query state in persisted snapshots
- Add an error-accumulating Validated container
- Add a persistent analysis cache to Vulture
- Add multi-module memory snapshots to wazero
- Add JSONPath query APIs to orderedmap and Starlark modules
- Add a per-origin circuit breaker to ofetch
- Add stepped slices for arrays and strings
- Add `matchEach` to ts-pattern
- Harden module loading, cache introspection, and script flags
- Add action pinning linting for actions and reusable workflows
- Add typed window function builders with OVER clauses
- Add typed blend range access and blend-if compositing
- Add deterministic map conflict detection to Y.Map writes
- Add dependency-aware async initialization to the container
- Add multipart response parsing to HTTPX
- Add scoped per-rule ignore markers to Obsidian Linter
- Partition report files by launcher and expand report templates
- Add keyset cursor pagination to `$find`
- Add ShapeIndex encoding and decoding
- Add entity snapshot and rollback APIs to Koota
- Restore RichLog follow-state parity and expand reflow behavior
- Add duration-aware sharding to Vitest
- Add grouped test phases with synchronized barriers
- Add bidirectional TOML table converters
- Add go:embed directive support for interpreted packages
- Add task snapshots, inspection, and diffing to aiomonitor
- Add drift detection and compliance baselines
- Expose accumulated streamed function-call args in SDK surfaces
- Add grouping-set and window-frame SQL helpers
- Add rule evaluation profiling to Rego
- Add typed variable bindings to Anko
- Add multiplexed ordered streams over KCP
- Add single-active-consumer priority and cancel tracking to virtual transports
- Reconstruct template strings in partial evaluation output
- Add bounded-memory spilling to SCC aggregation
- Fix isolated Go-side calls for Tengo callables and closures
- Add interprocedural taint checks for Bandit injection sinks
- Add deterministic multi-key sorting to fd
- Add task graph export with JSON, DOT, and text output
- Add default arguments to Anko function parameters
- Add deprecation, sunset, and successor headers to FastAPI routes
- Add a checker for broken doc comment links
- Add boundary modes to `@stencil`
- Add retry-aware publishing audit logs
- Add method declarations and interface dispatch to Scriggo
- Add partial structuring with error recovery to cattrs
- Add conditional required attributes to schemas
- Add worktree merge conflict handling
- Add session bundle recording and replay to IPython
- Add flattened dataclass fields to Mashumaro field options
- Add build-time grammar conflict analysis to participle
- Add XML diff, patch, and merge operations to etree
- Add streaming JSON iteration to HTTPX responses
- Validate daemon watch, status, and log lifecycle
- Add recursive schema composition to Valibot
- Add input key aliases to name mapping
- Add durability callbacks and wait APIs for sync writes
- Add duration encoding to TableVectorizer
- Add safe import checkpoints and invariant validation
- Add shorthand expansion and compression to the lexer
- Add RFC 5545 timezone interoperability to dateutil recurrence parsing
- Add atomic signal selectors to Kea
- Add consistent hash policy support to TrafficPolicy
- Fix PromQL label sorting across typed and untyped values
- Format CREATE TABLE DDL and add DDL parsing helpers
- Add trap coredump generation to wasmi
- Add hierarchical evaluation cancellation to Boa
- Add value-based query predicates to Koota
- Add request coalescing to `Runnable`
- Add explicit resource management declarations to the parser
- Add lazy recursive schemas with DTO and JSON Schema export
- Persist the fitted feature schema across evaluate, predict, serve, and export
- Add composite trait aspects to Koota
- Add tube multiplexing to pwntools
- Add scoped state data to state machine callbacks and history
- Add destructuring bindings to Tengo
- Add JSON Schema refs and dependency keywords
- Add link format conversion between wiki and markdown syntax
- Add implicit HEAD and automatic OPTIONS responses to FastAPI routes
- Add transparent encryption to dump uploads
- Add conditional option dependencies to Optique
- Preserve structure needed by stylesheet selectors
- Add policy-based alerting for failures, latency, and SSL expiry
- Add incremental cache controls to Bandit
- Coalesce qualifying choices into character classes
- Preserve ANSI resets during truncation and styling
- Add async autocomplete options and fetch lifecycle handling
- Add config file parsing to Cliffy commands
- Add HTML document format handling to Dasel
- Add CSS Grid layout to the Box component
- Add `\multicolumn` column spans to array-like environments
- Add a deferred mutation buffer to batch entity changes
- Add automatic table of contents generation for Obsidian linter
- Implement a deterministic IntersectionObserver in Happy DOM
- Add pair-level relation tracking modifiers
- Add error stack serialization to SuperJSON
- Complete Kitty keyboard phases and stable fallback key metadata
- Add structured nosec directives for regions and next line
- Add bail-on-test-failure handling to Testem
- Add try/catch error recovery to expr
- Add configurable array merge strategies to Helm value coalescing
- Implement recursive agent delegation through delegate_task tool calls
- Add SSE streaming endpoints to HttpApi
- Add transactional reload status and rollback tracking to Prometheus
- Add GraphQL incremental delivery with @defer and @stream
- Add dead-lettering, TTL, and overflow handling to virtual queues
- Reuse one toolbar across multiple Quill editors
- adaptive-rejection-sampler
- bn-fit-modify
- break-filter-js-from-html
- build-cython-ext
- build-pmars
- build-pov-ray
- caffe-cifar-10
- cancel-async-tasks
- chess-best-move
- circuit-fibsqrt
- cobol-modernization
- code-from-image
- compile-compcert
- configure-git-webserver
- constraints-scheduling
- count-dataset-tokens
- crack-7z-hash
- custom-memory-heap-crash
- db-wal-recovery
- distribution-search
- dna-assembly
- dna-insert
- extract-elf
- extract-moves-from-video
- feal-differential-cryptanalysis
- financial-document-processor
- fix-code-vulnerability
- fix-git
- fix-ocaml-gc
- gcode-to-text
- git-leak-recovery
- git-multibranch
- headless-terminal
- hf-model-inference
- install-windows-3.11
- kv-store-grpc
- large-scale-text-editing
- largest-eigenval
- llm-inference-batching-scheduler
- log-summary-date-ranges
- mailman
- make-mips-interpreter
- mcmc-sampling-stan
- merge-diff-arc-agi-task
- model-extraction-relu-logits
- modernize-scientific-stack
- mteb-leaderboard
- mteb-retrieve
- multi-source-data-merger
- nginx-request-logging
- openssl-selfsigned-cert
- overfull-hbox
- password-recovery
- path-tracing
- path-tracing-reverse
- polyglot-c-py
- polyglot-rust-c
- portfolio-optimization
- protein-assembly
- prove-plus-comm
- pypi-server
- pytorch-model-cli
- pytorch-model-recovery
- qemu-alpine-ssh
- qemu-startup
- query-optimize
- raman-fitting
- regex-chess
- regex-log
- reshard-c4-data
- rstan-to-pystan
- sam-cell-seg
- sanitize-git-repo
- schemelike-metacircular-eval
- sparql-university
- sqlite-db-truncate
- sqlite-with-gcov
- torch-pipeline-parallelism
- torch-tensor-parallelism
- train-fasttext
- tune-mjcf
- video-processing
- vulnerable-secret
- winning-avg-corewars
- 6905333b74f22949d97ba998
- 6905333b74f22949d97ba999
- 6905333b74f22949d97ba99a
- 6905333b74f22949d97ba99b
- 6905333b74f22949d97ba99d
- 6905333b74f22949d97ba99f
- 6905333b74f22949d97ba9a2
- 6905333b74f22949d97ba9a3
- 6905333b74f22949d97ba9a4
- 6905333b74f22949d97ba9a5
- 6905333b74f22949d97ba9a6
- 6905333b74f22949d97ba9a7
- 6905333b74f22949d97ba9a8
- 6905333b74f22949d97ba9a9
- 6905333b74f22949d97ba9aa
- 6905333b74f22949d97ba9ab
- 6905333b74f22949d97ba9ac
- 6905333b74f22949d97ba9ad
- 6905333b74f22949d97ba9ae
- 6905333b74f22949d97ba9af
- 6905333b74f22949d97ba9b1
- 6905333b74f22949d97ba9b2
- 6905333b74f22949d97ba9b3
- 6905333b74f22949d97ba9b5
- 6905333b74f22949d97ba9b6
- 6905333b74f22949d97ba9b7
- 6905333b74f22949d97ba9b8
- 6905333b74f22949d97ba9ba
- 6905333b74f22949d97ba9bb
- 6905333b74f22949d97ba9bc
- 6905333b74f22949d97ba9bd
- 6905333b74f22949d97ba9be
- 6905333b74f22949d97ba9bf
- 6905333b74f22949d97ba9c0
- 6905333b74f22949d97ba9c1
- 6905333b74f22949d97ba9c2
- 6905333b74f22949d97ba9c3
- 6905333b74f22949d97ba9c4
- 6905333b74f22949d97ba9c5
- 6905333b74f22949d97ba9c6
- 6905333b74f22949d97ba9c8
- 6905333b74f22949d97ba9c9
- 6905333b74f22949d97ba9ca
- 6905333b74f22949d97ba9cb
- 6905333b74f22949d97ba9cc
- 6905333b74f22949d97ba9cd
- 6905333b74f22949d97ba9ce
- 6905333b74f22949d97ba9cf
- 6905333b74f22949d97ba9d0
- 6905333b74f22949d97ba9d1
- 6905333b74f22949d97ba9d2
- 6905333b74f22949d97ba9d3
- 6905333b74f22949d97ba9d4
- 6905333b74f22949d97ba9d5
- 6905333b74f22949d97ba9d6
- 6905333b74f22949d97ba9d7
- 6905333b74f22949d97ba9d8
- 6905333b74f22949d97ba9d9
- 6905333b74f22949d97ba9db
- 6905333b74f22949d97ba9dc
- 6905333b74f22949d97ba9dd
- 6905333b74f22949d97ba9de
- 6905333b74f22949d97ba9e0
- 6905333b74f22949d97ba9e1
- 6905333b74f22949d97ba9e3
- 6905333b74f22949d97ba9e4
- 6905333b74f22949d97ba9e5
- 6905333b74f22949d97ba9e7
- 6905333b74f22949d97ba9e8
- 6905333b74f22949d97ba9e9
- 6905333b74f22949d97ba9eb
- 6905333b74f22949d97ba9ee
- 6905333b74f22949d97ba9f0
- 6905333b74f22949d97ba9f1
- 6905333b74f22949d97ba9f2
- 6905333b74f22949d97ba9f4
- 6905333b74f22949d97ba9f5
- 6905333b74f22949d97ba9f7
- 6905333b74f22949d97ba9f8
- 6905333b74f22949d97ba9f9
- 6905333b74f22949d97ba9fa
- 6905333b74f22949d97ba9fb
- 6905333b74f22949d97ba9fc
- 6905333b74f22949d97ba9fd
- 6905333b74f22949d97ba9ff
- 6905333b74f22949d97baa01
- 6905333b74f22949d97baa02
- 6905333b74f22949d97baa03
- 6905333b74f22949d97baa04
- 6905333b74f22949d97baa05
- 6905333b74f22949d97baa06
- 6905333b74f22949d97baa07
- 6905333b74f22949d97baa09
- 6905333b74f22949d97baa0b
- 6905333b74f22949d97baa0c
- 6905333b74f22949d97baa0d
- 6905333b74f22949d97baa0f
- 6905333b74f22949d97baa10
- 6905333b74f22949d97baa11
- 6905333b74f22949d97baa12
- 6905333b74f22949d97baa14
- 6905333b74f22949d97baa15
- 6905333b74f22949d97baa16
- 6905333b74f22949d97baa17
- 6905333b74f22949d97baa19
- 6905333b74f22949d97baa1a
- 6905333b74f22949d97baa1b
- 6905333b74f22949d97baa1c
- 6905333b74f22949d97baa1d
- 6905333b74f22949d97baa1e
- 6905333b74f22949d97baa1f
- 6905333b74f22949d97baa20
- 6905333b74f22949d97baa21
- 6905333b74f22949d97baa22
- 6905333b74f22949d97baa23
- 6905333b74f22949d97baa24
- 6905333b74f22949d97baa25
- 6905333b74f22949d97baa26
- 6905333b74f22949d97baa27
- 6905333b74f22949d97baa28
- 6905333b74f22949d97baa2a
- 6905333b74f22949d97baa2b
- 6905333b74f22949d97baa2c
- 6905333b74f22949d97baa2d
Indexの集約対象
Artificial Analysisでは、エージェントの各バリアントについて、含まれるベンチマーク構成要素ごとにpass@1スコアを計算し、それらのスコアを公開指数に集約します。
同じベンチマークスイートは、ベンチマークページに公開する実行コスト、トークン使用量、実行時間などの統合効率性指標の基礎にもなっています。つまり、性能と効率性のビューは、無関係な実行結果ではなく、同じ基礎ベンチマークの対象範囲に基づいています。
採点と結果
pass@1の結果
評価対象の各試行には、ベンチマーク評価器によるpass@1の結果が与えられます。SWE-Atlas-QnAでは、Scale AIのTask Resolve Rate評価手法に沿ったバイナリpass/fail採点を使用します。合格した試行は1、失敗またはエラーとなった試行は0です。
| 用語 | 定義 |
|---|---|
| バイナリpass@1 | タスクに合格なら1、不合格なら0を与えるテストスイート評価の結果。 |
評価ごとのスコア
各評価では、まずタスクごとに評価した3回の試行を平均し、次にタスクの重みがすべて等しくなるよう、タスク単位のpass@1スコアを平均します。
効率性指標
コスト、トークン使用量、実行時間は、現在公開しているコーディングエージェントベンチマークスイート全体で統合した、タスク試行あたりの平均値として示します。
- 実行コスト:コンシューマー向けプランではなく、プロバイダーのトークン料金に基づく、タスクあたりの従量制APIコストの平均。
- トークン使用量:タスクあたりの入力、キャッシュ、キャッシュ書き込み、推論、出力トークンの平均。
- 実行時間:タスクあたりの経過時間の平均。タスク全体の経過時間と、取得可能な場合はそのうちエージェントが動作していた時間を含みます。
特定の指標でテレメトリが欠けている場合、その値をゼロとして扱わず、該当する平均値から除外します。
コスト指標では、プロバイダーの料金体系で区別されている場合、キャッシュ済み入力と未キャッシュ入力を分けて扱います。また、プロバイダーがプロンプトのキャッシュ状態を作成する際に課金する場合は、キャッシュ書き込み料金を含めます。これは、一律のトークン単価による推定よりも、従量制APIのトークン料金を正確に反映することを目的としています。
エージェント設定
公開ベンチマークの各行は、モデル名だけでなく、エージェントのバリアントを表します。動作を変える可能性のある設定は、レポート上で個別に扱います。
特に記載がない限り、ベンチマークがデフォルトのユーザー体験を反映するよう、各エージェントのデフォルトの推論設定を使用します。
新しい評価やエージェントのバリアントが追加されるにつれて、ベンチマークの方法は変化する可能性があります。ただし、公開比較は、公開ベンチマークスイート内で同等条件のエージェントバリアントを反映することを意図しています。
バージョン履歴
バージョン1.3
2026年7月~現在
- SWE-Atlas-QnAのバイナリpass/fail採点を改善し、Scale AIのTask Resolve Rate評価手法に整合
バージョン1.2
2026年7月
- SWE-Atlas-QnAの採点をルーブリック報酬からバイナリpass/failに変更し、タスクを正解と判定するにはルーブリックの全基準への合格を必須化
バージョン1.1
2026年6月~2026年7月
- DeepSWE(長期的なソフトウェアエンジニアリング)を追加
- Coding Agent IndexからSWE-Bench-Pro-Hard-AAを削除
バージョン1.0
2026年5月~2026年6月
- SWE-Bench-Pro-Hard-AA(コード生成)、Terminal-Bench v2(エージェントによるターミナル操作)、SWE-Atlas-QnA(リポジトリに関するQ&A)を含む初回リリース