Coding Agent Indexの方法論

概要

Artificial Analysisでは、エンドツーエンドのソフトウェアエンジニアリングタスクでコーディングエージェントをベンチマークしています。現実的なコーディング作業をエージェントがどの程度完遂できるか、また成果、信頼性、トークン使用量、コスト、実行時間によって性能がどう異なるかを測定することが目的です。

Coding Agent Indexページの公開結果は、タスク単位のベンチマーク試行から構築され、評価ごとのスコア、統合した効率性指標、Artificial Analysis Coding Agent Indexに集約されます。

このページでは、公開されているArtificial Analysis Coding Agent Indexの構築方法、現在含まれているベンチマーク構成要素、公開されるpass@1、コスト、トークン使用量、実行時間の各指標の算出方法を説明します。

Artificial Analysis Coding Agent Index

現在公開しているArtificial Analysis Coding Agent Indexは、公開コーディングエージェントスイートに設定されたベンチマーク構成要素から算出する複合ベンチマークスコアです。

この指数の目的は、すべてのコーディング作業を1種類のベンチマークタスクに集約することではありません。コーディングエージェントの性能は、リポジトリに関するQ&A、実装やバグ修正のタスク、ターミナル操作が中心のワークフローで大きく異なる場合があります。指数は、ベンチマークごとの内訳を維持しながら、こうした異なるベンチマーク群を1つの上位の性能ビューにまとめるためのものです。

指数の構成要素

現在の公開指数には、次のベンチマーク構成要素が含まれています。

評価分野タスクタスクあたりの試行回数回答形式採点
DeepSWE長期的なソフトウェアエンジニアリング1133コードパッチ / リポジトリの変更プログラム検証のpass/fail、pass@1
Terminal-Bench v2エージェントによるターミナル操作84*3ターミナルベースのタスク実行テストスイートのpass/fail、pass@1
SWE-Atlas-QnAリポジトリに関するQ&A1243自由回答Scale AI Task Resolve Rate(バイナリpass/fail)、pass@1

* Terminal-Bench v2には本来89件のタスクがありますが、環境互換性の問題により5件を除外しています。

評価対象タスク

現在の公開指数は、3件のベンチマーク構成要素にわたる321件の評価対象タスクを網羅しています。

  • Add iterable collection combinators to true-myth
  • Abort pending body reads on shutdown
  • Add rolling min, max, median, and quantile methods
  • Format BigQuery pipe syntax queries correctly
  • Add unified manifest stream output across Helm commands
  • Add a deterministic CookieStore with modern Set-Cookie parsing
  • Preserve restored query state in persisted snapshots
  • Add an error-accumulating Validated container
  • Add a persistent analysis cache to Vulture
  • Add multi-module memory snapshots to wazero
  • Add JSONPath query APIs to orderedmap and Starlark modules
  • Add a per-origin circuit breaker to ofetch
  • Add stepped slices for arrays and strings
  • Add `matchEach` to ts-pattern
  • Harden module loading, cache introspection, and script flags
  • Add action pinning linting for actions and reusable workflows
  • Add typed window function builders with OVER clauses
  • Add typed blend range access and blend-if compositing
  • Add deterministic map conflict detection to Y.Map writes
  • Add dependency-aware async initialization to the container
  • Add multipart response parsing to HTTPX
  • Add scoped per-rule ignore markers to Obsidian Linter
  • Partition report files by launcher and expand report templates
  • Add keyset cursor pagination to `$find`
  • Add ShapeIndex encoding and decoding
  • Add entity snapshot and rollback APIs to Koota
  • Restore RichLog follow-state parity and expand reflow behavior
  • Add duration-aware sharding to Vitest
  • Add grouped test phases with synchronized barriers
  • Add bidirectional TOML table converters
  • Add go:embed directive support for interpreted packages
  • Add task snapshots, inspection, and diffing to aiomonitor
  • Add drift detection and compliance baselines
  • Expose accumulated streamed function-call args in SDK surfaces
  • Add grouping-set and window-frame SQL helpers
  • Add rule evaluation profiling to Rego
  • Add typed variable bindings to Anko
  • Add multiplexed ordered streams over KCP
  • Add single-active-consumer priority and cancel tracking to virtual transports
  • Reconstruct template strings in partial evaluation output
  • Add bounded-memory spilling to SCC aggregation
  • Fix isolated Go-side calls for Tengo callables and closures
  • Add interprocedural taint checks for Bandit injection sinks
  • Add deterministic multi-key sorting to fd
  • Add task graph export with JSON, DOT, and text output
  • Add default arguments to Anko function parameters
  • Add deprecation, sunset, and successor headers to FastAPI routes
  • Add a checker for broken doc comment links
  • Add boundary modes to `@stencil`
  • Add retry-aware publishing audit logs
  • Add method declarations and interface dispatch to Scriggo
  • Add partial structuring with error recovery to cattrs
  • Add conditional required attributes to schemas
  • Add worktree merge conflict handling
  • Add session bundle recording and replay to IPython
  • Add flattened dataclass fields to Mashumaro field options
  • Add build-time grammar conflict analysis to participle
  • Add XML diff, patch, and merge operations to etree
  • Add streaming JSON iteration to HTTPX responses
  • Validate daemon watch, status, and log lifecycle
  • Add recursive schema composition to Valibot
  • Add input key aliases to name mapping
  • Add durability callbacks and wait APIs for sync writes
  • Add duration encoding to TableVectorizer
  • Add safe import checkpoints and invariant validation
  • Add shorthand expansion and compression to the lexer
  • Add RFC 5545 timezone interoperability to dateutil recurrence parsing
  • Add atomic signal selectors to Kea
  • Add consistent hash policy support to TrafficPolicy
  • Fix PromQL label sorting across typed and untyped values
  • Format CREATE TABLE DDL and add DDL parsing helpers
  • Add trap coredump generation to wasmi
  • Add hierarchical evaluation cancellation to Boa
  • Add value-based query predicates to Koota
  • Add request coalescing to `Runnable`
  • Add explicit resource management declarations to the parser
  • Add lazy recursive schemas with DTO and JSON Schema export
  • Persist the fitted feature schema across evaluate, predict, serve, and export
  • Add composite trait aspects to Koota
  • Add tube multiplexing to pwntools
  • Add scoped state data to state machine callbacks and history
  • Add destructuring bindings to Tengo
  • Add JSON Schema refs and dependency keywords
  • Add link format conversion between wiki and markdown syntax
  • Add implicit HEAD and automatic OPTIONS responses to FastAPI routes
  • Add transparent encryption to dump uploads
  • Add conditional option dependencies to Optique
  • Preserve structure needed by stylesheet selectors
  • Add policy-based alerting for failures, latency, and SSL expiry
  • Add incremental cache controls to Bandit
  • Coalesce qualifying choices into character classes
  • Preserve ANSI resets during truncation and styling
  • Add async autocomplete options and fetch lifecycle handling
  • Add config file parsing to Cliffy commands
  • Add HTML document format handling to Dasel
  • Add CSS Grid layout to the Box component
  • Add `\multicolumn` column spans to array-like environments
  • Add a deferred mutation buffer to batch entity changes
  • Add automatic table of contents generation for Obsidian linter
  • Implement a deterministic IntersectionObserver in Happy DOM
  • Add pair-level relation tracking modifiers
  • Add error stack serialization to SuperJSON
  • Complete Kitty keyboard phases and stable fallback key metadata
  • Add structured nosec directives for regions and next line
  • Add bail-on-test-failure handling to Testem
  • Add try/catch error recovery to expr
  • Add configurable array merge strategies to Helm value coalescing
  • Implement recursive agent delegation through delegate_task tool calls
  • Add SSE streaming endpoints to HttpApi
  • Add transactional reload status and rollback tracking to Prometheus
  • Add GraphQL incremental delivery with @defer and @stream
  • Add dead-lettering, TTL, and overflow handling to virtual queues
  • Reuse one toolbar across multiple Quill editors

  • adaptive-rejection-sampler
  • bn-fit-modify
  • break-filter-js-from-html
  • build-cython-ext
  • build-pmars
  • build-pov-ray
  • caffe-cifar-10
  • cancel-async-tasks
  • chess-best-move
  • circuit-fibsqrt
  • cobol-modernization
  • code-from-image
  • compile-compcert
  • configure-git-webserver
  • constraints-scheduling
  • count-dataset-tokens
  • crack-7z-hash
  • custom-memory-heap-crash
  • db-wal-recovery
  • distribution-search
  • dna-assembly
  • dna-insert
  • extract-elf
  • extract-moves-from-video
  • feal-differential-cryptanalysis
  • financial-document-processor
  • fix-code-vulnerability
  • fix-git
  • fix-ocaml-gc
  • gcode-to-text
  • git-leak-recovery
  • git-multibranch
  • headless-terminal
  • hf-model-inference
  • install-windows-3.11
  • kv-store-grpc
  • large-scale-text-editing
  • largest-eigenval
  • llm-inference-batching-scheduler
  • log-summary-date-ranges
  • mailman
  • make-mips-interpreter
  • mcmc-sampling-stan
  • merge-diff-arc-agi-task
  • model-extraction-relu-logits
  • modernize-scientific-stack
  • mteb-leaderboard
  • mteb-retrieve
  • multi-source-data-merger
  • nginx-request-logging
  • openssl-selfsigned-cert
  • overfull-hbox
  • password-recovery
  • path-tracing
  • path-tracing-reverse
  • polyglot-c-py
  • polyglot-rust-c
  • portfolio-optimization
  • protein-assembly
  • prove-plus-comm
  • pypi-server
  • pytorch-model-cli
  • pytorch-model-recovery
  • qemu-alpine-ssh
  • qemu-startup
  • query-optimize
  • raman-fitting
  • regex-chess
  • regex-log
  • reshard-c4-data
  • rstan-to-pystan
  • sam-cell-seg
  • sanitize-git-repo
  • schemelike-metacircular-eval
  • sparql-university
  • sqlite-db-truncate
  • sqlite-with-gcov
  • torch-pipeline-parallelism
  • torch-tensor-parallelism
  • train-fasttext
  • tune-mjcf
  • video-processing
  • vulnerable-secret
  • winning-avg-corewars

  • 6905333b74f22949d97ba998
  • 6905333b74f22949d97ba999
  • 6905333b74f22949d97ba99a
  • 6905333b74f22949d97ba99b
  • 6905333b74f22949d97ba99d
  • 6905333b74f22949d97ba99f
  • 6905333b74f22949d97ba9a2
  • 6905333b74f22949d97ba9a3
  • 6905333b74f22949d97ba9a4
  • 6905333b74f22949d97ba9a5
  • 6905333b74f22949d97ba9a6
  • 6905333b74f22949d97ba9a7
  • 6905333b74f22949d97ba9a8
  • 6905333b74f22949d97ba9a9
  • 6905333b74f22949d97ba9aa
  • 6905333b74f22949d97ba9ab
  • 6905333b74f22949d97ba9ac
  • 6905333b74f22949d97ba9ad
  • 6905333b74f22949d97ba9ae
  • 6905333b74f22949d97ba9af
  • 6905333b74f22949d97ba9b1
  • 6905333b74f22949d97ba9b2
  • 6905333b74f22949d97ba9b3
  • 6905333b74f22949d97ba9b5
  • 6905333b74f22949d97ba9b6
  • 6905333b74f22949d97ba9b7
  • 6905333b74f22949d97ba9b8
  • 6905333b74f22949d97ba9ba
  • 6905333b74f22949d97ba9bb
  • 6905333b74f22949d97ba9bc
  • 6905333b74f22949d97ba9bd
  • 6905333b74f22949d97ba9be
  • 6905333b74f22949d97ba9bf
  • 6905333b74f22949d97ba9c0
  • 6905333b74f22949d97ba9c1
  • 6905333b74f22949d97ba9c2
  • 6905333b74f22949d97ba9c3
  • 6905333b74f22949d97ba9c4
  • 6905333b74f22949d97ba9c5
  • 6905333b74f22949d97ba9c6
  • 6905333b74f22949d97ba9c8
  • 6905333b74f22949d97ba9c9
  • 6905333b74f22949d97ba9ca
  • 6905333b74f22949d97ba9cb
  • 6905333b74f22949d97ba9cc
  • 6905333b74f22949d97ba9cd
  • 6905333b74f22949d97ba9ce
  • 6905333b74f22949d97ba9cf
  • 6905333b74f22949d97ba9d0
  • 6905333b74f22949d97ba9d1
  • 6905333b74f22949d97ba9d2
  • 6905333b74f22949d97ba9d3
  • 6905333b74f22949d97ba9d4
  • 6905333b74f22949d97ba9d5
  • 6905333b74f22949d97ba9d6
  • 6905333b74f22949d97ba9d7
  • 6905333b74f22949d97ba9d8
  • 6905333b74f22949d97ba9d9
  • 6905333b74f22949d97ba9db
  • 6905333b74f22949d97ba9dc
  • 6905333b74f22949d97ba9dd
  • 6905333b74f22949d97ba9de
  • 6905333b74f22949d97ba9e0
  • 6905333b74f22949d97ba9e1
  • 6905333b74f22949d97ba9e3
  • 6905333b74f22949d97ba9e4
  • 6905333b74f22949d97ba9e5
  • 6905333b74f22949d97ba9e7
  • 6905333b74f22949d97ba9e8
  • 6905333b74f22949d97ba9e9
  • 6905333b74f22949d97ba9eb
  • 6905333b74f22949d97ba9ee
  • 6905333b74f22949d97ba9f0
  • 6905333b74f22949d97ba9f1
  • 6905333b74f22949d97ba9f2
  • 6905333b74f22949d97ba9f4
  • 6905333b74f22949d97ba9f5
  • 6905333b74f22949d97ba9f7
  • 6905333b74f22949d97ba9f8
  • 6905333b74f22949d97ba9f9
  • 6905333b74f22949d97ba9fa
  • 6905333b74f22949d97ba9fb
  • 6905333b74f22949d97ba9fc
  • 6905333b74f22949d97ba9fd
  • 6905333b74f22949d97ba9ff
  • 6905333b74f22949d97baa01
  • 6905333b74f22949d97baa02
  • 6905333b74f22949d97baa03
  • 6905333b74f22949d97baa04
  • 6905333b74f22949d97baa05
  • 6905333b74f22949d97baa06
  • 6905333b74f22949d97baa07
  • 6905333b74f22949d97baa09
  • 6905333b74f22949d97baa0b
  • 6905333b74f22949d97baa0c
  • 6905333b74f22949d97baa0d
  • 6905333b74f22949d97baa0f
  • 6905333b74f22949d97baa10
  • 6905333b74f22949d97baa11
  • 6905333b74f22949d97baa12
  • 6905333b74f22949d97baa14
  • 6905333b74f22949d97baa15
  • 6905333b74f22949d97baa16
  • 6905333b74f22949d97baa17
  • 6905333b74f22949d97baa19
  • 6905333b74f22949d97baa1a
  • 6905333b74f22949d97baa1b
  • 6905333b74f22949d97baa1c
  • 6905333b74f22949d97baa1d
  • 6905333b74f22949d97baa1e
  • 6905333b74f22949d97baa1f
  • 6905333b74f22949d97baa20
  • 6905333b74f22949d97baa21
  • 6905333b74f22949d97baa22
  • 6905333b74f22949d97baa23
  • 6905333b74f22949d97baa24
  • 6905333b74f22949d97baa25
  • 6905333b74f22949d97baa26
  • 6905333b74f22949d97baa27
  • 6905333b74f22949d97baa28
  • 6905333b74f22949d97baa2a
  • 6905333b74f22949d97baa2b
  • 6905333b74f22949d97baa2c
  • 6905333b74f22949d97baa2d

Indexの集約対象

Artificial Analysisでは、エージェントの各バリアントについて、含まれるベンチマーク構成要素ごとにpass@1スコアを計算し、それらのスコアを公開指数に集約します。

同じベンチマークスイートは、ベンチマークページに公開する実行コスト、トークン使用量、実行時間などの統合効率性指標の基礎にもなっています。つまり、性能と効率性のビューは、無関係な実行結果ではなく、同じ基礎ベンチマークの対象範囲に基づいています。

採点と結果

pass@1の結果

評価対象の各試行には、ベンチマーク評価器によるpass@1の結果が与えられます。SWE-Atlas-QnAでは、Scale AIのTask Resolve Rate評価手法に沿ったバイナリpass/fail採点を使用します。合格した試行は1、失敗またはエラーとなった試行は0です。

用語定義
バイナリpass@1タスクに合格なら1、不合格なら0を与えるテストスイート評価の結果。

評価ごとのスコア

各評価では、まずタスクごとに評価した3回の試行を平均し、次にタスクの重みがすべて等しくなるよう、タスク単位のpass@1スコアを平均します。

効率性指標

コスト、トークン使用量、実行時間は、現在公開しているコーディングエージェントベンチマークスイート全体で統合した、タスク試行あたりの平均値として示します。

  • 実行コスト:コンシューマー向けプランではなく、プロバイダーのトークン料金に基づく、タスクあたりの従量制APIコストの平均。
  • トークン使用量:タスクあたりの入力、キャッシュ、キャッシュ書き込み、推論、出力トークンの平均。
  • 実行時間:タスクあたりの経過時間の平均。タスク全体の経過時間と、取得可能な場合はそのうちエージェントが動作していた時間を含みます。

特定の指標でテレメトリが欠けている場合、その値をゼロとして扱わず、該当する平均値から除外します。

コスト指標では、プロバイダーの料金体系で区別されている場合、キャッシュ済み入力と未キャッシュ入力を分けて扱います。また、プロバイダーがプロンプトのキャッシュ状態を作成する際に課金する場合は、キャッシュ書き込み料金を含めます。これは、一律のトークン単価による推定よりも、従量制APIのトークン料金を正確に反映することを目的としています。

エージェント設定

公開ベンチマークの各行は、モデル名だけでなく、エージェントのバリアントを表します。動作を変える可能性のある設定は、レポート上で個別に扱います。

特に記載がない限り、ベンチマークがデフォルトのユーザー体験を反映するよう、各エージェントのデフォルトの推論設定を使用します。

新しい評価やエージェントのバリアントが追加されるにつれて、ベンチマークの方法は変化する可能性があります。ただし、公開比較は、公開ベンチマークスイート内で同等条件のエージェントバリアントを反映することを意図しています。

バージョン履歴

バージョン1.3

2026年7月~現在

  • SWE-Atlas-QnAのバイナリpass/fail採点を改善し、Scale AIのTask Resolve Rate評価手法に整合

バージョン1.2

2026年7月

  • SWE-Atlas-QnAの採点をルーブリック報酬からバイナリpass/failに変更し、タスクを正解と判定するにはルーブリックの全基準への合格を必須化

バージョン1.1

2026年6月~2026年7月

  • DeepSWE(長期的なソフトウェアエンジニアリング)を追加
  • Coding Agent IndexからSWE-Bench-Pro-Hard-AAを削除

バージョン1.0

2026年5月~2026年6月

  • SWE-Bench-Pro-Hard-AA(コード生成)、Terminal-Bench v2(エージェントによるターミナル操作)、SWE-Atlas-QnA(リポジトリに関するQ&A)を含む初回リリース