Coding Agent Index 방법론

개요

Artificial Analysis는 종단 간 소프트웨어 엔지니어링 작업으로 코딩 에이전트를 벤치마킹합니다. 에이전트가 실제와 유사한 코딩 작업을 얼마나 잘 완료하는지, 그리고 결과, 신뢰성, 토큰 사용량, 비용, 실행 시간에 따라 성능이 어떻게 달라지는지를 측정하는 것이 목표입니다.

Coding Agent Index 페이지의 공개 결과는 작업 수준의 벤치마크 시도를 바탕으로 평가별 점수, 통합 효율성 지표, Artificial Analysis Coding Agent Index로 집계됩니다.

이 페이지에서는 공개 Artificial Analysis Coding Agent Index를 구성하는 방식, 현재 포함된 벤치마크 구성 요소, 공개 pass@1, 비용, 토큰 사용량, 실행 시간 지표의 산출 방식을 중점적으로 설명합니다.

Artificial Analysis Coding Agent Index

현재 공개된 Artificial Analysis Coding Agent Index는 공개 코딩 에이전트 스위트에 설정된 벤치마크 구성 요소로 만든 종합 벤치마크 점수입니다.

이 지수의 목적은 모든 코딩 작업을 한 가지 벤치마크 작업 유형으로 축약하는 데 있지 않습니다. 코딩 에이전트마다 저장소 질의응답, 구현 및 버그 수정 작업, 터미널 사용 비중이 높은 워크플로에서 성능이 크게 다를 수 있습니다. 이 지수는 벤치마크별 세부 내역을 그대로 제공하면서 서로 다른 벤치마크군을 하나의 상위 성능 보기로 요약합니다.

지수 구성 요소

현재 공개 지수에는 다음 벤치마크 구성 요소가 포함됩니다.

평가분야작업작업당 시도 횟수응답 유형점수 산정
DeepSWE장기 소프트웨어 엔지니어링1133코드 패치 / 저장소 변경프로그램 검증기 통과/실패, pass@1
Terminal-Bench v2에이전트 기반 터미널 사용84*3터미널 기반 작업 실행테스트 스위트 통과/실패, pass@1
SWE-Atlas-QnA저장소 질의응답1243자유 응답Scale AI Task Resolve Rate(이진 통과/실패), pass@1

* Terminal-Bench v2에는 원래 89개 작업이 포함되어 있으나, 환경 호환성 문제로 작업 5개를 제외합니다.

평가 작업

현재 공개 지수는 3개 벤치마크 구성 요소에서 평가한 총 321개 작업을 다룹니다.

  • Add iterable collection combinators to true-myth
  • Abort pending body reads on shutdown
  • Add rolling min, max, median, and quantile methods
  • Format BigQuery pipe syntax queries correctly
  • Add unified manifest stream output across Helm commands
  • Add a deterministic CookieStore with modern Set-Cookie parsing
  • Preserve restored query state in persisted snapshots
  • Add an error-accumulating Validated container
  • Add a persistent analysis cache to Vulture
  • Add multi-module memory snapshots to wazero
  • Add JSONPath query APIs to orderedmap and Starlark modules
  • Add a per-origin circuit breaker to ofetch
  • Add stepped slices for arrays and strings
  • Add `matchEach` to ts-pattern
  • Harden module loading, cache introspection, and script flags
  • Add action pinning linting for actions and reusable workflows
  • Add typed window function builders with OVER clauses
  • Add typed blend range access and blend-if compositing
  • Add deterministic map conflict detection to Y.Map writes
  • Add dependency-aware async initialization to the container
  • Add multipart response parsing to HTTPX
  • Add scoped per-rule ignore markers to Obsidian Linter
  • Partition report files by launcher and expand report templates
  • Add keyset cursor pagination to `$find`
  • Add ShapeIndex encoding and decoding
  • Add entity snapshot and rollback APIs to Koota
  • Restore RichLog follow-state parity and expand reflow behavior
  • Add duration-aware sharding to Vitest
  • Add grouped test phases with synchronized barriers
  • Add bidirectional TOML table converters
  • Add go:embed directive support for interpreted packages
  • Add task snapshots, inspection, and diffing to aiomonitor
  • Add drift detection and compliance baselines
  • Expose accumulated streamed function-call args in SDK surfaces
  • Add grouping-set and window-frame SQL helpers
  • Add rule evaluation profiling to Rego
  • Add typed variable bindings to Anko
  • Add multiplexed ordered streams over KCP
  • Add single-active-consumer priority and cancel tracking to virtual transports
  • Reconstruct template strings in partial evaluation output
  • Add bounded-memory spilling to SCC aggregation
  • Fix isolated Go-side calls for Tengo callables and closures
  • Add interprocedural taint checks for Bandit injection sinks
  • Add deterministic multi-key sorting to fd
  • Add task graph export with JSON, DOT, and text output
  • Add default arguments to Anko function parameters
  • Add deprecation, sunset, and successor headers to FastAPI routes
  • Add a checker for broken doc comment links
  • Add boundary modes to `@stencil`
  • Add retry-aware publishing audit logs
  • Add method declarations and interface dispatch to Scriggo
  • Add partial structuring with error recovery to cattrs
  • Add conditional required attributes to schemas
  • Add worktree merge conflict handling
  • Add session bundle recording and replay to IPython
  • Add flattened dataclass fields to Mashumaro field options
  • Add build-time grammar conflict analysis to participle
  • Add XML diff, patch, and merge operations to etree
  • Add streaming JSON iteration to HTTPX responses
  • Validate daemon watch, status, and log lifecycle
  • Add recursive schema composition to Valibot
  • Add input key aliases to name mapping
  • Add durability callbacks and wait APIs for sync writes
  • Add duration encoding to TableVectorizer
  • Add safe import checkpoints and invariant validation
  • Add shorthand expansion and compression to the lexer
  • Add RFC 5545 timezone interoperability to dateutil recurrence parsing
  • Add atomic signal selectors to Kea
  • Add consistent hash policy support to TrafficPolicy
  • Fix PromQL label sorting across typed and untyped values
  • Format CREATE TABLE DDL and add DDL parsing helpers
  • Add trap coredump generation to wasmi
  • Add hierarchical evaluation cancellation to Boa
  • Add value-based query predicates to Koota
  • Add request coalescing to `Runnable`
  • Add explicit resource management declarations to the parser
  • Add lazy recursive schemas with DTO and JSON Schema export
  • Persist the fitted feature schema across evaluate, predict, serve, and export
  • Add composite trait aspects to Koota
  • Add tube multiplexing to pwntools
  • Add scoped state data to state machine callbacks and history
  • Add destructuring bindings to Tengo
  • Add JSON Schema refs and dependency keywords
  • Add link format conversion between wiki and markdown syntax
  • Add implicit HEAD and automatic OPTIONS responses to FastAPI routes
  • Add transparent encryption to dump uploads
  • Add conditional option dependencies to Optique
  • Preserve structure needed by stylesheet selectors
  • Add policy-based alerting for failures, latency, and SSL expiry
  • Add incremental cache controls to Bandit
  • Coalesce qualifying choices into character classes
  • Preserve ANSI resets during truncation and styling
  • Add async autocomplete options and fetch lifecycle handling
  • Add config file parsing to Cliffy commands
  • Add HTML document format handling to Dasel
  • Add CSS Grid layout to the Box component
  • Add `\multicolumn` column spans to array-like environments
  • Add a deferred mutation buffer to batch entity changes
  • Add automatic table of contents generation for Obsidian linter
  • Implement a deterministic IntersectionObserver in Happy DOM
  • Add pair-level relation tracking modifiers
  • Add error stack serialization to SuperJSON
  • Complete Kitty keyboard phases and stable fallback key metadata
  • Add structured nosec directives for regions and next line
  • Add bail-on-test-failure handling to Testem
  • Add try/catch error recovery to expr
  • Add configurable array merge strategies to Helm value coalescing
  • Implement recursive agent delegation through delegate_task tool calls
  • Add SSE streaming endpoints to HttpApi
  • Add transactional reload status and rollback tracking to Prometheus
  • Add GraphQL incremental delivery with @defer and @stream
  • Add dead-lettering, TTL, and overflow handling to virtual queues
  • Reuse one toolbar across multiple Quill editors

  • adaptive-rejection-sampler
  • bn-fit-modify
  • break-filter-js-from-html
  • build-cython-ext
  • build-pmars
  • build-pov-ray
  • caffe-cifar-10
  • cancel-async-tasks
  • chess-best-move
  • circuit-fibsqrt
  • cobol-modernization
  • code-from-image
  • compile-compcert
  • configure-git-webserver
  • constraints-scheduling
  • count-dataset-tokens
  • crack-7z-hash
  • custom-memory-heap-crash
  • db-wal-recovery
  • distribution-search
  • dna-assembly
  • dna-insert
  • extract-elf
  • extract-moves-from-video
  • feal-differential-cryptanalysis
  • financial-document-processor
  • fix-code-vulnerability
  • fix-git
  • fix-ocaml-gc
  • gcode-to-text
  • git-leak-recovery
  • git-multibranch
  • headless-terminal
  • hf-model-inference
  • install-windows-3.11
  • kv-store-grpc
  • large-scale-text-editing
  • largest-eigenval
  • llm-inference-batching-scheduler
  • log-summary-date-ranges
  • mailman
  • make-mips-interpreter
  • mcmc-sampling-stan
  • merge-diff-arc-agi-task
  • model-extraction-relu-logits
  • modernize-scientific-stack
  • mteb-leaderboard
  • mteb-retrieve
  • multi-source-data-merger
  • nginx-request-logging
  • openssl-selfsigned-cert
  • overfull-hbox
  • password-recovery
  • path-tracing
  • path-tracing-reverse
  • polyglot-c-py
  • polyglot-rust-c
  • portfolio-optimization
  • protein-assembly
  • prove-plus-comm
  • pypi-server
  • pytorch-model-cli
  • pytorch-model-recovery
  • qemu-alpine-ssh
  • qemu-startup
  • query-optimize
  • raman-fitting
  • regex-chess
  • regex-log
  • reshard-c4-data
  • rstan-to-pystan
  • sam-cell-seg
  • sanitize-git-repo
  • schemelike-metacircular-eval
  • sparql-university
  • sqlite-db-truncate
  • sqlite-with-gcov
  • torch-pipeline-parallelism
  • torch-tensor-parallelism
  • train-fasttext
  • tune-mjcf
  • video-processing
  • vulnerable-secret
  • winning-avg-corewars

  • 6905333b74f22949d97ba998
  • 6905333b74f22949d97ba999
  • 6905333b74f22949d97ba99a
  • 6905333b74f22949d97ba99b
  • 6905333b74f22949d97ba99d
  • 6905333b74f22949d97ba99f
  • 6905333b74f22949d97ba9a2
  • 6905333b74f22949d97ba9a3
  • 6905333b74f22949d97ba9a4
  • 6905333b74f22949d97ba9a5
  • 6905333b74f22949d97ba9a6
  • 6905333b74f22949d97ba9a7
  • 6905333b74f22949d97ba9a8
  • 6905333b74f22949d97ba9a9
  • 6905333b74f22949d97ba9aa
  • 6905333b74f22949d97ba9ab
  • 6905333b74f22949d97ba9ac
  • 6905333b74f22949d97ba9ad
  • 6905333b74f22949d97ba9ae
  • 6905333b74f22949d97ba9af
  • 6905333b74f22949d97ba9b1
  • 6905333b74f22949d97ba9b2
  • 6905333b74f22949d97ba9b3
  • 6905333b74f22949d97ba9b5
  • 6905333b74f22949d97ba9b6
  • 6905333b74f22949d97ba9b7
  • 6905333b74f22949d97ba9b8
  • 6905333b74f22949d97ba9ba
  • 6905333b74f22949d97ba9bb
  • 6905333b74f22949d97ba9bc
  • 6905333b74f22949d97ba9bd
  • 6905333b74f22949d97ba9be
  • 6905333b74f22949d97ba9bf
  • 6905333b74f22949d97ba9c0
  • 6905333b74f22949d97ba9c1
  • 6905333b74f22949d97ba9c2
  • 6905333b74f22949d97ba9c3
  • 6905333b74f22949d97ba9c4
  • 6905333b74f22949d97ba9c5
  • 6905333b74f22949d97ba9c6
  • 6905333b74f22949d97ba9c8
  • 6905333b74f22949d97ba9c9
  • 6905333b74f22949d97ba9ca
  • 6905333b74f22949d97ba9cb
  • 6905333b74f22949d97ba9cc
  • 6905333b74f22949d97ba9cd
  • 6905333b74f22949d97ba9ce
  • 6905333b74f22949d97ba9cf
  • 6905333b74f22949d97ba9d0
  • 6905333b74f22949d97ba9d1
  • 6905333b74f22949d97ba9d2
  • 6905333b74f22949d97ba9d3
  • 6905333b74f22949d97ba9d4
  • 6905333b74f22949d97ba9d5
  • 6905333b74f22949d97ba9d6
  • 6905333b74f22949d97ba9d7
  • 6905333b74f22949d97ba9d8
  • 6905333b74f22949d97ba9d9
  • 6905333b74f22949d97ba9db
  • 6905333b74f22949d97ba9dc
  • 6905333b74f22949d97ba9dd
  • 6905333b74f22949d97ba9de
  • 6905333b74f22949d97ba9e0
  • 6905333b74f22949d97ba9e1
  • 6905333b74f22949d97ba9e3
  • 6905333b74f22949d97ba9e4
  • 6905333b74f22949d97ba9e5
  • 6905333b74f22949d97ba9e7
  • 6905333b74f22949d97ba9e8
  • 6905333b74f22949d97ba9e9
  • 6905333b74f22949d97ba9eb
  • 6905333b74f22949d97ba9ee
  • 6905333b74f22949d97ba9f0
  • 6905333b74f22949d97ba9f1
  • 6905333b74f22949d97ba9f2
  • 6905333b74f22949d97ba9f4
  • 6905333b74f22949d97ba9f5
  • 6905333b74f22949d97ba9f7
  • 6905333b74f22949d97ba9f8
  • 6905333b74f22949d97ba9f9
  • 6905333b74f22949d97ba9fa
  • 6905333b74f22949d97ba9fb
  • 6905333b74f22949d97ba9fc
  • 6905333b74f22949d97ba9fd
  • 6905333b74f22949d97ba9ff
  • 6905333b74f22949d97baa01
  • 6905333b74f22949d97baa02
  • 6905333b74f22949d97baa03
  • 6905333b74f22949d97baa04
  • 6905333b74f22949d97baa05
  • 6905333b74f22949d97baa06
  • 6905333b74f22949d97baa07
  • 6905333b74f22949d97baa09
  • 6905333b74f22949d97baa0b
  • 6905333b74f22949d97baa0c
  • 6905333b74f22949d97baa0d
  • 6905333b74f22949d97baa0f
  • 6905333b74f22949d97baa10
  • 6905333b74f22949d97baa11
  • 6905333b74f22949d97baa12
  • 6905333b74f22949d97baa14
  • 6905333b74f22949d97baa15
  • 6905333b74f22949d97baa16
  • 6905333b74f22949d97baa17
  • 6905333b74f22949d97baa19
  • 6905333b74f22949d97baa1a
  • 6905333b74f22949d97baa1b
  • 6905333b74f22949d97baa1c
  • 6905333b74f22949d97baa1d
  • 6905333b74f22949d97baa1e
  • 6905333b74f22949d97baa1f
  • 6905333b74f22949d97baa20
  • 6905333b74f22949d97baa21
  • 6905333b74f22949d97baa22
  • 6905333b74f22949d97baa23
  • 6905333b74f22949d97baa24
  • 6905333b74f22949d97baa25
  • 6905333b74f22949d97baa26
  • 6905333b74f22949d97baa27
  • 6905333b74f22949d97baa28
  • 6905333b74f22949d97baa2a
  • 6905333b74f22949d97baa2b
  • 6905333b74f22949d97baa2c
  • 6905333b74f22949d97baa2d

지수의 집계 대상

Artificial Analysis는 에이전트 변형별로 포함된 각 벤치마크 구성 요소의 pass@1 점수를 계산한 뒤, 구성 요소 점수를 공개 지수로 집계합니다.

동일한 벤치마크 스위트가 벤치마크 페이지의 실행 비용, 토큰 사용량, 실행 시간을 포함한 공개 통합 효율성 지표의 기반으로도 사용됩니다. 따라서 성능과 효율성 보기는 서로 관련 없는 실행 결과가 아니라 동일한 기반 벤치마크 범위에 맞춰져 있습니다.

점수 산정 및 결과

pass@1 결과

평가된 각 시도에는 벤치마크 평가기가 pass@1 결과를 부여합니다. SWE-Atlas-QnA는 Scale AI의 Task Resolve Rate 방법론에 맞춘 이진 통과/실패 점수 산정 방식을 사용합니다. 통과한 시도는 1점, 실패하거나 오류가 발생한 시도는 0점을 받습니다.

용어정의
이진 pass@1작업에 통과 시 1, 실패 시 0을 부여하는 테스트 스위트 평가 결과입니다.

평가별 점수

Artificial Analysis는 평가마다 먼저 각 작업에서 수행한 세 번의 평가 시도에 대한 평균을 계산한 다음, 모든 작업의 가중치가 같도록 작업 수준의 pass@1 점수를 평균합니다.

효율성 지표

비용, 토큰 사용량, 실행 시간은 현재 공개 코딩 에이전트 벤치마크 스위트 전체의 작업 시도당 통합 평균으로 표시합니다.

  • 실행 비용: 소비자용 요금제가 아닌 제공업체의 토큰 가격을 기준으로 한 작업당 평균 토큰 종량제 API 비용입니다.
  • 토큰 사용량: 작업당 평균 입력, 캐시, 캐시 쓰기, 추론, 출력 토큰 수입니다.
  • 실행 시간: 작업당 평균 실제 경과 시간으로, 전체 작업 경과 시간과 확인 가능한 경우 에이전트 경과 시간 구간을 포함합니다.

특정 지표의 텔레메트리가 누락된 경우 해당 누락값을 0으로 처리하지 않고 관련 평균에서 제외합니다.

비용 지표에서는 제공업체 가격이 이를 구분할 경우 캐시된 입력을 캐시되지 않은 입력과 별도로 처리하며, 제공업체가 프롬프트 캐시 상태 생성에 요금을 청구하면 캐시 쓰기 요금도 포함합니다. 이는 일률적인 토큰당 추정치보다 토큰 종량제 API 가격을 더 정확하게 반영하기 위한 것입니다.

에이전트 설정

공개 벤치마크의 각 행은 모델 이름만이 아니라 에이전트 변형을 나타냅니다. 동작을 바꿀 수 있는 설정은 보고할 때 별도로 구분합니다.

달리 명시하지 않는 한 벤치마크가 기본 사용자 경험을 반영하도록 각 에이전트의 기본 추론 설정을 사용합니다.

새로운 평가와 에이전트 변형이 추가됨에 따라 벤치마킹 방법론은 시간이 지나면서 변경될 수 있지만, 공개 비교는 게시된 벤치마크 스위트 내에서 동등한 조건의 에이전트 변형을 비교하기 위한 것입니다.

버전 기록

버전 1.3

2026년 7월—현재

  • Scale AI의 Task Resolve Rate 방법론에 맞게 SWE-Atlas-QnA 이진 통과/실패 점수 산정 방식을 개선

버전 1.2

2026년 7월

  • SWE-Atlas-QnA 점수 산정 방식을 루브릭 보상에서 이진 통과/실패 방식으로 변경하고, 작업이 정답으로 인정되려면 모든 루브릭 기준을 통과하도록 설정

버전 1.1

2026년 6월—2026년 7월

  • DeepSWE(장기 소프트웨어 엔지니어링) 추가
  • Coding Agent Index에서 SWE-Bench-Pro-Hard-AA 제거

버전 1.0

2026년 5월—2026년 6월

  • SWE-Bench-Pro-Hard-AA(코드 생성), Terminal-Bench v2(에이전트 기반 터미널 사용), SWE-Atlas-QnA(저장소 질의응답)로 최초 출시