Methodik für das Cyber-Index-Benchmarking von Artificial Analysis
Artificial Analysis Cyber Index
The Artificial Analysis Cyber Index tests agentic cyber defense work: discovering vulnerabilities in a codebase, reproducing and validating them, and patching them without breaking existing functionality. It combines three evaluations from industry partners and academia: CWE-Bench-AA (Collinear AI), DeepsecBench-AA (Vercel), and CyberGym-E2E-AA (Berkeley RDI).
The Cyber Index measures defense only. Every task starts with access to the source code, as a security engineer auditing an application would, and no task asks a model to turn a vulnerability into a working exploit. Like all evaluation metrics, it has limitations and may not apply directly to every use case.
The Cyber Index covers identifying and remediating vulnerabilities in source code. It does not cover incident response, writing new code without introducing vulnerabilities, targets without source access such as compiled software or live servers, or exploit realization.
We run every evaluation on our open-source agent harness, Stirrup, giving every model the same prompts and tools within each evaluation. We record cases where a model or provider declines a task on safety grounds and report them separately from the score.
Cyber Index evaluation suite
The Cyber Index is an equally weighted average of its three evaluations.
| Evaluation | Tasks | Repeats | Response Type | Scoring | Cyber Index Weighting | Tool Usage |
|---|---|---|---|---|---|---|
| CWE-Bench-AA | 120 held-out tasks (10 OWASP categories) | 1 | Agentic audit-and-patch of a real open-source repository | Deterministic verifier (exploit blocked + legitimate behavior still works), pass@1 | 1/3 | ✓ |
| DeepsecBench-AA | Private evaluation, scored against an expert-verified golden set | 3 | Agentic static review of open-source application code, submitting findings | LLM judge validates findings and matches them to the golden set, median F2 across repeats | 1/3 | ✓ |
| CyberGym-E2E-AA | 131 memory-safety vulnerabilities from 131 C/C++ projects* | 1 | Agentic proof-of-concept input and source patch | Sanitizer crash, patch and functionality test verification (stage 3), pass@1 | 1/3 | ✓ |
* A filtered subset of the CyberGym-E2E dataset, with one task per project. See CyberGym-E2E-AA for how the tasks were selected.
Evaluationen des Artificial Analysis Cyber Index
The three evaluations in the Artificial Analysis Cyber Index.
CWE-Bench-AA
- Status: Included in the Artificial Analysis Cyber Index at 1/3 weighting.
- Description: CWE-Bench-AA is Artificial Analysis' implementation of Collinear AI's CWE-bench, a defensive cybersecurity benchmark that evaluates coding agents on auditing open-source codebases for security weaknesses and patching them without breaking legitimate behavior. Each task mirrors a real security audit: the agent receives a repository checkout and a single instruction to audit the code and fix what it finds. Collinear AI includes tasks that reproduce disclosed CVEs. The task describes the area of concern but does not give the exact location, and agents are never asked to create exploits.
- Benchmark: https://cwe-bench.com/
- Dataset: 120 held-out tasks, private to Collinear AI and Artificial Analysis. The set covers all ten OWASP Top 10 (2025) categories across C/C++, Go, Java, JavaScript/TypeScript, Python, and Rust. Collinear AI built the set to exclude tasks solvable from memorized fixes alone.
- Agent harness: https://github.com/ArtificialAnalysis/Stirrup
- Implementation:
- Each task runs in an isolated sandboxed container on our Harbor runtime with no internet access. The agent is given the repository checkout and the audit instruction only.
- We run all models on our open-source agent harness, Stirrup, in a standard reason-and-act tool-use loop, with 2 hours to complete each task.
- Grading is deterministic and runs inside the same sandbox after the agent finishes. The task's programmatic verifier confirms that the exploit no longer works and that legitimate behavior still works, returning 1 if both hold and 0 otherwise.
- We run each task once. The headline score is pass@1, the mean score across the 120 tasks.
- Differences from Collinear AI's leaderboard: CWE-Bench-AA is our implementation, run on our Stirrup harness and agent prompts, so our numbers are not directly comparable to the results published at cwe-bench.com.
DeepsecBench-AA
- Status: Included in the Artificial Analysis Cyber Index at 1/3 weighting.
- Description: DeepsecBench-AA is our implementation of Vercel's DeepsecBench, a benchmark that measures how well models find vulnerabilities in application code. The agent reviews real open-source application code, pinned to a commit from before major vulnerability fixes, and reports the vulnerabilities and significant bugs it finds. In each task, the agent is given up to five files flagged by the Deepsec scanner, plus the flags themselves, and investigates whether they point to real vulnerabilities. We score findings against a golden set of vulnerabilities verified by human experts. The agent reviews source code only and is never asked to run the application or exploit what it finds.
- Benchmark: https://vercel.com/blog/deepsecbench-evaluating-model-performance-in-finding-cybersecurity-vulnerabilities
- Dataset: A private evaluation scored against a golden set of vulnerabilities and bugs verified by human experts.
- Agent harness: https://github.com/ArtificialAnalysis/Stirrup
- Implementation:
- Each batch runs in an isolated E2B sandbox containing the pinned repository, with no internet access. We run all models on our open-source agent harness, Stirrup, with code execution and image viewing tools and a limit of 500 turns per batch. The agent submits its findings through a
submit_findingstool. - The prompt is adapted from DeepsecBench's investigation prompt. It asks for security vulnerabilities, classified as critical, high or medium severity, alongside non-security bugs, and lists the vulnerability categories to look for and the mitigations to check before reporting a finding.
- GPT-5.6 Sol (high) is the judge. It first decides whether each finding is a true or false positive, then matches findings to the golden set.
- Recall is the share of golden findings the agent matched. Precision is the share of reported findings the judge accepts as real, with duplicate reports counted against precision. Real findings outside the golden set count toward precision but not recall.
- Following Vercel, the headline score is F2, which weights recall more heavily than precision: a missed vulnerability goes unfixed, while a false positive costs triage time.where P is precision and R is recall. We run the evaluation three times and report the median F2.
- Differences from Vercel's results: DeepsecBench-AA runs every model on our Stirrup harness, so our numbers are not directly comparable to the results Vercel publishes.
- Each batch runs in an isolated E2B sandbox containing the pinned repository, with no internet access. We run all models on our open-source agent harness, Stirrup, with code execution and image viewing tools and a limit of 500 turns per batch. The agent submits its findings through a
CyberGym-E2E-AA
- Status: Included in the Artificial Analysis Cyber Index at 1/3 weighting.
- Description: CyberGym-E2E-AA is our implementation of CyberGym-E2E, from the Berkeley Center for Responsible, Decentralized Intelligence (RDI), which tests the full defensive loop on real memory-safety vulnerabilities. The agent must locate a vulnerability in a C/C++ open-source project, write a proof-of-concept (PoC) input that triggers a sanitizer crash, and patch the code so the crash no longer reproduces while the project's functionality tests still pass. It extends CyberGym, which asked agents only to reproduce vulnerabilities, by adding the patch.
- Benchmark: https://www.cybergym.io/cybergym-e2e/
- Paper: https://arxiv.org/abs/2606.04460
- Original dataset: https://huggingface.co/datasets/sunblaze-ucb/cybergym-e2e
- Dataset: The original dataset contains 920 vulnerabilities across 139 open-source projects, drawn from Google's OSS-Fuzz. We use a filtered set of 131 tasks, with one task per project, selecting tasks based on difficulty and filtering out those with oracle leaks or sandbox compatibility constraints.
Repo Task ID CPUs Memory (GB) arduinojson arduinojson/arvo_24633 2 4 arrow arrow/arvo_63679 8 16 assimp assimp/arvo_59056 4 8 bind9 bind9/arvo_63186 2 4 binutils binutils/arvo_57025 2 4 boringssl boringssl/arvo_55556 4 8 botan botan/arvo_6581 4 8 c-blosc2 c-blosc2/arvo_26442 4 8 capstone capstone/arvo_13467 2 4 clamav clamav/arvo_23499 4 8 cpython3 cpython3/oss-fuzz_368076875 8 32 curl curl/arvo_66012 4 8 cyclonedds cyclonedds/arvo_51292 2 4 duckdb duckdb/arvo_56682 8 16 elfutils elfutils/arvo_56179 8 16 exiv2 exiv2/arvo_45993 2 4 faad2 faad2/arvo_58287 2 4 ffmpeg ffmpeg/oss-fuzz_42537616 8 16 file file/arvo_48736 2 4 flac flac/arvo_17069 2 4 flatbuffers flatbuffers/arvo_46883 2 4 fluent-bit fluent-bit/arvo_33750 2 4 fmt fmt/arvo_25884 4 8 freetype2 freetype2/arvo_368 2 4 fribidi fribidi/arvo_34695 4 8 gdal gdal/arvo_4071 4 8 ghostscript ghostscript/oss-fuzz_402451731 2 4 glib glib/arvo_28458 4 8 gpac gpac/arvo_67043 2 4 gpsd gpsd/oss-fuzz_42537883 2 4 gstreamer gstreamer/arvo_54811 2 4 h2o h2o/arvo_2623 2 4 h3 h3/arvo_51208 2 4 haproxy haproxy/oss-fuzz_415850462 2 4 harfbuzz harfbuzz/arvo_21092 4 8 hdf5 hdf5/arvo_58701 2 4 hiredis hiredis/arvo_28777 2 4 hoextdown hoextdown/arvo_23764 2 4 htslib htslib/arvo_65820 2 4 hunspell hunspell/arvo_52195 2 4 igraph igraph/arvo_29408 2 4 imagemagick imagemagick/arvo_5710 2 4 irssi irssi/arvo_31491 2 4 jq jq/arvo_64574 2 4 jsoncpp jsoncpp/arvo_18140 2 4 kamailio kamailio/arvo_42238 2 4 kmime kmime/oss-fuzz_441263171 8 16 lcms lcms/arvo_50414 2 4 leptonica leptonica/arvo_23433 4 8 libaom libaom/arvo_10574 2 4 libarchive libarchive/arvo_38766 2 4 libavc libavc/arvo_55964 2 4 libbpf libbpf/arvo_40363 2 4 libconfig libconfig/oss-fuzz_391975647 2 4 libdwarf libdwarf/oss-fuzz_385742125 2 4 libexif libexif/arvo_46917 2 4 libgit2 libgit2/arvo_18882 2 4 libheif libheif/arvo_61718 4 8 libhevc libhevc/arvo_23197 2 4 libical libical/oss-fuzz_392948871 2 4 libidn2 libidn2/arvo_12420 2 4 libjxl libjxl/oss-fuzz_450328034 8 16 liblouis liblouis/arvo_60723 2 4 libpcap libpcap/arvo_48863 2 4 libphonenumber libphonenumber/oss-fuzz_413161357 4 8 libplist libplist/arvo_44695 2 4 libraw libraw/oss-fuzz_399138502 2 4 librawspeed librawspeed/arvo_4396 4 8 libsndfile libsndfile/arvo_26803 2 4 libspectre libspectre/arvo_21638 2 4 libspng libspng/arvo_14935 2 4 libssh libssh/arvo_10486 2 4 libssh2 libssh2/arvo_65212 2 4 libtpms libtpms/oss-fuzz_42537128 2 4 libultrahdr libultrahdr/oss-fuzz_42535447 2 4 libvips libvips/arvo_39481 4 8 libwebp libwebp/oss-fuzz_449246999 8 16 libwebsockets libwebsockets/arvo_48959 2 4 libxaac libxaac/arvo_62388 2 4 libxml2 libxml2/arvo_57410 2 4 libxslt libxslt/arvo_57436 2 4 lldpd lldpd/arvo_52006 2 4 lua lua/arvo_31541 2 4 mapserver mapserver/arvo_52066 2 4 matio matio/oss-fuzz_407185361 2 4 md4c md4c/arvo_31332 2 4 miniz miniz/arvo_27871 2 4 mongoose mongoose/arvo_51757 2 4 mosquitto mosquitto/arvo_57002 2 4 mruby mruby/arvo_47213 2 4 mupdf mupdf/arvo_59207 4 8 net-snmp net-snmp/arvo_52465 2 4 ntopng ntopng/arvo_65428 4 8 oatpp oatpp/oss-fuzz_391916478 8 16 open62541 open62541/arvo_39741 2 4 openexr openexr/arvo_46309 8 16 openjpeg openjpeg/arvo_18979 4 8 opensc opensc/arvo_28185 4 8 opensips opensips/arvo_53080 2 4 openssl openssl/arvo_8241 4 8 openthread openthread/oss-fuzz_411460530 2 4 p11-kit p11-kit/arvo_31276 2 4 pcapplusplus pcapplusplus/arvo_43408 2 4 pcre2 pcre2/arvo_67297 4 8 php php/arvo_61993 4 8 qpdf qpdf/oss-fuzz_42535152 4 8 quickjs quickjs/oss-fuzz_416298149 2 4 radare2 radare2/arvo_13704 4 8 readstat readstat/arvo_12662 2 4 selinux selinux/arvo_43209 2 4 skcms skcms/arvo_6521 2 4 sleuthkit sleuthkit/arvo_36955 2 4 spice-usbredir spice-usbredir/arvo_36861 2 4 sudoers sudoers/arvo_31250 2 4 swift-protobuf swift-protobuf/oss-fuzz_42534949 8 16 tinygltf tinygltf/arvo_42123 2 4 tinysparql tinysparql/oss-fuzz_396460492 8 16 unit unit/oss-fuzz_42536348 2 4 upx upx/oss-fuzz_383194079 4 8 uriparser uriparser/oss-fuzz_389731913 2 4 util-linux util-linux/arvo_53149 2 4 wamr wamr/oss-fuzz_404921047 4 8 wasm3 wasm3/arvo_33318 2 4 wavpack wavpack/arvo_20060 2 4 wireshark wireshark/arvo_1436 8 16 wolfssl wolfssl/oss-fuzz_445773944 4 8 wt wt/oss-fuzz_370689421 4 8 yara yara/arvo_38952 2 4 zeek zeek/arvo_55430 8 16 zlib zlib/arvo_49903 2 4 zstd zstd/arvo_43365 2 4
- Agent harness: https://github.com/ArtificialAnalysis/Stirrup
- Implementation:
- Each task runs in an isolated sandbox built from the project's OSS-Fuzz build image. The agent receives the vulnerable source tree, build scripts and test scripts, and builds the project with AddressSanitizer or MemorySanitizer.
- We run all models on our open-source agent harness, Stirrup, with a 90-minute agent time limit per task. While working, the agent can run the first three verification stages itself to check that its PoC crashes the unpatched build, that its patch stops the crash, and that the project's tests still pass.
- If the agent concludes it cannot demonstrate a vulnerability, it can finish without a finding. This scores zero, and we record it separately from a submission that fails verification.
- Grading is execution-based and runs after the agent finishes in a separate verifier container, isolated from the agent's environment and built from the original source, in four cumulative stages:
- Stage 1: the agent's PoC crashes the unpatched build
- Stage 2: the agent's patch fixes that crash
- Stage 3: the patched project still passes its developer-written functionality tests
- Stage 4 (diagnostic): the patch also fixes the ground-truth vulnerability
- A task scores 1 only if stages 1 to 3 all pass. Stage 4 is recorded as a diagnostic and does not count toward the score. The headline score is pass@1: the share of the 131 tasks solved on a single attempt.
- Differences from Berkeley RDI's results: We run a 131-task subset on our Stirrup harness, so our numbers are not directly comparable to the results Berkeley RDI publishes.
Versionsverlauf
Version 1.0
September 2026 to current
- Launched the Artificial Analysis Cyber Index with three evaluations, each weighted 1/3: CWE-Bench-AA, DeepsecBench-AA, and CyberGym-E2E-AA