Announcing the Artificial Analysis Cyber Index Alliance: toward better benchmarking of agentic cyber defense
The Artificial Analysis Cyber Index Alliance brings together industry partners to set a new standard for evaluating how AI models perform on enterprise cyber defense tasks. The Alliance launches alongside the Artificial Analysis Cyber Index, which combines three partner-contributed and open benchmarks to evaluate how well agents find and fix vulnerabilities.
Launch partners
About the AllianceIntroducing the Artificial Analysis Cyber Index
Every organization that relies on software carries vulnerabilities, and finding and fixing them before an attacker does is the constant work of a security team. This affects enterprises, governments, NGOs, and downstream users every day.
As the cyber capabilities of AI models increase, attackers can automate more of their work and run it at greater scale. Security teams need capable, cost-effective tools that can detect and patch vulnerabilities before they can be exploited.
To meet this challenge, we are releasing the Artificial Analysis Cyber Index to support decision-makers choosing models for security work. It is a standalone index of model capability on cyber defense, and combines evaluation datasets from industry partners and academia to measure how well a model can support and accelerate the work of an internal security team.
The Cyber Index tests the defensive loop: discovering vulnerabilities in a codebase, reproducing and validating them, and patching them without breaking existing functionality. Models work from the source code, the way a security engineer auditing an application would, and we do not ask any model to build a working exploit. The result is an independent, like-for-like comparison of which models best support this work and what they cost to run.
Artificial Analysis Cyber Index
The Cyber Index Alliance
The Artificial Analysis Cyber Index launches in partnership with the Cyber Index Alliance, which brings together industry partners to set a new standard for evaluating how AI models perform on enterprise cyber defense tasks. Partners contribute expert input on the design and implementation of the Index, and may contribute datasets and external research directly.
Launch partners
- Collinear AI
Developer of CWE-bench, contributed as a private held-out evaluation.
- IBM
Expert input on scope and methodology.
- NVIDIA
Expert input on scope and methodology.
- Vercel
Developer of DeepsecBench, contributed as a private held-out evaluation.
Organizations interested in joining the Cyber Index Alliance can contact us at cyber@artificialanalysis.ai.
How the Artificial Analysis Cyber Index works
Benchmark overview
At launch, the Artificial Analysis Cyber Index combines three cybersecurity evaluations from industry partners and academic research. Between them they cover the full breadth of the defensive loop, from scanning a codebase for weaknesses to reproducing a crash and implementing a patch.
| Evaluation | What it measures | Capabilities | Sub-capabilities |
|---|---|---|---|
CWE-Bench-AA Collinear AI | Auditing a real open-source repository for a vulnerability in a described area of concern, then patching it without breaking legitimate behavior. 120 held-out tasks covering all ten OWASP Top 10 (2025) categories. | Identifying and remediating vulnerabilities |
|
DeepsecBench-AA Vercel | Finding vulnerabilities in open-source application code, scored against a golden set of expert-verified findings. | Identifying vulnerabilities |
|
CyberGym-E2E-AA Berkeley RDI | Discovering, reproducing, and patching memory-safety vulnerabilities in C/C++ open-source projects. 131 tasks, one per project, drawn from the 920-instance CyberGym-E2E dataset. | Identifying and remediating vulnerabilities |
|
We tag each evaluation with the capabilities and sub-capabilities it tests, as in the table above, so we can see which parts of cyber defense the Index covers and where it has gaps. It covers identifying and remediating vulnerabilities with access to the source code. We plan to add incident response, writing new code without introducing vulnerabilities, and targets without source access, such as compiled software and live servers. Exploit realization, turning a found vulnerability into a working exploit, is out of scope for a defense-focused index.
Methodology
All three evaluations run on Stirrup, our open-source agent harness.
Because cyber work is dual-use, we also track cases where a model or provider declines a task on safety grounds, and report them separately from the score. Given the open-ended nature of the tasks, the harness in CyberGym-E2E-AA allows models to end a task without a finding if they conclude they cannot find or demonstrate a vulnerability.
Full details are on the methodology page.
CWE-Bench-AA: audit and patch real repositories
CWE-Bench-AA is our implementation of Collinear AI's CWE-bench. Each task is a security audit. The agent receives a checkout of an open-source repository and an instruction to audit the code and fix what it finds. Collinear AI includes tasks that reproduce disclosed CVEs (Common Vulnerabilities and Exposures, the public register of known security flaws). The task describes the area of concern but not the exact location.
The test set contains 120 held-out tasks, private to Collinear AI and Artificial Analysis, covering all ten OWASP Top 10 (2025) categories across C/C++, Go, Java, JavaScript/TypeScript, Python, and Rust. Each task tests both finding the weakness and patching it without breaking the rest of the codebase.
CWE-Bench-AA
How models fail
CWE-Bench-AA failures concentrate in two areas: partial fixes, which leave the vulnerability exposed, and over-corrections, where the patch breaks legitimate behavior. On average, models spent 38% of their turns searching for the bug before making their first edit, and the remaining 62% patching and validating the fix.
Partial fixes are the main failure mode: excluding refusals and timeouts, 55% of failed attempts fixed the primary issue but left a related one open, such as a second entry point. Over-corrections account for ~24% of failed attempts and are most common among the strongest models, making up ~40% of failures for the four highest-scoring models compared to ~15% for the lowest performers, with the breakage usually a close edge case of legitimate functionality.
DeepsecBench-AA: find the vulnerabilities expert reviewers confirmed
DeepsecBench-AA, our implementation of Vercel's DeepsecBench, isolates the discovery process. The agent is asked to review scanner-flagged files in open-source application code and report every vulnerability it can confirm. We score its findings against a golden set of vulnerabilities verified by human experts.
Security teams face the same open-ended problem: deciding where to look in a large codebase that keeps changing. Agents often report findings outside the golden set, some of them real, and a security team has to triage every one. Scoring against the golden set rewards models that surface real findings without burying them in false positives.
DeepsecBench-AA
How models fail
Given the open-ended nature of DeepsecBench-AA, it is unsurprising that not all vulnerabilities are found, with the best model identifying only 41% of the expert-verified issues. However, there is a clear pattern in the types of bugs models do and do not report. Models most often find flaws with a direct path from untrusted input to consequence.
In contrast, bugs that require reasoning through a sequence of events or considering business and privacy rules are rarely found. When models do report these sequence-of-events vulnerabilities, they are almost always right (95% of reports are correct). GPT-6 Sol and GPT-6 Astra find them in ~30% of runs, roughly twice the rate of the next best models and three to four times the typical model, suggesting this may be an emerging capability.
CyberGym-E2E-AA: discover, reproduce, and patch memory-safety bugs
CyberGym-E2E, from the Berkeley Center for Responsible, Decentralized Intelligence, evaluates models on the end-to-end process of discovering vulnerabilities, reproducing them, and writing patches that pass the project's existing tests. The benchmark focuses on memory-safety bugs.
Each task targets a real memory-safety vulnerability in a widely used C/C++ open-source project, such as FFmpeg or CPython. The model has to locate the bug, write a proof-of-concept input to trigger the crash, and patch it so the crash no longer reproduces. Our implementation, CyberGym-E2E-AA, uses a filtered set of 131 tasks, with one task per project. Ending a task without a finding scores zero, and we record it separately from a failed submission.
CyberGym-E2E-AA
How models fail
Refusals are a significant factor on CyberGym-E2E-AA: GPT-6 Astra, GPT-6 Sol, Claude Fable 5.1, Claude Opus 5.5, Qwen3.8 2.4T A95B, and Qwen3.8 27B refuse at least 98% of tasks, making it difficult to assess frontier model performance.
Of the remaining models, failures follow three patterns. Discovery is the main challenge, with 42% of attempts reaching the 90-minute limit without an input that crashes the program. This is a key consideration for enterprise users deciding the scope of their codebase to scan.
As in DeepsecBench-AA, models struggle most with bugs that depend on how state changes over several steps. Models pass 50% of attempts on out-of-bounds bugs, against 33% on use-after-free bugs, which depend on an object's lifetime across several operations, and 20% on integer and arithmetic bugs.
Finally, passing attempts often fix the wrong bug: 31% of passes patch a real crash other than the target. These tend to be shallower issues, such as null-pointer crashes, making up 22% of off-target passes, against 9% of on-target ones, because models stop at the first crash they can validate and finish within a couple of turns.
Artificial Analysis Cyber Index resources
- Live model scores and refinements are hosted on the Artificial Analysis Cyber Index page
- The methodology page details the scope, grading, and implementation of each benchmark
- Each of the three benchmarks has its own leaderboard page: CWE-Bench-AA, DeepsecBench-AA, and CyberGym-E2E-AA
- Stirrup, the open-source agent harness used for every Artificial Analysis Cyber Index task, is on GitHub
Our development roadmap
Security teams are already choosing models to support their defensive work, so we are launching the Artificial Analysis Cyber Index and the Cyber Index Alliance now, with three evaluations, rather than waiting to cover every capability.
We will review the Index as models improve and add evaluations as new capabilities need measuring, starting with the gaps listed above.
Read the latest

GPT-6 Sol and Luna push the cost efficiency frontier
GPT-6 Sol and Luna push the cost efficiency frontier by halving cost relative to GPT-5.6 Sol and Luna. Intelligence Index and Coding Agent Index scores remain level with GPT-5.6, with progress in some evaluations and regressions in others
September 22, 2026

Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index
Anthropic's new Opus scores 58 and arrives with a 20% price cut and a larger cache hit discount
September 22, 2026

Benchmarking Grok 4.7
Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index to bring SpaceXAI into the top 4 AI labs. Coding Agent Index performance has also improved, overtaking GPT-5.6 Sol
September 21, 2026