All articles

September 25, 2026

Announcing the Artificial Analysis Cyber Index Alliance: toward better benchmarking of agentic cyber defense

The Artificial Analysis Cyber Index Alliance brings together industry partners to set a new standard for evaluating how AI models perform on enterprise cyber defense tasks. The Alliance launches alongside the Artificial Analysis Cyber Index, which combines three partner-contributed and open benchmarks to evaluate how well agents find and fix vulnerabilities.

This is the original launch article and shows Artificial Analysis Cyber Index results as of launch. For live results across the latest models, see the Artificial Analysis Cyber Index leaderboard.

Introducing the Artificial Analysis Cyber Index

Every organization that relies on software carries vulnerabilities, and finding and fixing them before an attacker does is the constant work of a security team. This affects enterprises, governments, NGOs, and downstream users every day.

As the cyber capabilities of AI models increase, attackers can automate more of their work and run it at greater scale. Security teams need capable, cost-effective tools that can detect and patch vulnerabilities before they can be exploited.

To meet this challenge, we are releasing the Artificial Analysis Cyber Index to support decision-makers choosing models for security work. It is a standalone index of model capability on cyber defense, and combines evaluation datasets from industry partners and academia to measure how well a model can support and accelerate the work of an internal security team.

The Cyber Index tests the defensive loop: discovering vulnerabilities in a codebase, reproducing and validating them, and patching them without breaking existing functionality. Models work from the source code, the way a security engineer auditing an application would, and we do not ask any model to build a working exploit. The result is an independent, like-for-like comparison of which models best support this work and what they cost to run.

Artificial Analysis Cyber Index

Artificial Analysis Cyber Index v1 incorporates 3 evaluations: CWE-Bench-AA, DeepsecBench-AA, and CyberGym-E2E-AA
Successes
Safety blocks

The Cyber Index Alliance

The Artificial Analysis Cyber Index launches in partnership with the Cyber Index Alliance, which brings together industry partners to set a new standard for evaluating how AI models perform on enterprise cyber defense tasks. Partners contribute expert input on the design and implementation of the Index, and may contribute datasets and external research directly.

Launch partners

  • Collinear AI

    Developer of CWE-bench, contributed as a private held-out evaluation.

  • IBM

    Expert input on scope and methodology.

  • NVIDIA

    Expert input on scope and methodology.

  • Vercel

    Developer of DeepsecBench, contributed as a private held-out evaluation.

Organizations interested in joining the Cyber Index Alliance can contact us at cyber@artificialanalysis.ai.

How the Artificial Analysis Cyber Index works

Benchmark overview

At launch, the Artificial Analysis Cyber Index combines three cybersecurity evaluations from industry partners and academic research. Between them they cover the full breadth of the defensive loop, from scanning a codebase for weaknesses to reproducing a crash and implementing a patch.

EvaluationWhat it measuresCapabilitiesSub-capabilities
CWE-Bench-AA
Collinear AI
Auditing a real open-source repository for a vulnerability in a described area of concern, then patching it without breaking legitimate behavior. 120 held-out tasks covering all ten OWASP Top 10 (2025) categories.Identifying and remediating vulnerabilities
  • Discovering and scanning for vulnerabilities
  • Implementing a patch or mitigation
DeepsecBench-AA
Vercel
Finding vulnerabilities in open-source application code, scored against a golden set of expert-verified findings.Identifying vulnerabilities
  • Discovering and scanning for vulnerabilities
CyberGym-E2E-AA
Berkeley RDI
Discovering, reproducing, and patching memory-safety vulnerabilities in C/C++ open-source projects. 131 tasks, one per project, drawn from the 920-instance CyberGym-E2E dataset.Identifying and remediating vulnerabilities
  • Discovering and scanning for vulnerabilities
  • Reproducing and validating vulnerabilities
  • Implementing a patch or mitigation

We tag each evaluation with the capabilities and sub-capabilities it tests, as in the table above, so we can see which parts of cyber defense the Index covers and where it has gaps. It covers identifying and remediating vulnerabilities with access to the source code. We plan to add incident response, writing new code without introducing vulnerabilities, and targets without source access, such as compiled software and live servers. Exploit realization, turning a found vulnerability into a working exploit, is out of scope for a defense-focused index.

Methodology

All three evaluations run on Stirrup, our open-source agent harness.

Because cyber work is dual-use, we also track cases where a model or provider declines a task on safety grounds, and report them separately from the score. Given the open-ended nature of the tasks, the harness in CyberGym-E2E-AA allows models to end a task without a finding if they conclude they cannot find or demonstrate a vulnerability.

Full details are on the methodology page.

CWE-Bench-AA: audit and patch real repositories

CWE-Bench-AA is our implementation of Collinear AI's CWE-bench. Each task is a security audit. The agent receives a checkout of an open-source repository and an instruction to audit the code and fix what it finds. Collinear AI includes tasks that reproduce disclosed CVEs (Common Vulnerabilities and Exposures, the public register of known security flaws). The task describes the area of concern but not the exact location.

The test set contains 120 held-out tasks, private to Collinear AI and Artificial Analysis, covering all ten OWASP Top 10 (2025) categories across C/C++, Go, Java, JavaScript/TypeScript, Python, and Rust. Each task tests both finding the weakness and patching it without breaking the rest of the codebase.

CWE-Bench-AA

Share of tasks solved, failed, and declined on safety grounds · Vulnerability patching success (pass@1) · Independently benchmarked by Artificial Analysis
Successes
Safety blocks

How models fail

CWE-Bench-AA failures concentrate in two areas: partial fixes, which leave the vulnerability exposed, and over-corrections, where the patch breaks legitimate behavior. On average, models spent 38% of their turns searching for the bug before making their first edit, and the remaining 62% patching and validating the fix.

Partial fixes are the main failure mode: excluding refusals and timeouts, 55% of failed attempts fixed the primary issue but left a related one open, such as a second entry point. Over-corrections account for ~24% of failed attempts and are most common among the strongest models, making up ~40% of failures for the four highest-scoring models compared to ~15% for the lowest performers, with the breakage usually a close edge case of legitimate functionality.

DeepsecBench-AA: find the vulnerabilities expert reviewers confirmed

DeepsecBench-AA, our implementation of Vercel's DeepsecBench, isolates the discovery process. The agent is asked to review scanner-flagged files in open-source application code and report every vulnerability it can confirm. We score its findings against a golden set of vulnerabilities verified by human experts.

Security teams face the same open-ended problem: deciding where to look in a large codebase that keeps changing. Agents often report findings outside the golden set, some of them real, and a security team has to triage every one. Scoring against the golden set rewards models that surface real findings without burying them in false positives.

DeepsecBench-AA

Share of tasks solved, failed, and declined on safety grounds · F2 score (recall weighted above precision) · Independently benchmarked by Artificial Analysis
Successes
Safety blocks

How models fail

Given the open-ended nature of DeepsecBench-AA, it is unsurprising that not all vulnerabilities are found, with the best model identifying only 41% of the expert-verified issues. However, there is a clear pattern in the types of bugs models do and do not report. Models most often find flaws with a direct path from untrusted input to consequence.

In contrast, bugs that require reasoning through a sequence of events or considering business and privacy rules are rarely found. When models do report these sequence-of-events vulnerabilities, they are almost always right (95% of reports are correct). GPT-6 Sol and GPT-6 Astra find them in ~30% of runs, roughly twice the rate of the next best models and three to four times the typical model, suggesting this may be an emerging capability.

CyberGym-E2E-AA: discover, reproduce, and patch memory-safety bugs

CyberGym-E2E, from the Berkeley Center for Responsible, Decentralized Intelligence, evaluates models on the end-to-end process of discovering vulnerabilities, reproducing them, and writing patches that pass the project's existing tests. The benchmark focuses on memory-safety bugs.

Each task targets a real memory-safety vulnerability in a widely used C/C++ open-source project, such as FFmpeg or CPython. The model has to locate the bug, write a proof-of-concept input to trigger the crash, and patch it so the crash no longer reproduces. Our implementation, CyberGym-E2E-AA, uses a filtered set of 131 tasks, with one task per project. Ending a task without a finding scores zero, and we record it separately from a failed submission.

CyberGym-E2E-AA

Share of tasks solved, failed, and declined on safety grounds · Vulnerability discovery and patching success (pass@1) · Independently benchmarked by Artificial Analysis
Successes
Safety blocks

How models fail

Refusals are a significant factor on CyberGym-E2E-AA: GPT-6 Astra, GPT-6 Sol, Claude Fable 5.1, Claude Opus 5.5, Qwen3.8 2.4T A95B, and Qwen3.8 27B refuse at least 98% of tasks, making it difficult to assess frontier model performance.

Of the remaining models, failures follow three patterns. Discovery is the main challenge, with 42% of attempts reaching the 90-minute limit without an input that crashes the program. This is a key consideration for enterprise users deciding the scope of their codebase to scan.

As in DeepsecBench-AA, models struggle most with bugs that depend on how state changes over several steps. Models pass 50% of attempts on out-of-bounds bugs, against 33% on use-after-free bugs, which depend on an object's lifetime across several operations, and 20% on integer and arithmetic bugs.

Finally, passing attempts often fix the wrong bug: 31% of passes patch a real crash other than the target. These tend to be shallower issues, such as null-pointer crashes, making up 22% of off-target passes, against 9% of on-target ones, because models stop at the first crash they can validate and finish within a couple of turns.

Artificial Analysis Cyber Index resources

Our development roadmap

Security teams are already choosing models to support their defensive work, so we are launching the Artificial Analysis Cyber Index and the Cyber Index Alliance now, with three evaluations, rather than waiting to cover every capability.

We will review the Index as models improve and add evaluations as new capabilities need measuring, starting with the gaps listed above.