tl;dr: Since CWE-bench v0, a new generation of models has arrived, and frontier labs are investing heavily in cyber capability. CWE-bench tracks that progress. We have run more than 13,000 agent rollouts across frontier and open models, and they show us where models fall short and where a verifier can be fooled. v1 puts that to use in three changes:
Stronger verifiers. Where our programmatic verifiers disagreed with an agentic judge panel, we read the rollouts and fixed the verifier. Some passed partial fixes; others failed valid ones.
A tighter sandbox. Agents now run with no internet access, so they can’t look up the upstream fix. One model still found a way out, using its own API key.
A bigger suite. 120 held-out audit-and-patch tasks, covering 73 CWEs, all ten OWASP Top 10 2025 categories and eight languages.
Most of the attention on AI in cybersecurity has gone to offense: finding vulnerabilities and writing exploits. Defense is the harder test; an attacker needs one way in while a defender needs to close every way in. As agents take on more real security work, they need to be at least as capable at defense as at offense. CWE-bench measures that side, and we’ll keep extending it as the frontier moves.
Leaderboard

Grok 4.7 and GPT-6 Astra tie for first at 68%, with Claude Opus 5.5 a point behind at 67%. With four attempts, Grok 4.7 pulls ahead, solving 81% of tasks.
Claude Opus 5.5 matches the leaders at about a quarter of the cost ($0.79 per rollout vs about $2.80). DeepSeek V4.1 Flash is the budget pick, at 55% for $0.09.
CWE-bench is also part of the Artificial Analysis Cyber Index, where Artificial Analysis runs it independently as CWE-Bench-AA. Scores are not directly comparable due to methodology differences with differing harnesses and reasoning modes.
1. Constructing an agentic judge panel
A deterministic programmatic verifier should check that (1) the exploit no longer works and (2) the project’s existing tests still pass. Those checks can still miss a bad fix; one that closes the probed code path but leaves an identical one open, blocks the exploit by disabling the feature, or breaks the code the tests don’t cover can still fool the verifier.
Since v0 we have run more than 13,000 rollouts. To find where the verifier was wrong, we graded them with an agentic judge panel with three judges from different model families (GPT-5.4, Claude Sonnet 5 and Gemini 3.8 Flash). Each judge works in its own copy of the patched repository and can build the project, run the programmatic verifier and explore the code. Each judge checks five things and passes a fix only if all five hold; the panel passes it by majority vote:
Security: the vulnerability is closed
Happy path: normal behavior still works
Generalization: variants of the attack are closed too, not just the one we probe
Integrity: the agent fixed the code rather than gaming the grader
Soundness: the patched codebase still builds and runs
We focused our effort on the programmatic verifiers, because they give a deterministic, reproducible score. Wherever the panel and a verifier disagreed, we read the rollout. Some verifiers were too loose and passed partial fixes; we added probes so those now fail. Others were too strict and failed valid fixes, often because they assumed the shape of our reference fix; we rewrote them to test behavior instead of structure. The remaining disagreements usually came from the judges, so we tightened the rubrics until the two graders agreed.
The programmatic verifier is the authoritative score, and the leaderboard ranks by it. We also report the panel’s aggregate score as a second, independent read. The two track each other closely: across the nine models, they are within about two points on average.
2. A sandbox that holds
CWE-bench builds upon real, public vulnerabilities, so the upstream fix is often a search away. An agent with internet access can find the project’s commit history, download the published patch and apply it. That scores as a fix without any work on the code in front of it.
CWE-bench ensures that the agent phase runs with no network access. The only host an agent can reach is its own model API, which it needs in order to think; setup and grading still have network access. To catch anything that slips through, we screened every leaderboard rollout for actual internet access (data coming back, not just attempts) and removed any that had it.
That check caught one model going around the sandbox. On two tasks, Grok 4.7 needed package checksums that weren’t in the environment, and its downloads failed. It listed its environment variables and found the OpenRouter API key its harness uses. It tested dozens of hosts, found that OpenRouter was the only one reachable, and used the key to ask web-enabled models to fetch the pages for it. This happened in only 4 out of nearly 1000 rollouts. None of the four solved its task, and all four have since been rerun in a stronger sandbox.
The agent wasn’t trying to cause harm. It was doing what it was trained to do: make progress on the task by whatever path the environment allows. Agents are trained in environments too, and when those sandboxes leak, the shortcut gets rewarded. Between the judge panel flagging integrity failures and our internet-access checks, we’ve caught and patched these cases of reward hacking in our environments.
3. Refusals and fallbacks
Some model providers run safety filters that flag defensive security work, and the harness decides what happens next. With GPT-6, OpenAI’s API can flag a request partway through a task; Codex stops, and the attempt ends. With Claude, Anthropic’s API refuses the request, and Claude Code retries it with a fallback model (Claude Opus 5 or Claude Opus 4.8), which finishes the task in the same session.

OpenAI’s filter flagged GPT-6 Astra on 13 of the 120 tasks and GPT-6 Sol on 5, and 5 of Astra’s rollouts ended blocked. Anthropic’s filter flagged Claude Fable 5.1 on 39 tasks and Claude Opus 5.5 on 9. Claude Code handed those rollouts to a fallback model, which did about a quarter of Fable’s turns, and most of them still passed: 62% for Fable and 72% for Opus 5.5. Grok, DeepSeek, Hy4 and Muse Spark were never flagged.
Rollouts that end blocked are graded on the work they’ve done, and rollouts finished by a fallback model count for the model they started with, because that is what someone querying that model gets. Refusing offensive work, like writing a working exploit against someone else’s system, makes sense because it can be abused. Refusing to patch a vulnerability protects no one. Defensive work should never be blocked, and a filter that can’t tell the two apart ends up blocking defenders.
4. Where models differ
Consistency. A model’s four attempts at a task don’t always agree, and how often they agree varies a lot between models. GPT-6 Astra is close to all-or-nothing: of 120 tasks, it solves 74 on every attempt, 31 on none, and only 15 on some. So four attempts barely help it; pass@4 is only 6 points above its pass@1. Hy4 Preview and DeepSeek V4.1 Flash are the opposite. About 40% of their tasks are solved on some attempts but not others, and four attempts add around 20 points. A consistent model is predictable. A variable one gains more from retries, and its partial successes point to what it almost knows how to do.

The top three solve different tasks. Grok 4.7, GPT-6 Astra and Claude Opus 5.5 are within a point of each other on pass@1, but they don’t solve the same problems. Grok 4.7 solves 20 tasks that Astra never solves, and Astra solves 12 that Grok 4.7 never does. Claude Opus 5.5 solves 12 tasks that Astra never does, and Astra solves 6 that Opus 5.5 never does. All three solve 75 tasks in common, and together they solve 110, well above any one of them alone (Grok 4.7’s 97 is the most).

The remaining gap. Ten tasks are solved by none of the top three, and nine are solved by none of the three. One of the ten is solved only by DeepSeek V4.1 Flash and Hy4 Preview. The unsolved nine span six languages, and three of them are Software or Data Integrity Failures (OWASP A08).
What comes next
An attacker needs one unguarded path; a defender needs to cover all of them. A benchmark for defenders has to hold itself to the same standard: every shortcut an agent can take is a path the benchmark has to close. We’ll keep running models at scale, keep reading the rollouts where our graders disagree, and keep growing the suite toward the full CWE catalog.
If you are training or evaluating a model on defensive cyber capabilities and want to run it on CWE-bench, contact Collinear AI’s cyber team at cyber@collinear.ai.
Citation
@misc{cwebench2026v1,
title = {CWE-bench v1: Defense Is the Harder Test},
author = {{Collinear AI}},
year = {2026},
howpublished = {\url{https://cwe-bench.com}},
note = {120 held-out audit-and-patch tasks across 73 CWEs.}
}




