Cibersegurança defensiva ganha índice próprio e números decepcionam
Defensive cybersecurity now has a dedicated benchmark with the launch of the Cyber Index, from Artificial Analysis. The announcement came out on September 28.

Defensive cybersecurity now has a dedicated benchmark with the launch of the Cyber Index, from Artificial Analysis. The announcement came out on September 28. The initiative is organized by the Cyber Index Alliance, with IBM, NVIDIA, Collinear AI, and Vercel as launch partners.
The test measures performance on production software repositories. In other words, the environment evaluates whether agents can audit raw code, isolate confirmed vulnerabilities, and write functional patches.
There's an important scope boundary. However, building offensive exploits is left out, since turning a vulnerability into a functional attack falls outside the defensive focus.
In addition, all evaluations run through Stirrup, an open-source platform created to standardize the tests.
The overall ranking shows a low ceiling
The composite index numbers stand out for their level. Grok 4.7 (xhigh) and MiMo V2.6 Pro lead with 56% each.
Next comes GPT 6 Luna, with 53%. However, GLM 5.3 Flash closed out the front group with 50%.
Still, note the meaning of this ceiling. Even the best result leaves nearly half of the tasks unsolved.
In CWE Bench AA, a private set from Collinear AI with 120 tasks, the picture improves a bit. Grok 4.7 and DeepSeek V4.1 Flash reached 68% success.
This set covers C, C++, Go, Java, JavaScript, TypeScript, Python, and Rust. In addition, models receive only the general area of concern.
Cybersecurity fails when the patch only half-fixes the issue
Here's the most useful data point for code reviewers. Partial remediations caused 55% of all failures that weren't refusals.
In these cases, the model fixed the main vulnerability. However, it left secondary attack paths open.
There's also a worrying side effect. The four top-performing systems broke existing functionality in about 40% of failed attempts.
Compare that with the worst-performing group, which stood at 15%. So a more capable model also produces more regression when it errs.
The engineers also measured effort distribution. On average, models spent 38% of steps searching for the bug before the first edit. The remaining 62% went to writing the patch and running tests.
Chained logic flaws still slip past cybersecurity
DeepsecBench AA, from Vercel, tests review of flagged files against expert-verified flaws. It uses the F2 score as its metric.
The result calls for caution. Even the best model found only 41% of the verified defects.
The accuracy pattern says a lot. Simple flaws, which link untrusted input directly to the execution point, were identified with ease.
Logic flaws that unfold over the course of execution, however, proved harder. GPT 6 Sol and GPT 6 Astra detected these cases in about 30% of runs.
Still, there was a bright side. When they did report a finding, these models were correct 95% of the time.
Memory-safety cybersecurity runs into safety refusals
CyberGym E2E AA, from Berkeley RDI, evaluates memory-safety fixes across 131 C and C++ projects. In this set, the obstacle came from somewhere else.
Frontier models refused at least 98% of the tasks on safety grounds. The list includes GPT 6 Astra, GPT 6 Sol, Claude Fable 5.1, and Claude Opus 5.5.
Among those that did execute, MiMo V2.6 Pro reached 79% success. GPT 6 Luna stood at 78% and Grok 4.7 at 74%.
Time also weighed heavily. About 42% of runs hit the 90-minute limit without producing a reproducible failure.
Results varied by bug type. Out-of-bounds access had 50% success, use-after-free dropped to 33%, and arithmetic flaws stood at 20%.
There's also a detail that dilutes the numbers. About 31% of approved solutions fixed unrelated issues, such as shallow null-pointer bugs.
What to do with this data in your workflow
First, treat a generated patch as a hypothesis. The partial-remediation rate shows that human review remains essential.
Second, always test for regressions. A 40% breakage rate among the best models justifies a full test suite before merging.
Third, combine tools. Static analysis still catches what the agent misses in long chains.
Also, plan for execution time. Tasks that exceed 90 minutes require a decision on cost and limits.
Finally, factor refusals into your design. Security workflows can stall on frontier models, and routing needs to account for that.
The alliance plans to expand its evaluations. Compiled software, real-time network targets, and incident-response workflows are on the roadmap.
Follow our profile on Instagram!
Translated from the Brazilian Portuguese original · Read the original
Anthropic launches Sonnet 5.5, and the leap in Terminal Bench impresses
Anthropic launched Claude Sonnet 5.5 on September 28. This is the second model in the Claude 5.5 family. The proposal addresses well-defined tasks.





