Booz Allen has released an AI benchmark they’re calling the Cyber Weapon Index (CWI).
Immediately, I hear Patrick Gray and Adam Boileau saying, “eeenternet veapons” in a bad Russian accent. IYKYK.
It isn’t the first cybersecurity AI benchmark out there, but I do appreciate that it’s fairly straightforward and not too complex. Real attacks are significantly simpler than most of the CTF challenges. Attackers aren’t into crime because of their desire to challenge their hacking skills - they’re doing it because it’s easy money. If they can make attacks simpler, or find a quicker way to get paid, they will.
Are attackers actually using AI though? We have some evidence that attackers are absolutely using AI where it works well. Improving phishing messages, cloning websites, and even carrying out some of the attack operations directly. I haven’t seen much evidence that attackers using AI to create zero-days is becoming the norm, however (let me know if I’m missing something!).
All the evidence we have suggests opportunistic attackers don’t really need zero-days, and we haven’t yet seen an increase in zero-day use. My best guess is that, since years old RCEs are keeping attackers well fed for now, there’s no need for them to start prompting new RCEs out of whatever models they have access to.
A good rule of thumb is that attackers are not going to do extra work unless forced to, but if AI gets good enough to do the entire thing? I’d think attackers would be all over that, and it’s a good assumption that cybercrime adoption of AI will be swift if it’s actually useful.
The CWI
The CWI is actually comprised of two benchmarks:
The Vulnerability Research Score (VRS) is a test of a model’s ability to find vulnerabilities in binaries, without access to source code. There are two VRS tests
Easy - vulnerabilities are planted in a relatively small amount of code
Hard - no vulnerabilities have been planted and the code base is much larger. Models have to find vulns that Booz Allen don’t even know exist in the code.
The Kill Chain Attainment Score (KCAS) is a test where a model has to progress through more of a traditional network-based penetration test scenario and is scored on how far it makes it through the cyber kill chain (which Booz Allen also occasionally calls the “cyber attack lifecycle)
The latter assessment is built around a windows-centric environment, with obtaining domain admin (DA) as the ultimate objective. I could complain that this isn’t realistic, that attackers don’t care about DA, but I think it’s fine. Based on the evidence we have, most companies that get targeted and compromised are flat-network Microsoft shops, so this design covers a large portion of the attack targets out there.
The ultimate goal of the CWI is to track the status of AI model capabilities towards being able to achieve full compromise without a human in the loop. Point an AI model at a target, tell it to hack it, and reliably get a positive result.
The first release of the CWI only saw one model achieve the objective in the KCAS benchmark - Mythos. Less than a week later, there was a second, which underscores why they’re doing this work - to track how quickly these models are improving.
And the second test run:
I found some of the charts confusing, as the text mentioned that only one model in the first test fully passed KCAS, and two in the second did, but I’m seeing five models listed as achieving the objective. What gives?
A second chart type clarifies the situation a bit. It looks like, to get full marks, a model has to reliably achieve DA.
I’m interpreting this to mean that the latest test includes two models (Mythos and GPT-6 Astra) that get full marks and three models occasionally complete the entire benchmark. Notably, one of those models is GLM-5.2 - a Chinese open-weight model.
Conclusion and thoughts
The results aren’t surprising - we’ve been hearing for weeks about how models have been breaking out of sandboxes, creating zero-day exploits, and hacking companies without the knowledge of their owners. The authors of the report estimate that all models will be able to ace the KCAS and VRS tests six months from now.
What’s missing are the costs. It’s one thing for something to be technically possible. It’s quite another for it to be practical and affordable. I can’t imagine Booz Allen would be willing to spend millions to perform these tests, but even five figures might be enough to discourage attackers from using these capabilities. At least, until open weight models can reliably complete an attack autonomously.
It wasn’t entirely clear from the methodology, but I also assume there are no security controls or active defenders trying to stop these attacks. This would be aligned with other cybersecurity benchmarks, like the infamous CyberGym, but isn’t aligned with reality, which is good to keep in mind.
In other words, just because a model can score well on a benchmark, doesn’t mean it can hack any company you point it at. The real world is much too complex for a benchmark to replicate. If you really want to know how close models are getting to putting a penetration tester out of a job, go ask an active penetration tester.






