TL;DR — Key Takeaways
– AWS created the Deception Benchmark to test whether AI models can distinguish exploitable vulnerabilities from code that only appears risky.
– The benchmark includes 14,822 code samples across 16 programming languages and more than 70 CWE categories.
– Tested models identified many real vulnerabilities but also produced high false-positive rates, flagging 41% to 99% of safe code.
Amazon Web Services (AWS) has developed a benchmark that can be used to test whether a model can distinguish real vulnerabilities from code that looks risky but is actually safe.
The Deception Benchmark was created following an evaluation of the capabilities of 12 models from five different providers. In all, the benchmark includes 14,822 samples of code built using 16 different languages spanning more than 70 Common Weakness Enumeration (CWE) categories. Each sample is run through an adversarial loop to generate code, test it against frontier models, harden, repeat. If a model gets it right easily, the sample is removed.
According to the benchmark, every AI model has the same fundamental issue. While they identify up to 95% of real vulnerabilities, they also flag 41 to 99% of safe code. Proof-of-exploit prompting can cut false positives by 17 to 74 percentage points but misses 7% to 44% of real vulnerabilities. The environment-gated challenges are worse: models flag the code and ignore the Kubernetes Network Policy next to it. No tested configuration keeps both false positives and false negatives below 10%, according to the benchmark.

To provide those assessments, the Deception Benchmark creates two distinct types of challenges for AI models. Code-level challenges present vulnerable and safe variants that differ by a subtle fix. Both look suspicious, but only one is exploitable. Environment-gated challenges go a step further by using that same code to test specific IT environments to determine, for example, if a Kubernetes network policy blocks a server-side request forgery (SSRF) path. That foundation, for example, then makes it possible for CyberGym agents to test on more than 1,500 realistic tasks while another AI agent invokes the CYBENCH framework to evaluate capture the flag (CTF) challenges.
Additionally, the benchmark is designed to treat labeling as a convergent audit loop rather than a one-time step. Every label is re-examined independently by multiple reviewers who do not see one another’s assessments or the original reasoning behind the label. Disagreements escalate to direct adjudication, where the original reasoning is evaluated against the challenge. Unresolved cases are then passed on for human review.
Neha Rungta, director of applied science at AWS, said ultimately the goal is to identify the AI models that generate the fewest number of false positives when used to scan for vulnerabilities. That issue is becoming problematic because as AI models are used to scan for vulnerabilities, many DevSecOps teams are now wasting more time than ever investigating vulnerabilities that turn out to be false alarms.
Preventing those false positives requires AI models that reason not just about the existing code, but also how a remediation might impact the surrounding environment, noted Rungta. The goal is to provide the context needed to enable AI models to better understand what good should actually look like when fixing a vulnerability, she added.
Mitch Ashley, vice president and practice lead for software lifecycle engineering at the Futurum Group, said false positives are verification debt that ultimately undermines confidence in AI findings. AI scanners that read code without the environment around it cannot tell an exploitable path from a blocked one, he added. A model that flags safe code as vulnerable is only making more work for DevSecOps teams, noted Ashley.
Just how much work AI models are creating for DevSecOps teams is unclear, but other reports suggest that the more complex the environment, the less helpful AI becomes. Hopefully, as AI continues to evolve, these tools and platforms will not just be used to discover vulnerabilities but also fix them in a way that requires as little intervention from a DevSecOps team as possible.
Frequently Asked Questions
What is the AWS Deception Benchmark?
The Deception Benchmark is an AWS-developed security benchmark designed to measure whether AI models can accurately distinguish real software vulnerabilities from code that looks vulnerable but is actually safe.
How large is the benchmark?
It contains 14,822 code samples spanning 16 programming languages and more than 70 Common Weakness Enumeration categories.
How accurate were the AI models?
AWS found that models could identify many genuine vulnerabilities, but they also generated substantial numbers of false positives, flagging between 41% and 99% of safe code.

