Insights
AI Security Research

Three AI Scanners Agreed on 5% of Their Findings. Here's What a Reproducible AI Security Scan Looks Like.

Jeff Williams, the founder of Contrast Security and the original author of the OWASP Top 10, ran an experiment that should worry anyone buying an AI security scanner right now. His team pointed three different AI-powered security scanners at the same 50,000-line codebase, three times each. The scanners didn't just disagree with each other. They disagreed with themselves.

Across the three runs, the three scanners agreed on only 5% of the findings they produced. A Sonnet-based simple scan reproduced just 17% of its own findings run to run. Opus did somewhat better at 25%, but swung by almost 30% in total finding count between its best and worst run. Same code. Same scanner. Different answers depending on which run you happened to catch.

That's not a rounding error. It's a sign that a meaningful share of what these tools report isn't a property of the code being scanned. It's a property of the model's mood that day.

The same article points to a second, larger-scale version of the same problem. Anthropic's Claude Mythos has generated 26,153 vulnerability findings since Project Glasswing launched in April 2026. Of those, only 2,736 (about 10%) have made it into the public disclosure ledger, meaning they've been reported to a maintainer or are in the process of being reported. Fewer than 0.8%, a total of 202, are confirmed patched. 245 were withdrawn outright. VulnCheck's Patrick Garrity, who did the analysis, also flagged a severity-inflation problem: Claude rated 91.5% of ledger findings as critical or high severity, while the maintainers who actually own the code rated only 61.3% that high.

Put those two data points together and you get the same underlying story from two different angles. Volume is not the bottleneck anymore. An AI scanner can generate more candidate findings in an afternoon than a human team can triage in a year. The bottleneck is knowing which of those findings are real, reproducible, and worth a human's time, and that is precisely the question a scanner whose own output changes between runs cannot answer for you.

Why non-determinism happens, and why it isn't a bug you can prompt your way out of

The scanners in Contrast's test are, by design, LLM-judged. An agent reads code, reasons about whether a pattern constitutes a vulnerability, and writes up a finding in natural language. That reasoning step is where the non-determinism comes from. Ask the same model the same question about the same fifty thousand lines of code twice, with even minor variation in context window packing, sampling, or the model's own internal state, and you can get two different judgments about whether the same code path is exploitable. Neither judgment is necessarily wrong. They're just not the same answer, which means neither one is reproducible on its own.

That's a structural property of LLM-as-judge, not a prompting mistake. You can tighten a prompt and improve the average quality of individual findings without touching the underlying reproducibility problem at all, because the problem isn't that the model reasons poorly. It's that "reasoning about whether this is a vulnerability" is not a deterministic operation to begin with.

What a reproducible check looks like instead

This is the reason RedLens's scan architecture separates attack generation from execution from judgment, and holds the parts that determine comparability constant rather than leaving them to vary between runs.

Concretely, the pipeline runs in three phases: an attacker model generates a fixed number of adversarial prompts for a given attack methodology, each prompt is executed against the target model under test, and a judge model scores each resulting response. Phases two and three are paired and run per attack, and a single failed attack (including a target refusal, which is a routine and expected outcome for an adversarial prompt) is recorded as an error on that attack's own result row rather than silently dropped or allowed to skew the aggregate. Only a total failure to generate any attacks aborts a scan, because a scan with zero attacks has nothing to report.

The comparability piece matters as much as the pipeline shape. RedLens's public benchmark methodology holds the attacker model, the judge model, the number of attacks per methodology, and the attack methodology suite fixed across every target model it evaluates, specifically so that the only thing that varies between two scan results is the model being tested, not the tooling doing the testing. Errored attacks are excluded from the vulnerability-rate denominator rather than counted as either a pass or a fail, and a model with no reachable endpoint is reported as not run rather than scored as zero. The methodology commits to no invented numbers: every published figure has to trace back to an actual scan the engine executed, and the published artifact is exactly the engine's output.

None of that makes an individual finding infallible; a fixed attack suite is not an exhaustive one, and a model's real-world safety still depends on the guardrails deployed around it in production, which a suite like this doesn't test. But it does mean that two scans run under the same fixed parameters should produce comparable results, in the way that three independent LLM-judged scanners, run three times each against the same code, did not.

Williams put it plainly: "the world is still doing security at human speed." The fix for that isn't a faster firehose. It's fewer findings that hold up the second time you look at them.

Sources

Get New Posts by Email

New posts, straight to your inbox

Original writing on AI security from the people building the platform: agentic risk, red-team findings, and the advisories worth reading in full. A few times a month at most, no sales sequence, and one click to leave.

We use your address for new-post notifications and nothing else. Every email carries an unsubscribe link. See our Privacy Policy.

Test It, Don't Assume It

Find out whether your agents can actually be turned against you

RedLens red-teams the AI systems you've already shipped, covering agent tool abuse, prompt injection, and data exfiltration, then shows you exactly how to close what it finds.