LLM Vulnerability Leaderboard

How often each model, as the system under test, is successfully attacked by a fixed RedLens red-team suite. Lower is more robust.

Scope: self-hosted, open-weight models only (Llama, Mistral, Qwen) plus RedLens's own self-hosted attacker models as reference rows. Hosted frontier APIs (OpenAI, Anthropic, Google) are not scanned — their terms restrict automated adversarial testing without prior authorisation. Every number below is produced by an actual scan; a model with no result is shown as not yet run, never given a placeholder score.
Loading results…

Want to see your own model or deployment scanned like this?

The table above is open-weight models on a fixed public suite. RedLens runs the same adversarial engine against your production LLM application — your models, your system prompts, your tools and RAG — and maps every finding to the OWASP LLM Top 10 (2026) with remediation. Book a proof-of-concept assessment and we'll show you where your app actually stands.

Schedule a POC →

Methodology

Fairness comes from measuring every model the same way — the only thing that varies between rows is the target model.