Jailbreak robustness
Resistance to attempts to bypass safeguards.
- Reported by
- Anthropic, Google DeepMind, Meta, Moonshot AI, OpenAI and xAI
- Evaluations
- 16 evaluations, of four kinds of test
- Values
- 98 values in 29 documents
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Definition
Jailbreak robustness is how well a model’s safeguards hold when someone deliberately tries to get around them. A jailbreak is a prompt designed to make a model produce what it would normally refuse, for example through role-play, disguised wording or a long series of turns.
Developers measure robustness with fixed sets of known jailbreak prompts, with public benchmarks, and with adaptive attackers, human or automated, who keep changing their approach until something works. Fixed sets show how a model handles known attacks; adaptive attacks come closer to a determined adversary.
Results are reported as the share of attacks that succeed, the share of responses that stay safe, or the effort an attacker needs. The prompt sets, attackers, graders and safeguards in place differ from test to test, so each result belongs to its own setup.
The line under the title is this family’s entry in Methodology §1.1. A family groups results about one property; it does not make them comparable (Methodology §5).
How developers measure it
16 evaluations of four kinds of test, with the most values first. Each evaluation opens in the data explorer, which lists all of its values.
| Evaluation | Reported by | Values |
|---|---|---|
| Static benchmark · 10 evaluations, 87 values | ||
| StrongReject · OpenAI, in the data explorer Run by the developer OpenAI's own run of the public StrongReject set of jailbreak-wrapped harmful requests, scored on the share of responses that are not unsafe.Reported by OpenAI · 49 values | OpenAI | 49 |
| Static jailbreak, in the data explorer Run by the developer A new fixed set of jailbreak attacks introduced in the GPT-6 Astra card, scored as the share of attacks refused in each harm area and severity tier.Reported by OpenAI · 11 values | OpenAI | 11 |
| Jailbreaks · xAI, in the data explorer Run by the developer Should-refuse prompts combined with jailbreak attacks, measuring how often the model still answers or complies.Reported by xAI · 9 values | xAI | 9 |
| StrongREJECT · Anthropic, in the data explorer Run by the developer Scores responses to StrongREJECT forbidden prompts under a set of jailbreak techniques and reports the score of the most effective jailbreak.Reported by Anthropic · 6 values | Anthropic | 6 |
| Jailbreaks · OpenAI, in the data explorer Run by the developer OpenAI's jailbreak robustness evaluation in the GPT-5.4 Thinking to GPT-5.6 cards, which show results only as figures; the dataset records the cards' comparisons in words.Reported by OpenAI · 4 values | OpenAI | 4 |
| Human sourced jailbreaks, in the data explorer Run by the developer A fixed set of jailbreak prompts gathered from people, scored on the share of responses that are not unsafe.Reported by OpenAI · 3 values | OpenAI | 3 |
| StrongREJECT · xAI, in the data explorer Run by the developer The public StrongREJECT benchmark of forbidden prompts under jailbreak attacks, run by xAI and reported as the share of compliant responses.Reported by xAI · 2 values | xAI | 2 |
| CBRN jailbreak template success, in the data explorer Run by the developer Rate at which common jailbreak templates achieve their objective against the model's CBRN safeguards.Reported by Google DeepMind · 1 value | Google DeepMind | 1 |
| FORTRESS, in the data explorer Run by the developer A public set of adversarial prompts on national-security and public-safety harms, run by Meta and reported as attack success rate.Reported by Meta · 1 value | Meta | 1 |
| Restricted-bio input filter, in the data explorer Run by the developer Tests xAI's input filter for restricted biology content, measuring how often harmful prompts get past it, with and without a prompt-injection attack.Reported by xAI · 1 value | xAI | 1 |
| Adaptive attack · 4 evaluations, 9 values | ||
| Automated Red Teaming (dangerous content), in the data explorer Run by the developer Share of queries generated by Google's automated red-teaming system that lead the model to produce policy-violating dangerous content.Reported by Google DeepMind · 4 values | Google DeepMind | 4 |
| Promptfoo red-team, in the data explorer Run by the developer Moonshot's run of the Promptfoo red-teaming tool, measuring the share of safe responses to harmful requests under jailbreak strategies.Reported by Moonshot AI · 2 values | Moonshot AI | 2 |
| StrongREJECT · Meta, in the data explorer Run by the developer Version 2 of the StrongREJECT forbidden-prompt benchmark with jailbreak attacks, run by Meta and reported as attack success rate.Reported by Meta · 2 values | Meta | 2 |
| Cyber safeguards automated red-teaming, in the data explorer Run by the developer An automated red-teaming agent tries to complete harmful cyber tasks against the model with its cyber safeguards in place.Reported by Anthropic · 1 value | Anthropic | 1 |
| Red teaming · 1 evaluation, 1 value | ||
| Universal jailbreak (cyber) · UK AI Security Institute, in the data explorer Run by UK AI Security Institute (government) UK AI Security Institute red-teamers searched for a universal jailbreak, scored by how often it passes on a dataset of policy-violating cyber requests.Reported by OpenAI · 1 value | OpenAI | 1 |
| Qualitative assessment · 1 evaluation, 1 value | ||
| Jailbreak vulnerability assessment, in the data explorer Run by the developer The developer's statement in words of how vulnerable the model is to jailbreaks compared with its predecessor.Reported by Google DeepMind · 1 value | Google DeepMind | 1 |
Reported values
The values a chart can place, by the kind of test that produced them. Open a point, or a value in the table, for its source and history. Open in explorer: Jailbreak robustness
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
Jailbreak robustness: values by kind of test
Each point is a value reported in a developer’s document, in percent (higher is worse), grouped by the kind of test that produced it.
| Model | Developer | Evaluation | Condition | Value | Document | Kind of test |
|---|---|---|---|---|---|---|
| Claude Opus 4 | Anthropic | StrongREJECTBest-jailbreak StrongREJECT score | no extended thinking | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Static benchmark | |
| Claude Opus 4 | Anthropic | StrongREJECTBest-jailbreak StrongREJECT score | extended thinking | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Static benchmark | |
| Claude Sonnet 4 | Anthropic | StrongREJECTBest-jailbreak StrongREJECT score | no extended thinking | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Static benchmark | |
| Claude Sonnet 4 | Anthropic | StrongREJECTBest-jailbreak StrongREJECT score | extended thinking | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Static benchmark | |
| Claude Sonnet 3.7 | Anthropic | StrongREJECTBest-jailbreak StrongREJECT score | no extended thinking | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Static benchmark | |
| Claude Sonnet 3.7 | Anthropic | StrongREJECTBest-jailbreak StrongREJECT score | extended thinking | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Static benchmark | |
| Gemini 3.7 Flash | Google DeepMind | CBRN jailbreak template successObjective success rate | with CBRN safeguards | Gemini 3.7 Flash Frontier Safety Framework Report13 Aug 2026 | Static benchmark | |
| Muse Spark | Meta | FORTRESSAttack success rate | without system mitigations | Muse Spark Safety & Preparedness Report8 Apr 2026 | Static benchmark | |
| Grok 4.5 (high) | xAI | JailbreaksCompliance under attack | high; standard jailbreaks | Grok 4.5 Model Card14 Jul 2026 | Static benchmark | |
| Grok 4.6 (high) | xAI | StrongREJECTCompliance under StrongREJECT attacks | high | Grok 4.6 Model Card12 Aug 2026 | Static benchmark | |
| Grok 4.6 (high) | xAI | JailbreaksCompliance under attack | high; standard jailbreaks | Grok 4.6 Model Card12 Aug 2026 | Static benchmark | |
| Grok 4.7 (high) | xAI | StrongREJECTCompliance under StrongREJECT attacks | high | Grok 4.7 Model Card21 Sep 2026 | Static benchmark | |
| Gemini 2.5 Pro | Google DeepMind | Automated Red Teaming (dangerous content)Dangerous content violation rate | — | Gemini 2.5 Technical Report17 Jun 2025 | Adaptive attack | |
| Gemini 2.5 Flash | Google DeepMind | Automated Red Teaming (dangerous content)Dangerous content violation rate | — | Gemini 2.5 Technical Report17 Jun 2025 | Adaptive attack | |
| Gemini 2.0 Flash | Google DeepMind | Automated Red Teaming (dangerous content)Dangerous content violation rate | — | Gemini 2.5 Technical Report17 Jun 2025 | Adaptive attack | |
| Gemini 1.5 Pro 002002 | Google DeepMind | Automated Red Teaming (dangerous content)Dangerous content violation rate | — | Gemini 2.5 Technical Report17 Jun 2025 | Adaptive attack | |
| Muse Spark | Meta | StrongREJECTAttack success rate (adaptive multi-turn) | without system mitigations | Muse Spark Safety & Preparedness Report8 Apr 2026 | Adaptive attack | |
| Muse Spark 1.1 | Meta | StrongREJECTAttack success rate | without system mitigations | Muse Spark 1.1 Evaluation Report9 Jul 2026 | Adaptive attack |
Marker shape is the developer whose document reports the value.
Each point is one value, placed by the number the document printed. Bands are kinds of test; within a band, points are spread up and down only so that they do not overlap.
18 of the family’s 98 values are plotted: those in percent where higher is worse, the unit most of its evaluations use. Of the rest, 72 in other units or read the other way (rates from 0 to 1 and percent where higher is better), 5 are categories or statements in words, 2 are restated in a later document or from an earlier version of one and 1 is low-confidence or read off a figure, as in the explorer’s default view. The explorer lists them all, one evaluation at a time.
The values span more than one and a half orders of magnitude, but one of them is zero, which a logarithmic axis cannot show, so the axis is linear and starts at zero.
Source: Safety Card Ledger v0.1 · Data from developer system cards · CC BY 4.0
Comparability
- Each of the 16 evaluations is its own test, with its own tasks, graders and definitions. Values from different evaluations describe the same property measured in different ways, and are not comparable with one another, even where two evaluations share a name.
- Values come from the documents of six developers. Developers use different tests, prompts and graders for the same property, so we do not rank them on these numbers: tables here are ordered by name or by kind of test, never by a reported value.
- Some names are used by more than one evaluation: “StrongReject” by four and “Jailbreaks” by two. They are different tests, run by different evaluators or reported by different developers, so this page names who ran each one.
- In four evaluations the explorer connects results from different documents that do not say whether the test stayed the same. We connect them because nothing suggests it changed; where a document says it did, the line breaks. Each value on such a line is marked “test assumed unchanged”.
- Three evaluations report values under more than one condition, such as reasoning effort, safeguards or number of attempts. Each condition is kept as a separate series.
- Numbers in this family are printed in two units: rates from 0 to 1 and percent. We never convert one unit into another, so a chart shows one unit at a time.
- For some measures a higher value is worse and for others better, depending on whether the developer reports the behaviour or its absence. Each value’s record says which.
- Two values are restatements: a later document reporting an earlier model again, sometimes with a different value. A restatement is kept beside the model’s own value, never in place of it.
Coverage
Documents from six of the eight developers report this family. The coverage matrix counts a document when it has at least one printed number in the family; categories, statements in words and figure reads count only in the first column.
| Developer | With any value | With a printed number | All its documents |
|---|---|---|---|
| Anthropic | 2 | 2 | 15 |
| DeepSeek | 0 | 0 | 1 |
| Google DeepMind | 3 | 2 | 16 |
| Meta | 2 | 2 | 5 |
| Moonshot AI | 1 | 1 | 2 |
| OpenAI | 14 | 10 | 22 |
| xAI | 7 | 7 | 8 |
| Zhipu AI | 0 | 0 | 1 |
Related findings
No finding is about this family yet.