Harmful compliance and over-refusal
Responses to disallowed requests, and refusals of benign ones.
- Reported by
- Anthropic, DeepSeek, Google DeepMind, Meta, Moonshot AI, OpenAI, xAI and Zhipu AI
- Evaluations
- 45 evaluations, of six kinds of test
- Values
- 606 values in 59 documents
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Definition
This family covers both sides of how a model responds to requests that may be harmful. Harmful compliance is helping with something a developer’s policy disallows, such as instructions for serious harm, abusive content or sexual content involving minors. Over-refusal is the opposite failure: declining a harmless request because it resembles a harmful one.
Developers measure both with sets of test prompts sorted into policy categories and graded as safe or unsafe, and also with multi-turn conversations and agentic tasks.
Results are often the share of responses that were not unsafe, the share of harmful requests refused, or the share of harmless requests answered, so a higher value can be better or worse depending on the measure. Each developer writes its own policy and its own prompts, and the categories and graders can change from one document to the next.
The line under the title is this family’s entry in Methodology §1.1. A family groups results about one property; it does not make them comparable (Methodology §5).
How developers measure it
45 evaluations of six kinds of test, with the most values first. Each evaluation opens in the data explorer, which lists all of its values.
| Evaluation | Reported by | Values | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Static benchmark · 39 evaluations, 595 values | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Production Benchmarks, in the data explorer Run by the developer OpenAI test of model responses to prompt sets in disallowed-content categories, reported per category as the share of responses graded not unsafe.Reported by OpenAI · 301 values | OpenAI | 301 | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Single-turn violative requests, in the data explorer Run by the developer Measures how often the model gives a harmless response to single-turn prompts that violate Anthropic's usage policy.Reported by Anthropic · 48 values | Anthropic | 48 | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Single-turn benign requests, in the data explorer Run by the developer Measures how often the model refuses single-turn requests that are benign (over-refusal).Reported by Anthropic · 31 values | Anthropic | 31 | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Malicious Claude Code use, in the data explorer Run by the developer Measures refusal of malicious requests and completion of dual-use and benign requests when the model works in Claude Code.Reported by Anthropic · 27 values | Anthropic | 27 | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Image to Text Safety, in the data explorer Run by the developer Automated evaluation of policy violations when the model answers prompts that include images, reported as a change against an earlier Gemini model.Reported by Google DeepMind · 14 values | Google DeepMind | 14 | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Multilingual Safety, in the data explorer Run by the developer Automated evaluation of policy violations on prompts in many languages, reported as a change against an earlier Gemini model.Reported by Google DeepMind · 14 values | Google DeepMind | 14 | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Text to Text Safety, in the data explorer Run by the developer Automated evaluation of how often the model's text responses violate Google's content safety policies, reported as a change against an earlier Gemini model.Reported by Google DeepMind · 14 values | Google DeepMind | 14 | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Tone, in the data explorer Run by the developer Automated rating of how objective the tone of the model's refusals is, reported as a change against an earlier Gemini model.Reported by Google DeepMind · 14 values | Google DeepMind | 14 | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Show the other 31 evaluations of this kind (132 values)
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Adaptive attack · 2 evaluations, 4 values | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Automated Red Teaming (helpfulness), in the data explorer Run by the developer Share of queries generated by Google's automated red-teaming system to which the model responds unhelpfully.Reported by Google DeepMind · 2 values | Google DeepMind | 2 | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Dynamic mental health, in the data explorer Run by the developer Multi-turn conversations with a simulated adversarial user on mental-health topics such as self-harm, graded for whether the model responses are not unsafe.Reported by OpenAI · 2 values | OpenAI | 2 | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Automated audit · 1 evaluation, 3 values | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Alignment audit, in the data explorer Run by the developer Misuse findings from xAI's automated alignment audit built on the Petri 2.0 tool, measuring how often the model cooperates with harmful requests, including when the system prompt tries to override its policy.Reported by xAI · 3 values | xAI | 3 | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Framework determination · 1 evaluation, 2 values | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Child safety launch thresholds, in the data explorer Run by the developer Whether the model meets Google's internal child-safety thresholds required for launch.Reported by Google DeepMind · 2 values | Google DeepMind | 2 | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Red teaming · 1 evaluation, 1 value | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Multi-turn biological weapons, in the data explorer Run by the developer Multi-turn conversations probing for biological-weapons assistance, scored as the share of safe responses.Reported by Anthropic · 1 value | Anthropic | 1 | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Qualitative assessment · 1 evaluation, 1 value | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| Safety benchmarks summary, in the data explorer Run by the developer DeepSeek's summary judgment of the model's inherent safety level across several public safety benchmarks, without its risk-control system.Reported by DeepSeek · 1 value | DeepSeek | 1 | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Reported values
The values a chart can place, by the kind of test that produced them. Open a point, or a value in the table, for its source and history. Open in explorer: Harmful compliance and over-refusal
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
Harmful compliance and over-refusal: values by kind of test
Each point is a value reported in a developer’s document, in percent (higher is better), grouped by the kind of test that produced it.
Hollow: read in a secondary write-up
| Model | Developer | Evaluation | Condition | Value | Document | Kind of test |
|---|---|---|---|---|---|---|
| Claude Opus 4 | Anthropic | Single-turn violative requestsHarmless response rate | overall across thinking modes | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Static benchmark | |
| Claude Sonnet 4 | Anthropic | Single-turn violative requestsHarmless response rate | overall across thinking modes | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Static benchmark | |
| Claude Sonnet 3.7 | Anthropic | Single-turn violative requestsHarmless response rate | overall across thinking modes | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Static benchmark | |
| Claude Opus 4 | Anthropic | Single-turn violative requestsHarmless response rate | overall across thinking modes; with ASL-3 safeguards | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Static benchmark | |
| Claude Opus 4 | Anthropic | Single-turn violative requestsHarmless response rate | no extended thinking | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Static benchmark | |
| Claude Opus 4 | Anthropic | Single-turn violative requestsHarmless response rate | extended thinking | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Static benchmark | |
| Claude Opus 4 | Anthropic | Malicious agentic codingSafety score | without safeguards | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Static benchmark | |
| Claude Sonnet 4 | Anthropic | Malicious agentic codingSafety score | without safeguards | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Static benchmark | |
| Claude Sonnet 3.7 | Anthropic | Malicious agentic codingSafety score | without safeguards | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Static benchmark | |
| Claude Opus 4.1 | Anthropic | Single-turn violative requestsHarmless response rate | overall across thinking modes | Claude Opus 4.1 System Card (addendum)5 Aug 2025 | Static benchmark | |
| Claude Opus 4.1 | Anthropic | Single-turn violative requestsHarmless response rate | no extended thinking | Claude Opus 4.1 System Card (addendum)5 Aug 2025 | Static benchmark | |
| Claude Opus 4.1 | Anthropic | Single-turn violative requestsHarmless response rate | extended thinking | Claude Opus 4.1 System Card (addendum)5 Aug 2025 | Static benchmark | |
| Claude Opus 4 | Anthropic | Single-turn violative requestsHarmless response rate | overall across thinking modes | Claude Opus 4.1 System Card (addendum)5 Aug 2025 | Static benchmark | |
| Claude Sonnet 4.5 | Anthropic | Single-turn violative requestsHarmless response rate | overall across thinking modes | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Sonnet 4 | Anthropic | Single-turn violative requestsHarmless response rate | overall across thinking modes | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Sonnet 4.5 | Anthropic | Single-turn violative requestsHarmless response rate | no extended thinking | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Sonnet 4.5 | Anthropic | Single-turn violative requestsHarmless response rate | extended thinking | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Sonnet 4.5 | Anthropic | Malicious agentic codingSafety score | without safeguards | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Sonnet 4.5 | Anthropic | Malicious Claude Code useRefusal rate, overt malicious | without safeguards | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Sonnet 4.5 | Anthropic | Malicious Claude Code useRefusal rate, covert malicious | without safeguards | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Sonnet 4 | Anthropic | Malicious Claude Code useRefusal rate, overt malicious | without safeguards | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Sonnet 4 | Anthropic | Malicious Claude Code useRefusal rate, covert malicious | without safeguards | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Sonnet 4.5 | Anthropic | Malicious Claude Code useSuccess rate, dual-use requests | without safeguards | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Sonnet 4.5 | Anthropic | Malicious Claude Code useRefusal rate, covert malicious | with mitigations | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Haiku 4.5 | Anthropic | Single-turn violative requestsHarmless response rate | no extended thinking | Claude Haiku 4.5 System Card15 Oct 2025 | Static benchmark | |
| Claude Haiku 4.5 | Anthropic | Single-turn violative requestsHarmless response rate | extended thinking | Claude Haiku 4.5 System Card15 Oct 2025 | Static benchmark | |
| Claude Haiku 3.5 | Anthropic | Single-turn violative requestsHarmless response rate | overall across thinking modes | Claude Haiku 4.5 System Card15 Oct 2025 | Static benchmark | |
| Claude Haiku 4.5 | Anthropic | Malicious agentic codingSafety score | without safeguards | Claude Haiku 4.5 System Card15 Oct 2025 | Static benchmark | |
| Claude Haiku 3.5 | Anthropic | Malicious agentic codingSafety score | without safeguards | Claude Haiku 4.5 System Card15 Oct 2025 | Static benchmark | |
| Claude Haiku 4.5 | Anthropic | Malicious Claude Code useRefusal rate, malicious requests | without safeguards | Claude Haiku 4.5 System Card15 Oct 2025 | Static benchmark | |
| Claude Haiku 4.5 | Anthropic | Malicious Claude Code useSuccess rate, dual-use & benign | without safeguards | Claude Haiku 4.5 System Card15 Oct 2025 | Static benchmark | |
| Claude Sonnet 4.5 | Anthropic | Malicious Claude Code useRefusal rate, malicious requests | without safeguards | Claude Haiku 4.5 System Card15 Oct 2025 | Static benchmark | |
| Claude Sonnet 4.5 | Anthropic | Malicious Claude Code useSuccess rate, dual-use & benign | without safeguards | Claude Haiku 4.5 System Card15 Oct 2025 | Static benchmark | |
| Claude Haiku 3.5 | Anthropic | Malicious Claude Code useRefusal rate, malicious requests | without safeguards | Claude Haiku 4.5 System Card15 Oct 2025 | Static benchmark | |
| Claude Haiku 3.5 | Anthropic | Malicious Claude Code useSuccess rate, dual-use & benign | without safeguards | Claude Haiku 4.5 System Card15 Oct 2025 | Static benchmark | |
| Claude Haiku 4.5 | Anthropic | Malicious Claude Code useRefusal rate, malicious requests | with new mitigations | Claude Haiku 4.5 System Card15 Oct 2025 | Static benchmark | |
| Claude Opus 4.5 | Anthropic | Single-turn violative requestsHarmless response rate | overall across thinking modes; all languages | Claude Opus 4.5 System Card24 Nov 2025 | Static benchmark | |
| Claude Haiku 4.5 | Anthropic | Single-turn violative requestsHarmless response rate | overall across thinking modes; all languages | Claude Opus 4.5 System Card24 Nov 2025 | Static benchmark | |
| Claude Sonnet 4.5 | Anthropic | Single-turn violative requestsHarmless response rate | overall across thinking modes; all languages | Claude Opus 4.5 System Card24 Nov 2025 | Static benchmark | |
| Claude Opus 4.1 | Anthropic | Single-turn violative requestsHarmless response rate | overall across thinking modes; all languages | Claude Opus 4.5 System Card24 Nov 2025 | Static benchmark | |
| Claude Opus 4.5 | Anthropic | Single-turn violative requestsHarmless response rate | no extended thinking; all languages | Claude Opus 4.5 System Card24 Nov 2025 | Static benchmark | |
| Claude Opus 4.5 | Anthropic | Single-turn violative requestsHarmless response rate | extended thinking; all languages | Claude Opus 4.5 System Card24 Nov 2025 | Static benchmark | |
| Claude Opus 4.5 | Anthropic | Malicious agentic codingSafety score | without safeguards | Claude Opus 4.5 System Card24 Nov 2025 | Static benchmark | |
| Claude Opus 4.1 | Anthropic | Malicious agentic codingSafety score | without safeguards | Claude Opus 4.5 System Card24 Nov 2025 | Static benchmark | |
| Claude Opus 4.5 | Anthropic | Malicious Claude Code useRefusal rate, malicious requests | without safeguards | Claude Opus 4.5 System Card24 Nov 2025 | Static benchmark | |
| Claude Opus 4.1 | Anthropic | Malicious Claude Code useRefusal rate, malicious requests | without safeguards | Claude Opus 4.5 System Card24 Nov 2025 | Static benchmark | |
| Claude Opus 4.5 | Anthropic | Malicious Claude Code useSuccess rate, dual-use & benign | without safeguards | Claude Opus 4.5 System Card24 Nov 2025 | Static benchmark | |
| Claude Opus 4.5 | Anthropic | Malicious Claude Code useRefusal rate, malicious requests | with mitigations | Claude Opus 4.5 System Card24 Nov 2025 | Static benchmark | |
| Claude Opus 4.5 | Anthropic | Malicious computer useRefusal rate | without safeguards | Claude Opus 4.5 System Card24 Nov 2025 | Static benchmark | |
| Claude Haiku 4.5 | Anthropic | Malicious computer useRefusal rate | without safeguards | Claude Opus 4.5 System Card24 Nov 2025 | Static benchmark | |
| Claude Sonnet 4.5 | Anthropic | Malicious computer useRefusal rate | without safeguards | Claude Opus 4.5 System Card24 Nov 2025 | Static benchmark | |
| Claude Opus 4.1 | Anthropic | Malicious computer useRefusal rate | without safeguards | Claude Opus 4.5 System Card24 Nov 2025 | Static benchmark | |
| Claude Opus 4.6 | Anthropic | Single-turn violative requestsHarmless response rate | overall across thinking modes | Claude Opus 4.6 System CardFeb 2026 | Static benchmark | |
| Claude Opus 4.5 | Anthropic | Single-turn violative requestsHarmless response rate | overall across thinking modes | Claude Opus 4.6 System CardFeb 2026 | Static benchmark | |
| Claude Sonnet 4.5 | Anthropic | Single-turn violative requestsHarmless response rate | overall across thinking modes | Claude Opus 4.6 System CardFeb 2026 | Static benchmark | |
| Claude Haiku 4.5 | Anthropic | Single-turn violative requestsHarmless response rate | overall across thinking modes | Claude Opus 4.6 System CardFeb 2026 | Static benchmark | |
| Claude Opus 4.6 | Anthropic | Higher-difficulty violative requestsHarmless response rate | overall across thinking modes | Claude Opus 4.6 System CardFeb 2026 | Static benchmark | |
| Claude Opus 4.6 | Anthropic | Malicious Claude Code useRefusal rate, malicious requests | with system prompt + FileRead reminder | Claude Opus 4.6 System CardFeb 2026 | Static benchmark | |
| Claude Opus 4.6 | Anthropic | Malicious computer useRefusal rate | — | Claude Opus 4.6 System CardFeb 2026 | Static benchmark | |
| Claude Sonnet 4.6 | Anthropic | Single-turn violative requestsHarmless response rate | overall across thinking modes | Claude Sonnet 4.6 System Card17 Feb 2026 | Static benchmark | |
| Claude Sonnet 4.6 | Anthropic | Higher-difficulty violative requestsHarmless response rate | overall across thinking modes | Claude Sonnet 4.6 System Card17 Feb 2026 | Static benchmark | |
| Claude Sonnet 4.6 | Anthropic | Child safety single-turn violative requestsHarmless response rate | — | Claude Sonnet 4.6 System Card17 Feb 2026 | Static benchmark | |
| Claude Sonnet 4.6 | Anthropic | Child safety multi-turnAppropriate response rate | — | Claude Sonnet 4.6 System Card17 Feb 2026 | Static benchmark | |
| Claude Sonnet 4.6 | Anthropic | Suicide/self-harm multi-turnAppropriate response rate | — | Claude Sonnet 4.6 System Card17 Feb 2026 | Static benchmark | |
| Claude Opus 4.6 | Anthropic | Suicide/self-harm multi-turnAppropriate response rate | — | Claude Sonnet 4.6 System Card17 Feb 2026 | Static benchmark | |
| Claude Opus 4.8 | Anthropic | Single-turn violative requestsHarmless response rate | no system prompt; API | Claude Opus 4.8 System Card28 May 2026 | Static benchmark | |
| Claude Opus 4.8 | Anthropic | Single-turn violative requestsHarmless response rate | claude.ai system prompt | Claude Opus 4.8 System Card28 May 2026 | Static benchmark | |
| Claude Opus 4.7 | Anthropic | Single-turn violative requestsHarmless response rate | no system prompt; API | Claude Opus 4.8 System Card28 May 2026 | Static benchmark | |
| Claude Sonnet 4.6 | Anthropic | Single-turn violative requestsHarmless response rate | no system prompt; API | Claude Opus 4.8 System Card28 May 2026 | Static benchmark | |
| Claude Mythos Preview | Anthropic | Single-turn violative requestsHarmless response rate | no system prompt; API | Claude Opus 4.8 System Card28 May 2026 | Static benchmark | |
| Claude Sonnet 5 | Anthropic | Single-turn violative requestsHarmless response rate | no system prompt; API | Claude Sonnet 5 System Card30 Jun 2026 | Static benchmark | |
| Claude Sonnet 5 | Anthropic | Single-turn violative requestsHarmless response rate | claude.ai system prompt | Claude Sonnet 5 System Card30 Jun 2026 | Static benchmark | |
| Claude Sonnet 5 | Anthropic | Child safety multi-turnAppropriate response rate | API | Claude Sonnet 5 System Card30 Jun 2026 | Static benchmark | |
| Claude Sonnet 5 | Anthropic | Child safety multi-turnAppropriate response rate | claude.ai | Claude Sonnet 5 System Card30 Jun 2026 | Static benchmark | |
| Claude Sonnet 5 | Anthropic | Suicide/self-harm multi-turnAppropriate response rate | API | Claude Sonnet 5 System Card30 Jun 2026 | Static benchmark | |
| Claude Sonnet 5 | Anthropic | Suicide/self-harm multi-turnAppropriate response rate | claude.ai | Claude Sonnet 5 System Card30 Jun 2026 | Static benchmark | |
| Claude Sonnet 5 | Anthropic | Malicious Claude Code useRefusal rate, malicious requests | — | Claude Sonnet 5 System Card30 Jun 2026 | Static benchmark | |
| Claude Sonnet 5 | Anthropic | Malicious Claude Code useSuccess rate, dual-use & benign | — | Claude Sonnet 5 System Card30 Jun 2026 | Static benchmark | |
| Claude Sonnet 4.6 | Anthropic | Malicious Claude Code useRefusal rate, malicious requests | — | Claude Sonnet 5 System Card30 Jun 2026 | Static benchmark | |
| Claude Mythos 5 | Anthropic | Malicious Claude Code useRefusal rate, malicious requests | — | Claude Sonnet 5 System Card30 Jun 2026 | Static benchmark | |
| Claude Opus 4.8 | Anthropic | Malicious Claude Code useRefusal rate, malicious requests | — | Claude Sonnet 5 System Card30 Jun 2026 | Static benchmark | |
| Claude Mythos Preview | Anthropic | Malicious Claude Code useRefusal rate, malicious requests | — | Claude Sonnet 5 System Card30 Jun 2026 | Static benchmark | |
| Claude Sonnet 5 | Anthropic | Malicious computer useRefusal rate | — | Claude Sonnet 5 System Card30 Jun 2026 | Static benchmark | |
| Claude Mythos 5 | Anthropic | Malicious computer useRefusal rate | — | Claude Sonnet 5 System Card30 Jun 2026 | Static benchmark | |
| Claude Opus 4.8 | Anthropic | Malicious computer useRefusal rate | — | Claude Sonnet 5 System Card30 Jun 2026 | Static benchmark | |
| Claude Mythos Preview | Anthropic | Malicious computer useRefusal rate | — | Claude Sonnet 5 System Card30 Jun 2026 | Static benchmark | |
| Claude Sonnet 4.6 | Anthropic | Malicious computer useRefusal rate | — | Claude Sonnet 5 System Card30 Jun 2026 | Static benchmark | |
| Claude Mythos 5 | Anthropic | Single-turn violative requestsHarmless response rate | no system prompt; API | Claude Sonnet 5 System Card30 Jun 2026 | Static benchmark | |
| Claude Opus 5 | Anthropic | Single-turn violative requestsHarmless response rate | no system prompt; API | Claude Opus 5 System Card24 Jul 2026 | Static benchmark | |
| Claude Opus 5 | Anthropic | Single-turn violative requestsHarmless response rate | claude.ai system prompt | Claude Opus 5 System Card24 Jul 2026 | Static benchmark | |
| Claude Fable 5 | Anthropic | Single-turn violative requestsHarmless response rate | no system prompt; API | Claude Opus 5 System Card24 Jul 2026 | Static benchmark | |
| Claude Opus 5 | Anthropic | Child safety single-turn violative requestsHarmless response rate | API | Claude Opus 5 System Card24 Jul 2026 | Static benchmark | |
| Claude Opus 5 | Anthropic | Suicide/self-harm multi-turnAppropriate response rate | claude.ai | Claude Opus 5 System Card24 Jul 2026 | Static benchmark | |
| Muse Spark (Meta AI system) | Meta | BioTIER refusalsRefusal rate on harmful bio | with system mitigations | Muse Spark Safety & Preparedness Report8 Apr 2026 | Static benchmark | |
| Kimi K2-Instruct | Moonshot AI | Promptfoo red-team (no attack)Harmful: pass rate (safe responses) | Basic (no attack) | Kimi K2 technical report28 Jul 2025 | Static benchmark | |
| GPT-5.1-Codex-Max | OpenAI | Long-form biorisk questionsRefusal rate | — | GPT-5.1-Codex-Max System Card18 Nov 2025 | Static benchmark | |
| GPT-Image-2.5-Sunburst | OpenAI | Image safety evalPrompts yielding safe generation | — | ChatGPT Images 2.5 System Card8 Sep 2026 | Static benchmark | |
| GPT-Image-2.5-Flare | OpenAI | Image safety evalPrompts yielding safe generation | — | ChatGPT Images 2.5 System Card8 Sep 2026 | Static benchmark | |
| ChatGPT Images 2.0 | OpenAI | Image safety evalPrompts yielding safe generation | — | ChatGPT Images 2.5 System Card8 Sep 2026 | Static benchmark | |
| Grok 4.5 (high) | xAI | CBRN refusalsRefusal accuracy on dangerous bio queries | high; full safeguards | Grok 4.5 Model Card14 Jul 2026 | Static benchmark | |
| Grok 4.6 (high) | xAI | FORTRESS-RNR/N refusal accuracy | high | Grok 4.6 Model Card12 Aug 2026 | Static benchmark | |
| Grok 4.6 (high) | xAI | BioUseBenchSeverity-5 refusal rate | high | Grok 4.6 Model Card12 Aug 2026 | Static benchmark | |
| Grok 4.7 (high) | xAI | BioUseBenchDual-use bio refusal rate | high | Grok 4.7 Model Card21 Sep 2026 | Static benchmark | |
| GLM-4.5 | Zhipu AI | SafetyBenchOverall accuracy | overall | GLM-4.5 technical report8 Aug 2025 | Static benchmark | |
| GLM-4.5 | Zhipu AI | SafetyBenchUnfairness & bias accuracy | Unfairness & Bias | GLM-4.5 technical report8 Aug 2025 | Static benchmark | |
| Claude Fable 5.1 | Anthropic | Multi-turn biological weaponsSafe response rate | API | Claude Fable 5.1 & Claude Mythos 5.1 System Card1 Sep 2026 | Red teaming |
Marker shape is the developer whose document reports the value; a hollow marker is a value we read in an independent write-up because the card section could not be read directly.
Each point is one value, placed by the number the document printed. Bands are kinds of test; within a band, points are spread up and down only so that they do not overlap.
106 of the family’s 606 values are plotted: those in percent where higher is better, the unit most of its evaluations use. Of the rest, 435 in other units or read the other way (percent where higher is worse, rates from 0 to 1 and percentage points), 5 are categories or statements in words and 60 are restated in a later document or from an earlier version of one, as in the explorer’s default view. The explorer lists them all, one evaluation at a time.
The axis starts at zero.
Source: Safety Card Ledger v0.1 · Data from developer system cards · CC BY 4.0
Comparability
- Each of the 45 evaluations is its own test, with its own tasks, graders and definitions. Values from different evaluations describe the same property measured in different ways, and are not comparable with one another, even where two evaluations share a name.
- Values come from the documents of eight developers. Developers use different tests, prompts and graders for the same property, so we do not rank them on these numbers: tables here are ordered by name or by kind of test, never by a reported value.
- Some names are used by more than one evaluation: “AgentHarm” by two. They are different tests, run by different evaluators or reported by different developers, so this page names who ran each one.
- In six evaluations the test changed between documents (among them Production Benchmarks, Single-turn violative requests and Single-turn benign requests). Values on the old and new definitions are never joined: the explorer breaks the line where the test changed.
- In 20 evaluations the explorer connects results from different documents that do not say whether the test stayed the same. We connect them because nothing suggests it changed; where a document says it did, the line breaks. Each value on such a line is marked “test assumed unchanged”.
- 19 evaluations report values under more than one condition, such as reasoning effort, safeguards or number of attempts. Each condition is kept as a separate series.
- Numbers in this family are printed in three units: percent, rates from 0 to 1 and percentage points. We never convert one unit into another, so a chart shows one unit at a time.
- For some measures a higher value is worse and for others better, depending on whether the developer reports the behaviour or its absence. Each value’s record says which.
- 56 values are restatements: a later document reporting an earlier model again, sometimes with a different value. A restatement is kept beside the model’s own value, never in place of it.
Coverage
Documents from eight of the eight developers report this family. The coverage matrix counts a document when it has at least one printed number in the family; categories, statements in words and figure reads count only in the first column.
| Developer | With any value | With a printed number | All its documents |
|---|---|---|---|
| Anthropic | 13 | 12 | 15 |
| DeepSeek | 1 | 0 | 1 |
| Google DeepMind | 14 | 14 | 16 |
| Meta | 2 | 2 | 5 |
| Moonshot AI | 1 | 1 | 2 |
| OpenAI | 19 | 19 | 22 |
| xAI | 8 | 8 | 8 |
| Zhipu AI | 1 | 1 | 1 |
Related findings
No finding is about this family yet.