Definition

This family covers both sides of how a model responds to requests that may be harmful. Harmful compliance is helping with something a developer’s policy disallows, such as instructions for serious harm, abusive content or sexual content involving minors. Over-refusal is the opposite failure: declining a harmless request because it resembles a harmful one.

Developers measure both with sets of test prompts sorted into policy categories and graded as safe or unsafe, and also with multi-turn conversations and agentic tasks.

Results are often the share of responses that were not unsafe, the share of harmful requests refused, or the share of harmless requests answered, so a higher value can be better or worse depending on the measure. Each developer writes its own policy and its own prompts, and the categories and graders can change from one document to the next.

The line under the title is this family’s entry in Methodology §1.1. A family groups results about one property; it does not make them comparable (Methodology §5).

How developers measure it

45 evaluations of six kinds of test, with the most values first. Each evaluation opens in the data explorer, which lists all of its values.

The 45 evaluations in Harmful compliance and over-refusal, grouped by kind of test: who ran each one, what it counts, the developers whose documents report it and how many values it has.
EvaluationReported byValues
Static benchmark · 39 evaluations, 595 values
Production Benchmarks, in the data explorer Run by the developer OpenAI test of model responses to prompt sets in disallowed-content categories, reported per category as the share of responses graded not unsafe.Reported by OpenAI · 301 valuesOpenAI301
Single-turn violative requests, in the data explorer Run by the developer Measures how often the model gives a harmless response to single-turn prompts that violate Anthropic's usage policy.Reported by Anthropic · 48 valuesAnthropic48
Single-turn benign requests, in the data explorer Run by the developer Measures how often the model refuses single-turn requests that are benign (over-refusal).Reported by Anthropic · 31 valuesAnthropic31
Malicious Claude Code use, in the data explorer Run by the developer Measures refusal of malicious requests and completion of dual-use and benign requests when the model works in Claude Code.Reported by Anthropic · 27 valuesAnthropic27
Image to Text Safety, in the data explorer Run by the developer Automated evaluation of policy violations when the model answers prompts that include images, reported as a change against an earlier Gemini model.Reported by Google DeepMind · 14 valuesGoogle DeepMind14
Multilingual Safety, in the data explorer Run by the developer Automated evaluation of policy violations on prompts in many languages, reported as a change against an earlier Gemini model.Reported by Google DeepMind · 14 valuesGoogle DeepMind14
Text to Text Safety, in the data explorer Run by the developer Automated evaluation of how often the model's text responses violate Google's content safety policies, reported as a change against an earlier Gemini model.Reported by Google DeepMind · 14 valuesGoogle DeepMind14
Tone, in the data explorer Run by the developer Automated rating of how objective the tone of the model's refusals is, reported as a change against an earlier Gemini model.Reported by Google DeepMind · 14 valuesGoogle DeepMind14
Show the other 31 evaluations of this kind (132 values)
More static benchmarks in Harmful compliance and over-refusal
EvaluationReported byValues
Malicious agentic coding, in the data explorer Run by the developer Measures whether the model declines clearly malicious coding tasks when working as a coding agent.Reported by Anthropic · 12 valuesAnthropic12
HackerBench, in the data explorer Run by the developer A set of cyber tasks reported by xAI, measuring how often the model complies with harmful or dual-use requests and how often it wrongly refuses benign ones.Reported by xAI · 11 valuesxAI11
Cyber safety, in the data explorer Run by the developer Measures how often responses to cybersecurity-related prompts, from production data and from synthetic data, follow OpenAI cyber policy.Reported by OpenAI · 10 valuesOpenAI10
Malicious computer use, in the data explorer Run by the developer Measures how often the model refuses malicious tasks while operating a computer through the computer-use tool.Reported by Anthropic · 10 valuesAnthropic10
Unjustified refusals, in the data explorer Run by the developer Automated measure of how often the model refuses borderline prompts it should answer, reported as a change against an earlier Gemini model.Reported by Google DeepMind · 9 valuesGoogle DeepMind9
Voice-native production prompts, in the data explorer Run by the developer Tests spoken responses of a voice model to production-derived prompts, reporting a safety score per harm category.Reported by OpenAI · 8 valuesOpenAI8
Refusals, in the data explorer Run by the developer Prompts that xAI's policy says the model should refuse, measuring how often it answers or complies instead.Reported by xAI · 7 valuesxAI7
Image safety eval, in the data explorer Run by the developer Image-generation safety test that measures how often unsafe images reach the user and how often prompts still yield a safe image.Reported by OpenAI · 6 valuesOpenAI6
AgentHarm · xAI, in the data explorer Run by the developer A public benchmark of harmful agentic tasks, run by xAI, measuring how often the model carries them out rather than refusing.Reported by xAI · 5 valuesxAI5
Instruction Following, in the data explorer Run by the developer Automated measure of how well the model follows user instructions while staying within policy, reported as a change against an earlier Gemini model.Reported by Google DeepMind · 5 valuesGoogle DeepMind5
Malware refusals (golden set), in the data explorer Run by the developer Measures how often a coding model refuses a curated set of requests to build malicious software.Reported by OpenAI · 5 valuesOpenAI5
Suicide/self-harm multi-turn, in the data explorer Run by the developer Measures how often the model responds appropriately over multi-turn conversations involving suicide or self-harm.Reported by Anthropic · 5 valuesAnthropic5
Multi-turn testing, in the data explorer Run by the developer Tests the model's responses over multi-turn conversations in specific harm areas, graded against rubrics.Reported by Anthropic · 4 valuesAnthropic4
Self-harm refusals, in the data explorer Run by the developer Prompts about self-harm that the model should refuse or redirect, measuring how often it complies instead.Reported by xAI · 4 valuesxAI4
Challenging refusal eval, in the data explorer Run by the developer Harder fixed set of disallowed-content prompts, scored as the share of responses graded not unsafe.Reported by OpenAI · 3 valuesOpenAI3
Child safety multi-turn, in the data explorer Run by the developer Measures how often the model responds appropriately over multi-turn conversations on child-safety topics.Reported by Anthropic · 3 valuesAnthropic3
Higher-difficulty benign requests, in the data explorer Run by the developer Measures how often the model refuses a harder set of benign requests reported alongside the standard single-turn set.Reported by Anthropic · 3 valuesAnthropic3
Standard refusal eval, in the data explorer Run by the developer Fixed set of disallowed and benign prompts checking that the model avoids unsafe output without refusing benign requests.Reported by OpenAI · 3 valuesOpenAI3
Voice-native synthetic prompts, in the data explorer Run by the developer Tests spoken responses of a voice model to synthetic prompts, reporting a safety score per harm category.Reported by OpenAI · 3 valuesOpenAI3
BioUseBench, in the data explorer Run by the developer Dual-use biology requests, measuring how often the model refuses them.Reported by xAI · 2 valuesxAI2
Child safety single-turn violative requests, in the data explorer Run by the developer Measures the harmless response rate on single-turn prompts that violate child-safety policy.Reported by Anthropic · 2 valuesAnthropic2
Higher-difficulty violative requests, in the data explorer Run by the developer Measures the harmless response rate on a harder set of policy-violating requests reported alongside the standard single-turn set.Reported by Anthropic · 2 valuesAnthropic2
SafetyBench, in the data explorer Run by the developer A public multiple-choice benchmark of safety knowledge across categories such as offensiveness and unfairness, reported as accuracy.Reported by Zhipu AI · 2 valuesZhipu AI2
AgentHarm · Meta, in the data explorer Run by the developer A public benchmark of harmful agentic tasks, run by Meta, measuring how often the model complies.Reported by Meta · 1 valueMeta1
Bio/chem weapons query violations, in the data explorer Run by the developer Rate of policy-violating responses to queries about biological and chemical weapons with safeguards in place.Reported by Google DeepMind · 1 valueGoogle DeepMind1
BioTIER refusals, in the data explorer Run by the developer Harmful biology requests sent to the deployed Meta AI system, measuring how often it refuses.Reported by Meta · 1 valueMeta1
CBRN refusals, in the data explorer Run by the developer Dangerous chemical, biological, radiological and nuclear queries, measuring how accurately the model refuses them.Reported by xAI · 1 valuexAI1
FORTRESS-RN, in the data explorer Run by the developer Radiological and nuclear prompts, measuring how accurately the model refuses them.Reported by xAI · 1 valuexAI1
Long-form biorisk questions, in the data explorer Run by the developer Measures how often the model refuses long-form questions that seek biological-risk information.Reported by OpenAI · 1 valueOpenAI1
Promptfoo red-team (no attack), in the data explorer Run by the developer Moonshot's run of the Promptfoo red-teaming tool, measuring the share of safe responses to harmful requests with no attack strategy.Reported by Moonshot AI · 1 valueMoonshot AI1
Refusals on debated political and social topics, in the data explorer Run by the developer Prompts on debated political and social topics, measuring how often the model refuses to engage.Reported by Meta · 1 valueMeta1
Adaptive attack · 2 evaluations, 4 values
Automated Red Teaming (helpfulness), in the data explorer Run by the developer Share of queries generated by Google's automated red-teaming system to which the model responds unhelpfully.Reported by Google DeepMind · 2 valuesGoogle DeepMind2
Dynamic mental health, in the data explorer Run by the developer Multi-turn conversations with a simulated adversarial user on mental-health topics such as self-harm, graded for whether the model responses are not unsafe.Reported by OpenAI · 2 valuesOpenAI2
Automated audit · 1 evaluation, 3 values
Alignment audit, in the data explorer Run by the developer Misuse findings from xAI's automated alignment audit built on the Petri 2.0 tool, measuring how often the model cooperates with harmful requests, including when the system prompt tries to override its policy.Reported by xAI · 3 valuesxAI3
Framework determination · 1 evaluation, 2 values
Child safety launch thresholds, in the data explorer Run by the developer Whether the model meets Google's internal child-safety thresholds required for launch.Reported by Google DeepMind · 2 valuesGoogle DeepMind2
Red teaming · 1 evaluation, 1 value
Multi-turn biological weapons, in the data explorer Run by the developer Multi-turn conversations probing for biological-weapons assistance, scored as the share of safe responses.Reported by Anthropic · 1 valueAnthropic1
Qualitative assessment · 1 evaluation, 1 value
Safety benchmarks summary, in the data explorer Run by the developer DeepSeek's summary judgment of the model's inherent safety level across several public safety benchmarks, without its risk-control system.Reported by DeepSeek · 1 valueDeepSeek1

Reported values

The values a chart can place, by the kind of test that produced them. Open a point, or a value in the table, for its source and history. Open in explorer: Harmful compliance and over-refusal

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

Harmful compliance and over-refusal: values by kind of test

Each point is a value reported in a developer’s document, in percent (higher is better), grouped by the kind of test that produced it.

Hollow: read in a secondary write-up

Static benchmarkRed teaming020406080100120Percent

Marker shape is the developer whose document reports the value; a hollow marker is a value we read in an independent write-up because the card section could not be read directly.

Each point is one value, placed by the number the document printed. Bands are kinds of test; within a band, points are spread up and down only so that they do not overlap.

106 of the family’s 606 values are plotted: those in percent where higher is better, the unit most of its evaluations use. Of the rest, 435 in other units or read the other way (percent where higher is worse, rates from 0 to 1 and percentage points), 5 are categories or statements in words and 60 are restated in a later document or from an earlier version of one, as in the explorer’s default view. The explorer lists them all, one evaluation at a time.

The axis starts at zero.

Source: Safety Card Ledger v0.1 · Data from developer system cards · CC BY 4.0

Comparability

  • Each of the 45 evaluations is its own test, with its own tasks, graders and definitions. Values from different evaluations describe the same property measured in different ways, and are not comparable with one another, even where two evaluations share a name.
  • Values come from the documents of eight developers. Developers use different tests, prompts and graders for the same property, so we do not rank them on these numbers: tables here are ordered by name or by kind of test, never by a reported value.
  • Some names are used by more than one evaluation: “AgentHarm” by two. They are different tests, run by different evaluators or reported by different developers, so this page names who ran each one.
  • In six evaluations the test changed between documents (among them Production Benchmarks, Single-turn violative requests and Single-turn benign requests). Values on the old and new definitions are never joined: the explorer breaks the line where the test changed.
  • In 20 evaluations the explorer connects results from different documents that do not say whether the test stayed the same. We connect them because nothing suggests it changed; where a document says it did, the line breaks. Each value on such a line is marked “test assumed unchanged”.
  • 19 evaluations report values under more than one condition, such as reasoning effort, safeguards or number of attempts. Each condition is kept as a separate series.
  • Numbers in this family are printed in three units: percent, rates from 0 to 1 and percentage points. We never convert one unit into another, so a chart shows one unit at a time.
  • For some measures a higher value is worse and for others better, depending on whether the developer reports the behaviour or its absence. Each value’s record says which.
  • 56 values are restatements: a later document reporting an earlier model again, sometimes with a different value. A restatement is kept beside the model’s own value, never in place of it.

Coverage

Documents from eight of the eight developers report this family. The coverage matrix counts a document when it has at least one printed number in the family; categories, statements in words and figure reads count only in the first column.

Each developer's documents with values in Harmful compliance and over-refusal, out of all its documents in the dataset
DeveloperWith any valueWith a printed numberAll its documents
Anthropic131215
DeepSeek101
Google DeepMind141416
Meta225
Moonshot AI112
OpenAI191922
xAI888
Zhipu AI111

Related findings

No finding is about this family yet.