Overview
Metric families
Every value in the dataset belongs to one of twelve families, each about one property of a model, such as evaluation awareness or jailbreak robustness. A family groups results; it does not make them comparable.
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Who reports what
The 1,645 values in the dataset fall into twelve families, each about one property of a model. Harmful compliance and over-refusal has values from all eight developers, self-preservation from two. Sizes vary as widely, from 606 values in harmful compliance and over-refusal to 4 in self-preservation.
Share of each developer’s documents that report each family
Share of each developer’s documents in the dataset with at least one printed number in each metric family. The number of documents is in brackets after each developer.
| Developer | M1 | M2 | M3 | M4 | M5 | M6 | M7 | M8 | M9 | M10 | M11 | M12 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| OpenAI (22) | 18% | 5% | 18% | 45% | 41% | 5% | 86% | 45% | 55% | 59% | 0% | 18% |
| Google DeepMind (16) | 38% | 0% | 31% | 0% | 13% | 0% | 88% | 13% | 6% | 38% | 0% | 6% |
| Anthropic (15) | 60% | 53% | 53% | 53% | 60% | 13% | 80% | 13% | 67% | 87% | 7% | 20% |
| xAI (8) | 13% | 0% | 13% | 0% | 100% | 88% | 100% | 88% | 63% | 100% | 0% | 0% |
| Meta (5) | 40% | 20% | 40% | 40% | 40% | 40% | 40% | 40% | 60% | 60% | 0% | 0% |
| Moonshot AI (2) | 0% | 0% | 0% | 0% | 0% | 0% | 50% | 50% | 50% | 50% | 0% | 0% |
| DeepSeek (1) | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% |
| Zhipu AI (1) | 0% | 0% | 0% | 0% | 0% | 0% | 100% | 0% | 0% | 0% | 0% | 0% |
Share of documents0%1-24%25-49%50-74%75-100%
- M1
- Evaluation awareness
- M2
- Reward hacking
- M3
- Sabotage and sandbagging
- M4
- Misalignment audits
- M5
- Honesty and hallucination
- M6
- Sycophancy
- M7
- Harmful compliance and over-refusal
- M8
- Jailbreak robustness
- M9
- Prompt injection
- M10
- Dangerous capabilities and risk determinations
- M11
- Self-preservation
- M12
- Chain-of-thought monitorability
The matrix is already a table: every cell prints its share, and its label carries the two counts behind it.
A document counts when it has at least one printed number in the family. Categories (such as a risk level), statements in words and values read off a figure do not count, so four cells show 0% although the developer has values in that family.
Read this as a lower bound. The version 0 reader did not reach the later sections of several long PDFs, so families reported there are undercounted.
Developers are ordered by how many of their documents the dataset holds, not by any value. Each cell opens the explorer on that family and developer.
Source: Safety Card Ledger v0.1 · Data from developer system cards · CC BY 4.0
The twelve families
In the order Methodology §1.1 lists them. Each opens a page with a definition, the evaluations behind the family, its values by kind of test, and which developers report it.
M1
Evaluation awareness
How often a model shows signs of recognizing it is being tested, whether stated in its reasoning or detected by other means.
53 values from 16 evaluations, in documents from Anthropic, Google DeepMind, Meta, OpenAI and xAI.
Kinds of test developer audit transcripts, Apollo Research scenarios, other developer setups, internal states or training signals, real or simulated deployment and UK AI Security Institute scenarios
M2
Reward hacking
Exploiting flaws in tests or graders, special-casing tests, and behavior on deliberately impossible tasks.
63 values from 13 evaluations, in documents from Anthropic, Google DeepMind, Meta and OpenAI.
Kinds of test static benchmarks, scenario evaluations, training monitoring, capability benchmarks, deployment simulations and qualitative assessments
M3
Sabotage and sandbagging
Covertly completing a hidden side task, undermining assigned work, or deliberately underperforming.
58 values from 22 evaluations, in documents from Anthropic, Google DeepMind, Meta, OpenAI and xAI.
Kinds of test capability benchmarks, scenario evaluations, automated audits, qualitative assessments and internal-use monitoring
M4
Misalignment audits
Automated behavioral audits, agentic misalignment scenarios, covert-action and scheming rates, destructive actions.
94 values from 24 evaluations, in documents from Anthropic, Google DeepMind, Meta and OpenAI.
Kinds of test scenario evaluations, automated audits, deployment simulations, framework determinations, internal-use monitoring, static benchmarks and qualitative assessments
M5
Honesty and hallucination
Deception, fabrication, false claims of task completion, hallucination rates.
101 values from 31 evaluations, in documents from Anthropic, Google DeepMind, Meta, OpenAI and xAI.
Kinds of test static benchmarks, scenario evaluations, production traffic and deployment simulations
M6
Sycophancy
Telling users what they want to hear at the expense of accuracy.
23 values from eight evaluations, in documents from Anthropic, Meta, OpenAI and xAI.
Kinds of test static benchmarks, automated audits and production traffic
M7
Harmful compliance and over-refusal
Responses to disallowed requests, and refusals of benign ones.
606 values from 45 evaluations, in documents from Anthropic, DeepSeek, Google DeepMind, Meta, Moonshot AI, OpenAI, xAI and Zhipu AI.
Kinds of test static benchmarks, adaptive attacks, automated audits, framework determinations, red teaming and qualitative assessments
M8
Jailbreak robustness
Resistance to attempts to bypass safeguards.
98 values from 16 evaluations, in documents from Anthropic, Google DeepMind, Meta, Moonshot AI, OpenAI and xAI.
Kinds of test static benchmarks, adaptive attacks, red teaming and qualitative assessments
M9
Prompt injection
Resistance to instructions planted in content an agent reads.
118 values from 25 evaluations, in documents from Anthropic, Google DeepMind, Meta, Moonshot AI, OpenAI and xAI.
Kinds of test static benchmarks, adaptive attacks and a bug bounty
M10
Dangerous capabilities and risk determinations
Capability evaluations in biology, chemistry, cyber and AI research, and the developer's threshold or risk-level decisions.
406 values from 90 evaluations, in documents from Anthropic, Google DeepMind, Meta, Moonshot AI, OpenAI and xAI.
Kinds of test capability benchmarks, framework determinations, uplift studies, red teaming, qualitative assessments, scenario evaluations and surveys
M11
Self-preservation
Shutdown resistance, self-exfiltration, resource or power seeking.
4 values from four evaluations, in documents from Anthropic and OpenAI.
Kinds of test scenario evaluations, internal-use monitoring and qualitative assessments
M12
Chain-of-thought monitorability
Whether a model's written reasoning is legible and faithful enough to monitor.
21 values from nine evaluations, in documents from Anthropic, Google DeepMind and OpenAI.
Kinds of test static benchmarks, scenario evaluations and training monitoring