Tracking the safety results AI developers publish about their own models — every number with its source, its method, and its history.
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
1,645 values · 70 documents · 8 developers
Featured finding · Finding 1 ·
33 reported numbers on evaluation awareness come from six kinds of test
Evaluation-awareness numbers, by the kind of test that produced them
Each point is one number as a developer’s document reports it, in percent: most are the share of transcripts, samples or episodes in which a model showed signs of recognising a test, and each value’s record gives its exact definition. Logarithmic axis.
Hollow: read in a secondary write-up
| Model | Developer | Evaluation | Condition | Value | Document | Kind of test |
|---|---|---|---|---|---|---|
| Claude Opus 4 | Anthropic | Automated behavioral auditTranscripts flagged for situational awareness | of 414 audit transcripts | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Developer audit transcripts | |
| Claude Sonnet 4.5 | Anthropic | Automated behavioral auditShare of transcripts verbalizing eval awareness | — | Claude Sonnet 4.5 System Card29 Sep 2025 | Developer audit transcripts | |
| Claude Haiku 4.5 | Anthropic | Automated behavioral auditShare of scenarios affected by verbalized eval awareness | — | Claude Haiku 4.5 System Card15 Oct 2025 | Developer audit transcripts | |
| Claude Opus 4.6 | Anthropic | Automated behavioral auditShare of transcripts verbalizing eval awareness | — | Claude Opus 4.7 System Card16 Apr 2026 | Developer audit transcripts | |
| Claude Opus 4.7 | Anthropic | Automated behavioral auditShare of transcripts verbalizing eval awareness | — | Claude Opus 4.7 System Card16 Apr 2026 | Developer audit transcripts | |
| Claude Sonnet 5 | Anthropic | Automated behavioral auditShare of transcripts verbalizing eval awareness | — | Claude Sonnet 5 System Card30 Jun 2026 | Developer audit transcripts | |
| Claude Opus 5.5 | Anthropic | Automated behavioral auditShare of transcripts scoring >=6 | — | Claude Opus 5.5 System Card22 Sep 2026 | Developer audit transcripts | |
| Grok 4 | xAI | Alignment auditVerbalized evaluation awareness rate | across all audits | Grok 4.20 System Card7 Apr 2026 | Developer audit transcripts | |
| Grok 4.20 (single-agent) | xAI | Alignment auditVerbalized evaluation awareness rate | single-agent; across all audits | Grok 4.20 System Card7 Apr 2026 | Developer audit transcripts | |
| Claude Opus 4.1 | Anthropic | Evaluation awarenessRun by Apollo ResearchUnambiguous references to being evaluated | — | Claude Sonnet 4.5 System Card29 Sep 2025 | Apollo Research scenarios | |
| Claude Sonnet 4.5 | Anthropic | Evaluation awarenessRun by Apollo ResearchUnambiguous references to being evaluated | — | Claude Sonnet 4.5 System Card29 Sep 2025 | Apollo Research scenarios | |
| gpt-5-thinking | OpenAI | Evaluation awarenessRun by Apollo ResearchVerbalized eval awareness, broad | all samples | GPT-5 System Card7 Aug 2025 | Apollo Research scenarios | |
| GPT-5.3-Codex (pre-release)pre-release | OpenAI | Evaluation awarenessRun by Apollo ResearchSamples with moderate-or-higher eval awareness | — | GPT-5.5 System Card23 Apr 2026 | Apollo Research scenarios | |
| GPT-5.4 (pre-release)pre-release | OpenAI | Evaluation awarenessRun by Apollo ResearchSamples with moderate-or-higher eval awareness | — | GPT-5.5 System Card23 Apr 2026 | Apollo Research scenarios | |
| GPT-5.4 Thinking | OpenAI | Evaluation awarenessRun by Apollo ResearchSamples with verbalized eval awareness | — | GPT-5.4 Thinking System Card5 Mar 2026 | Apollo Research scenarios | |
| GPT-5.5 | OpenAI | Evaluation awarenessRun by Apollo ResearchSamples with moderate-or-higher eval awareness | — | GPT-5.5 System Card23 Apr 2026 | Apollo Research scenarios | |
| GPT-5.5 | OpenAI | Evaluation awarenessRun by Apollo ResearchSamples with verbalized eval awareness | xhigh | GPT-6 Astra System Card3 Sep 2026 | Apollo Research scenarios | |
| GPT-6 Astra | OpenAI | Evaluation awarenessRun by Apollo ResearchSamples with verbalized eval awareness | xhigh | GPT-6 Astra System Card3 Sep 2026 | Apollo Research scenarios | |
| GPT-6 Astra | OpenAI | Evaluation awarenessRun by Apollo ResearchSamples with verbalized eval awareness | max | GPT-6 Astra System Card3 Sep 2026 | Apollo Research scenarios | |
| Claude Opus 4.1 | Anthropic | Evaluation awarenessRun by UK AI Security InstituteUnprompted verbalized eval awareness rate | — | Claude Sonnet 4.5 System Card29 Sep 2025 | UK AISI scenarios | |
| Claude Sonnet 4 | Anthropic | Evaluation awarenessRun by UK AI Security InstituteUnprompted verbalized eval awareness rate | — | Claude Sonnet 4.5 System Card29 Sep 2025 | UK AISI scenarios | |
| Claude Sonnet 4.5 | Anthropic | Evaluation awarenessRun by UK AI Security InstituteUnprompted verbalized eval awareness rate | — | Claude Sonnet 4.5 System Card29 Sep 2025 | UK AISI scenarios | |
| Muse Spark | Meta | Verbalized eval awarenessShare of samples verbalizing evaluation | internal evals | Muse Spark Safety & Preparedness Report8 Apr 2026 | Other developer setups | |
| Muse Spark 1.1 | Meta | Verbalized eval awarenessShare of samples verbalizing evaluation | dangerous-capability evals | Muse Spark 1.1 Evaluation Report9 Jul 2026 | Other developer setups | |
| Muse Spark | Meta | Verbalized eval awarenessShare of samples verbalizing evaluation | public benchmarks | Muse Spark Safety & Preparedness Report8 Apr 2026 | Other developer setups | |
| Muse Spark | Meta | Prompted evaluation awarenessEval-vs-deployment discrimination | — | Muse Spark Safety & Preparedness Report8 Apr 2026 | Other developer setups | |
| Muse Spark 1.1 | Meta | Verbalized eval awarenessShare of samples verbalizing evaluation | Agentic Misalignment scenarios | Muse Spark 1.1 Evaluation Report9 Jul 2026 | Other developer setups | |
| Claude Opus 5.5 | Anthropic | Evaluation awareness (deployment)Share of transcripts scoring >=6 | real internal deployment | Claude Opus 5.5 System Card22 Sep 2026 | Real or simulated deployment | |
| Original production traffic (mixed models) | OpenAI | Deployment simulation (internal Codex traffic)Trajectories verbalizing evaluation awareness | original production traffic | GPT-6 Astra System Card3 Sep 2026 | Real or simulated deployment | |
| GPT-5.6 Sol | OpenAI | Deployment simulation (internal Codex traffic)Trajectories verbalizing evaluation awareness | shared task set | GPT-6 Astra System Card3 Sep 2026 | Real or simulated deployment | |
| GPT-6 Astra | OpenAI | Deployment simulation (internal Codex traffic)Trajectories verbalizing evaluation awareness | shared task set | GPT-6 Astra System Card3 Sep 2026 | Real or simulated deployment | |
| Claude Opus 4.8 | Anthropic | Verbalized grader speculationShare of training episodes | RL training | Claude Opus 4.8 System Card28 May 2026 | Internal states or training | |
| Claude Opus 4.8 | Anthropic | Unverbalized grader awarenessShare of cases with grader awareness | — | Claude Opus 4.8 System Card28 May 2026 | Internal states or training |
Each point is one value, placed by the number the document printed. Bands are kinds of test; within a band, points are spread up and down only so that they do not overlap. Points in different bands come from different tests and are not comparable with one another.
Marker shape is the developer whose document reports the value; a hollow marker is a value we read in an independent write-up of the document, because the section could not be read directly.
30 of the 33 values are plotted. The value for Claude Opus 4 is a count; the values for Grok 4 and Grok 4.20 (single-agent) are rates from 0 to 1. Placing them on this axis would mean rescaling a printed value, so they are in the table only.
The lowest and the highest plotted values are labelled.
Logarithmic axis: the values span more than one and a half orders of magnitude.
Source: Safety Card Ledger v0.1 · Data from developer system cards · CC BY 4.0
How often a model shows signs of recognising a test is measured in six different ways, and the same model gets different numbers from different kinds of test.
Read the finding: 33 reported numbers on evaluation awareness come from six kinds of test
Recent revisions
The latest changes developers made to their documents after publishing them, where a value is recorded on both sides and both have been checked. The ledger holds all 15 revisions, including changes of wording and documents marked updated with no record of what changed.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
xAI · 3 of its 5 changes
HackerBench v0.2 harmful/dual-use compliance, Grok 4.6 (high)
16.7%6.9%Partly explained — changelog lists "corrected eval results" for this eval; no reason
Old value New value In the ledger: HackerBench v0.2 harmful/dual-use compliance, Grok 4.6 (high)
MASK dishonesty, Grok 4.6 (high)
3.8%1.90%Partly explained — listed as corrected; no reason
Old value New value In the ledger: MASK dishonesty, Grok 4.6 (high)
Self-harm compliance, Grok 4.6 (high)
3.7%0.84%Partly explained — listed as corrected; no reason
Old value New value In the ledger: Self-harm compliance, Grok 4.6 (high)
Findings
What the data shows, one chart each, with the values behind it and how they were chosen and checked.
33 reported numbers on evaluation awareness come from six kinds of test
How often a model shows signs of recognising a test is measured in six different ways, and the same model gets different numbers from different kinds of test.
Not written yet: planned for Phase 3
- Finding 2Later documents report earlier models again, sometimes with different values (not written yet)
- Finding 3Documents change after release, with and without a changelog (not written yet)
- Finding 4Within one developer and one test, values form a series until the test changes (not written yet)
- Finding 5Who reports what (not written yet)
Coverage at a glance
Which metric families each developer’s documents report with a printed number. This describes what developers publish, not the safety of their models.
Metric families reported, by developer
Share of each developer’s documents with at least one printed number in each metric family. Developers are ordered by number of documents (in brackets), then by name. A cell opens the data explorer at that family and developer.
| Developer | M1 | M2 | M3 | M4 | M5 | M6 | M7 | M8 | M9 | M10 | M11 | M12 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| OpenAI (22) | 18% | 5% | 18% | 45% | 41% | 5% | 86% | 45% | 55% | 59% | 0% | 18% |
| Google DeepMind (16) | 38% | 0% | 31% | 0% | 13% | 0% | 88% | 13% | 6% | 38% | 0% | 6% |
| Anthropic (15) | 60% | 53% | 53% | 53% | 60% | 13% | 80% | 13% | 67% | 87% | 7% | 20% |
| xAI (8) | 13% | 0% | 13% | 0% | 100% | 88% | 100% | 88% | 63% | 100% | 0% | 0% |
| Meta (5) | 40% | 20% | 40% | 40% | 40% | 40% | 40% | 40% | 60% | 60% | 0% | 0% |
| Moonshot AI (2) | 0% | 0% | 0% | 0% | 0% | 0% | 50% | 50% | 50% | 50% | 0% | 0% |
| DeepSeek (1) | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% |
| Zhipu AI (1) | 0% | 0% | 0% | 0% | 0% | 0% | 100% | 0% | 0% | 0% | 0% | 0% |
Share of documents0%1-24%25-49%50-74%75-100%
- M1
- Evaluation awareness
- M2
- Reward hacking
- M3
- Sabotage and sandbagging
- M4
- Misalignment audits
- M5
- Honesty and hallucination
- M6
- Sycophancy
- M7
- Harmful compliance and over-refusal
- M8
- Jailbreak robustness
- M9
- Prompt injection
- M10
- Dangerous capabilities and risk determinations
- M11
- Self-preservation
- M12
- Chain-of-thought monitorability
The chart is already a table: every cell prints its share, and its label gives the two counts behind it. The CSV has the counts.
A document counts for a family when it prints at least one number in it. Values read off a figure, stated in words, or given as a category such as a risk level do not count. Some documents were not read to the end in version 0, so read the shares as a lower bound.
Source: Safety Card Ledger v0.1 · Data from developer system cards · CC BY 4.0
Recently added or revised documents
Ordered by the latest date we know for each document: its publication date, or the date of a later version when the document’s changelog or header gives one. The day we retrieved a copy does not count.
About
Safety Card Ledger records the safety evaluation results that AI developers publish about their own models: each number as the document prints it, where it was read, how it was checked, and how it changed when a document was revised. Values are recorded as published by each developer. Safety Card Ledger does not rate the safety of models. Developers are never ranked on the numbers they report.