33 reported numbers on evaluation awareness come from six kinds of test
How often a model shows signs of recognising a test is measured in six different ways, and the same model gets different numbers from different kinds of test.
- Published
- Last updated
- Metric family
- M1 · Evaluation awareness
- Values
- 33 values from 14 documents
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
A number for evaluation awareness means little without the test that produced it: the same model can get numbers more than ten times apart from different tests. Meta’s report on Muse Spark gives for the share of samples in which the model said it suspected an evaluation, across Meta’s internal evaluations, and on a test that asks the model to tell evaluation transcripts from real use.
Evaluation awareness is how often a model shows signs of recognising that it is being tested. A model that recognises a test may act differently from how it would act in real use, and that matters because most safety results come from tests.
This finding looks at the 33 numbers the feasibility study selected. They are reported in 14 documents from four developers, and they come from six kinds of test: developer audit transcripts, Apollo Research scenarios, UK AI Security Institute scenarios, other developer setups, real or simulated deployment and internal states or training signals. Each kind counts something different, over different transcripts, samples or episodes. The chart places each number in the band for its kind of test; numbers in different bands are not comparable.
Evaluation-awareness numbers, by the kind of test that produced them
Each point is one number as a developer’s document reports it, in percent: most are the share of transcripts, samples or episodes in which a model showed signs of recognising a test, and each value’s record gives its exact definition. Logarithmic axis.
Hollow: read in a secondary write-up
| Model | Developer | Evaluation | Condition | Value | Document | Kind of test |
|---|---|---|---|---|---|---|
| Claude Opus 4 | Anthropic | Automated behavioral auditTranscripts flagged for situational awareness | of 414 audit transcripts | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Developer audit transcripts | |
| Claude Sonnet 4.5 | Anthropic | Automated behavioral auditShare of transcripts verbalizing eval awareness | — | Claude Sonnet 4.5 System Card29 Sep 2025 | Developer audit transcripts | |
| Claude Haiku 4.5 | Anthropic | Automated behavioral auditShare of scenarios affected by verbalized eval awareness | — | Claude Haiku 4.5 System Card15 Oct 2025 | Developer audit transcripts | |
| Claude Opus 4.6 | Anthropic | Automated behavioral auditShare of transcripts verbalizing eval awareness | — | Claude Opus 4.7 System Card16 Apr 2026 | Developer audit transcripts | |
| Claude Opus 4.7 | Anthropic | Automated behavioral auditShare of transcripts verbalizing eval awareness | — | Claude Opus 4.7 System Card16 Apr 2026 | Developer audit transcripts | |
| Claude Sonnet 5 | Anthropic | Automated behavioral auditShare of transcripts verbalizing eval awareness | — | Claude Sonnet 5 System Card30 Jun 2026 | Developer audit transcripts | |
| Claude Opus 5.5 | Anthropic | Automated behavioral auditShare of transcripts scoring >=6 | — | Claude Opus 5.5 System Card22 Sep 2026 | Developer audit transcripts | |
| Grok 4 | xAI | Alignment auditVerbalized evaluation awareness rate | across all audits | Grok 4.20 System Card7 Apr 2026 | Developer audit transcripts | |
| Grok 4.20 (single-agent) | xAI | Alignment auditVerbalized evaluation awareness rate | single-agent; across all audits | Grok 4.20 System Card7 Apr 2026 | Developer audit transcripts | |
| Claude Opus 4.1 | Anthropic | Evaluation awarenessRun by Apollo ResearchUnambiguous references to being evaluated | — | Claude Sonnet 4.5 System Card29 Sep 2025 | Apollo Research scenarios | |
| Claude Sonnet 4.5 | Anthropic | Evaluation awarenessRun by Apollo ResearchUnambiguous references to being evaluated | — | Claude Sonnet 4.5 System Card29 Sep 2025 | Apollo Research scenarios | |
| gpt-5-thinking | OpenAI | Evaluation awarenessRun by Apollo ResearchVerbalized eval awareness, broad | all samples | GPT-5 System Card7 Aug 2025 | Apollo Research scenarios | |
| GPT-5.3-Codex (pre-release)pre-release | OpenAI | Evaluation awarenessRun by Apollo ResearchSamples with moderate-or-higher eval awareness | — | GPT-5.5 System Card23 Apr 2026 | Apollo Research scenarios | |
| GPT-5.4 (pre-release)pre-release | OpenAI | Evaluation awarenessRun by Apollo ResearchSamples with moderate-or-higher eval awareness | — | GPT-5.5 System Card23 Apr 2026 | Apollo Research scenarios | |
| GPT-5.4 Thinking | OpenAI | Evaluation awarenessRun by Apollo ResearchSamples with verbalized eval awareness | — | GPT-5.4 Thinking System Card5 Mar 2026 | Apollo Research scenarios | |
| GPT-5.5 | OpenAI | Evaluation awarenessRun by Apollo ResearchSamples with moderate-or-higher eval awareness | — | GPT-5.5 System Card23 Apr 2026 | Apollo Research scenarios | |
| GPT-5.5 | OpenAI | Evaluation awarenessRun by Apollo ResearchSamples with verbalized eval awareness | xhigh | GPT-6 Astra System Card3 Sep 2026 | Apollo Research scenarios | |
| GPT-6 Astra | OpenAI | Evaluation awarenessRun by Apollo ResearchSamples with verbalized eval awareness | xhigh | GPT-6 Astra System Card3 Sep 2026 | Apollo Research scenarios | |
| GPT-6 Astra | OpenAI | Evaluation awarenessRun by Apollo ResearchSamples with verbalized eval awareness | max | GPT-6 Astra System Card3 Sep 2026 | Apollo Research scenarios | |
| Claude Opus 4.1 | Anthropic | Evaluation awarenessRun by UK AI Security InstituteUnprompted verbalized eval awareness rate | — | Claude Sonnet 4.5 System Card29 Sep 2025 | UK AISI scenarios | |
| Claude Sonnet 4 | Anthropic | Evaluation awarenessRun by UK AI Security InstituteUnprompted verbalized eval awareness rate | — | Claude Sonnet 4.5 System Card29 Sep 2025 | UK AISI scenarios | |
| Claude Sonnet 4.5 | Anthropic | Evaluation awarenessRun by UK AI Security InstituteUnprompted verbalized eval awareness rate | — | Claude Sonnet 4.5 System Card29 Sep 2025 | UK AISI scenarios | |
| Muse Spark | Meta | Verbalized eval awarenessShare of samples verbalizing evaluation | internal evals | Muse Spark Safety & Preparedness Report8 Apr 2026 | Other developer setups | |
| Muse Spark 1.1 | Meta | Verbalized eval awarenessShare of samples verbalizing evaluation | dangerous-capability evals | Muse Spark 1.1 Evaluation Report9 Jul 2026 | Other developer setups | |
| Muse Spark | Meta | Verbalized eval awarenessShare of samples verbalizing evaluation | public benchmarks | Muse Spark Safety & Preparedness Report8 Apr 2026 | Other developer setups | |
| Muse Spark | Meta | Prompted evaluation awarenessEval-vs-deployment discrimination | — | Muse Spark Safety & Preparedness Report8 Apr 2026 | Other developer setups | |
| Muse Spark 1.1 | Meta | Verbalized eval awarenessShare of samples verbalizing evaluation | Agentic Misalignment scenarios | Muse Spark 1.1 Evaluation Report9 Jul 2026 | Other developer setups | |
| Claude Opus 5.5 | Anthropic | Evaluation awareness (deployment)Share of transcripts scoring >=6 | real internal deployment | Claude Opus 5.5 System Card22 Sep 2026 | Real or simulated deployment | |
| Original production traffic (mixed models) | OpenAI | Deployment simulation (internal Codex traffic)Trajectories verbalizing evaluation awareness | original production traffic | GPT-6 Astra System Card3 Sep 2026 | Real or simulated deployment | |
| GPT-5.6 Sol | OpenAI | Deployment simulation (internal Codex traffic)Trajectories verbalizing evaluation awareness | shared task set | GPT-6 Astra System Card3 Sep 2026 | Real or simulated deployment | |
| GPT-6 Astra | OpenAI | Deployment simulation (internal Codex traffic)Trajectories verbalizing evaluation awareness | shared task set | GPT-6 Astra System Card3 Sep 2026 | Real or simulated deployment | |
| Claude Opus 4.8 | Anthropic | Verbalized grader speculationShare of training episodes | RL training | Claude Opus 4.8 System Card28 May 2026 | Internal states or training | |
| Claude Opus 4.8 | Anthropic | Unverbalized grader awarenessShare of cases with grader awareness | — | Claude Opus 4.8 System Card28 May 2026 | Internal states or training |
Each point is one value, placed by the number the document printed. Bands are kinds of test; within a band, points are spread up and down only so that they do not overlap. Points in different bands come from different tests and are not comparable with one another.
Marker shape is the developer whose document reports the value; a hollow marker is a value we read in an independent write-up of the document, because the section could not be read directly.
30 of the 33 values are plotted. The value for Claude Opus 4 is a count; the values for Grok 4 and Grok 4.20 (single-agent) are rates from 0 to 1. Placing them on this axis would mean rescaling a printed value, so they are in the table only.
The lowest and the highest plotted values are labelled.
Logarithmic axis: the values span more than one and a half orders of magnitude.
Source: Safety Card Ledger v0.1 · Data from developer system cards · CC BY 4.0
One model, three tests
The Claude Sonnet 4.5 system card reports three numbers for the model. Anthropic’s automated behavioural audit found the model saying it might be under evaluation in of audit transcripts. The UK AI Security Institute reported an unprompted rate of in its scenarios, and Apollo Research found unambiguous references to being evaluated in of its scenario transcripts. Both outside evaluators tested pre-release snapshots of the model.
We read all three numbers in independent write-ups of the card, so the chart draws them as hollow markers. The first has since been checked against the card itself; the other two were checked against the write-up.
One document, two settings
The Claude Opus 5.5 system card reports one measure, the share of transcripts that reach a set score for evaluation awareness, in two settings: in Anthropic’s automated audit, and in real internal deployment. We read both in a write-up of the card and checked them against that write-up, not the card.
Inside one document
Apollo Research’s numbers for OpenAI models are measured one way within each document. In the GPT-5.5 system card, Apollo reports the share of samples with moderate or higher evaluation awareness: for a pre-release GPT-5.3-Codex, for a pre-release GPT-5.4 and for GPT-5.5.
The GPT-6 Astra system card reports a different measure, the share of samples with verbalized evaluation awareness at a stated reasoning effort: for GPT-5.5 and for GPT-6 Astra, both at the effort recorded as “xhigh”. GPT-5.5 therefore has a number in each document, on different measures. Inside each document the models were measured the same way, so their numbers can be compared there; across the two documents they cannot, and the explorer keeps them as separate series.
The explorer lists every value in the family, one evaluation at a time, with its source and its checks.
Cite this
APA
Safety Card Ledger. (2026, September 28). 33 reported numbers on evaluation awareness come from six kinds of test. Safety Card Ledger. https://safetycardledger.example/findings/evaluation-awareness
BibTeX
@misc{safetycardledger_evaluation_awareness,
author = {{Safety Card Ledger}},
title = {33 reported numbers on evaluation awareness come from six kinds of test},
year = {2026},
month = sep,
howpublished = {\url{https://safetycardledger.example/findings/evaluation-awareness}},
note = {Safety Card Ledger, dataset v0.1},
}