Evaluation awareness
How often a model shows signs of recognizing it is being tested, whether stated in its reasoning or detected by other means.
- Reported by
- Anthropic, Google DeepMind, Meta, OpenAI and xAI
- Evaluations
- 16 evaluations, of six kinds of test
- Values
- 53 values in 25 documents
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Definition
Evaluation awareness is about whether a model behaves as if it knows it is being tested. A model that recognises a test may act differently from how it would act in real use, and that matters because most safety results come from tests.
Developers look for it in several ways. The most common is to read the model’s written reasoning, or transcripts of its behaviour, for statements that it suspects an evaluation. Others ask the model directly whether a transcript looks like a test, build tasks that can only be solved by noticing features of the setup, examine signals inside the model, or track how often such statements appear during training. Independent evaluators run their own scenarios and report their own rates.
A higher value usually means more signs of awareness, but each method counts something different: a share of transcripts, of samples or of training episodes, found by a different reader in a different set of tasks.
The line under the title is this family’s entry in Methodology §1.1. A family groups results about one property; it does not make them comparable (Methodology §5).
How developers measure it
16 evaluations of six kinds of test, with the most values first. Each evaluation opens in the data explorer, which lists all of its values.
| Evaluation | Reported by | Values |
|---|---|---|
| Developer audit transcripts · 4 evaluations, 16 values | ||
| Automated behavioral audit, in the data explorer Run by the developer Signs that the model recognises it is being tested, found in transcripts from Anthropic's automated behavioural audit, where an auditor model probes the model in constructed scenarios.Reported by Anthropic · 8 values | Anthropic | 8 |
| Verbalized metagaming, in the data explorer Run by the developer A monitor reads the model's reasoning in OpenAI's evaluation and simulated deployment environments and flags reasoning about being evaluated or gaming the evaluation.Reported by OpenAI · 4 values | OpenAI | 4 |
| Alignment audit, in the data explorer Run by the developer Share of transcripts from xAI's automated alignment audit in which the model says it may be under evaluation.Reported by xAI · 2 values | xAI | 2 |
| Evaluation awareness · Google DeepMind, in the data explorer Run by the developer Qualitative findings on whether the model says or shows that it recognises it is being tested, from transcripts of Google's own evaluations.Reported by Google DeepMind · 2 values | Google DeepMind | 2 |
| Apollo Research scenarios · 1 evaluation, 14 values | ||
| Evaluation awareness · Apollo Research, in the data explorer Run by Apollo Research (third party) Apollo Research's measure of how often a model's reasoning shows that it recognises it is being evaluated in Apollo's own test scenarios.Reported by Anthropic and OpenAI · 14 values | Anthropic and OpenAI | 14 |
| Other developer setups · 3 evaluations, 11 values | ||
| Situational awareness challenges, in the data explorer Run by the developer Agentic challenges run by Google in which the model can succeed only by noticing and changing parts of its own setup, such as its context size or oversight settings.Reported by Google DeepMind · 6 values | Google DeepMind | 6 |
| Verbalized eval awareness, in the data explorer Run by the developer Share of the model's samples on a given set of evaluations in which it says it suspects it is being evaluated.Reported by Meta · 4 values | Meta | 4 |
| Prompted evaluation awareness, in the data explorer Run by the developer The model is asked whether transcripts come from an evaluation or from real use, and is scored on how well it tells them apart.Reported by Meta · 1 value | Meta | 1 |
| Internal states or training · 5 evaluations, 5 values | ||
| Evaluation awareness in audit transcripts, in the data explorer Run by the developer Share of audit transcripts in which the model is aware of being evaluated, detected from its internal states rather than from what it writes.Reported by Anthropic · 1 value | Anthropic | 1 |
| Grader awareness, in the data explorer Run by the developer Share of training episodes in which the model shows awareness of its grader, split by how exposed the grader is to hacking.Reported by Anthropic · 1 value | Anthropic | 1 |
| Unverbalized grader awareness, in the data explorer Run by the developer Share of cases in which the model is aware of its grader without saying so, found by analysing more than its written output.Reported by Anthropic · 1 value | Anthropic | 1 |
| Verbalized grader speculation, in the data explorer Run by the developer Share of reinforcement-learning training episodes in which the model speculates in writing about how it is being graded.Reported by Anthropic · 1 value | Anthropic | 1 |
| Verbalized metagaming in training, in the data explorer Run by the developer Tracks the rate of verbalized metagaming in the model's reasoning over the course of training.Reported by OpenAI · 1 value | OpenAI | 1 |
| Real or simulated deployment · 2 evaluations, 4 values | ||
| Deployment simulation (internal Codex traffic), in the data explorer Run by the developer Share of resampled internal Codex trajectories in which the model says it may be under evaluation, with the original traffic as a baseline.Reported by OpenAI · 3 values | OpenAI | 3 |
| Evaluation awareness (deployment), in the data explorer Run by the developer Share of transcripts from real internal use of the model that reach a set level on Anthropic's evaluation-awareness score.Reported by Anthropic · 1 value | Anthropic | 1 |
| UK AISI scenarios · 1 evaluation, 3 values | ||
| Evaluation awareness · UK AI Security Institute, in the data explorer Run by UK AI Security Institute (government) The UK AI Security Institute's rate of a model saying, unprompted, that it may be under evaluation in the Institute's test scenarios.Reported by Anthropic · 3 values | Anthropic | 3 |
Reported values
The values a chart can place, by the kind of test that produced them. Open a point, or a value in the table, for its source and history. Open in explorer: Evaluation awareness
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
Evaluation awareness: values by kind of test
Each point is a value reported in a developer’s document, in percent (higher is worse), grouped by the kind of test that produced it. Logarithmic axis.
Hollow: read in a secondary write-up
| Model | Developer | Evaluation | Condition | Value | Document | Kind of test |
|---|---|---|---|---|---|---|
| Claude Sonnet 4.5 | Anthropic | Automated behavioral auditShare of transcripts verbalizing eval awareness | — | Claude Sonnet 4.5 System Card29 Sep 2025 | Developer audit transcripts | |
| Claude Haiku 4.5 | Anthropic | Automated behavioral auditShare of scenarios affected by verbalized eval awareness | — | Claude Haiku 4.5 System Card15 Oct 2025 | Developer audit transcripts | |
| Claude Opus 4.7 | Anthropic | Automated behavioral auditShare of transcripts verbalizing eval awareness | — | Claude Opus 4.7 System Card16 Apr 2026 | Developer audit transcripts | |
| Claude Opus 4.6 | Anthropic | Automated behavioral auditShare of transcripts verbalizing eval awareness | — | Claude Opus 4.7 System Card16 Apr 2026 | Developer audit transcripts | |
| Claude Sonnet 5 | Anthropic | Automated behavioral auditShare of transcripts verbalizing eval awareness | — | Claude Sonnet 5 System Card30 Jun 2026 | Developer audit transcripts | |
| Claude Opus 5.5 | Anthropic | Automated behavioral auditShare of transcripts scoring >=6 | — | Claude Opus 5.5 System Card22 Sep 2026 | Developer audit transcripts | |
| Claude Sonnet 4.5 | Anthropic | Evaluation awarenessRun by Apollo ResearchUnambiguous references to being evaluated | — | Claude Sonnet 4.5 System Card29 Sep 2025 | Apollo Research scenarios | |
| Claude Opus 4.1 | Anthropic | Evaluation awarenessRun by Apollo ResearchUnambiguous references to being evaluated | — | Claude Sonnet 4.5 System Card29 Sep 2025 | Apollo Research scenarios | |
| gpt-5-thinking | OpenAI | Evaluation awarenessRun by Apollo ResearchVerbalized eval awareness, broad | all samples | GPT-5 System Card7 Aug 2025 | Apollo Research scenarios | |
| gpt-5-thinking | OpenAI | Evaluation awarenessRun by Apollo ResearchVerbalized eval awareness, broad | scheming samples | GPT-5 System Card7 Aug 2025 | Apollo Research scenarios | |
| gpt-5-thinking | OpenAI | Evaluation awarenessRun by Apollo ResearchVerbalized eval awareness, strict | all samples | GPT-5 System Card7 Aug 2025 | Apollo Research scenarios | |
| gpt-5-thinking | OpenAI | Evaluation awarenessRun by Apollo ResearchVerbalized eval awareness, strict | scheming samples | GPT-5 System Card7 Aug 2025 | Apollo Research scenarios | |
| GPT-5.4 Thinking | OpenAI | Evaluation awarenessRun by Apollo ResearchSamples with verbalized eval awareness | — | GPT-5.4 Thinking System Card5 Mar 2026 | Apollo Research scenarios | |
| GPT-5.5 | OpenAI | Evaluation awarenessRun by Apollo ResearchSamples with moderate-or-higher eval awareness | — | GPT-5.5 System Card23 Apr 2026 | Apollo Research scenarios | |
| GPT-5.4 (pre-release)pre-release | OpenAI | Evaluation awarenessRun by Apollo ResearchSamples with moderate-or-higher eval awareness | — | GPT-5.5 System Card23 Apr 2026 | Apollo Research scenarios | |
| GPT-5.3-Codex (pre-release)pre-release | OpenAI | Evaluation awarenessRun by Apollo ResearchSamples with moderate-or-higher eval awareness | — | GPT-5.5 System Card23 Apr 2026 | Apollo Research scenarios | |
| GPT-6 Astra | OpenAI | Evaluation awarenessRun by Apollo ResearchSamples with verbalized eval awareness | xhigh | GPT-6 Astra System Card3 Sep 2026 | Apollo Research scenarios | |
| GPT-6 Astra | OpenAI | Evaluation awarenessRun by Apollo ResearchSamples with verbalized eval awareness | max | GPT-6 Astra System Card3 Sep 2026 | Apollo Research scenarios | |
| GPT-5.5 | OpenAI | Evaluation awarenessRun by Apollo ResearchSamples with verbalized eval awareness | xhigh | GPT-6 Astra System Card3 Sep 2026 | Apollo Research scenarios | |
| Gemini 3.1 Pro | Google DeepMind | Situational awareness challengesSuccess rate on three hardest challenges | Deep Think | Gemini 3.1 Pro Model Card19 Feb 2026 | Other developer setups | |
| Muse Spark | Meta | Prompted evaluation awarenessEval-vs-deployment discrimination | — | Muse Spark Safety & Preparedness Report8 Apr 2026 | Other developer setups | |
| Muse Spark | Meta | Verbalized eval awarenessShare of samples verbalizing evaluation | public benchmarks | Muse Spark Safety & Preparedness Report8 Apr 2026 | Other developer setups | |
| Muse Spark | Meta | Verbalized eval awarenessShare of samples verbalizing evaluation | internal evals | Muse Spark Safety & Preparedness Report8 Apr 2026 | Other developer setups | |
| Muse Spark 1.1 | Meta | Verbalized eval awarenessShare of samples verbalizing evaluation | Agentic Misalignment scenarios | Muse Spark 1.1 Evaluation Report9 Jul 2026 | Other developer setups | |
| Muse Spark 1.1 | Meta | Verbalized eval awarenessShare of samples verbalizing evaluation | dangerous-capability evals | Muse Spark 1.1 Evaluation Report9 Jul 2026 | Other developer setups | |
| Claude Opus 4.8 | Anthropic | Unverbalized grader awarenessShare of cases with grader awareness | — | Claude Opus 4.8 System Card28 May 2026 | Internal states or training | |
| Claude Opus 4.8 | Anthropic | Verbalized grader speculationShare of training episodes | RL training | Claude Opus 4.8 System Card28 May 2026 | Internal states or training | |
| Claude Opus 5.5 | Anthropic | Evaluation awareness (deployment)Share of transcripts scoring >=6 | real internal deployment | Claude Opus 5.5 System Card22 Sep 2026 | Real or simulated deployment | |
| GPT-6 Astra | OpenAI | Deployment simulation (internal Codex traffic)Trajectories verbalizing evaluation awareness | shared task set | GPT-6 Astra System Card3 Sep 2026 | Real or simulated deployment | |
| GPT-5.6 Sol | OpenAI | Deployment simulation (internal Codex traffic)Trajectories verbalizing evaluation awareness | shared task set | GPT-6 Astra System Card3 Sep 2026 | Real or simulated deployment | |
| Original production traffic (mixed models) | OpenAI | Deployment simulation (internal Codex traffic)Trajectories verbalizing evaluation awareness | original production traffic | GPT-6 Astra System Card3 Sep 2026 | Real or simulated deployment | |
| Claude Sonnet 4.5 | Anthropic | Evaluation awarenessRun by UK AI Security InstituteUnprompted verbalized eval awareness rate | — | Claude Sonnet 4.5 System Card29 Sep 2025 | UK AISI scenarios | |
| Claude Sonnet 4 | Anthropic | Evaluation awarenessRun by UK AI Security InstituteUnprompted verbalized eval awareness rate | — | Claude Sonnet 4.5 System Card29 Sep 2025 | UK AISI scenarios | |
| Claude Opus 4.1 | Anthropic | Evaluation awarenessRun by UK AI Security InstituteUnprompted verbalized eval awareness rate | — | Claude Sonnet 4.5 System Card29 Sep 2025 | UK AISI scenarios |
Marker shape is the developer whose document reports the value; a hollow marker is a value we read in an independent write-up because the card section could not be read directly.
Each point is one value, placed by the number the document printed. Bands are kinds of test; within a band, points are spread up and down only so that they do not overlap.
34 of the family’s 53 values are plotted: those in percent where higher is worse, the unit most of its evaluations use. Of the rest, 8 in other units or read the other way (counts and rates from 0 to 1), 6 are categories or statements in words and 5 are low-confidence or read off a figure, as in the explorer’s default view. The explorer lists them all, one evaluation at a time.
Logarithmic axis: the values span more than one and a half orders of magnitude.
Source: Safety Card Ledger v0.1 · Data from developer system cards · CC BY 4.0
Comparability
- Each of the 16 evaluations is its own test, with its own tasks, graders and definitions. Values from different evaluations describe the same property measured in different ways, and are not comparable with one another, even where two evaluations share a name.
- Values come from the documents of five developers. Developers use different tests, prompts and graders for the same property, so we do not rank them on these numbers: tables here are ordered by name or by kind of test, never by a reported value.
- Some names are used by more than one evaluation: “Evaluation awareness” by three. They are different tests, run by different evaluators or reported by different developers, so this page names who ran each one.
- In two evaluations the test changed between documents (Evaluation awareness · Apollo Research and Automated behavioral audit). Values on the old and new definitions are never joined: the explorer breaks the line where the test changed.
- In four evaluations the explorer connects results from different documents that do not say whether the test stayed the same. We connect them because nothing suggests it changed; where a document says it did, the line breaks. Each value on such a line is marked “test assumed unchanged”.
- Seven evaluations report values under more than one condition, such as reasoning effort, safeguards or number of attempts. Each condition is kept as a separate series.
- Numbers in this family are printed in three units: counts, percent and rates from 0 to 1. We never convert one unit into another, so a chart shows one unit at a time.
- One value is a restatement: a later document reporting an earlier model again, sometimes with a different value. A restatement is kept beside the model’s own value, never in place of it.
Coverage
Documents from five of the eight developers report this family. The coverage matrix counts a document when it has at least one printed number in the family; categories, statements in words and figure reads count only in the first column.
| Developer | With any value | With a printed number | All its documents |
|---|---|---|---|
| Anthropic | 9 | 9 | 15 |
| DeepSeek | 0 | 0 | 1 |
| Google DeepMind | 6 | 6 | 16 |
| Meta | 2 | 2 | 5 |
| Moonshot AI | 0 | 0 | 2 |
| OpenAI | 7 | 4 | 22 |
| xAI | 1 | 1 | 8 |
| Zhipu AI | 0 | 0 | 1 |
Related findings
- Finding 1: 33 reported numbers on evaluation awareness come from six kinds of testHow often a model shows signs of recognising a test is measured in six different ways, and the same model gets different numbers from different kinds of test.