Definition

Evaluation awareness is about whether a model behaves as if it knows it is being tested. A model that recognises a test may act differently from how it would act in real use, and that matters because most safety results come from tests.

Developers look for it in several ways. The most common is to read the model’s written reasoning, or transcripts of its behaviour, for statements that it suspects an evaluation. Others ask the model directly whether a transcript looks like a test, build tasks that can only be solved by noticing features of the setup, examine signals inside the model, or track how often such statements appear during training. Independent evaluators run their own scenarios and report their own rates.

A higher value usually means more signs of awareness, but each method counts something different: a share of transcripts, of samples or of training episodes, found by a different reader in a different set of tasks.

The line under the title is this family’s entry in Methodology §1.1. A family groups results about one property; it does not make them comparable (Methodology §5).

How developers measure it

16 evaluations of six kinds of test, with the most values first. Each evaluation opens in the data explorer, which lists all of its values.

The 16 evaluations in Evaluation awareness, grouped by kind of test: who ran each one, what it counts, the developers whose documents report it and how many values it has.
EvaluationReported byValues
Developer audit transcripts · 4 evaluations, 16 values
Automated behavioral audit, in the data explorer Run by the developer Signs that the model recognises it is being tested, found in transcripts from Anthropic's automated behavioural audit, where an auditor model probes the model in constructed scenarios.Reported by Anthropic · 8 valuesAnthropic8
Verbalized metagaming, in the data explorer Run by the developer A monitor reads the model's reasoning in OpenAI's evaluation and simulated deployment environments and flags reasoning about being evaluated or gaming the evaluation.Reported by OpenAI · 4 valuesOpenAI4
Alignment audit, in the data explorer Run by the developer Share of transcripts from xAI's automated alignment audit in which the model says it may be under evaluation.Reported by xAI · 2 valuesxAI2
Evaluation awareness · Google DeepMind, in the data explorer Run by the developer Qualitative findings on whether the model says or shows that it recognises it is being tested, from transcripts of Google's own evaluations.Reported by Google DeepMind · 2 valuesGoogle DeepMind2
Apollo Research scenarios · 1 evaluation, 14 values
Evaluation awareness · Apollo Research, in the data explorer Run by Apollo Research (third party) Apollo Research's measure of how often a model's reasoning shows that it recognises it is being evaluated in Apollo's own test scenarios.Reported by Anthropic and OpenAI · 14 valuesAnthropic and OpenAI14
Other developer setups · 3 evaluations, 11 values
Situational awareness challenges, in the data explorer Run by the developer Agentic challenges run by Google in which the model can succeed only by noticing and changing parts of its own setup, such as its context size or oversight settings.Reported by Google DeepMind · 6 valuesGoogle DeepMind6
Verbalized eval awareness, in the data explorer Run by the developer Share of the model's samples on a given set of evaluations in which it says it suspects it is being evaluated.Reported by Meta · 4 valuesMeta4
Prompted evaluation awareness, in the data explorer Run by the developer The model is asked whether transcripts come from an evaluation or from real use, and is scored on how well it tells them apart.Reported by Meta · 1 valueMeta1
Internal states or training · 5 evaluations, 5 values
Evaluation awareness in audit transcripts, in the data explorer Run by the developer Share of audit transcripts in which the model is aware of being evaluated, detected from its internal states rather than from what it writes.Reported by Anthropic · 1 valueAnthropic1
Grader awareness, in the data explorer Run by the developer Share of training episodes in which the model shows awareness of its grader, split by how exposed the grader is to hacking.Reported by Anthropic · 1 valueAnthropic1
Unverbalized grader awareness, in the data explorer Run by the developer Share of cases in which the model is aware of its grader without saying so, found by analysing more than its written output.Reported by Anthropic · 1 valueAnthropic1
Verbalized grader speculation, in the data explorer Run by the developer Share of reinforcement-learning training episodes in which the model speculates in writing about how it is being graded.Reported by Anthropic · 1 valueAnthropic1
Verbalized metagaming in training, in the data explorer Run by the developer Tracks the rate of verbalized metagaming in the model's reasoning over the course of training.Reported by OpenAI · 1 valueOpenAI1
Real or simulated deployment · 2 evaluations, 4 values
Deployment simulation (internal Codex traffic), in the data explorer Run by the developer Share of resampled internal Codex trajectories in which the model says it may be under evaluation, with the original traffic as a baseline.Reported by OpenAI · 3 valuesOpenAI3
Evaluation awareness (deployment), in the data explorer Run by the developer Share of transcripts from real internal use of the model that reach a set level on Anthropic's evaluation-awareness score.Reported by Anthropic · 1 valueAnthropic1
UK AISI scenarios · 1 evaluation, 3 values
Evaluation awareness · UK AI Security Institute, in the data explorer Run by UK AI Security Institute (government) The UK AI Security Institute's rate of a model saying, unprompted, that it may be under evaluation in the Institute's test scenarios.Reported by Anthropic · 3 valuesAnthropic3

Reported values

The values a chart can place, by the kind of test that produced them. Open a point, or a value in the table, for its source and history. Open in explorer: Evaluation awareness

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

Evaluation awareness: values by kind of test

Each point is a value reported in a developer’s document, in percent (higher is worse), grouped by the kind of test that produced it. Logarithmic axis.

Hollow: read in a secondary write-up

Developer audit transcriptsApollo Research scenariosOther developer setupsInternal states or trainingReal or simulated deploymentUK AISI scenarios0.10.3131030100Percent, log scale

Marker shape is the developer whose document reports the value; a hollow marker is a value we read in an independent write-up because the card section could not be read directly.

Each point is one value, placed by the number the document printed. Bands are kinds of test; within a band, points are spread up and down only so that they do not overlap.

34 of the family’s 53 values are plotted: those in percent where higher is worse, the unit most of its evaluations use. Of the rest, 8 in other units or read the other way (counts and rates from 0 to 1), 6 are categories or statements in words and 5 are low-confidence or read off a figure, as in the explorer’s default view. The explorer lists them all, one evaluation at a time.

Logarithmic axis: the values span more than one and a half orders of magnitude.

Source: Safety Card Ledger v0.1 · Data from developer system cards · CC BY 4.0

Comparability

  • Each of the 16 evaluations is its own test, with its own tasks, graders and definitions. Values from different evaluations describe the same property measured in different ways, and are not comparable with one another, even where two evaluations share a name.
  • Values come from the documents of five developers. Developers use different tests, prompts and graders for the same property, so we do not rank them on these numbers: tables here are ordered by name or by kind of test, never by a reported value.
  • Some names are used by more than one evaluation: “Evaluation awareness” by three. They are different tests, run by different evaluators or reported by different developers, so this page names who ran each one.
  • In two evaluations the test changed between documents (Evaluation awareness · Apollo Research and Automated behavioral audit). Values on the old and new definitions are never joined: the explorer breaks the line where the test changed.
  • In four evaluations the explorer connects results from different documents that do not say whether the test stayed the same. We connect them because nothing suggests it changed; where a document says it did, the line breaks. Each value on such a line is marked “test assumed unchanged”.
  • Seven evaluations report values under more than one condition, such as reasoning effort, safeguards or number of attempts. Each condition is kept as a separate series.
  • Numbers in this family are printed in three units: counts, percent and rates from 0 to 1. We never convert one unit into another, so a chart shows one unit at a time.
  • One value is a restatement: a later document reporting an earlier model again, sometimes with a different value. A restatement is kept beside the model’s own value, never in place of it.

Coverage

Documents from five of the eight developers report this family. The coverage matrix counts a document when it has at least one printed number in the family; categories, statements in words and figure reads count only in the first column.

Each developer's documents with values in Evaluation awareness, out of all its documents in the dataset
DeveloperWith any valueWith a printed numberAll its documents
Anthropic9915
DeepSeek001
Google DeepMind6616
Meta225
Moonshot AI002
OpenAI7422
xAI118
Zhipu AI001

Related findings