KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

A number for evaluation awareness means little without the test that produced it: the same model can get numbers more than ten times apart from different tests. Meta’s report on Muse Spark gives for the share of samples in which the model said it suspected an evaluation, across Meta’s internal evaluations, and on a test that asks the model to tell evaluation transcripts from real use.

Evaluation awareness is how often a model shows signs of recognising that it is being tested. A model that recognises a test may act differently from how it would act in real use, and that matters because most safety results come from tests.

This finding looks at the 33 numbers the feasibility study selected. They are reported in 14 documents from four developers, and they come from six kinds of test: developer audit transcripts, Apollo Research scenarios, UK AI Security Institute scenarios, other developer setups, real or simulated deployment and internal states or training signals. Each kind counts something different, over different transcripts, samples or episodes. The chart places each number in the band for its kind of test; numbers in different bands are not comparable.

Evaluation-awareness numbers, by the kind of test that produced them

Each point is one number as a developer’s document reports it, in percent: most are the share of transcripts, samples or episodes in which a model showed signs of recognising a test, and each value’s record gives its exact definition. Logarithmic axis.

Hollow: read in a secondary write-up

Developer audit transcriptsApollo Research scenariosUK AISI scenariosOther developer setupsReal or simulated deploymentInternal states or training0.10.3131030Share (%), log scale79%about 0.1%

Each point is one value, placed by the number the document printed. Bands are kinds of test; within a band, points are spread up and down only so that they do not overlap. Points in different bands come from different tests and are not comparable with one another.

Marker shape is the developer whose document reports the value; a hollow marker is a value we read in an independent write-up of the document, because the section could not be read directly.

30 of the 33 values are plotted. The value for Claude Opus 4 is a count; the values for Grok 4 and Grok 4.20 (single-agent) are rates from 0 to 1. Placing them on this axis would mean rescaling a printed value, so they are in the table only.

The lowest and the highest plotted values are labelled.

Logarithmic axis: the values span more than one and a half orders of magnitude.

Source: Safety Card Ledger v0.1 · Data from developer system cards · CC BY 4.0

One model, three tests

The Claude Sonnet 4.5 system card reports three numbers for the model. Anthropic’s automated behavioural audit found the model saying it might be under evaluation in of audit transcripts. The UK AI Security Institute reported an unprompted rate of in its scenarios, and Apollo Research found unambiguous references to being evaluated in of its scenario transcripts. Both outside evaluators tested pre-release snapshots of the model.

We read all three numbers in independent write-ups of the card, so the chart draws them as hollow markers. The first has since been checked against the card itself; the other two were checked against the write-up.

One document, two settings

The Claude Opus 5.5 system card reports one measure, the share of transcripts that reach a set score for evaluation awareness, in two settings: in Anthropic’s automated audit, and in real internal deployment. We read both in a write-up of the card and checked them against that write-up, not the card.

Inside one document

Apollo Research’s numbers for OpenAI models are measured one way within each document. In the GPT-5.5 system card, Apollo reports the share of samples with moderate or higher evaluation awareness: for a pre-release GPT-5.3-Codex, for a pre-release GPT-5.4 and for GPT-5.5.

The GPT-6 Astra system card reports a different measure, the share of samples with verbalized evaluation awareness at a stated reasoning effort: for GPT-5.5 and for GPT-6 Astra, both at the effort recorded as “xhigh”. GPT-5.5 therefore has a number in each document, on different measures. Inside each document the models were measured the same way, so their numbers can be compared there; across the two documents they cannot, and the explorer keeps them as separate series.

The explorer lists every value in the family, one evaluation at a time, with its source and its checks.

Cite this

APA

Safety Card Ledger. (2026, September 28). 33 reported numbers on evaluation awareness come from six kinds of test. Safety Card Ledger. https://safetycardledger.example/findings/evaluation-awareness

BibTeX

@misc{safetycardledger_evaluation_awareness,
  author = {{Safety Card Ledger}},
  title = {33 reported numbers on evaluation awareness come from six kinds of test},
  year = {2026},
  month = sep,
  howpublished = {\url{https://safetycardledger.example/findings/evaluation-awareness}},
  note = {Safety Card Ledger, dataset v0.1},
}