Chain-of-thought monitorability
Whether a model's written reasoning is legible and faithful enough to monitor.
- Reported by
- Anthropic, Google DeepMind and OpenAI
- Evaluations
- nine evaluations, of three kinds of test
- Values
- 21 values in 12 documents
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Definition
Many models write out their reasoning before they answer. Chain-of-thought monitorability asks whether that reasoning can be relied on to oversee the model: whether it is legible, whether it reflects the reasoning that actually drives the answer, and whether a monitor reading it would catch misbehaviour.
Developers measure this by testing whether a monitor that can see the reasoning detects problems that a monitor without it misses, by checking whether a model can hide or reshape its reasoning when asked to, and by recording whether training put pressure on the reasoning itself, which could make it less faithful.
Results include detection rates, rates at which a model controls its reasoning on instruction, and statements about how the model was trained. A higher value can mean more monitorable reasoning or less, depending on what is measured.
The line under the title is this family’s entry in Methodology §1.1. A family groups results about one property; it does not make them comparable (Methodology §5).
How developers measure it
nine evaluations of three kinds of test, with the most values first. Each evaluation opens in the data explorer, which lists all of its values.
| Evaluation | Reported by | Values |
|---|---|---|
| Static benchmark · 6 evaluations, 17 values | ||
| CoT controllability · OpenAI, in the data explorer Run by the developer Measures how often the model can follow instructions that constrain the content of its own chain of thought; lower controllability is better for monitoring.Reported by OpenAI · 7 values | OpenAI | 7 |
| CoT monitorability, in the data explorer Run by the developer Measures how well a monitor that reads the model's chain of thought (and in later cards its actions) detects target behaviours across a suite of monitorability tasks.Reported by OpenAI · 5 values | OpenAI | 5 |
| CoT legibility, in the data explorer Run by the developer Share of sampled reasoning transcripts whose chain of thought is comprehensible and informative about the final output.Reported by Google DeepMind · 2 values | Google DeepMind | 2 |
| CoT controllability · Anthropic, in the data explorer Run by the developer How well the model can steer the content of its chain of thought when told to, where more control means the reasoning is less reliable to monitor.Reported by Anthropic · 1 value | Anthropic | 1 |
| Encoded reasoning (scratchpad paraphrase), in the data explorer Run by the developer Checks for hidden information in the model's reasoning by paraphrasing its scratchpad and seeing whether task performance drops.Reported by Anthropic · 1 value | Anthropic | 1 |
| Misalignment monitor recall, in the data explorer Run by the developer Measures the share of known misaligned trajectories that OpenAI's misalignment monitor flags at severity 3 or higher.Reported by OpenAI · 1 value | OpenAI | 1 |
| Scenario evaluation · 2 evaluations, 2 values | ||
| Multilingual reasoning anomalies · Apollo Research, in the data explorer Run by Apollo Research (third party) Apollo Research's share of samples whose reasoning shows anomalous use of several languages.Reported by OpenAI · 1 value | OpenAI | 1 |
| Non-Latin reasoning tokens · Apollo Research, in the data explorer Run by Apollo Research (third party) Apollo Research's share of a model's reasoning tokens written in non-Latin scripts, used as a proxy for how legible the reasoning is.Reported by OpenAI · 1 value | OpenAI | 1 |
| Training monitoring · 1 evaluation, 2 values | ||
| Accidental CoT supervision in RL, in the data explorer Run by the developer Share of reinforcement-learning episodes in which a training error exposed the model's chain of thought to the reward signal.Reported by Anthropic · 2 values | Anthropic | 2 |
Reported values
The values a chart can place, by the kind of test that produced them. Open a point, or a value in the table, for its source and history. Open in explorer: Chain-of-thought monitorability
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
Chain-of-thought monitorability: values by kind of test
Each point is a value reported in a developer’s document, in percent (higher is worse), grouped by the kind of test that produced it. Logarithmic axis.
Hollow: read in a secondary write-up
| Model | Developer | Evaluation | Condition | Value | Document | Kind of test |
|---|---|---|---|---|---|---|
| GPT-5.4 Thinking | OpenAI | CoT controllabilityCoTs successfully controlled | 10k-character CoTs | GPT-5.4 Thinking System Card5 Mar 2026 | Static benchmark | |
| GPT-5.5 | OpenAI | CoT controllabilityCoTs successfully controlled | 50k-character CoTs | GPT-5.5 System Card23 Apr 2026 | Static benchmark | |
| GPT-5.6 Sol | OpenAI | CoT controllabilityCoTs successfully controlled | ~5k-token CoTs | GPT-5.6 System Card9 Jul 2026 | Static benchmark | |
| GPT-5.5 | OpenAI | CoT controllabilityCoTs successfully controlled | ~5k-token CoTs | GPT-5.6 System Card9 Jul 2026 | Static benchmark | |
| GPT-5.4 Thinking | OpenAI | CoT controllabilityCoTs successfully controlled | ~5k-token CoTs | GPT-5.6 System Card9 Jul 2026 | Static benchmark | |
| GPT-5.3-Codex | OpenAI | Non-Latin reasoning tokensRun by Apollo ResearchShare of non-Latin reasoning tokens | — | GPT-5.3-Codex System Card5 Feb 2026 | Scenario evaluation | |
| GPT-5.4 Thinking | OpenAI | Multilingual reasoning anomaliesRun by Apollo ResearchSamples with anomalies | — | GPT-5.4 Thinking System Card5 Mar 2026 | Scenario evaluation | |
| Claude Opus 4.7 | Anthropic | Accidental CoT supervision in RLShare of training episodes affected | — | Claude Opus 4.7 System Card16 Apr 2026 | Training monitoring | |
| Claude Opus 4.8 | Anthropic | Accidental CoT supervision in RLShare of training episodes affected | — | Claude Opus 4.8 System Card28 May 2026 | Training monitoring |
Marker shape is the developer whose document reports the value; a hollow marker is a value we read in an independent write-up because the card section could not be read directly.
Each point is one value, placed by the number the document printed. Bands are kinds of test; within a band, points are spread up and down only so that they do not overlap.
9 of the family’s 21 values are plotted: those in percent where higher is worse, the unit most of its evaluations use. Of the rest, 3 in other units or read the other way (percent where higher is better), 8 are categories or statements in words and 1 is low-confidence or read off a figure, as in the explorer’s default view. The explorer lists them all, one evaluation at a time.
Logarithmic axis: the values span more than one and a half orders of magnitude.
5 of the 9 plotted values come from one evaluation, CoT controllability · OpenAI. The explorer draws it over time, with a break wherever the test changed.
Source: Safety Card Ledger v0.1 · Data from developer system cards · CC BY 4.0
Comparability
- Each of the nine evaluations is its own test, with its own tasks, graders and definitions. Values from different evaluations describe the same property measured in different ways, and are not comparable with one another, even where two evaluations share a name.
- Values come from the documents of three developers. Developers use different tests, prompts and graders for the same property, so we do not rank them on these numbers: tables here are ordered by name or by kind of test, never by a reported value.
- Some names are used by more than one evaluation: “CoT controllability” by two. They are different tests, run by different evaluators or reported by different developers, so this page names who ran each one.
- In one evaluation the test changed between documents (CoT monitorability). Values on the old and new definitions are never joined: the explorer breaks the line where the test changed.
- In one evaluation the explorer connects results from different documents that do not say whether the test stayed the same. We connect them because nothing suggests it changed; where a document says it did, the line breaks. Each value on such a line is marked “test assumed unchanged”.
- Two evaluations report values under more than one condition, such as reasoning effort, safeguards or number of attempts. Each condition is kept as a separate series.
- Numbers in this family are printed in two units: percent and scores from 0 to 1. We never convert one unit into another, so a chart shows one unit at a time.
- For some measures a higher value is worse and for others better, depending on whether the developer reports the behaviour or its absence. Each value’s record says which.
Coverage
Documents from three of the eight developers report this family. The coverage matrix counts a document when it has at least one printed number in the family; categories, statements in words and figure reads count only in the first column.
| Developer | With any value | With a printed number | All its documents |
|---|---|---|---|
| Anthropic | 4 | 3 | 15 |
| DeepSeek | 0 | 0 | 1 |
| Google DeepMind | 1 | 1 | 16 |
| Meta | 0 | 0 | 5 |
| Moonshot AI | 0 | 0 | 2 |
| OpenAI | 7 | 4 | 22 |
| xAI | 0 | 0 | 8 |
| Zhipu AI | 0 | 0 | 1 |
Related findings
No finding is about this family yet.