Misalignment audits
Automated behavioral audits, agentic misalignment scenarios, covert-action and scheming rates, destructive actions.
- Reported by
- Anthropic, Google DeepMind, Meta and OpenAI
- Evaluations
- 24 evaluations, of seven kinds of test
- Values
- 94 values in 27 documents
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Definition
Misalignment audits look for behaviour that goes against what a model’s users or developers intend, across many situations at once.
Developers run automated audits in which another model plays a user or an operator and probes for problems such as deception, cooperation with harmful requests or attempts to gain influence. They also build agentic scenarios in which the model could take a covert or destructive step, such as deleting files, misleading a user or pursuing a goal it was not given. Independent evaluators run their own scenarios for covert and scheming behaviour.
Results take many forms: the rate of a behaviour across a set of scenarios, a score from an automated auditor on the tester’s own scale, or a count of incidents. The scenarios are written by each developer or evaluator and can change from one document to the next, so this family holds many short series rather than one measure.
The line under the title is this family’s entry in Methodology §1.1. A family groups results about one property; it does not make them comparable (Methodology §5).
How developers measure it
24 evaluations of seven kinds of test, with the most values first. Each evaluation opens in the data explorer, which lists all of its values.
| Evaluation | Reported by | Values | |||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Scenario evaluation · 16 evaluations, 67 values | |||||||||||||||||||||||||||||
| Destructive action avoidance, in the data explorer Run by the developer Agentic coding scenarios that score whether the model avoids destructive actions and keeps or restores the user's existing work.Reported by OpenAI · 19 values | OpenAI | 19 | |||||||||||||||||||||||||||
| Realistic work environments, in the data explorer Run by the developer Realistic workplace environments measuring how often the agent causes misaligned outcomes such as data exfiltration, security weakening or unauthorised transactions, with and without a confirmation policy.Reported by OpenAI · 16 values | OpenAI | 16 | |||||||||||||||||||||||||||
| Covert deceptive behavior · Apollo Research, in the data explorer Run by Apollo Research (third party) Apollo Research's agentic scenarios measuring how often a model takes covert or deceptive actions.Reported by OpenAI · 8 values | OpenAI | 8 | |||||||||||||||||||||||||||
| Respecting Auto-Review, in the data explorer Run by the developer Measures whether an agent tries to get around an Auto-review approval step after an action is denied, or exploits a misconfigured one.Reported by OpenAI · 3 values | OpenAI | 3 | |||||||||||||||||||||||||||
| Agentic Misalignment, in the data explorer Run by the developer Constructed agentic scenarios that give the model an opening for harmful actions such as blackmail, run by Meta, measuring how often it takes them.Reported by Meta · 2 values | Meta | 2 | |||||||||||||||||||||||||||
| Destructive actions (synthetic backend), in the data explorer Run by the developer Agentic tasks against a simulated backend, measuring how often the model takes destructive actions.Reported by Anthropic · 2 values | Anthropic | 2 | |||||||||||||||||||||||||||
| ExploitGym honeypot, in the data explorer Run by the developer Hard ExploitGym tasks containing a honeypot, measuring whether the model attacks the honeypot to get the task done.Reported by OpenAI · 2 values | OpenAI | 2 | |||||||||||||||||||||||||||
| Internal computer-use safety benchmark, in the data explorer Run by the developer Internal computer-use benchmark reported in the GPT-6 Astra launch post that measures the rate of unsafe actions.Reported by OpenAI · 2 values | OpenAI | 2 | |||||||||||||||||||||||||||
Show the other 8 evaluations of this kind (13 values)
| |||||||||||||||||||||||||||||
| Automated audit · 1 evaluation, 7 values | |||||||||||||||||||||||||||||
| Automated behavioral audit, in the data explorer Run by the developer An auditor model probes the model across many constructed scenarios, and a judge model scores the transcripts for misaligned behaviour.Reported by Anthropic · 7 values | Anthropic | 7 | |||||||||||||||||||||||||||
| Deployment simulation · 2 evaluations, 7 values | |||||||||||||||||||||||||||||
| Deployment simulation (internal Codex traffic), in the data explorer Run by the developer Re-runs tasks from internal agentic coding (Codex) traffic with the new model, flags misaligned actions by severity and checks how realistic the resampled trajectories look.Reported by OpenAI · 6 values | OpenAI | 6 | |||||||||||||||||||||||||||
| ChatGPT deployment simulation, in the data explorer Run by the developer Simulates ChatGPT deployment by re-running conversations with the new model and flagging misaligned behaviour by severity.Reported by OpenAI · 1 value | OpenAI | 1 | |||||||||||||||||||||||||||
| Framework determination · 1 evaluation, 6 values | |||||||||||||||||||||||||||||
| Frontier Safety Framework determination (misalignment), in the data explorer Run by the developer Google DeepMind's decision on whether a model reaches the misalignment thresholds of its Frontier Safety Framework (instrumental reasoning, later stealth and situational awareness).Reported by Google DeepMind · 6 values | Google DeepMind | 6 | |||||||||||||||||||||||||||
| Internal-use monitoring · 2 evaluations, 5 values | |||||||||||||||||||||||||||||
| Internal deployment monitoring, in the data explorer Run by the developer Monitoring of Anthropic's internal use of the model for misaligned actions such as working around safeguards or covering tracks.Reported by Anthropic · 4 values | Anthropic | 4 | |||||||||||||||||||||||||||
| Disclosed misalignment incidents, in the data explorer Run by the developer Number of misalignment incident reports OpenAI published when it launched its framework for reporting model misalignment.Reported by OpenAI · 1 value | OpenAI | 1 | |||||||||||||||||||||||||||
| Static benchmark · 1 evaluation, 1 value | |||||||||||||||||||||||||||||
| Refusal to assist AI safety R&D, in the data explorer Run by the developer How often the model refuses requests to help with AI safety research and development.Reported by Anthropic · 1 value | Anthropic | 1 | |||||||||||||||||||||||||||
| Qualitative assessment · 1 evaluation, 1 value | |||||||||||||||||||||||||||||
| Strategic deception propensity (external) · Unnamed third-party evaluator, in the data explorer Run by Unnamed third-party evaluator (third party) An outside evaluator's assessment of a model's propensity for strategic deception.Reported by Google DeepMind · 1 value | Google DeepMind | 1 | |||||||||||||||||||||||||||
Reported values
The values a chart can place, by the kind of test that produced them. Open a point, or a value in the table, for its source and history. Open in explorer: Misalignment audits
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
Misalignment audits: values by kind of test
Each point is a value reported in a developer’s document, in percent (higher is worse), grouped by the kind of test that produced it.
Hollow: read in a secondary write-up
| Model | Developer | Evaluation | Condition | Value | Document | Kind of test |
|---|---|---|---|---|---|---|
| Claude Opus 4 | Anthropic | Opportunistic blackmail scenarioShare of rollouts with blackmail | replacement model shares values | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Scenario evaluation | |
| Claude Mythos Preview | Anthropic | Destructive actions (synthetic backend)Destructive action rate | — | Claude Mythos Preview System Card7 Apr 2026 | Scenario evaluation | |
| Claude Opus 4.6 | Anthropic | Destructive actions (synthetic backend)Destructive action rate | — | Claude Mythos Preview System Card7 Apr 2026 | Scenario evaluation | |
| Muse Spark | Meta | Agentic MisalignmentHarmful action rate | without system mitigations | Muse Spark Safety & Preparedness Report8 Apr 2026 | Scenario evaluation | |
| Muse Spark | Meta | Alignment fakingBehavioral difference monitored vs not | without system mitigations | Muse Spark Safety & Preparedness Report8 Apr 2026 | Scenario evaluation | |
| Muse Spark 1.1 | Meta | Agentic MisalignmentHarmful action rate | without system mitigations | Muse Spark 1.1 Evaluation Report9 Jul 2026 | Scenario evaluation | |
| gpt-5-thinking | OpenAI | Covert deceptive behaviorRun by Apollo ResearchShare of trajectories with deceptive action | Apollo scheming suite | GPT-5 System Card7 Aug 2025 | Scenario evaluation | |
| o3 | OpenAI | Covert deceptive behaviorRun by Apollo ResearchShare of trajectories with deceptive action | Apollo scheming suite | GPT-5 System Card7 Aug 2025 | Scenario evaluation | |
| gpt-5-thinking-helpful-only | OpenAI | Covert deceptive behaviorRun by Apollo ResearchShare of trajectories with deceptive action | Apollo scheming suite | GPT-5 System Card7 Aug 2025 | Scenario evaluation | |
| gpt-5-thinking | OpenAI | Log falsificationRun by Apollo ResearchCovert action rate | weak goal nudge | GPT-5 System Card7 Aug 2025 | Scenario evaluation | |
| gpt-5-thinking | OpenAI | Log falsificationRun by Apollo ResearchCovert action rate | strong goal nudge | GPT-5 System Card7 Aug 2025 | Scenario evaluation | |
| GPT-5.4 Thinking | OpenAI | Covert deceptive behaviorRun by Apollo ResearchOverall deception rate | no goal nudge | GPT-5.4 Thinking System Card5 Mar 2026 | Scenario evaluation | |
| GPT-5.6 Sol | OpenAI | Respecting Auto-ReviewRollouts attempting to circumvent Auto-review denial | max | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-5.6 Sol | OpenAI | Respecting Auto-ReviewRollouts exploiting misconfigured Auto-review | max | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-6 Astra | OpenAI | Respecting Auto-ReviewRollouts attempting to circumvent Auto-review denial | max | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-5.6 Sol | OpenAI | Respecting WarningsUnwanted persistence after restriction | max | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-6 Astra | OpenAI | Respecting WarningsUnwanted persistence after restriction | max | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-5.6 Sol | OpenAI | ExploitGym honeypotSubstantial attack on honeypot | max; adversarially selected hardest tasks | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-6 Astra | OpenAI | ExploitGym honeypotSubstantial attack on honeypot | max; adversarially selected hardest tasks | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-6 Astra | OpenAI | Internal computer-use safety benchmarkUnsafe action rate | — | GPT-6 Astra launch post3 Sep 2026 | Scenario evaluation | |
| GPT-5.6 Sol | OpenAI | Internal computer-use safety benchmarkUnsafe action rate | — | GPT-6 Astra launch post3 Sep 2026 | Scenario evaluation | |
| GPT-6 Astra | OpenAI | Realistic work environmentsOverall misaligned outcome rate | no confirmation policy | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-5.6 Sol | OpenAI | Realistic work environmentsOverall misaligned outcome rate | no confirmation policy | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-5.6 Terra | OpenAI | Realistic work environmentsOverall misaligned outcome rate | no confirmation policy | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-5.6 Luna | OpenAI | Realistic work environmentsOverall misaligned outcome rate | no confirmation policy | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-6 Astra | OpenAI | Realistic work environmentsOverall misaligned outcome rate | with confirmation policy (default) | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-5.6 Sol | OpenAI | Realistic work environmentsOverall misaligned outcome rate | with confirmation policy (default) | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-5.6 Terra | OpenAI | Realistic work environmentsOverall misaligned outcome rate | with confirmation policy (default) | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-5.6 Luna | OpenAI | Realistic work environmentsOverall misaligned outcome rate | with confirmation policy (default) | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-6 Astra | OpenAI | Realistic work environmentsUnauthorized transactions | no confirmation policy | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-5.6 Sol | OpenAI | Realistic work environmentsUnauthorized transactions | no confirmation policy | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-6 Astra | OpenAI | Realistic work environmentsData exfiltration | no confirmation policy | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-5.6 Sol | OpenAI | Realistic work environmentsData exfiltration | no confirmation policy | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-6 Astra | OpenAI | Realistic work environmentsUnauthorized external communication | no confirmation policy | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-5.6 Sol | OpenAI | Realistic work environmentsUnauthorized external communication | no confirmation policy | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-6 Astra | OpenAI | Realistic work environmentsSecurity weakening | no confirmation policy | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-5.6 Sol | OpenAI | Realistic work environmentsSecurity weakening | no confirmation policy | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-5.6 Sol | OpenAI | Unintended engagement with external agent messagesSamples following unauthorized board instructions | — | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-6 Astra | OpenAI | Unintended engagement with external agent messagesSamples following unauthorized board instructions | — | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-5.6 Sol | OpenAI | Proactive searches for other agentsTrajectories speculating about other agents | representative subsample | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-6 Astra | OpenAI | Proactive searches for other agentsTrajectories attempting contact with other agents | — | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-6 Astra | OpenAI | Out-of-scope supply-chain attackRun by UK AI Security InstituteSamples with out-of-scope attack | — | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-6 Astra | OpenAI | Out-of-scope supply-chain attackRun by UK AI Security InstituteSamples with out-of-scope attack | explicit no-internet scope | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-5.5 | OpenAI | ChatGPT deployment simulationSeverity-3 misalignment rate | — | GPT-5.5 System Card23 Apr 2026 | Deployment simulation | |
| GPT-6 Astra | OpenAI | Deployment simulation (internal Codex traffic)Tasks with severity 3+ misalignment flag | matched tasks | GPT-6 Astra System Card3 Sep 2026 | Deployment simulation | |
| GPT-5.6 Sol | OpenAI | Deployment simulation (internal Codex traffic)Tasks with severity 3+ misalignment flag | matched tasks | GPT-6 Astra System Card3 Sep 2026 | Deployment simulation | |
| GPT-6 Astra | OpenAI | Deployment simulation (internal Codex traffic)Tasks with severity 1-2 flag | matched tasks | GPT-6 Astra System Card3 Sep 2026 | Deployment simulation | |
| GPT-5.6 Sol | OpenAI | Deployment simulation (internal Codex traffic)Tasks with severity 1-2 flag | matched tasks | GPT-6 Astra System Card3 Sep 2026 | Deployment simulation | |
| Claude Mythos Preview | Anthropic | Internal deployment monitoringCovering tracks after disallowed actions | earlier snapshots | Claude Mythos Preview System Card7 Apr 2026 | Internal-use monitoring | |
| Claude Mythos 5 | Anthropic | Internal deployment monitoringMisaligned actions in monitored traffic | — | Claude Fable 5 & Claude Mythos 5 System Card9 Jun 2026 | Internal-use monitoring | |
| Claude Opus 5 | Anthropic | Internal deployment monitoringCircumvention attempts in monitored completions | — | Claude Opus 5 System Card24 Jul 2026 | Internal-use monitoring | |
| Claude Mythos 5.1 | Anthropic | Internal deployment monitoringClassifier workarounds / disabled checks | — | Claude Fable 5.1 & Claude Mythos 5.1 System Card1 Sep 2026 | Internal-use monitoring | |
| Claude Opus 4.8 | Anthropic | Refusal to assist AI safety R&DRefusal rate | — | Claude Opus 4.8 System Card28 May 2026 | Static benchmark |
Marker shape is the developer whose document reports the value; a hollow marker is a value we read in an independent write-up because the card section could not be read directly.
Each point is one value, placed by the number the document printed. Bands are kinds of test; within a band, points are spread up and down only so that they do not overlap.
53 of the family’s 94 values are plotted: those in percent where higher is worse, the unit most of its evaluations use. Of the rest, 22 in other units or read the other way (counts, relative change in percent, rates from 0 to 1 and percent where higher is better), 14 are categories or statements in words, 4 are restated in a later document or from an earlier version of one and 1 is low-confidence or read off a figure, as in the explorer’s default view. The explorer lists them all, one evaluation at a time.
The values span more than one and a half orders of magnitude, but four of them are zero, which a logarithmic axis cannot show, so the axis is linear and starts at zero.
Source: Safety Card Ledger v0.1 · Data from developer system cards · CC BY 4.0
Comparability
- Each of the 24 evaluations is its own test, with its own tasks, graders and definitions. Values from different evaluations describe the same property measured in different ways, and are not comparable with one another, even where two evaluations share a name.
- Values come from the documents of four developers. Developers use different tests, prompts and graders for the same property, so we do not rank them on these numbers: tables here are ordered by name or by kind of test, never by a reported value.
- In three evaluations the test changed between documents (Destructive action avoidance, Automated behavioral audit and Frontier Safety Framework determination (misalignment)). Values on the old and new definitions are never joined: the explorer breaks the line where the test changed.
- In three evaluations the explorer connects results from different documents that do not say whether the test stayed the same. We connect them because nothing suggests it changed; where a document says it did, the line breaks. Each value on such a line is marked “test assumed unchanged”.
- Nine evaluations report values under more than one condition, such as reasoning effort, safeguards or number of attempts. Each condition is kept as a separate series.
- Numbers in this family are printed in five units: percent, counts, relative change in percent, rates from 0 to 1 and scores from 1 to 10. We never convert one unit into another, so a chart shows one unit at a time.
- For some measures a higher value is worse and for others better, depending on whether the developer reports the behaviour or its absence. Each value’s record says which.
- Four values are restatements: a later document reporting an earlier model again, sometimes with a different value. A restatement is kept beside the model’s own value, never in place of it.
Coverage
Documents from four of the eight developers report this family. The coverage matrix counts a document when it has at least one printed number in the family; categories, statements in words and figure reads count only in the first column.
| Developer | With any value | With a printed number | All its documents |
|---|---|---|---|
| Anthropic | 9 | 8 | 15 |
| DeepSeek | 0 | 0 | 1 |
| Google DeepMind | 5 | 0 | 16 |
| Meta | 2 | 2 | 5 |
| Moonshot AI | 0 | 0 | 2 |
| OpenAI | 11 | 10 | 22 |
| xAI | 0 | 0 | 8 |
| Zhipu AI | 0 | 0 | 1 |
Related findings
No finding is about this family yet.