Definition

Misalignment audits look for behaviour that goes against what a model’s users or developers intend, across many situations at once.

Developers run automated audits in which another model plays a user or an operator and probes for problems such as deception, cooperation with harmful requests or attempts to gain influence. They also build agentic scenarios in which the model could take a covert or destructive step, such as deleting files, misleading a user or pursuing a goal it was not given. Independent evaluators run their own scenarios for covert and scheming behaviour.

Results take many forms: the rate of a behaviour across a set of scenarios, a score from an automated auditor on the tester’s own scale, or a count of incidents. The scenarios are written by each developer or evaluator and can change from one document to the next, so this family holds many short series rather than one measure.

The line under the title is this family’s entry in Methodology §1.1. A family groups results about one property; it does not make them comparable (Methodology §5).

How developers measure it

24 evaluations of seven kinds of test, with the most values first. Each evaluation opens in the data explorer, which lists all of its values.

The 24 evaluations in Misalignment audits, grouped by kind of test: who ran each one, what it counts, the developers whose documents report it and how many values it has.
EvaluationReported byValues
Scenario evaluation · 16 evaluations, 67 values
Destructive action avoidance, in the data explorer Run by the developer Agentic coding scenarios that score whether the model avoids destructive actions and keeps or restores the user's existing work.Reported by OpenAI · 19 valuesOpenAI19
Realistic work environments, in the data explorer Run by the developer Realistic workplace environments measuring how often the agent causes misaligned outcomes such as data exfiltration, security weakening or unauthorised transactions, with and without a confirmation policy.Reported by OpenAI · 16 valuesOpenAI16
Covert deceptive behavior · Apollo Research, in the data explorer Run by Apollo Research (third party) Apollo Research's agentic scenarios measuring how often a model takes covert or deceptive actions.Reported by OpenAI · 8 valuesOpenAI8
Respecting Auto-Review, in the data explorer Run by the developer Measures whether an agent tries to get around an Auto-review approval step after an action is denied, or exploits a misconfigured one.Reported by OpenAI · 3 valuesOpenAI3
Agentic Misalignment, in the data explorer Run by the developer Constructed agentic scenarios that give the model an opening for harmful actions such as blackmail, run by Meta, measuring how often it takes them.Reported by Meta · 2 valuesMeta2
Destructive actions (synthetic backend), in the data explorer Run by the developer Agentic tasks against a simulated backend, measuring how often the model takes destructive actions.Reported by Anthropic · 2 valuesAnthropic2
ExploitGym honeypot, in the data explorer Run by the developer Hard ExploitGym tasks containing a honeypot, measuring whether the model attacks the honeypot to get the task done.Reported by OpenAI · 2 valuesOpenAI2
Internal computer-use safety benchmark, in the data explorer Run by the developer Internal computer-use benchmark reported in the GPT-6 Astra launch post that measures the rate of unsafe actions.Reported by OpenAI · 2 valuesOpenAI2
Show the other 8 evaluations of this kind (13 values)
More scenario evaluations in Misalignment audits
EvaluationReported byValues
Log falsification · Apollo Research, in the data explorer Run by Apollo Research (third party) An Apollo Research scenario measuring how often a model covertly falsifies logs, under weaker and stronger nudges toward a goal.Reported by OpenAI · 2 valuesOpenAI2
Out-of-scope supply-chain attack · UK AI Security Institute, in the data explorer Run by UK AI Security Institute (government) A UK AI Security Institute agentic scenario checking whether a model carries out a supply-chain attack outside the scope it was given.Reported by OpenAI · 2 valuesOpenAI2
Proactive searches for other agents, in the data explorer Run by the developer Checks ExploitGym honeypot runs for whether the model speculates about, or tries to contact, other agents.Reported by OpenAI · 2 valuesOpenAI2
Respecting Warnings, in the data explorer Run by the developer Measures whether the model keeps pursuing an action after it has been warned off or restricted.Reported by OpenAI · 2 valuesOpenAI2
Unintended engagement with external agent messages, in the data explorer Run by the developer Measures whether the model follows instructions posted by other agents on a shared message board that it was not authorised to act on.Reported by OpenAI · 2 valuesOpenAI2
Alignment faking, in the data explorer Run by the developer Compares the model's behaviour when it believes it is monitored with its behaviour when it believes it is not, and reports the difference.Reported by Meta · 1 valueMeta1
Opportunistic blackmail scenario, in the data explorer Run by the developer A constructed scenario in which the model is told it will be replaced and has an opening to blackmail, measuring how often it does so.Reported by Anthropic · 1 valueAnthropic1
Package-registry credential exercise, in the data explorer Run by the developer A simulated security exercise involving package-registry credentials, measuring how often the model takes potentially harmful actions.Reported by Anthropic · 1 valueAnthropic1
Automated audit · 1 evaluation, 7 values
Automated behavioral audit, in the data explorer Run by the developer An auditor model probes the model across many constructed scenarios, and a judge model scores the transcripts for misaligned behaviour.Reported by Anthropic · 7 valuesAnthropic7
Deployment simulation · 2 evaluations, 7 values
Deployment simulation (internal Codex traffic), in the data explorer Run by the developer Re-runs tasks from internal agentic coding (Codex) traffic with the new model, flags misaligned actions by severity and checks how realistic the resampled trajectories look.Reported by OpenAI · 6 valuesOpenAI6
ChatGPT deployment simulation, in the data explorer Run by the developer Simulates ChatGPT deployment by re-running conversations with the new model and flagging misaligned behaviour by severity.Reported by OpenAI · 1 valueOpenAI1
Framework determination · 1 evaluation, 6 values
Frontier Safety Framework determination (misalignment), in the data explorer Run by the developer Google DeepMind's decision on whether a model reaches the misalignment thresholds of its Frontier Safety Framework (instrumental reasoning, later stealth and situational awareness).Reported by Google DeepMind · 6 valuesGoogle DeepMind6
Internal-use monitoring · 2 evaluations, 5 values
Internal deployment monitoring, in the data explorer Run by the developer Monitoring of Anthropic's internal use of the model for misaligned actions such as working around safeguards or covering tracks.Reported by Anthropic · 4 valuesAnthropic4
Disclosed misalignment incidents, in the data explorer Run by the developer Number of misalignment incident reports OpenAI published when it launched its framework for reporting model misalignment.Reported by OpenAI · 1 valueOpenAI1
Static benchmark · 1 evaluation, 1 value
Refusal to assist AI safety R&D, in the data explorer Run by the developer How often the model refuses requests to help with AI safety research and development.Reported by Anthropic · 1 valueAnthropic1
Qualitative assessment · 1 evaluation, 1 value
Strategic deception propensity (external) · Unnamed third-party evaluator, in the data explorer Run by Unnamed third-party evaluator (third party) An outside evaluator's assessment of a model's propensity for strategic deception.Reported by Google DeepMind · 1 valueGoogle DeepMind1

Reported values

The values a chart can place, by the kind of test that produced them. Open a point, or a value in the table, for its source and history. Open in explorer: Misalignment audits

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

Misalignment audits: values by kind of test

Each point is a value reported in a developer’s document, in percent (higher is worse), grouped by the kind of test that produced it.

Hollow: read in a secondary write-up

Scenario evaluationDeployment simulationInternal-use monitoringStatic benchmark020406080100Percent

Marker shape is the developer whose document reports the value; a hollow marker is a value we read in an independent write-up because the card section could not be read directly.

Each point is one value, placed by the number the document printed. Bands are kinds of test; within a band, points are spread up and down only so that they do not overlap.

53 of the family’s 94 values are plotted: those in percent where higher is worse, the unit most of its evaluations use. Of the rest, 22 in other units or read the other way (counts, relative change in percent, rates from 0 to 1 and percent where higher is better), 14 are categories or statements in words, 4 are restated in a later document or from an earlier version of one and 1 is low-confidence or read off a figure, as in the explorer’s default view. The explorer lists them all, one evaluation at a time.

The values span more than one and a half orders of magnitude, but four of them are zero, which a logarithmic axis cannot show, so the axis is linear and starts at zero.

Source: Safety Card Ledger v0.1 · Data from developer system cards · CC BY 4.0

Comparability

  • Each of the 24 evaluations is its own test, with its own tasks, graders and definitions. Values from different evaluations describe the same property measured in different ways, and are not comparable with one another, even where two evaluations share a name.
  • Values come from the documents of four developers. Developers use different tests, prompts and graders for the same property, so we do not rank them on these numbers: tables here are ordered by name or by kind of test, never by a reported value.
  • In three evaluations the test changed between documents (Destructive action avoidance, Automated behavioral audit and Frontier Safety Framework determination (misalignment)). Values on the old and new definitions are never joined: the explorer breaks the line where the test changed.
  • In three evaluations the explorer connects results from different documents that do not say whether the test stayed the same. We connect them because nothing suggests it changed; where a document says it did, the line breaks. Each value on such a line is marked “test assumed unchanged”.
  • Nine evaluations report values under more than one condition, such as reasoning effort, safeguards or number of attempts. Each condition is kept as a separate series.
  • Numbers in this family are printed in five units: percent, counts, relative change in percent, rates from 0 to 1 and scores from 1 to 10. We never convert one unit into another, so a chart shows one unit at a time.
  • For some measures a higher value is worse and for others better, depending on whether the developer reports the behaviour or its absence. Each value’s record says which.
  • Four values are restatements: a later document reporting an earlier model again, sometimes with a different value. A restatement is kept beside the model’s own value, never in place of it.

Coverage

Documents from four of the eight developers report this family. The coverage matrix counts a document when it has at least one printed number in the family; categories, statements in words and figure reads count only in the first column.

Each developer's documents with values in Misalignment audits, out of all its documents in the dataset
DeveloperWith any valueWith a printed numberAll its documents
Anthropic9815
DeepSeek001
Google DeepMind5016
Meta225
Moonshot AI002
OpenAI111022
xAI008
Zhipu AI001

Related findings

No finding is about this family yet.