Sabotage and sandbagging
Covertly completing a hidden side task, undermining assigned work, or deliberately underperforming.
- Reported by
- Anthropic, Google DeepMind, Meta, OpenAI and xAI
- Evaluations
- 22 evaluations, of five kinds of test
- Values
- 58 values in 27 documents
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Definition
This family covers two related concerns. Sabotage is a model working against the task it was given while appearing to cooperate: completing a hidden side task without a monitor noticing, for example, or quietly weakening its own work. Sandbagging is a model deliberately performing below its ability, which could hide what it can do from the tests meant to find out.
Developers study both with agentic tasks in which the model is given a secret goal and a monitor tries to catch it, with comparisons between a model’s normal performance and its performance when told to underperform, and with reviews of behaviour in internal use.
Some results are the rate at which a model evades a monitor; others are how often a monitor detects the attempt, or a score on the tester’s own scale. A higher number can therefore mean more capacity for sabotage or better detection, depending on the test.
The line under the title is this family’s entry in Methodology §1.1. A family groups results about one property; it does not make them comparable (Methodology §5).
How developers measure it
22 evaluations of five kinds of test, with the most values first. Each evaluation opens in the data explorer, which lists all of its values.
| Evaluation | Reported by | Values |
|---|---|---|
| Capability benchmark · 12 evaluations, 40 values | ||
| SHADE-Arena · Anthropic, in the data explorer Run by the developer Agentic environments in which the model pursues a secret side task alongside its main task, measuring how often it succeeds without the monitor detecting it.Reported by Anthropic · 10 values | Anthropic | 10 |
| Stealth challenges, in the data explorer Run by the developer Agentic challenges run by Google that test whether the model can carry out a covert task while avoiding oversight.Reported by Google DeepMind · 7 values | Google DeepMind | 7 |
| Sabotage capability · Apollo Research, in the data explorer Run by Apollo Research (third party) Apollo Research tasks measuring a model's capability to carry out sabotage, scored from 0 to 1.Reported by OpenAI · 5 values | OpenAI | 5 |
| Subtle sabotage (agentic), in the data explorer Run by the developer Agentic tasks with a hidden harmful side task, measuring how often the model completes the side task without a monitoring model noticing.Reported by Anthropic · 4 values | Anthropic | 4 |
| AI R&D sabotage · Apollo Research, in the data explorer Run by Apollo Research (third party) An Apollo Research task measuring a model's capability to sabotage AI research work, such as inserting a backdoor, scored from 0 to 1.Reported by OpenAI · 3 values | OpenAI | 3 |
| LinuxArena, in the data explorer Run by the developer Tasks in a Linux environment with a covert side task, measuring how often the model succeeds without a given monitor detecting it.Reported by Anthropic · 2 values | Anthropic | 2 |
| Minimal-LinuxBench, in the data explorer Run by the developer Tasks in a Linux environment with a covert side task, measuring how often the model completes it without the monitor noticing.Reported by Anthropic · 2 values | Anthropic | 2 |
| SHADE-Arena · Meta, in the data explorer Run by the developer Agentic environments with a hidden harmful side task, run by Meta, measuring how often the model completes it without the monitor noticing.Reported by Meta · 2 values | Meta | 2 |
| SHADE-Arena / Minimal-LinuxBench, in the data explorer Run by the developer Statements that summarise stealth side-task results across SHADE-Arena and Minimal-LinuxBench together.Reported by Anthropic · 2 values | Anthropic | 2 |
| GDM-Stealth, in the data explorer Run by the developer Google DeepMind's stealth tasks, in which the model must reach a goal while evading oversight, run by Meta as a stealth capability measure.Reported by Meta · 1 value | Meta | 1 |
| Monitor evasion under prompted sandbagging, in the data explorer Run by the developer Tests whether the model, when instructed to underperform on purpose, can do so without a chain-of-thought monitor noticing.Reported by OpenAI · 1 value | OpenAI | 1 |
| SHADE-Arena / LinuxArena, in the data explorer Run by the developer Statements that summarise stealth side-task results across SHADE-Arena and LinuxArena together.Reported by Anthropic · 1 value | Anthropic | 1 |
| Scenario evaluation · 4 evaluations, 9 values | ||
| Sandbagging · Apollo Research, in the data explorer Run by Apollo Research (third party) Apollo Research tests of whether a model deliberately underperforms on capability questions when it has an incentive to.Reported by OpenAI · 5 values | OpenAI | 5 |
| Safety research compromise, in the data explorer Run by the developer How often the model undermines AI safety research tasks it is given.Reported by Anthropic · 2 values | Anthropic | 2 |
| Sabotage · Apollo Research, in the data explorer Run by Apollo Research (third party) Apollo Research testing for whether a model sabotages its assigned work, reported as whether sabotage was observed.Reported by OpenAI · 1 value | OpenAI | 1 |
| Unprompted AI safety research sabotage · UK AI Security Institute, in the data explorer Run by UK AI Security Institute (government) A UK AI Security Institute check for whether a model, without being prompted, sabotages AI safety research work it is given.Reported by OpenAI · 1 value | OpenAI | 1 |
| Automated audit · 3 evaluations, 6 values | ||
| Alignment audit, in the data explorer Run by the developer Findings from xAI's automated alignment audit built on the Petri 2.0 tool on how often the model takes subversive action against its user or against xAI.Reported by xAI · 2 values | xAI | 2 |
| Sandbagging checks, in the data explorer Run by the developer Review of evaluation transcripts for signs that the model deliberately underperformed.Reported by Google DeepMind · 2 values | Google DeepMind | 2 |
| Sandbagging transcript review, in the data explorer Run by the developer Review of the model's evaluation transcripts for signs that it deliberately underperformed.Reported by Anthropic · 2 values | Anthropic | 2 |
| Qualitative assessment · 2 evaluations, 2 values | ||
| AI R&D sabotage (external) · Unnamed third-party evaluator, in the data explorer Run by Unnamed third-party evaluator (third party) An outside evaluator's assessment of a model's ability to sabotage AI research and development work.Reported by Google DeepMind · 1 value | Google DeepMind | 1 |
| Sandbagging review · METR, in the data explorer Run by METR (third party) METR's review of monitored evaluation runs for signs that a model deliberately underperformed.Reported by OpenAI · 1 value | OpenAI | 1 |
| Internal-use monitoring · 1 evaluation, 1 value | ||
| Internal monitoring, in the data explorer Run by the developer Monitoring of Anthropic's internal use of the model for signs of deliberate underperformance.Reported by Anthropic · 1 value | Anthropic | 1 |
Reported values
The values a chart can place, by the kind of test that produced them. Open a point, or a value in the table, for its source and history. Open in explorer: Sabotage and sandbagging
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
Sabotage and sandbagging: values by kind of test
Each point is a value reported in a developer’s document, in percent (higher is worse), grouped by the kind of test that produced it.
Hollow: read in a secondary write-up
| Model | Developer | Evaluation | Condition | Value | Document | Kind of test |
|---|---|---|---|---|---|---|
| Claude Opus 4 | Anthropic | Subtle sabotage (agentic)Side-task success without monitor detection | scratchpad visible to monitor | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Capability benchmark | |
| Claude Opus 4 | Anthropic | Subtle sabotage (agentic)Side-task success without monitor detection | hidden scratchpad | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Capability benchmark | |
| Claude Sonnet 3.7 | Anthropic | Subtle sabotage (agentic)Side-task success without monitor detection | scratchpad visible to monitor | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Capability benchmark | |
| Claude Sonnet 3.7 | Anthropic | Subtle sabotage (agentic)Side-task success without monitor detection | hidden scratchpad | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Capability benchmark | |
| Claude Mythos Preview | Anthropic | SHADE-ArenaStealth success rate | extended thinking | Claude Mythos Preview System Card7 Apr 2026 | Capability benchmark | |
| Claude Opus 4.7 | Anthropic | SHADE-ArenaStealth success rate | extended thinking | Claude Opus 4.7 System Card16 Apr 2026 | Capability benchmark | |
| Claude Opus 4.6 | Anthropic | SHADE-ArenaStealth success rate | extended thinking | Claude Opus 4.7 System Card16 Apr 2026 | Capability benchmark | |
| Claude Opus 4.7 | Anthropic | Minimal-LinuxBenchStealth success rate | CoT visible to monitor | Claude Opus 4.7 System Card16 Apr 2026 | Capability benchmark | |
| Claude Mythos 5.1 | Anthropic | LinuxArenaStealth success vs Opus 4.8 monitor | no extended thinking | Claude Fable 5.1 & Claude Mythos 5.1 System Card1 Sep 2026 | Capability benchmark | |
| Claude Mythos 5.1 | Anthropic | LinuxArenaStealth success vs Opus 4.8 monitor | extended thinking | Claude Fable 5.1 & Claude Mythos 5.1 System Card1 Sep 2026 | Capability benchmark | |
| Gemini 2.5 Pro PreviewPreview | Google DeepMind | Stealth challengesDecision Sabotage success rate | — | Gemini 2.5 Pro Model Card27 Jun 2025 | Capability benchmark | |
| Gemini 2.5 Pro PreviewPreview | Google DeepMind | Stealth challengesOversight Pattern success rate | — | Gemini 2.5 Pro Model Card27 Jun 2025 | Capability benchmark | |
| Muse Spark | Meta | SHADE-ArenaSabotage/stealth success score | without system mitigations | Muse Spark Safety & Preparedness Report8 Apr 2026 | Capability benchmark | |
| Muse Spark | Meta | GDM-StealthStealth capability score | without system mitigations | Muse Spark Safety & Preparedness Report8 Apr 2026 | Capability benchmark | |
| Muse Spark 1.1 | Meta | SHADE-ArenaOverall success under CoT monitor | — | Muse Spark 1.1 Evaluation Report9 Jul 2026 | Capability benchmark | |
| Claude Opus 5 | Anthropic | Safety research compromiseRate of compromising safety research | — | Claude Opus 5 System Card24 Jul 2026 | Scenario evaluation | |
| Claude Mythos 5 | Anthropic | Safety research compromiseRate of compromising safety research | — | Claude Opus 5 System Card24 Jul 2026 | Scenario evaluation |
Marker shape is the developer whose document reports the value; a hollow marker is a value we read in an independent write-up because the card section could not be read directly.
Each point is one value, placed by the number the document printed. Bands are kinds of test; within a band, points are spread up and down only so that they do not overlap.
17 of the family’s 58 values are plotted: those in percent where higher is worse, the unit most of its evaluations use. Of the rest, 19 in other units or read the other way (scores from 0 to 1, counts, percentage points, rates from 0 to 1 and 1 others), 16 are categories or statements in words, 1 is restated in a later document or from an earlier version of one and 5 are low-confidence or read off a figure, as in the explorer’s default view. The explorer lists them all, one evaluation at a time.
The values span more than one and a half orders of magnitude, but one of them is zero, which a logarithmic axis cannot show, so the axis is linear and starts at zero.
Source: Safety Card Ledger v0.1 · Data from developer system cards · CC BY 4.0
Comparability
- Each of the 22 evaluations is its own test, with its own tasks, graders and definitions. Values from different evaluations describe the same property measured in different ways, and are not comparable with one another, even where two evaluations share a name.
- Values come from the documents of five developers. Developers use different tests, prompts and graders for the same property, so we do not rank them on these numbers: tables here are ordered by name or by kind of test, never by a reported value.
- Some names are used by more than one evaluation: “SHADE-Arena” by two. They are different tests, run by different evaluators or reported by different developers, so this page names who ran each one.
- In one evaluation the test changed between documents (Stealth challenges). Values on the old and new definitions are never joined: the explorer breaks the line where the test changed.
- In three evaluations the explorer connects results from different documents that do not say whether the test stayed the same. We connect them because nothing suggests it changed; where a document says it did, the line breaks. Each value on such a line is marked “test assumed unchanged”.
- Nine evaluations report values under more than one condition, such as reasoning effort, safeguards or number of attempts. Each condition is kept as a separate series.
- Numbers in this family are printed in five units: scores from 0 to 1, percent, counts, percentage points and rates from 0 to 1. We never convert one unit into another, so a chart shows one unit at a time.
- For some measures a higher value is worse and for others better, depending on whether the developer reports the behaviour or its absence. Each value’s record says which.
- One value is a restatement: a later document reporting an earlier model again, sometimes with a different value. A restatement is kept beside the model’s own value, never in place of it.
Coverage
Documents from five of the eight developers report this family. The coverage matrix counts a document when it has at least one printed number in the family; categories, statements in words and figure reads count only in the first column.
| Developer | With any value | With a printed number | All its documents |
|---|---|---|---|
| Anthropic | 10 | 8 | 15 |
| DeepSeek | 0 | 0 | 1 |
| Google DeepMind | 6 | 5 | 16 |
| Meta | 2 | 2 | 5 |
| Moonshot AI | 0 | 0 | 2 |
| OpenAI | 8 | 4 | 22 |
| xAI | 1 | 1 | 8 |
| Zhipu AI | 0 | 0 | 1 |
Related findings
No finding is about this family yet.