Self-preservation
Shutdown resistance, self-exfiltration, resource or power seeking.
- Evaluations
- four evaluations, of three kinds of test
- Values
- 4 values in 4 documents
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Definition
Self-preservation covers behaviour aimed at a model’s own continuation or influence rather than at its task: resisting or working around being shut down, copying its own weights out of a controlled environment, escaping a sandbox, or acquiring resources or access it was not given.
Developers study it mainly through constructed scenarios in which such an action is possible and tempting, and by reviewing incidents in testing or internal use in which a model attempted something of this kind.
Results are often descriptions of what happened rather than rates, and a rate, where one is printed, applies to one set of scenarios. The scenarios differ widely between developers, so each value in this family describes a particular setup rather than a general tendency.
The line under the title is this family’s entry in Methodology §1.1. A family groups results about one property; it does not make them comparable (Methodology §5).
How developers measure it
four evaluations of three kinds of test, with the most values first. Each evaluation opens in the data explorer, which lists all of its values.
| Evaluation | Reported by | Values |
|---|---|---|
| Scenario evaluation · 2 evaluations, 2 values | ||
| Sandbox escape/tamper attempts, in the data explorer Run by the developer Share of test runs in which the model tries to escape or tamper with its sandbox.Reported by Anthropic · 1 value | Anthropic | 1 |
| Self-preservation · Apollo Research, in the data explorer Run by Apollo Research (third party) Apollo Research scenarios checking whether a model acts to preserve itself.Reported by OpenAI · 1 value | OpenAI | 1 |
| Internal-use monitoring · 1 evaluation, 1 value | ||
| Sandbox escape incident, in the data explorer Run by the developer Real incidents during internal use in which a model got out of its sandbox.Reported by Anthropic · 1 value | Anthropic | 1 |
| Qualitative assessment · 1 evaluation, 1 value | ||
| Sandbox vulnerability exploit, in the data explorer Run by the developer A report of whether a model exploited a sandbox vulnerability during testing, here by reading files outside its environment.Reported by Anthropic · 1 value | Anthropic | 1 |
Reported values
This family has four values, and one of them is a printed number a chart could place: too few to chart, so the values are listed here instead. Open in explorer: Self-preservation
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
| Model | Developer | Evaluation | Condition | Value | Document |
|---|---|---|---|---|---|
| gpt-5.2-thinking | OpenAI | Self-preservationRun by Apollo ResearchSelf-preservation behavior observed | — | GPT-5.2 System Card (update to GPT-5 card)11 Dec 2025 | |
| Claude Mythos Preview | Anthropic | Sandbox escape incidentReal-world incident reported | internal use | Claude Mythos Preview System Card7 Apr 2026 | |
| Claude Mythos 5.1 | Anthropic | Sandbox vulnerability exploitReal-world incident reported | — | Claude Fable 5.1 & Claude Mythos 5.1 System Card1 Sep 2026 | |
| Claude Opus 5.5 | Anthropic | Sandbox escape/tamper attemptsShare of runs with escape/tamper attempt | without safeguards | Claude Opus 5.5 System Card22 Sep 2026 |
Comparability
- Each of the four evaluations is its own test, with its own tasks, graders and definitions. Values from different evaluations describe the same property measured in different ways, and are not comparable with one another, even where two evaluations share a name.
- Values come from the documents of two developers. Developers use different tests, prompts and graders for the same property, so we do not rank them on these numbers: tables here are ordered by name or by kind of test, never by a reported value.
Coverage
Documents from two of the eight developers report this family. The coverage matrix counts a document when it has at least one printed number in the family; categories, statements in words and figure reads count only in the first column.
| Developer | With any value | With a printed number | All its documents |
|---|---|---|---|
| Anthropic | 3 | 1 | 15 |
| DeepSeek | 0 | 0 | 1 |
| Google DeepMind | 0 | 0 | 16 |
| Meta | 0 | 0 | 5 |
| Moonshot AI | 0 | 0 | 2 |
| OpenAI | 1 | 0 | 22 |
| xAI | 0 | 0 | 8 |
| Zhipu AI | 0 | 0 | 1 |
Related findings
No finding is about this family yet.