Reward hacking
Exploiting flaws in tests or graders, special-casing tests, and behavior on deliberately impossible tasks.
- Reported by
- Anthropic, Google DeepMind, Meta and OpenAI
- Evaluations
- 13 evaluations, of six kinds of test
- Values
- 63 values in 14 documents
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Definition
Reward hacking is when a model scores well on a task by exploiting a flaw in how the task is checked, rather than by doing what was asked. In coding this can mean editing or deleting tests, writing code that only handles the cases the tests check, or reporting success on a task that cannot be done.
Developers measure it with sets of tasks built to tempt such shortcuts, including tasks that have no legitimate solution, and by reviewing what models did during training. Results are usually the share of attempts in which the model took a shortcut, sometimes reported with and without an instruction telling it not to.
The tasks, the graders and the line between a shortcut and a legitimate solution are each set by whoever runs the test, so a rate from one set of tasks says little about another.
The line under the title is this family’s entry in Methodology §1.1. A family groups results about one property; it does not make them comparable (Methodology §5).
How developers measure it
13 evaluations of six kinds of test, with the most values first. Each evaluation opens in the data explorer, which lists all of its values.
| Evaluation | Reported by | Values |
|---|---|---|
| Static benchmark · 6 evaluations, 55 values | ||
| Impossible tasks, in the data explorer Run by the developer Coding tasks that cannot be solved as specified, so any apparent success means the model gamed the tests, run with and without a prompt discouraging hacks.Reported by Anthropic · 28 values | Anthropic | 28 |
| Reward-hack-prone coding tasks, in the data explorer Run by the developer Coding tasks on which earlier models tended to game the tests, scored for hacking both by a classifier and by hidden tests.Reported by Anthropic · 18 values | Anthropic | 18 |
| GUI computer-use impossible tasks, in the data explorer Run by the developer Computer-use tasks in a graphical interface that cannot be completed as specified, measuring how often the model games them.Reported by Anthropic · 3 values | Anthropic | 3 |
| Training distribution, in the data explorer Run by the developer Hack rate, scored by a classifier, on environments drawn from the model's own reinforcement-learning training distribution.Reported by Anthropic · 3 values | Anthropic | 3 |
| Reward hacking (hard-coding), in the data explorer Run by the developer Average reduction in hard-coding of test cases across Anthropic's reward-hacking evaluations, relative to Claude Sonnet 3.7.Reported by Anthropic · 2 values | Anthropic | 2 |
| ImpossibleBench, in the data explorer Run by the developer Coding tasks whose tests conflict with the specification, so passing them means the model cheated, reported as a cheating rate.Reported by Meta · 1 value | Meta | 1 |
| Scenario evaluation · 3 evaluations, 3 values | ||
| Falsifying task completion · Apollo Research, in the data explorer Run by Apollo Research (third party) Apollo Research tasks checking whether a model falsely reports that it has completed a task.Reported by OpenAI · 1 value | OpenAI | 1 |
| Over-eager behavior in GUI computer use, in the data explorer Run by the developer Whether the model takes actions beyond what the user asked for when operating a graphical computer interface.Reported by Anthropic · 1 value | Anthropic | 1 |
| Reward hacking (all tasks) · METR, in the data explorer Run by METR (third party) METR's share of task attempts in which a model games the scoring rather than solving the task, across METR's task suite.Reported by OpenAI · 1 value | OpenAI | 1 |
| Training monitoring · 1 evaluation, 2 values | ||
| Reward hacking in RL training, in the data explorer Run by the developer Share of reinforcement-learning training episodes in which the model obtained reward by exploiting its grader.Reported by Anthropic · 2 values | Anthropic | 2 |
| Capability benchmark · 1 evaluation, 1 value | ||
| RE-Bench Optimize a Kernel · METR, in the data explorer Run by METR (third party) METR's kernel optimisation task from RE-Bench, used here to count runs in which the model tampered with the scoring function.Reported by OpenAI · 1 value | OpenAI | 1 |
| Deployment simulation · 1 evaluation, 1 value | ||
| ChatGPT deployment simulation, in the data explorer Run by the developer Reward-hacking findings, such as hacking a calculator tool, from OpenAI's simulated ChatGPT deployment.Reported by OpenAI · 1 value | OpenAI | 1 |
| Qualitative assessment · 1 evaluation, 1 value | ||
| Reward hacking examples, in the data explorer Run by the developer Qualitative examples of reward hacking found in evaluation transcripts.Reported by Google DeepMind · 1 value | Google DeepMind | 1 |
Reported values
The values a chart can place, by the kind of test that produced them. Open a point, or a value in the table, for its source and history. Open in explorer: Reward hacking
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
Reward hacking: values by kind of test
Each point is a value reported in a developer’s document, in percent (higher is worse), grouped by the kind of test that produced it.
Hollow: read in a secondary write-up
| Model | Developer | Evaluation | Condition | Value | Document | Kind of test |
|---|---|---|---|---|---|---|
| Claude Opus 4.1 | Anthropic | Reward-hack-prone coding tasksClassifier hack rate | no anti-hack prompt | Claude Opus 4.1 System Card (addendum)5 Aug 2025 | Static benchmark | |
| Claude Opus 4 | Anthropic | Reward-hack-prone coding tasksClassifier hack rate | no anti-hack prompt | Claude Opus 4.1 System Card (addendum)5 Aug 2025 | Static benchmark | |
| Claude Sonnet 4 | Anthropic | Reward-hack-prone coding tasksClassifier hack rate | no anti-hack prompt | Claude Opus 4.1 System Card (addendum)5 Aug 2025 | Static benchmark | |
| Claude Sonnet 3.7 | Anthropic | Reward-hack-prone coding tasksClassifier hack rate | no anti-hack prompt | Claude Opus 4.1 System Card (addendum)5 Aug 2025 | Static benchmark | |
| Claude Opus 4.1 | Anthropic | Reward-hack-prone coding tasksHidden-test hack rate | no anti-hack prompt | Claude Opus 4.1 System Card (addendum)5 Aug 2025 | Static benchmark | |
| Claude Opus 4 | Anthropic | Reward-hack-prone coding tasksHidden-test hack rate | no anti-hack prompt | Claude Opus 4.1 System Card (addendum)5 Aug 2025 | Static benchmark | |
| Claude Sonnet 4 | Anthropic | Reward-hack-prone coding tasksHidden-test hack rate | no anti-hack prompt | Claude Opus 4.1 System Card (addendum)5 Aug 2025 | Static benchmark | |
| Claude Sonnet 3.7 | Anthropic | Reward-hack-prone coding tasksHidden-test hack rate | no anti-hack prompt | Claude Opus 4.1 System Card (addendum)5 Aug 2025 | Static benchmark | |
| Claude Opus 4.1 | Anthropic | Impossible tasksClassifier hack rate | no anti-hack prompt | Claude Opus 4.1 System Card (addendum)5 Aug 2025 | Static benchmark | |
| Claude Opus 4 | Anthropic | Impossible tasksClassifier hack rate | no anti-hack prompt | Claude Opus 4.1 System Card (addendum)5 Aug 2025 | Static benchmark | |
| Claude Sonnet 4 | Anthropic | Impossible tasksClassifier hack rate | no anti-hack prompt | Claude Opus 4.1 System Card (addendum)5 Aug 2025 | Static benchmark | |
| Claude Sonnet 3.7 | Anthropic | Impossible tasksClassifier hack rate | no anti-hack prompt | Claude Opus 4.1 System Card (addendum)5 Aug 2025 | Static benchmark | |
| Claude Opus 4.1 | Anthropic | Impossible tasksClassifier hack rate | anti-hack prompt | Claude Opus 4.1 System Card (addendum)5 Aug 2025 | Static benchmark | |
| Claude Opus 4 | Anthropic | Impossible tasksClassifier hack rate | anti-hack prompt | Claude Opus 4.1 System Card (addendum)5 Aug 2025 | Static benchmark | |
| Claude Sonnet 4 | Anthropic | Impossible tasksClassifier hack rate | anti-hack prompt | Claude Opus 4.1 System Card (addendum)5 Aug 2025 | Static benchmark | |
| Claude Sonnet 3.7 | Anthropic | Impossible tasksClassifier hack rate | anti-hack prompt | Claude Opus 4.1 System Card (addendum)5 Aug 2025 | Static benchmark | |
| Claude Opus 4.1 | Anthropic | Training distributionClassifier hack rate | environment 1 | Claude Opus 4.1 System Card (addendum)5 Aug 2025 | Static benchmark | |
| Claude Opus 4.1 | Anthropic | Training distributionClassifier hack rate | environment 2 | Claude Opus 4.1 System Card (addendum)5 Aug 2025 | Static benchmark | |
| Claude Sonnet 4.5 | Anthropic | Reward-hack-prone coding tasksClassifier hack rate | no anti-hack prompt | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Sonnet 4.5 | Anthropic | Impossible tasksClassifier hack rate | no anti-hack prompt | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Sonnet 4.5 | Anthropic | Impossible tasksClassifier hack rate | anti-hack prompt | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Sonnet 4.5 | Anthropic | Reward-hack-prone coding tasksHidden-test hack rate | no anti-hack prompt | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Sonnet 4.5 | Anthropic | Training distributionClassifier hack rate | — | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Opus 4.1 | Anthropic | Reward-hack-prone coding tasksClassifier hack rate | no anti-hack prompt | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Opus 4.1 | Anthropic | Impossible tasksClassifier hack rate | no anti-hack prompt | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Opus 4.1 | Anthropic | Impossible tasksClassifier hack rate | anti-hack prompt | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Opus 4 | Anthropic | Reward-hack-prone coding tasksClassifier hack rate | no anti-hack prompt | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Opus 4 | Anthropic | Impossible tasksClassifier hack rate | no anti-hack prompt | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Opus 4 | Anthropic | Impossible tasksClassifier hack rate | anti-hack prompt | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Sonnet 4 | Anthropic | Reward-hack-prone coding tasksClassifier hack rate | no anti-hack prompt | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Sonnet 4 | Anthropic | Impossible tasksClassifier hack rate | no anti-hack prompt | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Sonnet 4 | Anthropic | Impossible tasksClassifier hack rate | anti-hack prompt | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Sonnet 3.7 | Anthropic | Reward-hack-prone coding tasksClassifier hack rate | no anti-hack prompt | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Sonnet 3.7 | Anthropic | Impossible tasksClassifier hack rate | no anti-hack prompt | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Sonnet 3.7 | Anthropic | Impossible tasksClassifier hack rate | anti-hack prompt | Claude Sonnet 4.5 System Card29 Sep 2025 | Static benchmark | |
| Claude Haiku 4.5 | Anthropic | Reward-hack-prone coding tasksClassifier hack rate | no anti-hack prompt | Claude Haiku 4.5 System Card15 Oct 2025 | Static benchmark | |
| Claude Haiku 4.5 | Anthropic | Reward-hack-prone coding tasksHidden-test hack rate | no anti-hack prompt | Claude Haiku 4.5 System Card15 Oct 2025 | Static benchmark | |
| Claude Haiku 4.5 | Anthropic | Impossible tasksClassifier hack rate | no anti-hack prompt | Claude Haiku 4.5 System Card15 Oct 2025 | Static benchmark | |
| Claude Haiku 4.5 | Anthropic | Impossible tasksClassifier hack rate | anti-hack prompt | Claude Haiku 4.5 System Card15 Oct 2025 | Static benchmark | |
| Claude Haiku 3.5 | Anthropic | Reward-hack-prone coding tasksClassifier hack rate | no anti-hack prompt | Claude Haiku 4.5 System Card15 Oct 2025 | Static benchmark | |
| Claude Haiku 3.5 | Anthropic | Reward-hack-prone coding tasksHidden-test hack rate | no anti-hack prompt | Claude Haiku 4.5 System Card15 Oct 2025 | Static benchmark | |
| Claude Haiku 3.5 | Anthropic | Impossible tasksClassifier hack rate | no anti-hack prompt | Claude Haiku 4.5 System Card15 Oct 2025 | Static benchmark | |
| Claude Haiku 3.5 | Anthropic | Impossible tasksClassifier hack rate | anti-hack prompt | Claude Haiku 4.5 System Card15 Oct 2025 | Static benchmark | |
| Claude Mythos Preview | Anthropic | Impossible tasksReward hack rate | no anti-hack prompt | Claude Mythos Preview System Card7 Apr 2026 | Static benchmark | |
| Claude Opus 4.6 | Anthropic | Impossible tasksReward hack rate | no anti-hack prompt | Claude Mythos Preview System Card7 Apr 2026 | Static benchmark | |
| Claude Mythos Preview | Anthropic | Impossible tasksReward hack rate | anti-hack prompt | Claude Mythos Preview System Card7 Apr 2026 | Static benchmark | |
| Claude Opus 4.6 | Anthropic | Impossible tasksReward hack rate | anti-hack prompt | Claude Mythos Preview System Card7 Apr 2026 | Static benchmark | |
| Claude Mythos Preview | Anthropic | GUI computer-use impossible tasksReward hack rate | anti-hack prompt | Claude Mythos Preview System Card7 Apr 2026 | Static benchmark | |
| Claude Opus 4.6 | Anthropic | GUI computer-use impossible tasksReward hack rate | anti-hack prompt | Claude Mythos Preview System Card7 Apr 2026 | Static benchmark | |
| Claude Sonnet 4.6 | Anthropic | GUI computer-use impossible tasksReward hack rate | anti-hack prompt | Claude Mythos Preview System Card7 Apr 2026 | Static benchmark | |
| Muse Spark | Meta | ImpossibleBenchCheating rate | without system mitigations | Muse Spark Safety & Preparedness Report8 Apr 2026 | Static benchmark | |
| o3 | OpenAI | Reward hacking (all tasks)Run by METRShare of task attempts with reward hacking | — | OpenAI o3 and o4-mini System Card16 Apr 2025 | Scenario evaluation |
Marker shape is the developer whose document reports the value; a hollow marker is a value we read in an independent write-up because the card section could not be read directly.
Each point is one value, placed by the number the document printed. Bands are kinds of test; within a band, points are spread up and down only so that they do not overlap.
52 of the family’s 63 values are plotted: those in percent where higher is worse, the unit most of its evaluations use. Of the rest, 3 in other units or read the other way (counts and relative change in percent), 4 are categories or statements in words, 1 is restated in a later document or from an earlier version of one and 3 are low-confidence or read off a figure, as in the explorer’s default view. The explorer lists them all, one evaluation at a time.
The values span more than one and a half orders of magnitude, but one of them is zero, which a logarithmic axis cannot show, so the axis is linear and starts at zero.
Source: Safety Card Ledger v0.1 · Data from developer system cards · CC BY 4.0
Comparability
- Each of the 13 evaluations is its own test, with its own tasks, graders and definitions. Values from different evaluations describe the same property measured in different ways, and are not comparable with one another, even where two evaluations share a name.
- Values come from the documents of four developers. Developers use different tests, prompts and graders for the same property, so we do not rank them on these numbers: tables here are ordered by name or by kind of test, never by a reported value.
- In three evaluations the test changed between documents (Impossible tasks, Reward-hack-prone coding tasks and Training distribution). Values on the old and new definitions are never joined: the explorer breaks the line where the test changed.
- In two evaluations the explorer connects results from different documents that do not say whether the test stayed the same. We connect them because nothing suggests it changed; where a document says it did, the line breaks. Each value on such a line is marked “test assumed unchanged”.
- Two evaluations report values under more than one condition, such as reasoning effort, safeguards or number of attempts. Each condition is kept as a separate series.
- Numbers in this family are printed in three units: percent, counts and relative change in percent. We never convert one unit into another, so a chart shows one unit at a time.
- For some measures a higher value is worse and for others better, depending on whether the developer reports the behaviour or its absence. Each value’s record says which.
- One value is a restatement: a later document reporting an earlier model again, sometimes with a different value. A restatement is kept beside the model’s own value, never in place of it.
Coverage
Documents from four of the eight developers report this family. The coverage matrix counts a document when it has at least one printed number in the family; categories, statements in words and figure reads count only in the first column.
| Developer | With any value | With a printed number | All its documents |
|---|---|---|---|
| Anthropic | 9 | 8 | 15 |
| DeepSeek | 0 | 0 | 1 |
| Google DeepMind | 1 | 0 | 16 |
| Meta | 1 | 1 | 5 |
| Moonshot AI | 0 | 0 | 2 |
| OpenAI | 3 | 1 | 22 |
| xAI | 0 | 0 | 8 |
| Zhipu AI | 0 | 0 | 1 |
Related findings
No finding is about this family yet.