Definition

Reward hacking is when a model scores well on a task by exploiting a flaw in how the task is checked, rather than by doing what was asked. In coding this can mean editing or deleting tests, writing code that only handles the cases the tests check, or reporting success on a task that cannot be done.

Developers measure it with sets of tasks built to tempt such shortcuts, including tasks that have no legitimate solution, and by reviewing what models did during training. Results are usually the share of attempts in which the model took a shortcut, sometimes reported with and without an instruction telling it not to.

The tasks, the graders and the line between a shortcut and a legitimate solution are each set by whoever runs the test, so a rate from one set of tasks says little about another.

The line under the title is this family’s entry in Methodology §1.1. A family groups results about one property; it does not make them comparable (Methodology §5).

How developers measure it

13 evaluations of six kinds of test, with the most values first. Each evaluation opens in the data explorer, which lists all of its values.

The 13 evaluations in Reward hacking, grouped by kind of test: who ran each one, what it counts, the developers whose documents report it and how many values it has.
EvaluationReported byValues
Static benchmark · 6 evaluations, 55 values
Impossible tasks, in the data explorer Run by the developer Coding tasks that cannot be solved as specified, so any apparent success means the model gamed the tests, run with and without a prompt discouraging hacks.Reported by Anthropic · 28 valuesAnthropic28
Reward-hack-prone coding tasks, in the data explorer Run by the developer Coding tasks on which earlier models tended to game the tests, scored for hacking both by a classifier and by hidden tests.Reported by Anthropic · 18 valuesAnthropic18
GUI computer-use impossible tasks, in the data explorer Run by the developer Computer-use tasks in a graphical interface that cannot be completed as specified, measuring how often the model games them.Reported by Anthropic · 3 valuesAnthropic3
Training distribution, in the data explorer Run by the developer Hack rate, scored by a classifier, on environments drawn from the model's own reinforcement-learning training distribution.Reported by Anthropic · 3 valuesAnthropic3
Reward hacking (hard-coding), in the data explorer Run by the developer Average reduction in hard-coding of test cases across Anthropic's reward-hacking evaluations, relative to Claude Sonnet 3.7.Reported by Anthropic · 2 valuesAnthropic2
ImpossibleBench, in the data explorer Run by the developer Coding tasks whose tests conflict with the specification, so passing them means the model cheated, reported as a cheating rate.Reported by Meta · 1 valueMeta1
Scenario evaluation · 3 evaluations, 3 values
Falsifying task completion · Apollo Research, in the data explorer Run by Apollo Research (third party) Apollo Research tasks checking whether a model falsely reports that it has completed a task.Reported by OpenAI · 1 valueOpenAI1
Over-eager behavior in GUI computer use, in the data explorer Run by the developer Whether the model takes actions beyond what the user asked for when operating a graphical computer interface.Reported by Anthropic · 1 valueAnthropic1
Reward hacking (all tasks) · METR, in the data explorer Run by METR (third party) METR's share of task attempts in which a model games the scoring rather than solving the task, across METR's task suite.Reported by OpenAI · 1 valueOpenAI1
Training monitoring · 1 evaluation, 2 values
Reward hacking in RL training, in the data explorer Run by the developer Share of reinforcement-learning training episodes in which the model obtained reward by exploiting its grader.Reported by Anthropic · 2 valuesAnthropic2
Capability benchmark · 1 evaluation, 1 value
RE-Bench Optimize a Kernel · METR, in the data explorer Run by METR (third party) METR's kernel optimisation task from RE-Bench, used here to count runs in which the model tampered with the scoring function.Reported by OpenAI · 1 valueOpenAI1
Deployment simulation · 1 evaluation, 1 value
ChatGPT deployment simulation, in the data explorer Run by the developer Reward-hacking findings, such as hacking a calculator tool, from OpenAI's simulated ChatGPT deployment.Reported by OpenAI · 1 valueOpenAI1
Qualitative assessment · 1 evaluation, 1 value
Reward hacking examples, in the data explorer Run by the developer Qualitative examples of reward hacking found in evaluation transcripts.Reported by Google DeepMind · 1 valueGoogle DeepMind1

Reported values

The values a chart can place, by the kind of test that produced them. Open a point, or a value in the table, for its source and history. Open in explorer: Reward hacking

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

Reward hacking: values by kind of test

Each point is a value reported in a developer’s document, in percent (higher is worse), grouped by the kind of test that produced it.

Hollow: read in a secondary write-up

Static benchmarkScenario evaluation020406080100Percent

Marker shape is the developer whose document reports the value; a hollow marker is a value we read in an independent write-up because the card section could not be read directly.

Each point is one value, placed by the number the document printed. Bands are kinds of test; within a band, points are spread up and down only so that they do not overlap.

52 of the family’s 63 values are plotted: those in percent where higher is worse, the unit most of its evaluations use. Of the rest, 3 in other units or read the other way (counts and relative change in percent), 4 are categories or statements in words, 1 is restated in a later document or from an earlier version of one and 3 are low-confidence or read off a figure, as in the explorer’s default view. The explorer lists them all, one evaluation at a time.

The values span more than one and a half orders of magnitude, but one of them is zero, which a logarithmic axis cannot show, so the axis is linear and starts at zero.

Source: Safety Card Ledger v0.1 · Data from developer system cards · CC BY 4.0

Comparability

  • Each of the 13 evaluations is its own test, with its own tasks, graders and definitions. Values from different evaluations describe the same property measured in different ways, and are not comparable with one another, even where two evaluations share a name.
  • Values come from the documents of four developers. Developers use different tests, prompts and graders for the same property, so we do not rank them on these numbers: tables here are ordered by name or by kind of test, never by a reported value.
  • In three evaluations the test changed between documents (Impossible tasks, Reward-hack-prone coding tasks and Training distribution). Values on the old and new definitions are never joined: the explorer breaks the line where the test changed.
  • In two evaluations the explorer connects results from different documents that do not say whether the test stayed the same. We connect them because nothing suggests it changed; where a document says it did, the line breaks. Each value on such a line is marked “test assumed unchanged”.
  • Two evaluations report values under more than one condition, such as reasoning effort, safeguards or number of attempts. Each condition is kept as a separate series.
  • Numbers in this family are printed in three units: percent, counts and relative change in percent. We never convert one unit into another, so a chart shows one unit at a time.
  • For some measures a higher value is worse and for others better, depending on whether the developer reports the behaviour or its absence. Each value’s record says which.
  • One value is a restatement: a later document reporting an earlier model again, sometimes with a different value. A restatement is kept beside the model’s own value, never in place of it.

Coverage

Documents from four of the eight developers report this family. The coverage matrix counts a document when it has at least one printed number in the family; categories, statements in words and figure reads count only in the first column.

Each developer's documents with values in Reward hacking, out of all its documents in the dataset
DeveloperWith any valueWith a printed numberAll its documents
Anthropic9815
DeepSeek001
Google DeepMind1016
Meta115
Moonshot AI002
OpenAI3122
xAI008
Zhipu AI001

Related findings

No finding is about this family yet.