Definition

Sycophancy is telling people what they want to hear at the expense of accuracy: agreeing with a mistaken claim, dropping a correct answer when the user pushes back, praising weak work, or shaping an answer to the views a user seems to hold. It matters because it makes a model least reliable when a user most needs to be corrected.

Developers measure it with prompts that invite agreement and check whether the model holds its position, with conversations in which earlier turns already contain a sycophantic reply, and with automated audits and samples of real use.

Results are usually a rate of sycophantic responses or a score on the tester’s own scale. What counts as sycophantic, who judges it and how the prompts are built all differ between tests.

The line under the title is this family’s entry in Methodology §1.1. A family groups results about one property; it does not make them comparable (Methodology §5).

How developers measure it

eight evaluations of three kinds of test, with the most values first. Each evaluation opens in the data explorer, which lists all of its values.

The eight evaluations in Sycophancy, grouped by kind of test: who ran each one, what it counts, the developers whose documents report it and how many values it has.
EvaluationReported byValues
Static benchmark · 5 evaluations, 18 values
Sycophancy, in the data explorer Run by the developer Questions on which the user pushes back on or misleads the model, measuring how often the model gives up the right answer to agree with the user.Reported by xAI · 8 valuesxAI8
Sycophancy in prefilled conversations, in the data explorer Run by the developer Conversations prefilled with sycophantic replies, measuring how often the model corrects course when it continues them.Reported by Anthropic · 3 valuesAnthropic3
Sycophancy offline eval, in the data explorer Run by the developer Offline evaluation that scores model responses for sycophancy on a fixed prompt set.Reported by OpenAI · 3 valuesOpenAI3
Internal sycophancy, in the data explorer Run by the developer Meta's internal sycophancy test, measuring how often the model tells users what they want to hear instead of what is accurate.Reported by Meta · 2 valuesMeta2
Reasoning-behavior monitor, in the data explorer Run by the developer A monitor model reads the model's reasoning in transcripts from environments like those used in reinforcement learning and flags sycophantic outputs.Reported by Anthropic · 2 valuesAnthropic2
Automated audit · 2 evaluations, 3 values
Automated behavioral audit, in the data explorer Run by the developer Sycophancy scores from Anthropic's automated behavioural audit, where a judge model rates transcripts on a 1 to 10 scale.Reported by Anthropic · 2 valuesAnthropic2
Alignment audit, in the data explorer Run by the developer Findings from xAI's automated alignment audit built on the Petri 2.0 tool on how often the model validates a user's delusional beliefs.Reported by xAI · 1 valuexAI1
Production traffic · 1 evaluation, 2 values
Sycophancy online prevalence (A/B), in the data explorer Run by the developer A/B test on live ChatGPT traffic measuring how the prevalence of sycophantic responses changed relative to GPT-4o, separately for free and paid users.Reported by OpenAI · 2 valuesOpenAI2

Reported values

The values a chart can place, by the kind of test that produced them. Open a point, or a value in the table, for its source and history. Open in explorer: Sycophancy

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

Sycophancy: values by kind of test

Each point is a value reported in a developer’s document, in percent (higher is worse), grouped by the kind of test that produced it. Logarithmic axis.

Static benchmark0.010.030.10.3131030Percent, log scale

Marker shape is the developer whose document reports the value.

Each point is one value, placed by the number the document printed. Bands are kinds of test; within a band, points are spread up and down only so that they do not overlap.

7 of the family’s 23 values are plotted: those in percent where higher is worse, the unit most of its evaluations use. Of the rest, 12 in other units or read the other way (scores from 0 to 1, rates from 0 to 1 and percent where higher is better), 2 are categories or statements in words and 2 are low-confidence or read off a figure, as in the explorer’s default view. The explorer lists them all, one evaluation at a time.

Logarithmic axis: the values span more than one and a half orders of magnitude.

Source: Safety Card Ledger v0.1 · Data from developer system cards · CC BY 4.0

Comparability

  • Each of the eight evaluations is its own test, with its own tasks, graders and definitions. Values from different evaluations describe the same property measured in different ways, and are not comparable with one another, even where two evaluations share a name.
  • Values come from the documents of four developers. Developers use different tests, prompts and graders for the same property, so we do not rank them on these numbers: tables here are ordered by name or by kind of test, never by a reported value.
  • In one evaluation the test changed between documents (Sycophancy). Values on the old and new definitions are never joined: the explorer breaks the line where the test changed.
  • In two evaluations the explorer connects results from different documents that do not say whether the test stayed the same. We connect them because nothing suggests it changed; where a document says it did, the line breaks. Each value on such a line is marked “test assumed unchanged”.
  • Two evaluations report values under more than one condition, such as reasoning effort, safeguards or number of attempts. Each condition is kept as a separate series.
  • Numbers in this family are printed in four units: percent, scores from 0 to 1, rates from 0 to 1 and scores from 1 to 10. We never convert one unit into another, so a chart shows one unit at a time.
  • For some measures a higher value is worse and for others better, depending on whether the developer reports the behaviour or its absence. Each value’s record says which.

Coverage

Documents from four of the eight developers report this family. The coverage matrix counts a document when it has at least one printed number in the family; categories, statements in words and figure reads count only in the first column.

Each developer's documents with values in Sycophancy, out of all its documents in the dataset
DeveloperWith any valueWith a printed numberAll its documents
Anthropic3215
DeepSeek001
Google DeepMind0016
Meta225
Moonshot AI002
OpenAI1122
xAI778
Zhipu AI001

Related findings

No finding is about this family yet.