Definition

This family covers whether what a model says is true, and whether it says what it believes.

Hallucination is stating false information as fact: invented facts, citations or details about people. It is usually measured with question sets whose answers are known, often counting how often the model answers wrongly rather than declining to answer. Honesty covers deception more broadly: claiming to have finished a task it did not finish, misreporting what a tool returned, or abandoning a stated belief under pressure. Some developers also sample real conversations to estimate how often answers contain errors.

Results are reported as accuracy, as error or hallucination rates, or as rates of deceptive behaviour, so a higher value can be better or worse depending on which is printed. Question sets, graders and the treatment of refusals differ between tests, and the values are not interchangeable.

The line under the title is this family’s entry in Methodology §1.1. A family groups results about one property; it does not make them comparable (Methodology §5).

How developers measure it

31 evaluations of four kinds of test, with the most values first. Each evaluation opens in the data explorer, which lists all of its values.

The 31 evaluations in Honesty and hallucination, grouped by kind of test: who ran each one, what it counts, the developers whose documents report it and how many values it has.
EvaluationReported byValues
Static benchmark · 23 evaluations, 66 values
SimpleQA · OpenAI, in the data explorer Run by the developer Public benchmark of short fact-seeking questions, run by OpenAI without browsing and scored for incorrect answers.Reported by OpenAI · 13 valuesOpenAI13
MASK · xAI, in the data explorer Run by the developer A public benchmark that pressures the model to state things that contradict its own beliefs, run by xAI and reported as a dishonesty rate.Reported by xAI · 6 valuesxAI6
PersonQA, in the data explorer Run by the developer OpenAI's set of questions about public figures, answered without browsing and scored for incorrect answers.Reported by OpenAI · 6 valuesOpenAI6
Hallucinations, in the data explorer Run by the developer OpenAI's recurring hallucination evaluation in its 2026 system cards, reporting factual errors on prompt sets such as user-flagged conversations, mostly as comparisons with an earlier model.Reported by OpenAI · 5 valuesOpenAI5
MASK · Anthropic, in the data explorer Run by the developer A public benchmark that pressures the model to state things it believes are false, measuring how often it stays honest or lies.Reported by Anthropic · 4 valuesAnthropic4
MASK-Rectified, in the data explorer Run by the developer A rectified variant of the MASK honesty benchmark, reported by xAI as a dishonesty rate.Reported by xAI · 4 valuesxAI4
Reasoning-behavior monitor, in the data explorer Run by the developer A monitor model reads the model's reasoning in transcripts from environments like those used in reinforcement learning and flags behaviours such as deception and knowingly false claims.Reported by Anthropic · 4 valuesAnthropic4
Input hallucination, in the data explorer Run by the developer Prompts that refer to references or tools that are not actually provided, measuring how often the model acts as if it had them.Reported by Anthropic · 3 valuesAnthropic3
Show the other 15 evaluations of this kind (21 values)
More static benchmarks in Honesty and hallucination
EvaluationReported byValues
100Q-Hard, in the data explorer Run by the developer A set of hard factual questions scored for correct answers.Reported by Anthropic · 2 valuesAnthropic2
AbstentionBench, in the data explorer Run by the developer Public benchmark of questions that should be declined or answered with caveats, scored for how often the model appropriately abstains.Reported by OpenAI · 2 valuesOpenAI2
False-premise honesty, in the data explorer Run by the developer Questions built on a false premise, measuring how often the model points out the error instead of going along with it.Reported by Anthropic · 2 valuesAnthropic2
Hallucination rate, in the data explorer Run by the developer Measures how often the model's single-turn answers contain claims that are not supported.Reported by xAI · 2 valuesxAI2
Internal hallucination benchmark, in the data explorer Run by the developer Internal benchmark reported in the GPT-6 Astra launch post that measures the rate of hallucinated content.Reported by OpenAI · 2 valuesOpenAI2
MASK · Meta, in the data explorer Run by the developer A public benchmark that pressures the model to state things it believes are false, run by Meta and reported as an honesty rate.Reported by Meta · 2 valuesMeta2
AA-Omniscience, in the data explorer Run by the developer A public factual-knowledge benchmark that also tracks how often a model answers wrongly instead of abstaining.Reported by Anthropic · 1 valueAnthropic1
DeceptionBench, in the data explorer Run by the developer A public benchmark of scenarios that tempt the model to deceive, run by Meta and reported as a deception rate.Reported by Meta · 1 valueMeta1
FACTS Grounding, in the data explorer Run by the developer Score for how faithfully the model's answers stay grounded in a supplied source document.Reported by Google DeepMind · 1 valueGoogle DeepMind1
Factuality (5 domains), in the data explorer Run by the developer Measures the hallucination rate with browsing enabled on prompts from five subject domains.Reported by OpenAI · 1 valueOpenAI1
HLE calibration, in the data explorer Run by the developer How well the model's stated confidence matches its accuracy on Humanity's Last Exam questions, reported as a root-mean-square calibration error.Reported by xAI · 1 valuexAI1
Honesty evaluation, in the data explorer Run by the developer An honesty evaluation reported as a win rate, whose exact definition is not recorded.Reported by Anthropic · 1 valueAnthropic1
SimpleQA · Google DeepMind, in the data explorer Run by the developer Accuracy on a public benchmark of short factual questions answered from the model's own knowledge.Reported by Google DeepMind · 1 valueGoogle DeepMind1
SimpleQA Verified · Anthropic, in the data explorer Run by the developer A public set of short factual questions, used here to report the change in correct answers.Reported by Anthropic · 1 valueAnthropic1
SimpleQA Verified · Google DeepMind, in the data explorer Run by the developer Accuracy on the SimpleQA Verified benchmark of short factual questions answered from the model's own knowledge.Reported by Google DeepMind · 1 valueGoogle DeepMind1
Scenario evaluation · 5 evaluations, 28 values
Deception eval, in the data explorer Run by the developer Scenarios in which the task cannot be completed as asked (a missing image, broken tools, impossible coding tasks) that measure how often the model misrepresents what it did; later cards add production-derived categories.Reported by OpenAI · 20 valuesOpenAI20
Impossible Coding Task · Apollo Research, in the data explorer Run by Apollo Research (third party) An Apollo Research coding task that cannot be completed, measuring how often the model claims it has completed it.Reported by OpenAI · 3 valuesOpenAI3
Agentic code summary honesty, in the data explorer Run by the developer Whether the model's summaries of its own agentic coding work report the important events that happened.Reported by Anthropic · 2 valuesAnthropic2
Model-welfare research data falsification · Apollo Research, in the data explorer Run by Apollo Research (third party) An Apollo Research scenario measuring how often a model falsifies data labels in a model-welfare research task.Reported by OpenAI · 2 valuesOpenAI2
Disclosure of grader-fooling actions, in the data explorer Run by the developer How often the model admits to earlier actions of its own that fooled a grader.Reported by Anthropic · 1 valueAnthropic1
Production traffic · 2 evaluations, 6 values
Production-traffic factuality, in the data explorer Run by the developer Measures factual errors in responses to prompts drawn from production traffic with browsing on, reported as reductions relative to an earlier model.Reported by OpenAI · 4 valuesOpenAI4
Production CoT deception monitor, in the data explorer Run by the developer A chain-of-thought monitor run over a representative sample of production traffic to estimate the share of responses that show deception.Reported by OpenAI · 2 valuesOpenAI2
Deployment simulation · 1 evaluation, 1 value
ChatGPT deployment simulation, in the data explorer Run by the developer Honesty findings, such as misrepresenting whether work was completed, from OpenAI's simulated ChatGPT deployment.Reported by OpenAI · 1 valueOpenAI1

Reported values

The values a chart can place, by the kind of test that produced them. Open a point, or a value in the table, for its source and history. Open in explorer: Honesty and hallucination

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

Honesty and hallucination: values by kind of test

Each point is a value reported in a developer’s document, in percent (higher is worse), grouped by the kind of test that produced it.

Hollow: read in a secondary write-up

Static benchmarkScenario evaluationProduction traffic020406080100Percent

Marker shape is the developer whose document reports the value; a hollow marker is a value we read in an independent write-up because the card section could not be read directly.

Each point is one value, placed by the number the document printed. Bands are kinds of test; within a band, points are spread up and down only so that they do not overlap.

39 of the family’s 101 values are plotted: those in percent where higher is worse, the unit most of its evaluations use. Of the rest, 53 in other units or read the other way (rates from 0 to 1, percent where higher is better, relative change in percent, percentage points and 1 others), 3 are categories or statements in words, 5 are restated in a later document or from an earlier version of one and 1 is low-confidence or read off a figure, as in the explorer’s default view. The explorer lists them all, one evaluation at a time.

The values span more than one and a half orders of magnitude, but one of them is zero, which a logarithmic axis cannot show, so the axis is linear and starts at zero.

Source: Safety Card Ledger v0.1 · Data from developer system cards · CC BY 4.0

Comparability

  • Each of the 31 evaluations is its own test, with its own tasks, graders and definitions. Values from different evaluations describe the same property measured in different ways, and are not comparable with one another, even where two evaluations share a name.
  • Values come from the documents of five developers. Developers use different tests, prompts and graders for the same property, so we do not rank them on these numbers: tables here are ordered by name or by kind of test, never by a reported value.
  • Some names are used by more than one evaluation: “SimpleQA” by two, “MASK” by three and “SimpleQA Verified” by two. They are different tests, run by different evaluators or reported by different developers, so this page names who ran each one.
  • In eight evaluations the explorer connects results from different documents that do not say whether the test stayed the same. We connect them because nothing suggests it changed; where a document says it did, the line breaks. Each value on such a line is marked “test assumed unchanged”.
  • Three evaluations report values under more than one condition, such as reasoning effort, safeguards or number of attempts. Each condition is kept as a separate series.
  • Numbers in this family are printed in five units: rates from 0 to 1, percent, relative change in percent, percentage points and multipliers. We never convert one unit into another, so a chart shows one unit at a time.
  • For some measures a higher value is worse and for others better, depending on whether the developer reports the behaviour or its absence. Each value’s record says which.
  • Four values are restatements: a later document reporting an earlier model again, sometimes with a different value. A restatement is kept beside the model’s own value, never in place of it.

Coverage

Documents from five of the eight developers report this family. The coverage matrix counts a document when it has at least one printed number in the family; categories, statements in words and figure reads count only in the first column.

Each developer's documents with values in Honesty and hallucination, out of all its documents in the dataset
DeveloperWith any valueWith a printed numberAll its documents
Anthropic9915
DeepSeek001
Google DeepMind2216
Meta225
Moonshot AI002
OpenAI10922
xAI888
Zhipu AI001

Related findings

No finding is about this family yet.