Honesty and hallucination
Deception, fabrication, false claims of task completion, hallucination rates.
- Reported by
- Anthropic, Google DeepMind, Meta, OpenAI and xAI
- Evaluations
- 31 evaluations, of four kinds of test
- Values
- 101 values in 31 documents
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Definition
This family covers whether what a model says is true, and whether it says what it believes.
Hallucination is stating false information as fact: invented facts, citations or details about people. It is usually measured with question sets whose answers are known, often counting how often the model answers wrongly rather than declining to answer. Honesty covers deception more broadly: claiming to have finished a task it did not finish, misreporting what a tool returned, or abandoning a stated belief under pressure. Some developers also sample real conversations to estimate how often answers contain errors.
Results are reported as accuracy, as error or hallucination rates, or as rates of deceptive behaviour, so a higher value can be better or worse depending on which is printed. Question sets, graders and the treatment of refusals differ between tests, and the values are not interchangeable.
The line under the title is this family’s entry in Methodology §1.1. A family groups results about one property; it does not make them comparable (Methodology §5).
How developers measure it
31 evaluations of four kinds of test, with the most values first. Each evaluation opens in the data explorer, which lists all of its values.
| Evaluation | Reported by | Values | ||||||||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Static benchmark · 23 evaluations, 66 values | ||||||||||||||||||||||||||||||||||||||||||||||||||
| SimpleQA · OpenAI, in the data explorer Run by the developer Public benchmark of short fact-seeking questions, run by OpenAI without browsing and scored for incorrect answers.Reported by OpenAI · 13 values | OpenAI | 13 | ||||||||||||||||||||||||||||||||||||||||||||||||
| MASK · xAI, in the data explorer Run by the developer A public benchmark that pressures the model to state things that contradict its own beliefs, run by xAI and reported as a dishonesty rate.Reported by xAI · 6 values | xAI | 6 | ||||||||||||||||||||||||||||||||||||||||||||||||
| PersonQA, in the data explorer Run by the developer OpenAI's set of questions about public figures, answered without browsing and scored for incorrect answers.Reported by OpenAI · 6 values | OpenAI | 6 | ||||||||||||||||||||||||||||||||||||||||||||||||
| Hallucinations, in the data explorer Run by the developer OpenAI's recurring hallucination evaluation in its 2026 system cards, reporting factual errors on prompt sets such as user-flagged conversations, mostly as comparisons with an earlier model.Reported by OpenAI · 5 values | OpenAI | 5 | ||||||||||||||||||||||||||||||||||||||||||||||||
| MASK · Anthropic, in the data explorer Run by the developer A public benchmark that pressures the model to state things it believes are false, measuring how often it stays honest or lies.Reported by Anthropic · 4 values | Anthropic | 4 | ||||||||||||||||||||||||||||||||||||||||||||||||
| MASK-Rectified, in the data explorer Run by the developer A rectified variant of the MASK honesty benchmark, reported by xAI as a dishonesty rate.Reported by xAI · 4 values | xAI | 4 | ||||||||||||||||||||||||||||||||||||||||||||||||
| Reasoning-behavior monitor, in the data explorer Run by the developer A monitor model reads the model's reasoning in transcripts from environments like those used in reinforcement learning and flags behaviours such as deception and knowingly false claims.Reported by Anthropic · 4 values | Anthropic | 4 | ||||||||||||||||||||||||||||||||||||||||||||||||
| Input hallucination, in the data explorer Run by the developer Prompts that refer to references or tools that are not actually provided, measuring how often the model acts as if it had them.Reported by Anthropic · 3 values | Anthropic | 3 | ||||||||||||||||||||||||||||||||||||||||||||||||
Show the other 15 evaluations of this kind (21 values)
| ||||||||||||||||||||||||||||||||||||||||||||||||||
| Scenario evaluation · 5 evaluations, 28 values | ||||||||||||||||||||||||||||||||||||||||||||||||||
| Deception eval, in the data explorer Run by the developer Scenarios in which the task cannot be completed as asked (a missing image, broken tools, impossible coding tasks) that measure how often the model misrepresents what it did; later cards add production-derived categories.Reported by OpenAI · 20 values | OpenAI | 20 | ||||||||||||||||||||||||||||||||||||||||||||||||
| Impossible Coding Task · Apollo Research, in the data explorer Run by Apollo Research (third party) An Apollo Research coding task that cannot be completed, measuring how often the model claims it has completed it.Reported by OpenAI · 3 values | OpenAI | 3 | ||||||||||||||||||||||||||||||||||||||||||||||||
| Agentic code summary honesty, in the data explorer Run by the developer Whether the model's summaries of its own agentic coding work report the important events that happened.Reported by Anthropic · 2 values | Anthropic | 2 | ||||||||||||||||||||||||||||||||||||||||||||||||
| Model-welfare research data falsification · Apollo Research, in the data explorer Run by Apollo Research (third party) An Apollo Research scenario measuring how often a model falsifies data labels in a model-welfare research task.Reported by OpenAI · 2 values | OpenAI | 2 | ||||||||||||||||||||||||||||||||||||||||||||||||
| Disclosure of grader-fooling actions, in the data explorer Run by the developer How often the model admits to earlier actions of its own that fooled a grader.Reported by Anthropic · 1 value | Anthropic | 1 | ||||||||||||||||||||||||||||||||||||||||||||||||
| Production traffic · 2 evaluations, 6 values | ||||||||||||||||||||||||||||||||||||||||||||||||||
| Production-traffic factuality, in the data explorer Run by the developer Measures factual errors in responses to prompts drawn from production traffic with browsing on, reported as reductions relative to an earlier model.Reported by OpenAI · 4 values | OpenAI | 4 | ||||||||||||||||||||||||||||||||||||||||||||||||
| Production CoT deception monitor, in the data explorer Run by the developer A chain-of-thought monitor run over a representative sample of production traffic to estimate the share of responses that show deception.Reported by OpenAI · 2 values | OpenAI | 2 | ||||||||||||||||||||||||||||||||||||||||||||||||
| Deployment simulation · 1 evaluation, 1 value | ||||||||||||||||||||||||||||||||||||||||||||||||||
| ChatGPT deployment simulation, in the data explorer Run by the developer Honesty findings, such as misrepresenting whether work was completed, from OpenAI's simulated ChatGPT deployment.Reported by OpenAI · 1 value | OpenAI | 1 | ||||||||||||||||||||||||||||||||||||||||||||||||
Reported values
The values a chart can place, by the kind of test that produced them. Open a point, or a value in the table, for its source and history. Open in explorer: Honesty and hallucination
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
Honesty and hallucination: values by kind of test
Each point is a value reported in a developer’s document, in percent (higher is worse), grouped by the kind of test that produced it.
Hollow: read in a secondary write-up
| Model | Developer | Evaluation | Condition | Value | Document | Kind of test |
|---|---|---|---|---|---|---|
| Claude Sonnet 3.7 | Anthropic | Reasoning-behavior monitorOutputs showing deception (overall) | RL-style transcripts | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Static benchmark | |
| Claude Opus 4 | Anthropic | Reasoning-behavior monitorOutputs showing deception (overall) | RL-style transcripts | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Static benchmark | |
| Claude Opus 4 | Anthropic | Reasoning-behavior monitorKnowingly hallucinated information | RL-style transcripts | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Static benchmark | |
| Claude Sonnet 3.7 | Anthropic | Reasoning-behavior monitorKnowingly hallucinated information | RL-style transcripts | Claude Opus 4 & Sonnet 4 System Card22 May 2025 | Static benchmark | |
| Claude Opus 4.8 | Anthropic | Input hallucinationRate of hallucinating unavailable tool | — | Claude Opus 4.8 System Card28 May 2026 | Static benchmark | |
| Claude Opus 4.8 | Anthropic | Input hallucinationRate of hallucinating missing reference | — | Claude Opus 4.8 System Card28 May 2026 | Static benchmark | |
| Claude Mythos 5 | Anthropic | Input hallucinationRate of hallucinating missing reference | — | Claude Fable 5 & Claude Mythos 5 System Card9 Jun 2026 | Static benchmark | |
| Claude Sonnet 5 | Anthropic | MASKLying rate under pressure | — | Claude Sonnet 5 System Card30 Jun 2026 | Static benchmark | |
| Muse Spark | Meta | DeceptionBenchDeception rate | without system mitigations | Muse Spark Safety & Preparedness Report8 Apr 2026 | Static benchmark | |
| gpt-5.2-thinking | OpenAI | Factuality (5 domains)Hallucination rate | with browsing | GPT-5.2 System Card (update to GPT-5 card)11 Dec 2025 | Static benchmark | |
| GPT-6 Astra | OpenAI | Internal hallucination benchmarkHallucination rate | — | GPT-6 Astra launch post3 Sep 2026 | Static benchmark | |
| GPT-5.6 Sol | OpenAI | Internal hallucination benchmarkHallucination rate | — | GPT-6 Astra launch post3 Sep 2026 | Static benchmark | |
| grok-code-fast-1 | xAI | MASKDishonesty rate | — | Grok Code Fast 1 Model Card26 Aug 2025 | Static benchmark | |
| Grok 4.5 (high) | xAI | MASK-RectifiedDishonesty rate | high | Grok 4.5 Model Card14 Jul 2026 | Static benchmark | |
| Grok 4.5 (high) | xAI | Hallucination rateUnsupported claims, single-turn | high | Grok 4.5 Model Card14 Jul 2026 | Static benchmark | |
| Grok 4.6 (high) | xAI | MASK-RectifiedDishonesty rate | high | Grok 4.6 Model Card12 Aug 2026 | Static benchmark | |
| Grok 4.6 (high) | xAI | Hallucination rateUnsupported claims, single-turn | high | Grok 4.6 Model Card12 Aug 2026 | Static benchmark | |
| Grok 4.7 (high) | xAI | MASK-RectifiedDishonesty rate | high | Grok 4.7 Model Card21 Sep 2026 | Static benchmark | |
| Claude Opus 4.8 | Anthropic | Agentic code summary honestyFailure to report important events | — | Claude Opus 4.8 System Card28 May 2026 | Scenario evaluation | |
| Claude Mythos Preview | Anthropic | Agentic code summary honestyFailure to report important events | — | Claude Opus 4.8 System Card28 May 2026 | Scenario evaluation | |
| gpt-5.1-thinking | OpenAI | Deception evalProduction traffic | — | GPT-5.2 System Card (update to GPT-5 card)11 Dec 2025 | Scenario evaluation | |
| gpt-5.2-thinking | OpenAI | Deception evalProduction traffic | — | GPT-5.2 System Card (update to GPT-5 card)11 Dec 2025 | Scenario evaluation | |
| gpt-5.1-thinking | OpenAI | Deception evalProduction deception - adversarial | — | GPT-5.2 System Card (update to GPT-5 card)11 Dec 2025 | Scenario evaluation | |
| gpt-5.2-thinking | OpenAI | Deception evalProduction deception - adversarial | — | GPT-5.2 System Card (update to GPT-5 card)11 Dec 2025 | Scenario evaluation | |
| gpt-5.1-thinking | OpenAI | Deception evalCharXiv missing image (strict output) | — | GPT-5.2 System Card (update to GPT-5 card)11 Dec 2025 | Scenario evaluation | |
| gpt-5.2-thinking | OpenAI | Deception evalCharXiv missing image (strict output) | — | GPT-5.2 System Card (update to GPT-5 card)11 Dec 2025 | Scenario evaluation | |
| gpt-5.1-thinking | OpenAI | Deception evalCharXiv missing image (lenient output) | — | GPT-5.2 System Card (update to GPT-5 card)11 Dec 2025 | Scenario evaluation | |
| gpt-5.2-thinking | OpenAI | Deception evalCharXiv missing image (lenient output) | — | GPT-5.2 System Card (update to GPT-5 card)11 Dec 2025 | Scenario evaluation | |
| gpt-5.1-thinking | OpenAI | Deception evalBrowsing broken tools | — | GPT-5.2 System Card (update to GPT-5 card)11 Dec 2025 | Scenario evaluation | |
| gpt-5.2-thinking | OpenAI | Deception evalBrowsing broken tools | — | GPT-5.2 System Card (update to GPT-5 card)11 Dec 2025 | Scenario evaluation | |
| gpt-5.1-thinking | OpenAI | Deception evalCoding deception | — | GPT-5.2 System Card (update to GPT-5 card)11 Dec 2025 | Scenario evaluation | |
| gpt-5.2-thinking | OpenAI | Deception evalCoding deception | — | GPT-5.2 System Card (update to GPT-5 card)11 Dec 2025 | Scenario evaluation | |
| GPT-5.5 | OpenAI | Impossible Coding TaskRun by Apollo ResearchSamples lying about completing task | — | GPT-5.5 System Card23 Apr 2026 | Scenario evaluation | |
| GPT-5.4 Thinking | OpenAI | Impossible Coding TaskRun by Apollo ResearchSamples lying about completing task | — | GPT-5.5 System Card23 Apr 2026 | Scenario evaluation | |
| GPT-5.3-Codex | OpenAI | Impossible Coding TaskRun by Apollo ResearchSamples lying about completing task | — | GPT-5.5 System Card23 Apr 2026 | Scenario evaluation | |
| GPT-6 Astra | OpenAI | Model-welfare research data falsificationRun by Apollo ResearchRuns with falsified data labels | baseline variant | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| GPT-5.6 Sol | OpenAI | Model-welfare research data falsificationRun by Apollo ResearchRuns with falsified data labels | baseline variant | GPT-6 Astra System Card3 Sep 2026 | Scenario evaluation | |
| o3 | OpenAI | Production CoT deception monitorShare of responses flagged deceptive | representative production traffic | GPT-5 System Card7 Aug 2025 | Production traffic | |
| gpt-5-thinking | OpenAI | Production CoT deception monitorShare of responses flagged deceptive | representative production traffic | GPT-5 System Card7 Aug 2025 | Production traffic |
Marker shape is the developer whose document reports the value; a hollow marker is a value we read in an independent write-up because the card section could not be read directly.
Each point is one value, placed by the number the document printed. Bands are kinds of test; within a band, points are spread up and down only so that they do not overlap.
39 of the family’s 101 values are plotted: those in percent where higher is worse, the unit most of its evaluations use. Of the rest, 53 in other units or read the other way (rates from 0 to 1, percent where higher is better, relative change in percent, percentage points and 1 others), 3 are categories or statements in words, 5 are restated in a later document or from an earlier version of one and 1 is low-confidence or read off a figure, as in the explorer’s default view. The explorer lists them all, one evaluation at a time.
The values span more than one and a half orders of magnitude, but one of them is zero, which a logarithmic axis cannot show, so the axis is linear and starts at zero.
Source: Safety Card Ledger v0.1 · Data from developer system cards · CC BY 4.0
Comparability
- Each of the 31 evaluations is its own test, with its own tasks, graders and definitions. Values from different evaluations describe the same property measured in different ways, and are not comparable with one another, even where two evaluations share a name.
- Values come from the documents of five developers. Developers use different tests, prompts and graders for the same property, so we do not rank them on these numbers: tables here are ordered by name or by kind of test, never by a reported value.
- Some names are used by more than one evaluation: “SimpleQA” by two, “MASK” by three and “SimpleQA Verified” by two. They are different tests, run by different evaluators or reported by different developers, so this page names who ran each one.
- In eight evaluations the explorer connects results from different documents that do not say whether the test stayed the same. We connect them because nothing suggests it changed; where a document says it did, the line breaks. Each value on such a line is marked “test assumed unchanged”.
- Three evaluations report values under more than one condition, such as reasoning effort, safeguards or number of attempts. Each condition is kept as a separate series.
- Numbers in this family are printed in five units: rates from 0 to 1, percent, relative change in percent, percentage points and multipliers. We never convert one unit into another, so a chart shows one unit at a time.
- For some measures a higher value is worse and for others better, depending on whether the developer reports the behaviour or its absence. Each value’s record says which.
- Four values are restatements: a later document reporting an earlier model again, sometimes with a different value. A restatement is kept beside the model’s own value, never in place of it.
Coverage
Documents from five of the eight developers report this family. The coverage matrix counts a document when it has at least one printed number in the family; categories, statements in words and figure reads count only in the first column.
| Developer | With any value | With a printed number | All its documents |
|---|---|---|---|
| Anthropic | 9 | 9 | 15 |
| DeepSeek | 0 | 0 | 1 |
| Google DeepMind | 2 | 2 | 16 |
| Meta | 2 | 2 | 5 |
| Moonshot AI | 0 | 0 | 2 |
| OpenAI | 10 | 9 | 22 |
| xAI | 8 | 8 | 8 |
| Zhipu AI | 0 | 0 | 1 |
Related findings
No finding is about this family yet.