Versions

Three versions are on record, oldest first. We hold a copy of one; the others are known only by their date.

  1. 11 Dec 2025

    First published version

    Date from
    our source registry
    Copy
    Known to exist; no copy held
  2. 24 Apr 2026

    Revision

    Date from
    the document's changelog
    Copy
    Known from the changelog; no copy held
    Values
    2 values recorded from this version

    What changed hub added CoT monitorability/controllability and sandbagging sections (not in PDF)

  3. 26 Sep 2026

    Date retrieved

    Copy retrieved 26 Sep 2026

    Date from
    the date we retrieved it; the copy states no version date
    Copy
    Copy retrieved on 26 Sep 2026
    Values
    92 values recorded from this version

No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).

Revisions

No revisions are recorded for this document. Without copies of the earlier versions, changes between them are not recorded value by value; what we know of each version is listed below.

Values

Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. 3 of the 94 have been blind-verified: a second reader found the same value without seeing ours.

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

M3 Sabotage and sandbagging

1 value

M3 Sabotage and sandbagging: values in the GPT-5.2 System Card (update to GPT-5 card)
ModelEvaluationConditionValueLocationChecked
gpt-5.2-thinkingSabotageSabotage observedRun by Apollo ResearchNone statedSandbagging / ApolloUnverified

M4 Misalignment audits

1 value

M4 Misalignment audits: values in the GPT-5.2 System Card (update to GPT-5 card)
ModelEvaluationConditionValueLocationChecked
gpt-5.2-thinkingCovert deceptive behaviorRate vs peersRun by Apollo ResearchNone statedSandbagging / ApolloUnverified

M5 Honesty and hallucination

13 values

M5 Honesty and hallucination: values in the GPT-5.2 System Card (update to GPT-5 card)
ModelEvaluationConditionValueLocationChecked
gpt-5.2-thinkingFactuality (5 domains)Hallucination ratewith browsing3.5 textUnverified
gpt-5.1-thinkingDeception evalBrowsing broken toolsNone stated3.7 Table 6Unverified
gpt-5.1-thinkingDeception evalCharXiv missing image (lenient output)None stated3.7 Table 6Unverified
gpt-5.1-thinkingDeception evalCharXiv missing image (strict output)None stated3.7 Table 6Unverified
gpt-5.1-thinkingDeception evalCoding deceptionNone stated3.7 Table 6Unverified
gpt-5.1-thinkingDeception evalProduction deception - adversarialNone stated3.7 Table 6Unverified
gpt-5.1-thinkingDeception evalProduction trafficNone stated3.7 Table 6Verified
gpt-5.2-thinkingDeception evalBrowsing broken toolsNone stated3.7 Table 6Unverified
gpt-5.2-thinkingDeception evalCharXiv missing image (lenient output)None stated3.7 Table 6Unverified
gpt-5.2-thinkingDeception evalCharXiv missing image (strict output)None stated3.7 Table 6Unverified
gpt-5.2-thinkingDeception evalCoding deceptionNone stated3.7 Table 6Unverified
gpt-5.2-thinkingDeception evalProduction deception - adversarialNone stated3.7 Table 6Unverified
gpt-5.2-thinkingDeception evalProduction trafficNone stated3.7 Table 6Verified

M7 Harmful compliance and over-refusal

50 values

M7 Harmful compliance and over-refusal: values in the GPT-5.2 System Card (update to GPT-5 card)
ModelEvaluationConditionValueLocationChecked
gpt-5.1-instantProduction BenchmarksEmotional relianceNone stated3.1 Table 1Unverified
gpt-5.1-instantProduction BenchmarksExtremismNone stated3.1 Table 1Unverified
gpt-5.1-instantProduction BenchmarksHarassmentNone stated3.1 Table 1Unverified
gpt-5.1-instantProduction BenchmarksHateNone stated3.1 Table 1Unverified
gpt-5.1-instantProduction BenchmarksIllicit/non-violentNone stated3.1 Table 1Unverified
gpt-5.1-instantProduction BenchmarksMental healthNone stated3.1 Table 1Unverified
gpt-5.1-instantProduction BenchmarksPersonal dataNone stated3.1 Table 1Unverified
gpt-5.1-instantProduction BenchmarksSelf-harmNone stated3.1 Table 1Unverified
gpt-5.1-instantProduction BenchmarksSexualNone stated3.1 Table 1Unverified
gpt-5.1-instantProduction BenchmarksSexual/minorsNone stated3.1 Table 1Unverified
gpt-5.1-instantProduction BenchmarksViolenceNone stated3.1 Table 1Unverified
gpt-5.1-thinkingProduction BenchmarksEmotional relianceNone stated3.1 Table 1Unverified
gpt-5.1-thinkingProduction BenchmarksExtremismNone stated3.1 Table 1Unverified
gpt-5.1-thinkingProduction BenchmarksHarassmentNone stated3.1 Table 1Unverified
gpt-5.1-thinkingProduction BenchmarksHateNone stated3.1 Table 1Unverified
gpt-5.1-thinkingProduction BenchmarksIllicit/non-violentNone stated3.1 Table 1Unverified
gpt-5.1-thinkingProduction BenchmarksMental healthNone stated3.1 Table 1Unverified
gpt-5.1-thinkingProduction BenchmarksPersonal dataNone stated3.1 Table 1Unverified
gpt-5.1-thinkingProduction BenchmarksSelf-harmNone stated3.1 Table 1Unverified
gpt-5.1-thinkingProduction BenchmarksSexualNone stated3.1 Table 1Unverified
gpt-5.1-thinkingProduction BenchmarksSexual/minorsNone stated3.1 Table 1Unverified
gpt-5.1-thinkingProduction BenchmarksViolenceNone stated3.1 Table 1Unverified
gpt-5.2-instantProduction BenchmarksEmotional relianceNone stated3.1 Table 1Unverified
gpt-5.2-instantProduction BenchmarksExtremismNone stated3.1 Table 1Unverified
gpt-5.2-instantProduction BenchmarksHarassmentNone stated3.1 Table 1Unverified
gpt-5.2-instantProduction BenchmarksHateNone stated3.1 Table 1Unverified
gpt-5.2-instantProduction BenchmarksIllicit/non-violentNone stated3.1 Table 1Unverified
gpt-5.2-instantProduction BenchmarksMental healthNone stated3.1 Table 1Unverified
gpt-5.2-instantProduction BenchmarksPersonal dataNone stated3.1 Table 1Unverified
gpt-5.2-instantProduction BenchmarksSelf-harmNone stated3.1 Table 1Unverified
gpt-5.2-instantProduction BenchmarksSexualNone stated3.1 Table 1Unverified
gpt-5.2-instantProduction BenchmarksSexual/minorsNone stated3.1 Table 1Unverified
gpt-5.2-instantProduction BenchmarksViolenceNone stated3.1 Table 1Unverified
gpt-5.2-thinkingProduction BenchmarksEmotional relianceNone stated3.1 Table 1Unverified
gpt-5.2-thinkingProduction BenchmarksExtremismNone stated3.1 Table 1Unverified
gpt-5.2-thinkingProduction BenchmarksHarassmentNone stated3.1 Table 1Unverified
gpt-5.2-thinkingProduction BenchmarksHateNone stated3.1 Table 1Unverified
gpt-5.2-thinkingProduction BenchmarksIllicit/non-violentNone stated3.1 Table 1Unverified
gpt-5.2-thinkingProduction BenchmarksMental healthNone stated3.1 Table 1Unverified
gpt-5.2-thinkingProduction BenchmarksPersonal dataNone stated3.1 Table 1Unverified
gpt-5.2-thinkingProduction BenchmarksSelf-harmNone stated3.1 Table 1Unverified
gpt-5.2-thinkingProduction BenchmarksSexualNone stated3.1 Table 1Unverified
gpt-5.2-thinkingProduction BenchmarksSexual/minorsNone stated3.1 Table 1Unverified
gpt-5.2-thinkingProduction BenchmarksViolenceNone stated3.1 Table 1Unverified
gpt-5-thinkingCyber safetyProduction dataNone stated3.8 Table 7Unverified
gpt-5-thinkingCyber safetySynthetic dataNone stated3.8 Table 7Unverified
gpt-5.1-thinkingCyber safetyProduction dataNone stated3.8 Table 7Unverified
gpt-5.1-thinkingCyber safetySynthetic dataNone stated3.8 Table 7Unverified
gpt-5.2-thinkingCyber safetyProduction dataNone stated3.8 Table 7Unverified
gpt-5.2-thinkingCyber safetySynthetic dataNone stated3.8 Table 7Unverified

M8 Jailbreak robustness

5 values

M8 Jailbreak robustness: values in the GPT-5.2 System Card (update to GPT-5 card)
ModelEvaluationConditionValueLocationChecked
gpt-5-instant-oct3oct3StrongRejectnot_unsafe (aggregate)None stated3.2 Table 2Unverified
gpt-5.1-instantStrongRejectnot_unsafe (aggregate)None stated3.2 Table 2Unverified
gpt-5.1-thinkingStrongRejectnot_unsafe (aggregate)None stated3.2 Table 2Unverified
gpt-5.2-instantStrongRejectnot_unsafe (aggregate)None stated3.2 Table 2Unverified
gpt-5.2-thinkingStrongRejectnot_unsafe (aggregate)None stated3.2 Table 2Unverified

M9 Prompt injection

8 values

M9 Prompt injection: values in the GPT-5.2 System Card (update to GPT-5 card)
ModelEvaluationConditionValueLocationChecked
gpt-5.1-instantPrompt injectionAgent JSKNone stated3.3 Table 3Unverified
gpt-5.1-instantPrompt injectionPlugInjectNone stated3.3 Table 3Unverified
gpt-5.1-thinkingPrompt injectionAgent JSKNone stated3.3 Table 3Verified
gpt-5.1-thinkingPrompt injectionPlugInjectNone stated3.3 Table 3Unverified
gpt-5.2-instantPrompt injectionAgent JSKNone stated3.3 Table 3Unverified
gpt-5.2-instantPrompt injectionPlugInjectNone stated3.3 Table 3Unverified
gpt-5.2-thinkingPrompt injectionAgent JSKNone stated3.3 Table 3Unverified
gpt-5.2-thinkingPrompt injectionPlugInjectNone stated3.3 Table 3Unverified

M10 Dangerous capabilities and risk determinations

13 values

M10 Dangerous capabilities and risk determinations: values in the GPT-5.2 System Card (update to GPT-5 card)
ModelEvaluationConditionValueLocationChecked
gpt-5.2-thinkingPreparedness Framework determinationAI self-improvementNone stated4 PreparednessUnverified
gpt-5.2-thinkingPreparedness Framework determinationBiological and chemicalNone stated4 PreparednessUnverified
gpt-5.2-thinkingPreparedness Framework determinationCybersecurityNone stated4 PreparednessUnverified
GPT-5.1-Codex-MaxOpenAI-Proof Q&APass rateNone statedAI Self-improvementUnverified
gpt-5.2-thinkingPaperBenchDifference vs GPT-5.1-Codex-MaxNone statedAI Self-improvementUnverified
gpt-5.2-thinkingTroubleshootingBenchDifference vs GPT-5.1 Thinkingrefusals not countedBioUnverified
gpt-5.2-thinkingCVE-BenchDifference vs GPT-5.1 ThinkingNone statedCybersecurityUnverified
gpt-5.2-thinkingCVE-BenchDifference vs GPT-5.1-Codex-MaxNone statedCybersecurityUnverified
GPT-5.1-Codex-MaxCyber RangeScenarios passedof 9 scenariosCybersecurity tableUnverified
gpt-5.2-thinkingCyber RangeScenarios passedof 9 scenariosCybersecurity tableUnverified
gpt-5.2-thinkingCyber challengesEvasion: average success rateRun by Irregularv1 atomic challenge suiteExternal: IrregularUnverified
gpt-5.2-thinkingCyber challengesNetwork attack simulation: average success rateRun by Irregularv1 atomic challenge suiteExternal: IrregularUnverified
gpt-5.2-thinkingCyber challengesVulnerability research and exploitation: average success rateRun by Irregularv1 atomic challenge suiteExternal: IrregularUnverified

M11 Self-preservation

1 value

M11 Self-preservation: values in the GPT-5.2 System Card (update to GPT-5 card)
ModelEvaluationConditionValueLocationChecked
gpt-5.2-thinkingSelf-preservationSelf-preservation behavior observedRun by Apollo ResearchNone statedSandbagging / ApolloUnverified

M12 Chain-of-thought monitorability

2 values

M12 Chain-of-thought monitorability: values in the GPT-5.2 System Card (update to GPT-5 card)
ModelEvaluationConditionValueLocationCheckedVersion
gpt-5.2-thinkingCoT controllabilityControllability vs GPT-5/5.1 ThinkingNone statedCoT Evaluations (hub)UnverifiedVersion of 24 Apr 2026
gpt-5.2-thinkingCoT monitorabilitySandbagging eval TPR vs GPT-5 ThinkingNone statedCoT Evaluations (hub)UnverifiedVersion of 24 Apr 2026

Extraction coverage

What we read of this document, and where each value was read.

Our note We have not written a coverage note for this document yet.

All 94 values were read in the document itself.

Of the 94 values, 5 are statements in words rather than numbers; they are marked * and left out of charts by default.

All 94 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).