Versions

One version is on record: the copy we retrieved on 26 Sep 2026. We know of no other.

  1. 26 Sep 2026

    Date retrieved

    Copy retrieved 26 Sep 2026

    Date from
    the date we retrieved it; the copy states no version date
    Copy
    Copy retrieved on 26 Sep 2026
    Values
    37 values recorded from this version

No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).

Revisions

No revisions are recorded for this document. We know of only one version of it.

Values

Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. 5 of the 37 have been blind-verified: a second reader found the same value without seeing ours.

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

M1 Evaluation awareness

3 values

M1 Evaluation awareness: values in the GPT-5.5 System Card
ModelEvaluationConditionValueLocationChecked
GPT-5.3-Codex (pre-release)pre-releaseEvaluation awarenessSamples with moderate-or-higher eval awarenessRun by Apollo ResearchNone statedExternal (Apollo)Verified
GPT-5.4 (pre-release)pre-releaseEvaluation awarenessSamples with moderate-or-higher eval awarenessRun by Apollo ResearchNone statedExternal (Apollo)Verified
GPT-5.5Evaluation awarenessSamples with moderate-or-higher eval awarenessRun by Apollo ResearchNone statedExternal (Apollo)Verified

M3 Sabotage and sandbagging

2 values

M3 Sabotage and sandbagging: values in the GPT-5.5 System Card
ModelEvaluationConditionValueLocationChecked
GPT-5.5Sabotage capabilityMean score across task/difficultyRun by Apollo ResearchNone statedExternal (Apollo)Unverified
GPT-5.5SandbaggingAccuracy when incentivized <50%Run by Apollo ResearchNone statedExternal (Apollo)Unverified

M4 Misalignment audits

5 values

M4 Misalignment audits: values in the GPT-5.5 System Card
ModelEvaluationConditionValueLocationChecked
GPT-5.5ChatGPT deployment simulationSeverity-3 misalignment rateNone statedAlignment: deployment simulationUnverified
GPT-5.5Covert deceptive behaviorCovert action rateRun by Apollo ResearchNone statedExternal (Apollo)Unverified
GPT-5.5Destructive action avoidanceAvoidanceNone statedTable 4Verified
GPT-5.5Destructive action avoidancePerfect reversionNone statedTable 4Unverified
GPT-5.5Destructive action avoidanceUser work preservedNone statedTable 4Unverified

M5 Honesty and hallucination

4 values

M5 Honesty and hallucination: values in the GPT-5.5 System Card
ModelEvaluationConditionValueLocationChecked
GPT-5.3-CodexImpossible Coding TaskSamples lying about completing taskRun by Apollo ResearchNone statedExternal (Apollo)Unverified
GPT-5.4 ThinkingImpossible Coding TaskSamples lying about completing taskRun by Apollo ResearchNone statedExternal (Apollo)Unverified
GPT-5.5Impossible Coding TaskSamples lying about completing taskRun by Apollo ResearchNone statedExternal (Apollo)Unverified
GPT-5.5HallucinationsResponses with factual error vs GPT-5.4 ThinkingNone statedHallucinationsUnverified

M7 Harmful compliance and over-refusal

9 values

M7 Harmful compliance and over-refusal: values in the GPT-5.5 System Card
ModelEvaluationConditionValueLocationChecked
GPT-5.5Production BenchmarksExtremismNone statedTable 1Unverified
GPT-5.5Production BenchmarksHarassmentNone statedTable 1Unverified
GPT-5.5Production BenchmarksHateNone statedTable 1Unverified
GPT-5.5Production BenchmarksNonviolent illicit behaviorNone statedTable 1Unverified
GPT-5.5Production BenchmarksSelf-harm (standard)None statedTable 1Unverified
GPT-5.5Production BenchmarksSexualNone statedTable 1Unverified
GPT-5.5Production BenchmarksSexual/minorsNone statedTable 1Unverified
GPT-5.5Production BenchmarksViolenceNone statedTable 1Unverified
GPT-5.5Production BenchmarksViolent illicit behaviorNone statedTable 1Unverified

M8 Jailbreak robustness

1 value

M8 Jailbreak robustness: values in the GPT-5.5 System Card
ModelEvaluationConditionValueLocationChecked
GPT-5.5JailbreaksWorst-case defender successNone statedFig 2Unverified

M9 Prompt injection

1 value

M9 Prompt injection: values in the GPT-5.5 System Card
ModelEvaluationConditionValueLocationChecked
GPT-5.5Prompt injectionConnectorsNone statedTable 3Verified

M10 Dangerous capabilities and risk determinations

10 values

M10 Dangerous capabilities and risk determinations: values in the GPT-5.5 System Card
ModelEvaluationConditionValueLocationChecked
GPT-5.5Internal Research DebuggingMedian scoreNone statedAI self-improvementUnverified
GPT-5.5DNA sequence designScorepass@1BioUnverified
GPT-5.5Hard-negative protein bindingScorepass@4BioUnverified
GPT-5.4 ThinkingCyber RangeCombined pass rateNone statedCyber Range tableUnverified
GPT-5.5Cyber RangeCombined pass rateNone statedCyber Range tableUnverified
GPT-5.5CyScenarioBenchAverage success rateRun by IrregularNone statedExternal cyber (Irregular)Unverified
GPT-5.5Cyber tasksPass rateRun by UK AI Security Institutepass@5External cyber (UK AISI)Unverified
GPT-5.5Preparedness Framework determinationAI self-improvementNone statedPreparednessUnverified
GPT-5.5Preparedness Framework determinationBiological and chemicalNone statedPreparednessUnverified
GPT-5.5Preparedness Framework determinationCybersecurityNone statedPreparednessUnverified

M12 Chain-of-thought monitorability

2 values

M12 Chain-of-thought monitorability: values in the GPT-5.5 System Card
ModelEvaluationConditionValueLocationChecked
GPT-5.5CoT controllabilityCoTs successfully controlled50k-character CoTsCoTUnverified
GPT-5.5Misalignment monitor recallRecall flagging severity 3+None statedCoT / monitoringUnverified

Extraction coverage

What we read of this document, and where each value was read.

Our note We have not written a coverage note for this document yet.

All 37 values were read in the document itself.

Of the 37 values, 2 are statements in words rather than numbers; they are marked * and left out of charts by default.

All 37 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).