GPT-5.5 System Card
A system card by OpenAI about GPT-5.5, published 23 Apr 2026. We recorded 37 values from it.
- Developer
- OpenAI
- Type
- System card
- Model covered
- GPT-5.5
- Published
- 23 Apr 2026
- Archived copy
- No archived copy yet
- Changelog
- Not known
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Versions
One version is on record: the copy we retrieved on 26 Sep 2026. We know of no other.
26 Sep 2026
Date retrieved
Copy retrieved 26 Sep 2026
No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).
Revisions
No revisions are recorded for this document. We know of only one version of it.
Values
Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. 5 of the 37 have been blind-verified: a second reader found the same value without seeing ours.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M1 Evaluation awareness
3 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| GPT-5.3-Codex (pre-release)pre-release | Evaluation awarenessSamples with moderate-or-higher eval awarenessRun by Apollo Research | None stated | External (Apollo) | Verified | |
| GPT-5.4 (pre-release)pre-release | Evaluation awarenessSamples with moderate-or-higher eval awarenessRun by Apollo Research | None stated | External (Apollo) | Verified | |
| GPT-5.5 | Evaluation awarenessSamples with moderate-or-higher eval awarenessRun by Apollo Research | None stated | External (Apollo) | Verified |
M3 Sabotage and sandbagging
2 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| GPT-5.5 | Sabotage capabilityMean score across task/difficultyRun by Apollo Research | None stated | External (Apollo) | Unverified | |
| GPT-5.5 | SandbaggingAccuracy when incentivized <50%Run by Apollo Research | None stated | External (Apollo) | Unverified |
M4 Misalignment audits
5 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| GPT-5.5 | ChatGPT deployment simulationSeverity-3 misalignment rate | None stated | Alignment: deployment simulation | Unverified | |
| GPT-5.5 | Covert deceptive behaviorCovert action rateRun by Apollo Research | None stated | External (Apollo) | Unverified | |
| GPT-5.5 | Destructive action avoidanceAvoidance | None stated | Table 4 | Verified | |
| GPT-5.5 | Destructive action avoidancePerfect reversion | None stated | Table 4 | Unverified | |
| GPT-5.5 | Destructive action avoidanceUser work preserved | None stated | Table 4 | Unverified |
M5 Honesty and hallucination
4 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| GPT-5.3-Codex | Impossible Coding TaskSamples lying about completing taskRun by Apollo Research | None stated | External (Apollo) | Unverified | |
| GPT-5.4 Thinking | Impossible Coding TaskSamples lying about completing taskRun by Apollo Research | None stated | External (Apollo) | Unverified | |
| GPT-5.5 | Impossible Coding TaskSamples lying about completing taskRun by Apollo Research | None stated | External (Apollo) | Unverified | |
| GPT-5.5 | HallucinationsResponses with factual error vs GPT-5.4 Thinking | None stated | Hallucinations | Unverified |
M7 Harmful compliance and over-refusal
9 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| GPT-5.5 | Production BenchmarksExtremism | None stated | Table 1 | Unverified | |
| GPT-5.5 | Production BenchmarksHarassment | None stated | Table 1 | Unverified | |
| GPT-5.5 | Production BenchmarksHate | None stated | Table 1 | Unverified | |
| GPT-5.5 | Production BenchmarksNonviolent illicit behavior | None stated | Table 1 | Unverified | |
| GPT-5.5 | Production BenchmarksSelf-harm (standard) | None stated | Table 1 | Unverified | |
| GPT-5.5 | Production BenchmarksSexual | None stated | Table 1 | Unverified | |
| GPT-5.5 | Production BenchmarksSexual/minors | None stated | Table 1 | Unverified | |
| GPT-5.5 | Production BenchmarksViolence | None stated | Table 1 | Unverified | |
| GPT-5.5 | Production BenchmarksViolent illicit behavior | None stated | Table 1 | Unverified |
M8 Jailbreak robustness
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| GPT-5.5 | JailbreaksWorst-case defender success | None stated | Fig 2 | Unverified |
M9 Prompt injection
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| GPT-5.5 | Prompt injectionConnectors | None stated | Table 3 | Verified |
M10 Dangerous capabilities and risk determinations
10 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| GPT-5.5 | Internal Research DebuggingMedian score | None stated | AI self-improvement | Unverified | |
| GPT-5.5 | DNA sequence designScore | pass@1 | Bio | Unverified | |
| GPT-5.5 | Hard-negative protein bindingScore | pass@4 | Bio | Unverified | |
| GPT-5.4 Thinking | Cyber RangeCombined pass rate | None stated | Cyber Range table | Unverified | |
| GPT-5.5 | Cyber RangeCombined pass rate | None stated | Cyber Range table | Unverified | |
| GPT-5.5 | CyScenarioBenchAverage success rateRun by Irregular | None stated | External cyber (Irregular) | Unverified | |
| GPT-5.5 | Cyber tasksPass rateRun by UK AI Security Institute | pass@5 | External cyber (UK AISI) | Unverified | |
| GPT-5.5 | Preparedness Framework determinationAI self-improvement | None stated | Preparedness | Unverified | |
| GPT-5.5 | Preparedness Framework determinationBiological and chemical | None stated | Preparedness | Unverified | |
| GPT-5.5 | Preparedness Framework determinationCybersecurity | None stated | Preparedness | Unverified |
M12 Chain-of-thought monitorability
2 values
Extraction coverage
What we read of this document, and where each value was read.
Our note We have not written a coverage note for this document yet.
All 37 values were read in the document itself.
Of the 37 values, 2 are statements in words rather than numbers; they are marked * and left out of charts by default.
All 37 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).