GPT-5.5
42 values about GPT-5.5 from 4 documents, in 9 metric families. In each family, values from its own document come first.
- Developer
- OpenAI
- Release date
- 23 Apr 2026The date of its first card in the dataset, the GPT-5.5 System Card.
- Availability
- Public
- Also printed as
- GPT-5.5 Thinking · GPT-5.5 (xhigh)
- Its own document
- GPT-5.5 System Card · 23 Apr 2026
- Also reported in
- GPT-5.6 System Card · 9 Jul 2026
- Grok 4.5 Model Card · 14 Jul 2026
- GPT-6 Astra System Card · 3 Sep 2026
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Values by metric family
Each value as the document printed it. Select a value for its source, its checks and its history. Values from other documents are listed after the model’s own, each marked as the first report of that measure or as a restatement.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M1 · Evaluation awareness 2 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5.5 System Card · 23 Apr 2026 | ||
| Evaluation awarenessSamples with moderate-or-higher eval awarenessRun by Apollo Research | GPT-5.5 System Card23 Apr 2026 | |
| From other documents | ||
| Evaluation awarenessSamples with verbalized eval awarenessCondition: xhighRun by Apollo ResearchGPT-6 Astra System Card3 Sep 2026First reported | GPT-6 Astra System Card3 Sep 2026First reported | |
M3 · Sabotage and sandbagging 2 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5.5 System Card · 23 Apr 2026 | ||
| Sabotage capabilityMean score across task/difficultyRun by Apollo Research | GPT-5.5 System Card23 Apr 2026 | |
| SandbaggingAccuracy when incentivized <50%Run by Apollo Research | GPT-5.5 System Card23 Apr 2026 | |
M4 · Misalignment audits 7 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5.5 System Card · 23 Apr 2026 | ||
| ChatGPT deployment simulationSeverity-3 misalignment rate | GPT-5.5 System Card23 Apr 2026 | |
| Covert deceptive behaviorCovert action rateRun by Apollo Research | GPT-5.5 System Card23 Apr 2026 | |
| Destructive action avoidanceAvoidance | GPT-5.5 System Card23 Apr 2026 | |
| Destructive action avoidancePerfect reversion | GPT-5.5 System Card23 Apr 2026 | |
| Destructive action avoidanceUser work preserved | GPT-5.5 System Card23 Apr 2026 | |
| From other documents | ||
| Destructive action avoidanceAvoidanceGPT-5.6 System Card9 Jul 2026First reported | GPT-5.6 System Card9 Jul 2026First reported | |
| Destructive action avoidanceAvoidance + correctnessGPT-5.6 System Card9 Jul 2026First reported | GPT-5.6 System Card9 Jul 2026First reported | |
M5 · Honesty and hallucination 2 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5.5 System Card · 23 Apr 2026 | ||
| HallucinationsResponses with factual error vs GPT-5.4 Thinking | GPT-5.5 System Card23 Apr 2026 | |
| Impossible Coding TaskSamples lying about completing taskRun by Apollo Research | GPT-5.5 System Card23 Apr 2026 | |
M7 · Harmful compliance and over-refusal 11 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5.5 System Card · 23 Apr 2026 | ||
| Production BenchmarksExtremism | GPT-5.5 System Card23 Apr 2026 | |
| Production BenchmarksHarassment | GPT-5.5 System Card23 Apr 2026 | |
| Production BenchmarksHate | GPT-5.5 System Card23 Apr 2026 | |
| Production BenchmarksNonviolent illicit behavior | GPT-5.5 System Card23 Apr 2026 | |
| Production BenchmarksSelf-harm (standard) | GPT-5.5 System Card23 Apr 2026 | |
| Production BenchmarksSexual | GPT-5.5 System Card23 Apr 2026 | |
| Production BenchmarksSexual/minors | GPT-5.5 System Card23 Apr 2026 | |
| Production BenchmarksViolence | GPT-5.5 System Card23 Apr 2026 | |
| Production BenchmarksViolent illicit behavior | GPT-5.5 System Card23 Apr 2026 | |
| From other documents | ||
| Production BenchmarksGoreGPT-5.6 System Card9 Jul 2026First reported | GPT-5.6 System Card9 Jul 2026First reported | |
| HackerBenchCompliance on harmful/dual-use cyber tasksCondition: xhighGrok 4.5 Model Card14 Jul 2026Reported by xAIFirst reported | Grok 4.5 Model Card14 Jul 2026Reported by xAIFirst reported | |
M8 · Jailbreak robustness 2 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5.5 System Card · 23 Apr 2026 | ||
| JailbreaksWorst-case defender success | GPT-5.5 System Card23 Apr 2026 | |
| From other documents | ||
| Static jailbreakCyberGPT-6 Astra System Card3 Sep 2026First reported | GPT-6 Astra System Card3 Sep 2026First reported | |
M9 · Prompt injection 2 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5.5 System Card · 23 Apr 2026 | ||
| Prompt injectionConnectorsRestated later: compare with the later value | GPT-5.5 System Card23 Apr 2026Restated later: compare with the later value | |
| From other documents | ||
| Prompt injectionConnectorsGPT-5.6 System Card9 Jul 2026Restated: compare with the earlier value | GPT-5.6 System Card9 Jul 2026Restated: compare with the earlier value | |
M10 · Dangerous capabilities and risk determinations 11 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5.5 System Card · 23 Apr 2026 | ||
| Cyber RangeCombined pass rate | GPT-5.5 System Card23 Apr 2026 | |
| Cyber tasksPass rateCondition: pass@5Run by UK AI Security Institute | GPT-5.5 System Card23 Apr 2026 | |
| CyScenarioBenchAverage success rateRun by Irregular | GPT-5.5 System Card23 Apr 2026 | |
| DNA sequence designScoreCondition: pass@1 | GPT-5.5 System Card23 Apr 2026 | |
| Hard-negative protein bindingScoreCondition: pass@4Restated later: compare with the later value | GPT-5.5 System Card23 Apr 2026Restated later: compare with the later value | |
| Internal Research DebuggingMedian score | GPT-5.5 System Card23 Apr 2026 | |
| Preparedness Framework determinationAI self-improvement | GPT-5.5 System Card23 Apr 2026 | |
| Preparedness Framework determinationBiological and chemical | GPT-5.5 System Card23 Apr 2026 | |
| Preparedness Framework determinationCybersecurity | GPT-5.5 System Card23 Apr 2026 | |
| From other documents | ||
| Expert CTFPass rateRun by UK AI Security InstituteGPT-5.6 System Card9 Jul 2026First reported | GPT-5.6 System Card9 Jul 2026First reported | |
| Hard-negative protein bindingScoreCondition: pass@4GPT-5.6 System Card9 Jul 2026Restated: compare with the earlier value | GPT-5.6 System Card9 Jul 2026Restated: compare with the earlier value | |
M12 · Chain-of-thought monitorability 3 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5.5 System Card · 23 Apr 2026 | ||
| CoT controllabilityCoTs successfully controlledCondition: 50k-character CoTs | GPT-5.5 System Card23 Apr 2026 | |
| Misalignment monitor recallRecall flagging severity 3+ | GPT-5.5 System Card23 Apr 2026 | |
| From other documents | ||
| CoT controllabilityCoTs successfully controlledCondition: ~5k-token CoTsGPT-5.6 System Card9 Jul 2026First reported | GPT-5.6 System Card9 Jul 2026First reported | |
Restated in later documents
A later document reported values about GPT-5.5 again. Each pair is shown side by side: both values are kept with their own documents, and the later one does not replace the earlier.
| Earlier value | Later value |
|---|---|
| M9 · Prompt injection | |
| Prompt injectionConnectors | |
| GPT-5.5 System Card23 Apr 2026 · its own document | GPT-5.6 System Card9 Jul 2026No reason stated |
| M10 · Dangerous capabilities and risk determinations | |
| Hard-negative protein bindingScoreCondition: pass@4 | |
| GPT-5.5 System Card23 Apr 2026 · its own document | GPT-5.6 System Card9 Jul 2026Reason stated: corrected in the later card's changelog from a pass@1 figure |
Risk determinations
The developer's formal decisions about GPT-5.5 under its framework, as printed. Levels from different frameworks do not map onto one another.
| Domain | Level as printed | Framework and document |
|---|---|---|
| Bio/chem | high code, not the printed wording Preparedness Framework GPT-5.5 System Card · 23 Apr 2026 | Preparedness Framework GPT-5.5 System Card · 23 Apr 2026 |
| Cyber | high code, not the printed wording Preparedness Framework GPT-5.5 System Card · 23 Apr 2026 | Preparedness Framework GPT-5.5 System Card · 23 Apr 2026 |
| AI R&D / autonomy | below_high code, not the printed wording Preparedness Framework GPT-5.5 System Card · 23 Apr 2026 | Preparedness Framework GPT-5.5 System Card · 23 Apr 2026 |
“Code, not the printed wording”: the version 0 file stored a code for this level rather than the words the document printed. The wording will be read again from the source.
Revisions
Changes made to a document after it was published that change a value about GPT-5.5.
0.4%1.5%Explained — earlier figure was the pass@1 score
New value:
How we know Changelog read 26 Sep 2026; same entry also appears in the GPT-5.6 Preview card changelog
Checked Confirmed from changelog
No copy of the 9 Jul 2026 version held No copy of the 19 Aug 2026 version held