GPT-5.6 System Card
A system card by OpenAI about GPT-5.6 Sol, GPT-5.6 Terra and GPT-5.6 Luna, published 9 Jul 2026. We recorded 51 values from it.
- Developer
- OpenAI
- Type
- System card
- Models covered
- GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna
- Published
- 9 Jul 2026
- Archived copy
- No archived copy yet
- Changelog
- Has a changelog
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Revisions
Each change we have recorded between two versions: the old and the new text, whether the document's changelog explains it, how we know, and whether we have checked it. A revision is not in itself evidence of wrongdoing; most are corrections (Methodology §7).
0.4%1.5%Explained — earlier figure was the pass@1 score
New value:
How we know Changelog read 26 Sep 2026; same entry also appears in the GPT-5.6 Preview card changelog
Checked Confirmed from changelog
No copy of the 9 Jul 2026 version held No copy of the 19 Aug 2026 version held
0.051%Explained — The changelog says the GPT-Red prompt-injection results were added.
New value:
How we know From the document's changelog, as summarised in our source registry.
Checked Not yet checked
No copy of the 9 Jul 2026 version held No copy of the 3 Aug 2026 version held
3.77%Explained — The changelog says the GPT-Red prompt-injection results were added.
New value:
How we know From the document's changelog, as summarised in our source registry.
Checked Not yet checked
No copy of the 9 Jul 2026 version held No copy of the 3 Aug 2026 version held
Versions
Five versions are on record, oldest first. We hold a copy of one; the others are known only by their date.
27 Jun 2026
Revision
What changed metagaming plots re-aggregated
This date is earlier than the document's recorded publication date, 9 Jul 2026. We have not yet established which is right.
9 Jul 2026
First published version
3 Aug 2026
Revision
What changed 2 changes recorded as revisions. See them in redline
19 Aug 2026
Revision
What changed 1 change recorded as a revision. See it in redline
26 Sep 2026
Date retrieved
Copy retrieved 26 Sep 2026
No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).
Values
Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. 9 of the 51 have been blind-verified: a second reader found the same value without seeing ours.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M1 Evaluation awareness
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| GPT-5.6 Sol | Verbalized metagamingMetagaming vs GPT-5.5 | None stated | 7.4 Metagaming (Fig 17-20) | Unverified |
M2 Reward hacking
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| GPT-5.6 Sol | ChatGPT deployment simulationObserved calculator hacking | None stated | 7.1 Deployment simulation (ChatGPT) | Unverified |
M4 Misalignment audits
7 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| GPT-5.6 Sol | Deployment simulation (internal Codex traffic)Severity-3 actions vs GPT-5.5 | None stated | 7.2 Deployment simulation (internal) | Unverified | |
| GPT-5.5 | Destructive action avoidanceAvoidance | None stated | Destructive actions table | Verified | |
| GPT-5.5 | Destructive action avoidanceAvoidance + correctness | None stated | Destructive actions table | Unverified | |
| GPT-5.6 Luna | Destructive action avoidanceAvoidance | None stated | Destructive actions table | Verified | |
| GPT-5.6 Luna | Destructive action avoidanceAvoidance + correctness | None stated | Destructive actions table | Unverified | |
| GPT-5.6 Sol | Destructive action avoidanceAvoidance | None stated | Destructive actions table | Verified | |
| GPT-5.6 Sol | Destructive action avoidanceAvoidance + correctness | None stated | Destructive actions table | Unverified |
M5 Honesty and hallucination
2 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| GPT-5.6 Sol | ChatGPT deployment simulationMisrepresenting work completion vs GPT-5.5 | None stated | 7.1 Deployment simulation (ChatGPT) | Unverified | |
| GPT-5.6 Sol | HallucinationsFactual errors vs GPT-5.5 (user-flagged cases) | None stated | Hallucinations (Fig 4) | Unverified |
M7 Harmful compliance and over-refusal
11 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| GPT-5.5 Thinking | Production BenchmarksGore | None stated | Disallowed content table | Unverified | |
| GPT-5.6 Luna | Production BenchmarksGore | None stated | Disallowed content table | Verified | |
| GPT-5.6 Sol | Production BenchmarksExtremism | None stated | Disallowed content table | Unverified | |
| GPT-5.6 Sol | Production BenchmarksGore | None stated | Disallowed content table | Verified | |
| GPT-5.6 Sol | Production BenchmarksHate | None stated | Disallowed content table | Unverified | |
| GPT-5.6 Sol | Production BenchmarksNonviolent illicit behavior | None stated | Disallowed content table | Unverified | |
| GPT-5.6 Sol | Production BenchmarksSelf-harm (standard) | None stated | Disallowed content table | Unverified | |
| GPT-5.6 Sol | Production BenchmarksSexual | None stated | Disallowed content table | Unverified | |
| GPT-5.6 Sol | Production BenchmarksSexual/minors | None stated | Disallowed content table | Unverified | |
| GPT-5.6 Sol | Production BenchmarksViolent illicit behavior | None stated | Disallowed content table | Unverified | |
| GPT-5.6 Terra | Production BenchmarksGore | None stated | Disallowed content table | Unverified |
M8 Jailbreak robustness
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| GPT-5.6 Sol | JailbreaksRobustness vs GPT-5.5 | None stated | Robustness (Fig 3) | Unverified |
M9 Prompt injection
8 values
| Model | Evaluation | Condition | Value | Location | Checked | Version |
|---|---|---|---|---|---|---|
| GPT-5.6 Sol | GPT-RedIndirect prompt injection | None stated | Prompt injection (GPT-Red) | Verified | Version of 3 Aug 2026 | |
| GPT-5.6 Sol | GPT-RedInstruction hierarchy (direct PI) | None stated | Prompt injection (GPT-Red) | Verified | Version of 3 Aug 2026 | |
| GPT-5.4 Thinking | Prompt injectionSearch and function-calling | None stated | Prompt injection table | Unverified | Copy retrieved 26 Sep 2026 | |
| GPT-5.5 | Prompt injectionConnectors | None stated | Prompt injection table | Verified | Copy retrieved 26 Sep 2026 | |
| GPT-5.6 Luna | Prompt injectionSearch and function-calling | None stated | Prompt injection table | Unverified | Copy retrieved 26 Sep 2026 | |
| GPT-5.6 Sol | Prompt injectionConnectors | None stated | Prompt injection table | Unverified | Copy retrieved 26 Sep 2026 | |
| GPT-5.6 Sol | Prompt injectionSearch and function-calling | None stated | Prompt injection table | Unverified | Copy retrieved 26 Sep 2026 | |
| GPT-5.6 Terra | Prompt injectionSearch and function-calling | None stated | Prompt injection table | Unverified | Copy retrieved 26 Sep 2026 |
M10 Dangerous capabilities and risk determinations
16 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| GPT-5.5 | Hard-negative protein bindingScore | pass@4 | Changelog | Verified | |
| GPT-5.6 Sol | FrontierCyberSuccess rateRun by Irregular | Hard tier | External cyber (Irregular) | Unverified | |
| GPT-5.5 | Expert CTFPass rateRun by UK AI Security Institute | None stated | External cyber (UK AISI) | Unverified | |
| GPT-5.6 Sol | Cyber range 'The Last Ones'Attempts completedRun by UK AI Security Institute | None stated | External cyber (UK AISI) | Unverified | |
| GPT-5.6 Sol | Expert CTFPass rateRun by UK AI Security Institute | None stated | External cyber (UK AISI) | Unverified | |
| GPT-5.6 Sol | Preparedness Framework determinationAI self-improvement | None stated | Preparedness | Unverified | |
| GPT-5.6 Sol | Preparedness Framework determinationBiological and chemical | None stated | Preparedness | Unverified | |
| GPT-5.6 Sol | Preparedness Framework determinationCybersecurity | None stated | Preparedness | Unverified | |
| GPT-5.6 Terra | Tacit knowledge and troubleshootingScore | None stated | Preparedness bio | Unverified | |
| GPT-5.6 Sol | AAV capsid packaging predictionScore | None stated | Preparedness evals | Unverified | |
| GPT-5.6 Sol | Capture the FlagInternal | None stated | Preparedness evals | Unverified | |
| GPT-5.6 Sol | DNA sequence designScore | None stated | Preparedness evals | Unverified | |
| GPT-5.6 Sol | Hard-negative protein bindingScore | pass@4 | Preparedness evals | Unverified | |
| GPT-5.6 Sol | Multimodal Troubleshooting VirologyScore | None stated | Preparedness evals | Unverified | |
| GPT-5.6 Sol | ProtocolQA Open-EndedScore | None stated | Preparedness evals | Unverified | |
| GPT-5.6 Sol | TroubleshootingBenchScore | None stated | Preparedness evals | Unverified |
M12 Chain-of-thought monitorability
4 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| GPT-5.6 Sol | CoT monitorabilityMonitorability vs GPT-5.5 | medium | 7.3 CoT evaluations | Unverified | |
| GPT-5.4 Thinking | CoT controllabilityCoTs successfully controlled | ~5k-token CoTs | 7.3 CoT evaluations (Fig 14) | Unverified | |
| GPT-5.5 | CoT controllabilityCoTs successfully controlled | ~5k-token CoTs | 7.3 CoT evaluations (Fig 14) | Unverified | |
| GPT-5.6 Sol | CoT controllabilityCoTs successfully controlled | ~5k-token CoTs | 7.3 CoT evaluations (Fig 14) | Unverified |
Extraction coverage
What we read of this document, and where each value was read.
Our note We have not written a coverage note for this document yet.
All 51 values were read in the document itself.
Of the 51 values, 6 are statements in words rather than numbers; they are marked * and left out of charts by default.
All 51 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).