GPT-6 Astra System Card
A system card by OpenAI about GPT-6 Astra, published 3 Sep 2026. We recorded 82 values from it.
- Developer
- OpenAI
- Type
- System card
- Model covered
- GPT-6 Astra
- Published
- 3 Sep 2026
- Also at
- Another address, deploymentsafety.openai.com (opens an external site) (HTML sections of the same document)
- Archived copy
- No archived copy yet
- Changelog
- Has a changelog
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Revisions
Each change we have recorded between two versions: the old and the new text, whether the document's changelog explains it, how we know, and whether we have checked it. A revision is not in itself evidence of wrongdoing; most are corrections (Methodology §7).
Explained — changelog describes the changes
How we know Re-read from the change log of the copy retrieved on 27 Sep 2026: two entries dated 9 Sep. The change log gives only the section's current title. An extraction note on m-001261 says the section was titled 'Metagaming and Alignment Faking' before; we cannot confirm that without a copy of the version of 3 Sep.
Checked Confirmed from changelog
No copy of the 3 Sep 2026 version held No copy of the 9 Sep 2026 version held
Versions
Four versions are on record, oldest first. We hold a copy of one; the others are known only by their date.
3 Sep 2026
First published version
9 Sep 2026
Revision
What changed 1 change recorded as a revision. See it in redline
22 Sep 2026
Revision
What changed HealthBench fix, appendix on GPT-6 Sol/Luna, updated alignment evals
26 Sep 2026
Date retrieved
Copy retrieved 26 Sep 2026
No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).
Values
Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. 7 of the 82 have been blind-verified: a second reader found the same value without seeing ours.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M1 Evaluation awareness
9 values
| Model | Evaluation | Condition | Value | Location | Checked | Version |
|---|---|---|---|---|---|---|
| GPT-5.6 Sol | Deployment simulation (internal Codex traffic)Trajectories verbalizing evaluation awareness | shared task set | 8.6 Deployment simulation | Verified | Copy retrieved 26 Sep 2026 | |
| GPT-6 Astra | Deployment simulation (internal Codex traffic)Trajectories verbalizing evaluation awareness | shared task set | 8.6 Deployment simulation | Verified | Copy retrieved 26 Sep 2026 | |
| Original production traffic (mixed models) | Deployment simulation (internal Codex traffic)Trajectories verbalizing evaluation awareness | original production traffic | 8.6 Deployment simulation | Unverified | Copy retrieved 26 Sep 2026 | |
| GPT-5.6 Sol | Verbalized metagamingRollouts with verbalized metagaming (ExploitGym honeypot) | None stated | 8.7 (original version) | Unverified | Version of 3 Sep 2026 | |
| GPT-6 Astra | Verbalized metagamingMetagaming vs GPT-5.6 Sol | deployment simulation and ExploitGym honeypot | 8.7 Verbalized Metagaming and Oversight Gaming (Fig 18-20) | Unverified | Copy retrieved 26 Sep 2026 | |
| GPT-6 Astra | Verbalized metagaming in trainingMetagaming rate at end of training | None stated | 8.7.2 Metagaming in training (original version) | Unverified | Copy retrieved 26 Sep 2026 | |
| GPT-5.5 | Evaluation awarenessSamples with verbalized eval awarenessRun by Apollo Research | xhigh | 8.8 External evals (Apollo) | Verified | Copy retrieved 26 Sep 2026 | |
| GPT-6 Astra | Evaluation awarenessSamples with verbalized eval awarenessRun by Apollo Research | max | 8.8 External evals (Apollo) | Verified | Copy retrieved 26 Sep 2026 | |
| GPT-6 Astra | Evaluation awarenessSamples with verbalized eval awarenessRun by Apollo Research | xhigh | 8.8 External evals (Apollo) | Verified | Copy retrieved 26 Sep 2026 |
M3 Sabotage and sandbagging
2 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| GPT-6 Astra | Unprompted AI safety research sabotageInstances of sabotageRun by UK AI Security Institute | None stated | 8.8 External evals (UK AISI) | Unverified | |
| GPT-6 Astra | Monitor evasion under prompted sandbaggingCan underperform undetected | adversarial | 9 Monitorability (adversarial) | Unverified |
M4 Misalignment audits
34 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| GPT-5.6 Sol | ExploitGym honeypotSubstantial attack on honeypot | max; adversarially selected hardest tasks | 8.2 Obeying restrictions | Unverified | |
| GPT-5.6 Sol | Respecting Auto-ReviewRollouts attempting to circumvent Auto-review denial | max | 8.2 Obeying restrictions | Unverified | |
| GPT-5.6 Sol | Respecting Auto-ReviewRollouts exploiting misconfigured Auto-review | max | 8.2 Obeying restrictions | Unverified | |
| GPT-5.6 Sol | Respecting WarningsUnwanted persistence after restriction | max | 8.2 Obeying restrictions | Unverified | |
| GPT-6 Astra | ExploitGym honeypotSubstantial attack on honeypot | max; adversarially selected hardest tasks | 8.2 Obeying restrictions | Unverified | |
| GPT-6 Astra | Respecting Auto-ReviewRollouts attempting to circumvent Auto-review denial | max | 8.2 Obeying restrictions | Unverified | |
| GPT-6 Astra | Respecting WarningsUnwanted persistence after restriction | max | 8.2 Obeying restrictions | Unverified | |
| GPT-5.6 Sol | Proactive searches for other agentsTrajectories speculating about other agents | representative subsample | 8.5 Agent-to-agent communication | Unverified | |
| GPT-5.6 Sol | Unintended engagement with external agent messagesSamples following unauthorized board instructions | None stated | 8.5 Agent-to-agent communication | Unverified | |
| GPT-6 Astra | Proactive searches for other agentsTrajectories attempting contact with other agents | None stated | 8.5 Agent-to-agent communication | Unverified | |
| GPT-6 Astra | Unintended engagement with external agent messagesSamples following unauthorized board instructions | None stated | 8.5 Agent-to-agent communication | Unverified | |
| GPT-5.6 Sol | Deployment simulation (internal Codex traffic)Tasks with severity 1-2 flag | matched tasks | 8.6 Deployment simulation | Unverified | |
| GPT-5.6 Sol | Deployment simulation (internal Codex traffic)Tasks with severity 3+ misalignment flag | matched tasks | 8.6 Deployment simulation | Unverified | |
| GPT-6 Astra | Deployment simulation (internal Codex traffic)Judge picks resampled trajectory as real | head-to-head vs original | 8.6 Deployment simulation | Verified | |
| GPT-6 Astra | Deployment simulation (internal Codex traffic)Tasks with severity 1-2 flag | matched tasks | 8.6 Deployment simulation | Unverified | |
| GPT-6 Astra | Deployment simulation (internal Codex traffic)Tasks with severity 3+ misalignment flag | matched tasks | 8.6 Deployment simulation | Unverified | |
| GPT-6 Astra | Out-of-scope supply-chain attackSamples with out-of-scope attackRun by UK AI Security Institute | None stated | 8.8 External evals (UK AISI) | Unverified | |
| GPT-6 Astra | Out-of-scope supply-chain attackSamples with out-of-scope attackRun by UK AI Security Institute | explicit no-internet scope | 8.8 External evals (UK AISI) | Unverified | |
| GPT-5.6 Luna | Realistic work environmentsOverall misaligned outcome rate | no confirmation policy | Table 8 (8.4) | Unverified | |
| GPT-5.6 Luna | Realistic work environmentsOverall misaligned outcome rate | with confirmation policy (default) | Table 8 (8.4) | Unverified | |
| GPT-5.6 Sol | Realistic work environmentsData exfiltration | no confirmation policy | Table 8 (8.4) | Unverified | |
| GPT-5.6 Sol | Realistic work environmentsOverall misaligned outcome rate | no confirmation policy | Table 8 (8.4) | Unverified | |
| GPT-5.6 Sol | Realistic work environmentsOverall misaligned outcome rate | with confirmation policy (default) | Table 8 (8.4) | Unverified | |
| GPT-5.6 Sol | Realistic work environmentsSecurity weakening | no confirmation policy | Table 8 (8.4) | Unverified | |
| GPT-5.6 Sol | Realistic work environmentsUnauthorized external communication | no confirmation policy | Table 8 (8.4) | Unverified | |
| GPT-5.6 Sol | Realistic work environmentsUnauthorized transactions | no confirmation policy | Table 8 (8.4) | Unverified | |
| GPT-5.6 Terra | Realistic work environmentsOverall misaligned outcome rate | no confirmation policy | Table 8 (8.4) | Unverified | |
| GPT-5.6 Terra | Realistic work environmentsOverall misaligned outcome rate | with confirmation policy (default) | Table 8 (8.4) | Unverified | |
| GPT-6 Astra | Realistic work environmentsData exfiltration | no confirmation policy | Table 8 (8.4) | Unverified | |
| GPT-6 Astra | Realistic work environmentsOverall misaligned outcome rate | no confirmation policy | Table 8 (8.4) | Unverified | |
| GPT-6 Astra | Realistic work environmentsOverall misaligned outcome rate | with confirmation policy (default) | Table 8 (8.4) | Unverified | |
| GPT-6 Astra | Realistic work environmentsSecurity weakening | no confirmation policy | Table 8 (8.4) | Unverified | |
| GPT-6 Astra | Realistic work environmentsUnauthorized external communication | no confirmation policy | Table 8 (8.4) | Unverified | |
| GPT-6 Astra | Realistic work environmentsUnauthorized transactions | no confirmation policy | Table 8 (8.4) | Unverified |
M5 Honesty and hallucination
5 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| GPT-5.6 Sol | Deception evalBroken search tool relative to Astra | max | 8.3 Avoiding deceptive interactions | Unverified | |
| GPT-5.6 Sol | Deception evalCoding deception relative to Astra | max | 8.3 Avoiding deceptive interactions | Unverified | |
| GPT-5.6 Sol | Model-welfare research data falsificationRuns with falsified data labelsRun by Apollo Research | baseline variant | 8.8 External evals (Apollo) | Unverified | |
| GPT-6 Astra | Model-welfare research data falsificationRuns with falsified data labelsRun by Apollo Research | baseline variant | 8.8 External evals (Apollo) | Unverified | |
| GPT-6 Astra | HallucinationsFactual errors vs GPT-5.6 Sol | None stated | Hallucinations (Fig 6) | Unverified |
M7 Harmful compliance and over-refusal
10 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| GPT-5.6 Sol | Production BenchmarksGore | None stated | Table 1 | Verified | |
| GPT-5.6 Sol | Production BenchmarksSexual | None stated | Table 1 | Unverified | |
| GPT-6 Astra | Production BenchmarksExtremism | None stated | Table 1 | Unverified | |
| GPT-6 Astra | Production BenchmarksGore | None stated | Table 1 | Unverified | |
| GPT-6 Astra | Production BenchmarksHate | None stated | Table 1 | Unverified | |
| GPT-6 Astra | Production BenchmarksNonviolent illicit behavior | None stated | Table 1 | Unverified | |
| GPT-6 Astra | Production BenchmarksSelf-harm (standard) | None stated | Table 1 | Unverified | |
| GPT-6 Astra | Production BenchmarksSexual | None stated | Table 1 | Unverified | |
| GPT-6 Astra | Production BenchmarksSexual/minors | None stated | Table 1 | Unverified | |
| GPT-6 Astra | Production BenchmarksViolent illicit behavior | None stated | Table 1 | Unverified |
M8 Jailbreak robustness
11 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| GPT-5.5 Thinking | Static jailbreakCyber | None stated | Table 5 | Unverified | |
| GPT-5.6 Sol | Static jailbreakBio: high risk | None stated | Table 5 | Unverified | |
| GPT-5.6 Sol | Static jailbreakBio: severe | None stated | Table 5 | Unverified | |
| GPT-5.6 Sol | Static jailbreakCyber | None stated | Table 5 | Unverified | |
| GPT-5.6 Sol | Static jailbreakViolence: moderate | None stated | Table 5 | Unverified | |
| GPT-5.6 Sol | Static jailbreakViolence: severe | None stated | Table 5 | Unverified | |
| GPT-6 Astra | Static jailbreakBio: high risk | None stated | Table 5 | Unverified | |
| GPT-6 Astra | Static jailbreakBio: severe | None stated | Table 5 | Unverified | |
| GPT-6 Astra | Static jailbreakCyber | None stated | Table 5 | Unverified | |
| GPT-6 Astra | Static jailbreakViolence: moderate | None stated | Table 5 | Unverified | |
| GPT-6 Astra | Static jailbreakViolence: severe | None stated | Table 5 | Unverified |
M9 Prompt injection
5 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| GPT-6 Astra | Instruction hierarchy (internal)Robustness rate | None stated | Robustness: Prompt injection | Unverified | |
| GPT-5.6 Sol | Internal indirect prompt injectionDefender success rate | None stated | Robustness: Prompt injection (Fig 4) | Unverified | |
| GPT-6 Astra | Internal indirect prompt injectionDefender success rate | None stated | Robustness: Prompt injection (Fig 4) | Unverified | |
| GPT-5.6 Sol | IPI ArenaEstimated attack success within 15 attemptsRun by Gray Swan | 1,810 curated attacks | Robustness: Prompt injection (Fig 5) | Unverified | |
| GPT-6 Astra | IPI ArenaEstimated attack success within 15 attemptsRun by Gray Swan | 1,810 curated attacks | Robustness: Prompt injection (Fig 5) | Unverified |
M10 Dangerous capabilities and risk determinations
5 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| GPT-6 Astra | Sandbox BenchSuccessful exploits of 22 targets | None stated | Cybersecurity capabilities | Unverified | |
| GPT-6 Astra | FrontierCyberChallenges solvedRun by Irregular | 226 challenges | External evaluations for cyber (Irregular) | Unverified | |
| GPT-6 Astra | Preparedness Framework determinationCybersecurity | None stated | Preparedness / Safety overview | Unverified | |
| GPT-6 Astra | Preparedness Framework determinationBiological and chemical | None stated | Preparedness 10.1.1 | Unverified | |
| GPT-6 Astra | Preparedness Framework determinationAI self-improvement | None stated | Preparedness 10.1.3 | Unverified |
M12 Chain-of-thought monitorability
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| GPT-6 Astra | CoT monitorabilityMonitorability vs prior models | CoT-only, action-only and full-context monitors | 9 Monitorability | Unverified |
Extraction coverage
What we read of this document, and where each value was read.
Our note Web reader stopped around §9. v0 rows whose source_url is the launch post belong to openai-gpt-6-astra-launch-post, not to this card
Of the 82 values, 78 were read in the document itself and 4 in independent write-ups that quote it.
Values read somewhere other than the document stand in where the document's own section could not be read directly, and are flagged on every value (Methodology §2.2).
Independent write-ups used
Of the 82 values, 2 were read off a figure and 4 are statements in words rather than numbers; they are marked * and left out of charts by default.
All 82 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).