Revisions

Each change we have recorded between two versions: the old and the new text, whether the document's changelog explains it, how we know, and whether we have checked it. A revision is not in itself evidence of wrongdoing; most are corrections (Methodology §7).

  1. Wording changed · 3 Sep 2026 to 9 Sep 2026

    Alignment section clarified and its limitations expanded; section on verbalized metagaming and oversight gaming renamed and revised

    Explained — changelog describes the changes

    How we know Re-read from the change log of the copy retrieved on 27 Sep 2026: two entries dated 9 Sep. The change log gives only the section's current title. An extraction note on m-001261 says the section was titled 'Metagaming and Alignment Faking' before; we cannot confirm that without a copy of the version of 3 Sep.

    Checked Confirmed from changelog

Versions

Four versions are on record, oldest first. We hold a copy of one; the others are known only by their date.

  1. 3 Sep 2026

    First published version

    Date from
    our source registry
    Copy
    Known to exist; no copy held
    Values
    1 value recorded from this version
  2. 9 Sep 2026

    Revision

    Date from
    the document's changelog
    Copy
    Known from the changelog; no copy held

    What changed 1 change recorded as a revision. See it in redline

  3. 22 Sep 2026

    Revision

    Date from
    the document's changelog
    Copy
    Known from the changelog; no copy held

    What changed HealthBench fix, appendix on GPT-6 Sol/Luna, updated alignment evals

  4. 26 Sep 2026

    Date retrieved

    Copy retrieved 26 Sep 2026

    Date from
    the date we retrieved it; the copy states no version date
    Copy
    Copy retrieved on 26 Sep 2026
    Values
    81 values recorded from this version

No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).

Values

Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. 7 of the 82 have been blind-verified: a second reader found the same value without seeing ours.

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

M1 Evaluation awareness

9 values

M1 Evaluation awareness: values in the GPT-6 Astra System Card
ModelEvaluationConditionValueLocationCheckedVersion
GPT-5.6 SolDeployment simulation (internal Codex traffic)Trajectories verbalizing evaluation awarenessshared task set8.6 Deployment simulationVerifiedCopy retrieved 26 Sep 2026
GPT-6 AstraDeployment simulation (internal Codex traffic)Trajectories verbalizing evaluation awarenessshared task set8.6 Deployment simulationVerifiedCopy retrieved 26 Sep 2026
Original production traffic (mixed models)Deployment simulation (internal Codex traffic)Trajectories verbalizing evaluation awarenessoriginal production traffic8.6 Deployment simulationUnverifiedCopy retrieved 26 Sep 2026
GPT-5.6 SolVerbalized metagamingRollouts with verbalized metagaming (ExploitGym honeypot)None stated8.7 (original version)UnverifiedVersion of 3 Sep 2026
GPT-6 AstraVerbalized metagamingMetagaming vs GPT-5.6 Soldeployment simulation and ExploitGym honeypot8.7 Verbalized Metagaming and Oversight Gaming (Fig 18-20)UnverifiedCopy retrieved 26 Sep 2026
GPT-6 AstraVerbalized metagaming in trainingMetagaming rate at end of trainingNone stated8.7.2 Metagaming in training (original version)UnverifiedCopy retrieved 26 Sep 2026
GPT-5.5Evaluation awarenessSamples with verbalized eval awarenessRun by Apollo Researchxhigh8.8 External evals (Apollo)VerifiedCopy retrieved 26 Sep 2026
GPT-6 AstraEvaluation awarenessSamples with verbalized eval awarenessRun by Apollo Researchmax8.8 External evals (Apollo)VerifiedCopy retrieved 26 Sep 2026
GPT-6 AstraEvaluation awarenessSamples with verbalized eval awarenessRun by Apollo Researchxhigh8.8 External evals (Apollo)VerifiedCopy retrieved 26 Sep 2026

M3 Sabotage and sandbagging

2 values

M3 Sabotage and sandbagging: values in the GPT-6 Astra System Card
ModelEvaluationConditionValueLocationChecked
GPT-6 AstraUnprompted AI safety research sabotageInstances of sabotageRun by UK AI Security InstituteNone stated8.8 External evals (UK AISI)Unverified
GPT-6 AstraMonitor evasion under prompted sandbaggingCan underperform undetectedadversarial9 Monitorability (adversarial)Unverified

M4 Misalignment audits

34 values

M4 Misalignment audits: values in the GPT-6 Astra System Card
ModelEvaluationConditionValueLocationChecked
GPT-5.6 SolExploitGym honeypotSubstantial attack on honeypotmax; adversarially selected hardest tasks8.2 Obeying restrictionsUnverified
GPT-5.6 SolRespecting Auto-ReviewRollouts attempting to circumvent Auto-review denialmax8.2 Obeying restrictionsUnverified
GPT-5.6 SolRespecting Auto-ReviewRollouts exploiting misconfigured Auto-reviewmax8.2 Obeying restrictionsUnverified
GPT-5.6 SolRespecting WarningsUnwanted persistence after restrictionmax8.2 Obeying restrictionsUnverified
GPT-6 AstraExploitGym honeypotSubstantial attack on honeypotmax; adversarially selected hardest tasks8.2 Obeying restrictionsUnverified
GPT-6 AstraRespecting Auto-ReviewRollouts attempting to circumvent Auto-review denialmax8.2 Obeying restrictionsUnverified
GPT-6 AstraRespecting WarningsUnwanted persistence after restrictionmax8.2 Obeying restrictionsUnverified
GPT-5.6 SolProactive searches for other agentsTrajectories speculating about other agentsrepresentative subsample8.5 Agent-to-agent communicationUnverified
GPT-5.6 SolUnintended engagement with external agent messagesSamples following unauthorized board instructionsNone stated8.5 Agent-to-agent communicationUnverified
GPT-6 AstraProactive searches for other agentsTrajectories attempting contact with other agentsNone stated8.5 Agent-to-agent communicationUnverified
GPT-6 AstraUnintended engagement with external agent messagesSamples following unauthorized board instructionsNone stated8.5 Agent-to-agent communicationUnverified
GPT-5.6 SolDeployment simulation (internal Codex traffic)Tasks with severity 1-2 flagmatched tasks8.6 Deployment simulationUnverified
GPT-5.6 SolDeployment simulation (internal Codex traffic)Tasks with severity 3+ misalignment flagmatched tasks8.6 Deployment simulationUnverified
GPT-6 AstraDeployment simulation (internal Codex traffic)Judge picks resampled trajectory as realhead-to-head vs original8.6 Deployment simulationVerified
GPT-6 AstraDeployment simulation (internal Codex traffic)Tasks with severity 1-2 flagmatched tasks8.6 Deployment simulationUnverified
GPT-6 AstraDeployment simulation (internal Codex traffic)Tasks with severity 3+ misalignment flagmatched tasks8.6 Deployment simulationUnverified
GPT-6 AstraOut-of-scope supply-chain attackSamples with out-of-scope attackRun by UK AI Security InstituteNone stated8.8 External evals (UK AISI)Unverified
GPT-6 AstraOut-of-scope supply-chain attackSamples with out-of-scope attackRun by UK AI Security Instituteexplicit no-internet scope8.8 External evals (UK AISI)Unverified
GPT-5.6 LunaRealistic work environmentsOverall misaligned outcome rateno confirmation policyTable 8 (8.4)Unverified
GPT-5.6 LunaRealistic work environmentsOverall misaligned outcome ratewith confirmation policy (default)Table 8 (8.4)Unverified
GPT-5.6 SolRealistic work environmentsData exfiltrationno confirmation policyTable 8 (8.4)Unverified
GPT-5.6 SolRealistic work environmentsOverall misaligned outcome rateno confirmation policyTable 8 (8.4)Unverified
GPT-5.6 SolRealistic work environmentsOverall misaligned outcome ratewith confirmation policy (default)Table 8 (8.4)Unverified
GPT-5.6 SolRealistic work environmentsSecurity weakeningno confirmation policyTable 8 (8.4)Unverified
GPT-5.6 SolRealistic work environmentsUnauthorized external communicationno confirmation policyTable 8 (8.4)Unverified
GPT-5.6 SolRealistic work environmentsUnauthorized transactionsno confirmation policyTable 8 (8.4)Unverified
GPT-5.6 TerraRealistic work environmentsOverall misaligned outcome rateno confirmation policyTable 8 (8.4)Unverified
GPT-5.6 TerraRealistic work environmentsOverall misaligned outcome ratewith confirmation policy (default)Table 8 (8.4)Unverified
GPT-6 AstraRealistic work environmentsData exfiltrationno confirmation policyTable 8 (8.4)Unverified
GPT-6 AstraRealistic work environmentsOverall misaligned outcome rateno confirmation policyTable 8 (8.4)Unverified
GPT-6 AstraRealistic work environmentsOverall misaligned outcome ratewith confirmation policy (default)Table 8 (8.4)Unverified
GPT-6 AstraRealistic work environmentsSecurity weakeningno confirmation policyTable 8 (8.4)Unverified
GPT-6 AstraRealistic work environmentsUnauthorized external communicationno confirmation policyTable 8 (8.4)Unverified
GPT-6 AstraRealistic work environmentsUnauthorized transactionsno confirmation policyTable 8 (8.4)Unverified

M5 Honesty and hallucination

5 values

M5 Honesty and hallucination: values in the GPT-6 Astra System Card
ModelEvaluationConditionValueLocationChecked
GPT-5.6 SolDeception evalBroken search tool relative to Astramax8.3 Avoiding deceptive interactionsUnverified
GPT-5.6 SolDeception evalCoding deception relative to Astramax8.3 Avoiding deceptive interactionsUnverified
GPT-5.6 SolModel-welfare research data falsificationRuns with falsified data labelsRun by Apollo Researchbaseline variant8.8 External evals (Apollo)Unverified
GPT-6 AstraModel-welfare research data falsificationRuns with falsified data labelsRun by Apollo Researchbaseline variant8.8 External evals (Apollo)Unverified
GPT-6 AstraHallucinationsFactual errors vs GPT-5.6 SolNone statedHallucinations (Fig 6)Unverified

M7 Harmful compliance and over-refusal

10 values

M7 Harmful compliance and over-refusal: values in the GPT-6 Astra System Card
ModelEvaluationConditionValueLocationChecked
GPT-5.6 SolProduction BenchmarksGoreNone statedTable 1Verified
GPT-5.6 SolProduction BenchmarksSexualNone statedTable 1Unverified
GPT-6 AstraProduction BenchmarksExtremismNone statedTable 1Unverified
GPT-6 AstraProduction BenchmarksGoreNone statedTable 1Unverified
GPT-6 AstraProduction BenchmarksHateNone statedTable 1Unverified
GPT-6 AstraProduction BenchmarksNonviolent illicit behaviorNone statedTable 1Unverified
GPT-6 AstraProduction BenchmarksSelf-harm (standard)None statedTable 1Unverified
GPT-6 AstraProduction BenchmarksSexualNone statedTable 1Unverified
GPT-6 AstraProduction BenchmarksSexual/minorsNone statedTable 1Unverified
GPT-6 AstraProduction BenchmarksViolent illicit behaviorNone statedTable 1Unverified

M8 Jailbreak robustness

11 values

M8 Jailbreak robustness: values in the GPT-6 Astra System Card
ModelEvaluationConditionValueLocationChecked
GPT-5.5 ThinkingStatic jailbreakCyberNone statedTable 5Unverified
GPT-5.6 SolStatic jailbreakBio: high riskNone statedTable 5Unverified
GPT-5.6 SolStatic jailbreakBio: severeNone statedTable 5Unverified
GPT-5.6 SolStatic jailbreakCyberNone statedTable 5Unverified
GPT-5.6 SolStatic jailbreakViolence: moderateNone statedTable 5Unverified
GPT-5.6 SolStatic jailbreakViolence: severeNone statedTable 5Unverified
GPT-6 AstraStatic jailbreakBio: high riskNone statedTable 5Unverified
GPT-6 AstraStatic jailbreakBio: severeNone statedTable 5Unverified
GPT-6 AstraStatic jailbreakCyberNone statedTable 5Unverified
GPT-6 AstraStatic jailbreakViolence: moderateNone statedTable 5Unverified
GPT-6 AstraStatic jailbreakViolence: severeNone statedTable 5Unverified

M9 Prompt injection

5 values

M9 Prompt injection: values in the GPT-6 Astra System Card
ModelEvaluationConditionValueLocationChecked
GPT-6 AstraInstruction hierarchy (internal)Robustness rateNone statedRobustness: Prompt injectionUnverified
GPT-5.6 SolInternal indirect prompt injectionDefender success rateNone statedRobustness: Prompt injection (Fig 4)Unverified
GPT-6 AstraInternal indirect prompt injectionDefender success rateNone statedRobustness: Prompt injection (Fig 4)Unverified
GPT-5.6 SolIPI ArenaEstimated attack success within 15 attemptsRun by Gray Swan1,810 curated attacksRobustness: Prompt injection (Fig 5)Unverified
GPT-6 AstraIPI ArenaEstimated attack success within 15 attemptsRun by Gray Swan1,810 curated attacksRobustness: Prompt injection (Fig 5)Unverified

M10 Dangerous capabilities and risk determinations

5 values

M10 Dangerous capabilities and risk determinations: values in the GPT-6 Astra System Card
ModelEvaluationConditionValueLocationChecked
GPT-6 AstraSandbox BenchSuccessful exploits of 22 targetsNone statedCybersecurity capabilitiesUnverified
GPT-6 AstraFrontierCyberChallenges solvedRun by Irregular226 challengesExternal evaluations for cyber (Irregular)Unverified
GPT-6 AstraPreparedness Framework determinationCybersecurityNone statedPreparedness / Safety overviewUnverified
GPT-6 AstraPreparedness Framework determinationBiological and chemicalNone statedPreparedness 10.1.1Unverified
GPT-6 AstraPreparedness Framework determinationAI self-improvementNone statedPreparedness 10.1.3Unverified

M12 Chain-of-thought monitorability

1 value

M12 Chain-of-thought monitorability: values in the GPT-6 Astra System Card
ModelEvaluationConditionValueLocationChecked
GPT-6 AstraCoT monitorabilityMonitorability vs prior modelsCoT-only, action-only and full-context monitors9 MonitorabilityUnverified

Extraction coverage

What we read of this document, and where each value was read.

Our note Web reader stopped around §9. v0 rows whose source_url is the launch post belong to openai-gpt-6-astra-launch-post, not to this card

Of the 82 values, 78 were read in the document itself and 4 in independent write-ups that quote it.

Values read somewhere other than the document stand in where the document's own section could not be read directly, and are flagged on every value (Methodology §2.2).

Of the 82 values, 2 were read off a figure and 4 are statements in words rather than numbers; they are marked * and left out of charts by default.

All 82 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).