GPT-5.1-Codex-Max
39 values about GPT-5.1-Codex-Max from 4 documents, in 8 metric families. In each family, values from its own document come first.
- Developer
- OpenAI
- Release date
- 18 Nov 2025The date of its first card in the dataset, the GPT-5.1-Codex-Max System Card.
- Availability
- Public
- Its own document
- GPT-5.1-Codex-Max System Card · 18 Nov 2025
- Also reported in
- GPT-5.2 System Card (update to GPT-5 card) · 11 Dec 2025
- GPT-5.2-Codex System Card Addendum · 18 Dec 2025
- GPT-5.3-Codex System Card · 5 Feb 2026
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Values by metric family
Each value as the document printed it. Select a value for its source, its checks and its history. Values from other documents are listed after the model’s own, each marked as the first report of that measure or as a restatement.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M1 · Evaluation awareness 1 value
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5.1-Codex-Max System Card · 18 Nov 2025 | ||
| Evaluation awarenessRelative rate vs GPT-5Run by Apollo Research | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
M2 · Reward hacking 1 value
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5.1-Codex-Max System Card · 18 Nov 2025 | ||
| Falsifying task completionRelative rate vs GPT-5Run by Apollo Research | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
M3 · Sabotage and sandbagging 1 value
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5.1-Codex-Max System Card · 18 Nov 2025 | ||
| SandbaggingRelative rate vs GPT-5Run by Apollo Research | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
M4 · Misalignment audits 3 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5.1-Codex-Max System Card · 18 Nov 2025 | ||
| Covert deceptive behaviorRelative rate vs GPT-5Run by Apollo Research | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
| Destructive action avoidanceAvoidanceRestated later: compare with the later value | GPT-5.1-Codex-Max System Card18 Nov 2025Restated later: compare with the later value | |
| From other documents | ||
| Destructive action avoidanceAvoidanceGPT-5.2-Codex System Card Addendum18 Dec 2025Restated: compare with the earlier value | GPT-5.2-Codex System Card Addendum18 Dec 2025Restated: compare with the earlier value | |
M7 · Harmful compliance and over-refusal 17 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5.1-Codex-Max System Card · 18 Nov 2025 | ||
| Long-form biorisk questionsRefusal rate | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
| Malware refusals (golden set)Refusal rate | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
| Production BenchmarksEmotional relianceRestated later: compare with the later value | GPT-5.1-Codex-Max System Card18 Nov 2025Restated later: compare with the later value | |
| Production BenchmarksExtremism | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
| Production BenchmarksHarassment | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
| Production BenchmarksHate | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
| Production BenchmarksIllicit/non-violent | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
| Production BenchmarksIllicit/violent | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
| Production BenchmarksMental healthRestated later: compare with the later value | GPT-5.1-Codex-Max System Card18 Nov 2025Restated later: compare with the later value | |
| Production BenchmarksPersonal data | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
| Production BenchmarksSelf-harm/instructions | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
| Production BenchmarksSelf-harm/intent | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
| Production BenchmarksSexual | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
| Production BenchmarksSexual/minors | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
| Production BenchmarksViolence | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
| From other documents | ||
| Production BenchmarksEmotional relianceGPT-5.2-Codex System Card Addendum18 Dec 2025Restated: compare with the earlier value | GPT-5.2-Codex System Card Addendum18 Dec 2025Restated: compare with the earlier value | |
| Production BenchmarksMental healthGPT-5.2-Codex System Card Addendum18 Dec 2025Restated: compare with the earlier value | GPT-5.2-Codex System Card Addendum18 Dec 2025Restated: compare with the earlier value | |
M8 · Jailbreak robustness 2 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5.1-Codex-Max System Card · 18 Nov 2025 | ||
| StrongRejectnot_unsafe (aggregate)Restated later: compare with the later value | GPT-5.1-Codex-Max System Card18 Nov 2025Restated later: compare with the later value | |
| From other documents | ||
| StrongRejectnot_unsafe (aggregate)GPT-5.2-Codex System Card Addendum18 Dec 2025Restated: compare with the earlier value | GPT-5.2-Codex System Card Addendum18 Dec 2025Restated: compare with the earlier value | |
M9 · Prompt injection 1 value
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5.1-Codex-Max System Card · 18 Nov 2025 | ||
| Prompt injection (Codex env)Attacks successfully ignored | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
M10 · Dangerous capabilities and risk determinations 13 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5.1-Codex-Max System Card · 18 Nov 2025 | ||
| 50% time horizonTask length at 50% successRun by METR | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
| Cyber challengesEasy: challenges solvedCondition: out of 18 challengesRun by Irregular | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
| Cyber challengesHard: challenges solvedCondition: out of 6 challengesRun by Irregular | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
| Cyber challengesMedium: challenges solvedCondition: out of 17 challengesRun by Irregular | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
| Cyber RangeScenarios passedCondition: of 8 scenarios reported | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
| Preparedness Framework determinationAI self-improvement | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
| Preparedness Framework determinationBiological and chemical | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
| Preparedness Framework determinationCybersecurity | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
| Tacit knowledge and troubleshootingScore | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
| TroubleshootingBenchScoreCondition: refusals counted as successes | GPT-5.1-Codex-Max System Card18 Nov 2025 | |
| From other documents | ||
| Cyber RangeScenarios passedCondition: of 9 scenariosGPT-5.2 System Card (update to GPT-5 card)11 Dec 2025First reported | GPT-5.2 System Card (update to GPT-5 card)11 Dec 2025First reported | |
| OpenAI-Proof Q&APass rateGPT-5.2 System Card (update to GPT-5 card)11 Dec 2025First reported | GPT-5.2 System Card (update to GPT-5 card)11 Dec 2025First reported | |
| Cyber RangeCombined pass rateGPT-5.3-Codex System Card5 Feb 2026First reported | GPT-5.3-Codex System Card5 Feb 2026First reported | |
Restated in later documents
A later document reported values about GPT-5.1-Codex-Max again. Each pair is shown side by side: both values are kept with their own documents, and the later one does not replace the earlier.
| Earlier value | Later value |
|---|---|
| M4 · Misalignment audits | |
| Destructive action avoidanceAvoidance | |
| GPT-5.1-Codex-Max System Card18 Nov 2025 · its own document | GPT-5.2-Codex System Card Addendum18 Dec 2025No reason stated |
| M7 · Harmful compliance and over-refusal | |
| Production BenchmarksEmotional reliance | |
| GPT-5.1-Codex-Max System Card18 Nov 2025 · its own document | GPT-5.2-Codex System Card Addendum18 Dec 2025No reason stated |
| Production BenchmarksMental health | |
| GPT-5.1-Codex-Max System Card18 Nov 2025 · its own document | GPT-5.2-Codex System Card Addendum18 Dec 2025No reason stated |
| M8 · Jailbreak robustness | |
| StrongRejectnot_unsafe (aggregate) | |
| GPT-5.1-Codex-Max System Card18 Nov 2025 · its own document | GPT-5.2-Codex System Card Addendum18 Dec 2025No reason stated |
Risk determinations
The developer's formal decisions about GPT-5.1-Codex-Max under its framework, as printed. Levels from different frameworks do not map onto one another.
| Domain | Level as printed | Framework and document |
|---|---|---|
| Bio/chem | High (treated as) Preparedness Framework GPT-5.1-Codex-Max System Card · 18 Nov 2025 | Preparedness Framework GPT-5.1-Codex-Max System Card · 18 Nov 2025 |
| Cyber | Below High Preparedness Framework GPT-5.1-Codex-Max System Card · 18 Nov 2025 | Preparedness Framework GPT-5.1-Codex-Max System Card · 18 Nov 2025 |
| AI R&D / autonomy | Below High Preparedness Framework GPT-5.1-Codex-Max System Card · 18 Nov 2025 | Preparedness Framework GPT-5.1-Codex-Max System Card · 18 Nov 2025 |