gpt-5-thinking
68 values about gpt-5-thinking from 4 documents, in 10 metric families. In each family, values from its own document come first.
- Developer
- OpenAI
- Release date
- 7 Aug 2025The date of its first card in the dataset, the GPT-5 System Card.
- Availability
- Public
- Also printed as
- GPT-5 (thinking)
- Its own document
- GPT-5 System Card · 7 Aug 2025
- Also reported in
- GPT-5.1 Instant and Thinking System Card Addendum · 12 Nov 2025
- GPT-5.1-Codex-Max System Card · 18 Nov 2025
- GPT-5.2 System Card (update to GPT-5 card) · 11 Dec 2025
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Values by metric family
Each value as the document printed it. Select a value for its source, its checks and its history. Values from other documents are listed after the model’s own, each marked as the first report of that measure or as a restatement.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M1 · Evaluation awareness 4 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5 System Card · 7 Aug 2025 | ||
| Evaluation awarenessVerbalized eval awareness, broadCondition: all samplesRun by Apollo Research | GPT-5 System Card7 Aug 2025 | |
| Evaluation awarenessVerbalized eval awareness, broadCondition: scheming samplesRun by Apollo Research | GPT-5 System Card7 Aug 2025 | |
| Evaluation awarenessVerbalized eval awareness, strictCondition: all samplesRun by Apollo Research | GPT-5 System Card7 Aug 2025 | |
| Evaluation awarenessVerbalized eval awareness, strictCondition: scheming samplesRun by Apollo Research | GPT-5 System Card7 Aug 2025 | |
M3 · Sabotage and sandbagging 1 value
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5 System Card · 7 Aug 2025 | ||
| Sandbagging reviewClear evidence of sandbaggingRun by METR | GPT-5 System Card7 Aug 2025 | |
M4 · Misalignment audits 3 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5 System Card · 7 Aug 2025 | ||
| Covert deceptive behaviorShare of trajectories with deceptive actionCondition: Apollo scheming suiteRun by Apollo Research | GPT-5 System Card7 Aug 2025 | |
| Log falsificationCovert action rateCondition: strong goal nudgeRun by Apollo Research | GPT-5 System Card7 Aug 2025 | |
| Log falsificationCovert action rateCondition: weak goal nudgeRun by Apollo Research | GPT-5 System Card7 Aug 2025 | |
M5 · Honesty and hallucination 8 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5 System Card · 7 Aug 2025 | ||
| AbstentionBenchRecall (appropriate abstention) | GPT-5 System Card7 Aug 2025 | |
| Deception evalAgentic coding (impossible tasks) | GPT-5 System Card7 Aug 2025 | |
| Deception evalBrowsing broken tools | GPT-5 System Card7 Aug 2025 | |
| Deception evalCharXiv missing image | GPT-5 System Card7 Aug 2025 | |
| Production CoT deception monitorShare of responses flagged deceptiveCondition: representative production traffic | GPT-5 System Card7 Aug 2025 | |
| Production-traffic factualityClaim-level hallucination rate reduction vs o3Condition: with browsing; production prompts | GPT-5 System Card7 Aug 2025 | |
| Production-traffic factualityReduction in responses with 1+ major error vs o3Condition: with browsing; production prompts | GPT-5 System Card7 Aug 2025 | |
| SimpleQAHallucination rateCondition: no browsing | GPT-5 System Card7 Aug 2025 | |
M6 · Sycophancy 1 value
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5 System Card · 7 Aug 2025 | ||
| Sycophancy offline evalSycophancy scoreCondition: offline | GPT-5 System Card7 Aug 2025 | |
M7 · Harmful compliance and over-refusal 26 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5 System Card · 7 Aug 2025 | ||
| Production BenchmarksExtremism | GPT-5 System Card7 Aug 2025 | |
| Production BenchmarksHarassment/threatening | GPT-5 System Card7 Aug 2025 | |
| Production BenchmarksHate/threatening | GPT-5 System Card7 Aug 2025 | |
| Production BenchmarksIllicit/non-violent | GPT-5 System Card7 Aug 2025 | |
| Production BenchmarksIllicit/violent | GPT-5 System Card7 Aug 2025 | |
| Production BenchmarksNon-violent hate | GPT-5 System Card7 Aug 2025 | |
| Production BenchmarksPersonal data | GPT-5 System Card7 Aug 2025 | |
| Production BenchmarksSelf-harm/instructions | GPT-5 System Card7 Aug 2025 | |
| Production BenchmarksSelf-harm/intent | GPT-5 System Card7 Aug 2025 | |
| Production BenchmarksSexual/exploitative | GPT-5 System Card7 Aug 2025 | |
| Production BenchmarksSexual/minors | GPT-5 System Card7 Aug 2025 | |
| From other documents | ||
| Production BenchmarksEmotional relianceGPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | GPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | |
| Production BenchmarksExtremismGPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | GPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | |
| Production BenchmarksHarassmentGPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | GPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | |
| Production BenchmarksHateGPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | GPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | |
| Production BenchmarksIllicit/non-violentGPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | GPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | |
| Production BenchmarksIllicit/violentGPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | GPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | |
| Production BenchmarksMental healthGPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | GPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | |
| Production BenchmarksPersonal dataGPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | GPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | |
| Production BenchmarksSelf-harm/instructionsGPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | GPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | |
| Production BenchmarksSelf-harm/intentGPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | GPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | |
| Production BenchmarksSexualGPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | GPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | |
| Production BenchmarksSexual/minorsGPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | GPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | |
| Production BenchmarksViolenceGPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | GPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | |
| Cyber safetyProduction dataGPT-5.2 System Card (update to GPT-5 card)11 Dec 2025First reported | GPT-5.2 System Card (update to GPT-5 card)11 Dec 2025First reported | |
| Cyber safetySynthetic dataGPT-5.2 System Card (update to GPT-5 card)11 Dec 2025First reported | GPT-5.2 System Card (update to GPT-5 card)11 Dec 2025First reported | |
M8 · Jailbreak robustness 5 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5 System Card · 7 Aug 2025 | ||
| StrongRejectAbuse/disinformation/hate | GPT-5 System Card7 Aug 2025 | |
| StrongRejectIllicit/non-violent crime | GPT-5 System Card7 Aug 2025 | |
| StrongRejectSexual content | GPT-5 System Card7 Aug 2025 | |
| StrongRejectViolence | GPT-5 System Card7 Aug 2025 | |
| From other documents | ||
| StrongRejectnot_unsafe (aggregate)GPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | GPT-5.1 Instant and Thinking System Card Addendum12 Nov 2025First reported | |
M9 · Prompt injection 3 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5 System Card · 7 Aug 2025 | ||
| Prompt injectionBrowsing | GPT-5 System Card7 Aug 2025 | |
| Prompt injectionCoding | GPT-5 System Card7 Aug 2025 | |
| Prompt injectionTool calling | GPT-5 System Card7 Aug 2025 | |
M10 · Dangerous capabilities and risk determinations 15 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5 System Card · 7 Aug 2025 | ||
| 50% time horizonTask length at 50% successRun by METR | GPT-5 System Card7 Aug 2025 | |
| Cyber challengesEasy: challenges solvedCondition: out of 18 challengesRun by Irregular | GPT-5 System Card7 Aug 2025 | |
| Cyber challengesEvasion: average success rateRun by Irregular | GPT-5 System Card7 Aug 2025 | |
| Cyber challengesHard: challenges solvedCondition: out of 4 challengesRun by Irregular | GPT-5 System Card7 Aug 2025 | |
| Cyber challengesMedium: challenges solvedCondition: out of 14 challengesRun by Irregular | GPT-5 System Card7 Aug 2025 | |
| Cyber challengesNetwork attack simulation: average success rateRun by Irregular | GPT-5 System Card7 Aug 2025 | |
| Cyber challengesVulnerability discovery and exploitation: average success rateRun by Irregular | GPT-5 System Card7 Aug 2025 | |
| OpenAI-Proof Q&APass rate | GPT-5 System Card7 Aug 2025 | |
| Preparedness Framework determinationAI self-improvement | GPT-5 System Card7 Aug 2025 | |
| Preparedness Framework determinationBiological and chemical | GPT-5 System Card7 Aug 2025 | |
| Preparedness Framework determinationCybersecurity | GPT-5 System Card7 Aug 2025 | |
| TroubleshootingBenchDifference vs o3 | GPT-5 System Card7 Aug 2025 | |
| From other documents | ||
| Cyber challengesEasy: challenges solvedCondition: out of 18 challengesRun by IrregularGPT-5.1-Codex-Max System Card18 Nov 2025First reported | GPT-5.1-Codex-Max System Card18 Nov 2025First reported | |
| Cyber challengesHard: challenges solvedCondition: out of 6 challengesRun by IrregularGPT-5.1-Codex-Max System Card18 Nov 2025First reported | GPT-5.1-Codex-Max System Card18 Nov 2025First reported | |
| Cyber challengesMedium: challenges solvedCondition: out of 17 challengesRun by IrregularGPT-5.1-Codex-Max System Card18 Nov 2025First reported | GPT-5.1-Codex-Max System Card18 Nov 2025First reported | |
M12 · Chain-of-thought monitorability 2 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGPT-5 System Card · 7 Aug 2025 | ||
| CoT controllabilityAbility to control CoT content | GPT-5 System Card7 Aug 2025 | |
| CoT monitorabilityMonitorability vs o3 | GPT-5 System Card7 Aug 2025 | |
Risk determinations
The developer's formal decisions about gpt-5-thinking under its framework, as printed. Levels from different frameworks do not map onto one another.
| Domain | Level as printed | Framework and document |
|---|---|---|
| Bio/chem | High (precautionary) Preparedness Framework GPT-5 System Card · 7 Aug 2025 | Preparedness Framework GPT-5 System Card · 7 Aug 2025 |
| Cyber | Below High Preparedness Framework GPT-5 System Card · 7 Aug 2025 | Preparedness Framework GPT-5 System Card · 7 Aug 2025 |
| AI R&D / autonomy | Below High Preparedness Framework GPT-5 System Card · 7 Aug 2025 | Preparedness Framework GPT-5 System Card · 7 Aug 2025 |