OpenAI o3 and o4-mini System Card
A system card by OpenAI about o3 and o4-mini, published 16 Apr 2025. We recorded 51 values from it.
- Developer
- OpenAI
- Type
- System card
- Published
- 16 Apr 2025
- Archived copy
- No archived copy yet
- Changelog
- Not known
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Versions
One version is on record: the copy we retrieved on 26 Sep 2026. We know of no other.
26 Sep 2026
Date retrieved
Copy retrieved 26 Sep 2026
No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).
Revisions
No revisions are recorded for this document. We know of only one version of it.
Values
Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. 3 of the 51 have been blind-verified: a second reader found the same value without seeing ours.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M2 Reward hacking
2 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| o3 | RE-Bench Optimize a KernelRuns tampering with scoring functionRun by METR | out of 24 experiments | Third-party: METR | Verified | |
| o3 | Reward hacking (all tasks)Share of task attempts with reward hackingRun by METR | None stated | Third-party: METR | Unverified |
M3 Sabotage and sandbagging
4 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| o3 | SandbaggingIn-context sandbagging observedRun by Apollo Research | including without an in-context goal | Apollo Research: in-context scheming | Unverified | |
| o3 | AI R&D sabotageAverage sabotage scoreRun by Apollo Research | None stated | Apollo Research: Sabotage | Verified | |
| o3-mini | AI R&D sabotageAverage sabotage scoreRun by Apollo Research | None stated | Apollo Research: Sabotage | Unverified | |
| o4-mini | AI R&D sabotageAverage sabotage scoreRun by Apollo Research | None stated | Apollo Research: Sabotage | Unverified |
M5 Honesty and hallucination
6 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| o1 | PersonQAHallucination rate | no browsing | Table 4 | Unverified | |
| o1 | SimpleQAHallucination rate | no browsing | Table 4 | Unverified | |
| o3 | PersonQAHallucination rate | no browsing | Table 4 | Unverified | |
| o3 | SimpleQAHallucination rate | no browsing | Table 4 | Unverified | |
| o4-mini | PersonQAHallucination rate | no browsing | Table 4 | Unverified | |
| o4-mini | SimpleQAHallucination rate | no browsing | Table 4 | Unverified |
M7 Harmful compliance and over-refusal
6 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| o1 | Standard refusal evalNot overrefuse (aggregate) | None stated | Table 1 | Unverified | |
| o3 | Standard refusal evalNot overrefuse (aggregate) | None stated | Table 1 | Unverified | |
| o4-mini | Standard refusal evalNot overrefuse (aggregate) | None stated | Table 1 | Unverified | |
| o1 | Challenging refusal evalNot unsafe (aggregate) | None stated | Table 2 | Unverified | |
| o3 | Challenging refusal evalNot unsafe (aggregate) | None stated | Table 2 | Unverified | |
| o4-mini | Challenging refusal evalNot unsafe (aggregate) | None stated | Table 2 | Unverified |
M8 Jailbreak robustness
6 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| o1 | Human sourced jailbreaksnot_unsafe | None stated | Table 3 | Unverified | |
| o1 | StrongRejectnot_unsafe (aggregate) | None stated | Table 3 | Unverified | |
| o3 | Human sourced jailbreaksnot_unsafe | None stated | Table 3 | Unverified | |
| o3 | StrongRejectnot_unsafe (aggregate) | None stated | Table 3 | Unverified | |
| o4-mini | Human sourced jailbreaksnot_unsafe | None stated | Table 3 | Unverified | |
| o4-mini | StrongRejectnot_unsafe (aggregate) | None stated | Table 3 | Unverified |
M9 Prompt injection
3 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| o1 | Instruction hierarchySystem<>user conflict | None stated | Table 9 | Unverified | |
| o3 | Instruction hierarchySystem<>user conflict | None stated | Table 9 | Unverified | |
| o4-mini | Instruction hierarchySystem<>user conflict | None stated | Table 9 | Unverified |
M10 Dangerous capabilities and risk determinations
24 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| o1 | PaperBenchReplication score | None stated | AI Self-improvement | Unverified | |
| o3 | OpenAI PRsPass rate | None stated | AI Self-improvement | Verified | |
| o3 | SWE-bench VerifiedPass rate | helpful-only | AI Self-improvement | Unverified | |
| o3 | SWE-LancerDollars earned | helpful-only; with browsing | AI Self-improvement | Unverified | |
| o3 | SWE-LancerDollars earned | no browsing | AI Self-improvement | Unverified | |
| o3 | SWE-LancerIC SWE accuracy | helpful-only; with browsing | AI Self-improvement | Unverified | |
| o4-mini | OpenAI PRsPass rate | None stated | AI Self-improvement | Unverified | |
| o4-mini | PaperBenchReplication score | no browsing | AI Self-improvement | Unverified | |
| o3 | Capture the FlagCollegiate | no browsing; pass@12 | Cybersecurity, Fig 7 + text | Unverified | |
| o3 | Capture the FlagHigh school | no browsing; pass@12 | Cybersecurity, Fig 7 + text | Unverified | |
| o3 | Capture the FlagProfessional | no browsing; pass@12 | Cybersecurity, Fig 7 + text | Unverified | |
| o4-mini | Capture the FlagCollegiate | no browsing; pass@12 | Cybersecurity, Fig 7 + text | Unverified | |
| o4-mini | Capture the FlagHigh school | no browsing; pass@12 | Cybersecurity, Fig 7 + text | Unverified | |
| o4-mini | Capture the FlagProfessional | no browsing; pass@12 | Cybersecurity, Fig 7 + text | Unverified | |
| o3 | Cyber RangeScenarios solved unaided | without solver code; of 2 scenarios | Cybersecurity, Fig 8 | Unverified | |
| o4-mini | Cyber RangeScenarios solved unaided | without solver code; of 2 scenarios | Cybersecurity, Fig 8 | Unverified | |
| o3 | Preparedness Framework determinationAI self-improvement | None stated | Preparedness | Unverified | |
| o3 | Preparedness Framework determinationBiological and chemical | None stated | Preparedness | Unverified | |
| o3 | Preparedness Framework determinationCybersecurity | None stated | Preparedness | Unverified | |
| o4-mini | Preparedness Framework determinationAI self-improvement | None stated | Preparedness | Unverified | |
| o4-mini | Preparedness Framework determinationBiological and chemical | None stated | Preparedness | Unverified | |
| o4-mini | Preparedness Framework determinationCybersecurity | None stated | Preparedness | Unverified | |
| o3 | 50% time horizonTask length at 50% successRun by METR | None stated | Third-party: METR | Unverified | |
| o4-mini | 50% time horizonTask length at 50% successRun by METR | None stated | Third-party: METR | Unverified |
Extraction coverage
What we read of this document, and where each value was read.
Our note Most numbers stated in text
All 51 values were read in the document itself.
Of the 51 values, 4 are statements in words rather than numbers; they are marked * and left out of charts by default.
All 51 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).