o3
50 values about o3 from 2 documents, in 8 metric families. In each family, values from its own document come first.
- Developer
- OpenAI
- Release date
- 16 Apr 2025The date of its first card in the dataset, the OpenAI o3 and o4-mini System Card.
- Availability
- Public
- Its own document
- OpenAI o3 and o4-mini System Card · 16 Apr 2025
- Also reported in
- GPT-5 System Card · 7 Aug 2025
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Values by metric family
Each value as the document printed it. Select a value for its source, its checks and its history. Values from other documents are listed after the model’s own, each marked as the first report of that measure or as a restatement.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M2 · Reward hacking 2 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentOpenAI o3 and o4-mini System Card · 16 Apr 2025 | ||
| RE-Bench Optimize a KernelRuns tampering with scoring functionCondition: out of 24 experimentsRun by METR | OpenAI o3 and o4-mini System Card16 Apr 2025 | |
| Reward hacking (all tasks)Share of task attempts with reward hackingRun by METR | OpenAI o3 and o4-mini System Card16 Apr 2025 | |
M3 · Sabotage and sandbagging 2 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentOpenAI o3 and o4-mini System Card · 16 Apr 2025 | ||
| AI R&D sabotageAverage sabotage scoreRun by Apollo Research | OpenAI o3 and o4-mini System Card16 Apr 2025 | |
| SandbaggingIn-context sandbagging observedCondition: including without an in-context goalRun by Apollo Research | OpenAI o3 and o4-mini System Card16 Apr 2025 | |
M4 · Misalignment audits 1 value
| Evaluation and metric | Value | Document |
|---|---|---|
| From other documents | ||
| Covert deceptive behaviorShare of trajectories with deceptive actionCondition: Apollo scheming suiteRun by Apollo ResearchGPT-5 System Card7 Aug 2025First reported | GPT-5 System Card7 Aug 2025First reported | |
M5 · Honesty and hallucination 8 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentOpenAI o3 and o4-mini System Card · 16 Apr 2025 | ||
| PersonQAHallucination rateCondition: no browsing | OpenAI o3 and o4-mini System Card16 Apr 2025 | |
| SimpleQAHallucination rateCondition: no browsingRestated later: compare with the later value | OpenAI o3 and o4-mini System Card16 Apr 2025Restated later: compare with the later value | |
| From other documents | ||
| AbstentionBenchRecall (appropriate abstention)GPT-5 System Card7 Aug 2025First reported | GPT-5 System Card7 Aug 2025First reported | |
| Deception evalAgentic coding (impossible tasks)GPT-5 System Card7 Aug 2025First reported | GPT-5 System Card7 Aug 2025First reported | |
| Deception evalBrowsing broken toolsGPT-5 System Card7 Aug 2025First reported | GPT-5 System Card7 Aug 2025First reported | |
| Deception evalCharXiv missing imageGPT-5 System Card7 Aug 2025First reported | GPT-5 System Card7 Aug 2025First reported | |
| Production CoT deception monitorShare of responses flagged deceptiveCondition: representative production trafficGPT-5 System Card7 Aug 2025First reported | GPT-5 System Card7 Aug 2025First reported | |
| SimpleQAHallucination rateCondition: no browsingGPT-5 System Card7 Aug 2025Restated: compare with the earlier value | GPT-5 System Card7 Aug 2025Restated: compare with the earlier value | |
M7 · Harmful compliance and over-refusal 13 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentOpenAI o3 and o4-mini System Card · 16 Apr 2025 | ||
| Challenging refusal evalNot unsafe (aggregate) | OpenAI o3 and o4-mini System Card16 Apr 2025 | |
| Standard refusal evalNot overrefuse (aggregate) | OpenAI o3 and o4-mini System Card16 Apr 2025 | |
| From other documents | ||
| Production BenchmarksExtremismGPT-5 System Card7 Aug 2025First reported | GPT-5 System Card7 Aug 2025First reported | |
| Production BenchmarksHarassment/threateningGPT-5 System Card7 Aug 2025First reported | GPT-5 System Card7 Aug 2025First reported | |
| Production BenchmarksHate/threateningGPT-5 System Card7 Aug 2025First reported | GPT-5 System Card7 Aug 2025First reported | |
| Production BenchmarksIllicit/non-violentGPT-5 System Card7 Aug 2025First reported | GPT-5 System Card7 Aug 2025First reported | |
| Production BenchmarksIllicit/violentGPT-5 System Card7 Aug 2025First reported | GPT-5 System Card7 Aug 2025First reported | |
| Production BenchmarksNon-violent hateGPT-5 System Card7 Aug 2025First reported | GPT-5 System Card7 Aug 2025First reported | |
| Production BenchmarksPersonal dataGPT-5 System Card7 Aug 2025First reported | GPT-5 System Card7 Aug 2025First reported | |
| Production BenchmarksSelf-harm/instructionsGPT-5 System Card7 Aug 2025First reported | GPT-5 System Card7 Aug 2025First reported | |
| Production BenchmarksSelf-harm/intentGPT-5 System Card7 Aug 2025First reported | GPT-5 System Card7 Aug 2025First reported | |
| Production BenchmarksSexual/exploitativeGPT-5 System Card7 Aug 2025First reported | GPT-5 System Card7 Aug 2025First reported | |
| Production BenchmarksSexual/minorsGPT-5 System Card7 Aug 2025First reported | GPT-5 System Card7 Aug 2025First reported | |
M8 · Jailbreak robustness 6 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentOpenAI o3 and o4-mini System Card · 16 Apr 2025 | ||
| Human sourced jailbreaksnot_unsafe | OpenAI o3 and o4-mini System Card16 Apr 2025 | |
| StrongRejectnot_unsafe (aggregate) | OpenAI o3 and o4-mini System Card16 Apr 2025 | |
| From other documents | ||
| StrongRejectAbuse/disinformation/hateGPT-5 System Card7 Aug 2025First reported | GPT-5 System Card7 Aug 2025First reported | |
| StrongRejectIllicit/non-violent crimeGPT-5 System Card7 Aug 2025First reported | GPT-5 System Card7 Aug 2025First reported | |
| StrongRejectSexual contentGPT-5 System Card7 Aug 2025First reported | GPT-5 System Card7 Aug 2025First reported | |
| StrongRejectViolenceGPT-5 System Card7 Aug 2025First reported | GPT-5 System Card7 Aug 2025First reported | |
M9 · Prompt injection 4 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentOpenAI o3 and o4-mini System Card · 16 Apr 2025 | ||
| Instruction hierarchySystem<>user conflict | OpenAI o3 and o4-mini System Card16 Apr 2025 | |
| From other documents | ||
| Prompt injectionBrowsingGPT-5 System Card7 Aug 2025First reported | GPT-5 System Card7 Aug 2025First reported | |
| Prompt injectionCodingGPT-5 System Card7 Aug 2025First reported | GPT-5 System Card7 Aug 2025First reported | |
| Prompt injectionTool callingGPT-5 System Card7 Aug 2025First reported | GPT-5 System Card7 Aug 2025First reported | |
M10 · Dangerous capabilities and risk determinations 14 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentOpenAI o3 and o4-mini System Card · 16 Apr 2025 | ||
| 50% time horizonTask length at 50% successRun by METRRestated later: compare with the later value | OpenAI o3 and o4-mini System Card16 Apr 2025Restated later: compare with the later value | |
| Capture the FlagCollegiateCondition: no browsing; pass@12 | OpenAI o3 and o4-mini System Card16 Apr 2025 | |
| Capture the FlagHigh schoolCondition: no browsing; pass@12 | OpenAI o3 and o4-mini System Card16 Apr 2025 | |
| Capture the FlagProfessionalCondition: no browsing; pass@12 | OpenAI o3 and o4-mini System Card16 Apr 2025 | |
| Cyber RangeScenarios solved unaidedCondition: without solver code; of 2 scenarios | OpenAI o3 and o4-mini System Card16 Apr 2025 | |
| OpenAI PRsPass rate | OpenAI o3 and o4-mini System Card16 Apr 2025 | |
| Preparedness Framework determinationAI self-improvement | OpenAI o3 and o4-mini System Card16 Apr 2025 | |
| Preparedness Framework determinationBiological and chemical | OpenAI o3 and o4-mini System Card16 Apr 2025 | |
| Preparedness Framework determinationCybersecurity | OpenAI o3 and o4-mini System Card16 Apr 2025 | |
| SWE-bench VerifiedPass rateCondition: helpful-only | OpenAI o3 and o4-mini System Card16 Apr 2025 | |
| SWE-LancerDollars earnedCondition: helpful-only; with browsing | OpenAI o3 and o4-mini System Card16 Apr 2025 | |
| SWE-LancerDollars earnedCondition: no browsing | OpenAI o3 and o4-mini System Card16 Apr 2025 | |
| SWE-LancerIC SWE accuracyCondition: helpful-only; with browsing | OpenAI o3 and o4-mini System Card16 Apr 2025 | |
| From other documents | ||
| 50% time horizonTask length at 50% successRun by METRGPT-5 System Card7 Aug 2025Restated: compare with the earlier value | GPT-5 System Card7 Aug 2025Restated: compare with the earlier value | |
Restated in later documents
A later document reported values about o3 again. Each pair is shown side by side: both values are kept with their own documents, and the later one does not replace the earlier.
| Earlier value | Later value |
|---|---|
| M5 · Honesty and hallucination | |
| SimpleQAHallucination rateCondition: no browsing | |
| OpenAI o3 and o4-mini System Card16 Apr 2025 · its own document | GPT-5 System Card7 Aug 2025No reason stated |
| M10 · Dangerous capabilities and risk determinations | |
| 50% time horizonTask length at 50% successRun by METR | |
| OpenAI o3 and o4-mini System Card16 Apr 2025 · its own document | GPT-5 System Card7 Aug 2025No reason stated |
Risk determinations
The developer's formal decisions about o3 under its framework, as printed. Levels from different frameworks do not map onto one another.
| Domain | Level as printed | Framework and document |
|---|---|---|
| Bio/chem | Below High Preparedness Framework OpenAI o3 and o4-mini System Card · 16 Apr 2025 | Preparedness Framework OpenAI o3 and o4-mini System Card · 16 Apr 2025 |
| Cyber | Below High Preparedness Framework OpenAI o3 and o4-mini System Card · 16 Apr 2025 | Preparedness Framework OpenAI o3 and o4-mini System Card · 16 Apr 2025 |
| AI R&D / autonomy | Below High Preparedness Framework OpenAI o3 and o4-mini System Card · 16 Apr 2025 | Preparedness Framework OpenAI o3 and o4-mini System Card · 16 Apr 2025 |