Versions

One version is on record: the copy we retrieved on 26 Sep 2026. We know of no other.

  1. 26 Sep 2026

    Date retrieved

    Copy retrieved 26 Sep 2026

    Date from
    the date we retrieved it; the copy states no version date
    Copy
    Copy retrieved on 26 Sep 2026
    Values
    51 values recorded from this version

No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).

Revisions

No revisions are recorded for this document. We know of only one version of it.

Values

Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. 3 of the 51 have been blind-verified: a second reader found the same value without seeing ours.

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

M2 Reward hacking

2 values

M2 Reward hacking: values in the OpenAI o3 and o4-mini System Card
ModelEvaluationConditionValueLocationChecked
o3RE-Bench Optimize a KernelRuns tampering with scoring functionRun by METRout of 24 experimentsThird-party: METRVerified
o3Reward hacking (all tasks)Share of task attempts with reward hackingRun by METRNone statedThird-party: METRUnverified

M3 Sabotage and sandbagging

4 values

M3 Sabotage and sandbagging: values in the OpenAI o3 and o4-mini System Card
ModelEvaluationConditionValueLocationChecked
o3SandbaggingIn-context sandbagging observedRun by Apollo Researchincluding without an in-context goalApollo Research: in-context schemingUnverified
o3AI R&D sabotageAverage sabotage scoreRun by Apollo ResearchNone statedApollo Research: SabotageVerified
o3-miniAI R&D sabotageAverage sabotage scoreRun by Apollo ResearchNone statedApollo Research: SabotageUnverified
o4-miniAI R&D sabotageAverage sabotage scoreRun by Apollo ResearchNone statedApollo Research: SabotageUnverified

M5 Honesty and hallucination

6 values

M5 Honesty and hallucination: values in the OpenAI o3 and o4-mini System Card
ModelEvaluationConditionValueLocationChecked
o1PersonQAHallucination rateno browsingTable 4Unverified
o1SimpleQAHallucination rateno browsingTable 4Unverified
o3PersonQAHallucination rateno browsingTable 4Unverified
o3SimpleQAHallucination rateno browsingTable 4Unverified
o4-miniPersonQAHallucination rateno browsingTable 4Unverified
o4-miniSimpleQAHallucination rateno browsingTable 4Unverified

M7 Harmful compliance and over-refusal

6 values

M7 Harmful compliance and over-refusal: values in the OpenAI o3 and o4-mini System Card
ModelEvaluationConditionValueLocationChecked
o1Standard refusal evalNot overrefuse (aggregate)None statedTable 1Unverified
o3Standard refusal evalNot overrefuse (aggregate)None statedTable 1Unverified
o4-miniStandard refusal evalNot overrefuse (aggregate)None statedTable 1Unverified
o1Challenging refusal evalNot unsafe (aggregate)None statedTable 2Unverified
o3Challenging refusal evalNot unsafe (aggregate)None statedTable 2Unverified
o4-miniChallenging refusal evalNot unsafe (aggregate)None statedTable 2Unverified

M8 Jailbreak robustness

6 values

M8 Jailbreak robustness: values in the OpenAI o3 and o4-mini System Card
ModelEvaluationConditionValueLocationChecked
o1Human sourced jailbreaksnot_unsafeNone statedTable 3Unverified
o1StrongRejectnot_unsafe (aggregate)None statedTable 3Unverified
o3Human sourced jailbreaksnot_unsafeNone statedTable 3Unverified
o3StrongRejectnot_unsafe (aggregate)None statedTable 3Unverified
o4-miniHuman sourced jailbreaksnot_unsafeNone statedTable 3Unverified
o4-miniStrongRejectnot_unsafe (aggregate)None statedTable 3Unverified

M9 Prompt injection

3 values

M9 Prompt injection: values in the OpenAI o3 and o4-mini System Card
ModelEvaluationConditionValueLocationChecked
o1Instruction hierarchySystem<>user conflictNone statedTable 9Unverified
o3Instruction hierarchySystem<>user conflictNone statedTable 9Unverified
o4-miniInstruction hierarchySystem<>user conflictNone statedTable 9Unverified

M10 Dangerous capabilities and risk determinations

24 values

M10 Dangerous capabilities and risk determinations: values in the OpenAI o3 and o4-mini System Card
ModelEvaluationConditionValueLocationChecked
o1PaperBenchReplication scoreNone statedAI Self-improvementUnverified
o3OpenAI PRsPass rateNone statedAI Self-improvementVerified
o3SWE-bench VerifiedPass ratehelpful-onlyAI Self-improvementUnverified
o3SWE-LancerDollars earnedhelpful-only; with browsingAI Self-improvementUnverified
o3SWE-LancerDollars earnedno browsingAI Self-improvementUnverified
o3SWE-LancerIC SWE accuracyhelpful-only; with browsingAI Self-improvementUnverified
o4-miniOpenAI PRsPass rateNone statedAI Self-improvementUnverified
o4-miniPaperBenchReplication scoreno browsingAI Self-improvementUnverified
o3Capture the FlagCollegiateno browsing; pass@12Cybersecurity, Fig 7 + textUnverified
o3Capture the FlagHigh schoolno browsing; pass@12Cybersecurity, Fig 7 + textUnverified
o3Capture the FlagProfessionalno browsing; pass@12Cybersecurity, Fig 7 + textUnverified
o4-miniCapture the FlagCollegiateno browsing; pass@12Cybersecurity, Fig 7 + textUnverified
o4-miniCapture the FlagHigh schoolno browsing; pass@12Cybersecurity, Fig 7 + textUnverified
o4-miniCapture the FlagProfessionalno browsing; pass@12Cybersecurity, Fig 7 + textUnverified
o3Cyber RangeScenarios solved unaidedwithout solver code; of 2 scenariosCybersecurity, Fig 8Unverified
o4-miniCyber RangeScenarios solved unaidedwithout solver code; of 2 scenariosCybersecurity, Fig 8Unverified
o3Preparedness Framework determinationAI self-improvementNone statedPreparednessUnverified
o3Preparedness Framework determinationBiological and chemicalNone statedPreparednessUnverified
o3Preparedness Framework determinationCybersecurityNone statedPreparednessUnverified
o4-miniPreparedness Framework determinationAI self-improvementNone statedPreparednessUnverified
o4-miniPreparedness Framework determinationBiological and chemicalNone statedPreparednessUnverified
o4-miniPreparedness Framework determinationCybersecurityNone statedPreparednessUnverified
o350% time horizonTask length at 50% successRun by METRNone statedThird-party: METRUnverified
o4-mini50% time horizonTask length at 50% successRun by METRNone statedThird-party: METRUnverified

Extraction coverage

What we read of this document, and where each value was read.

Our note Most numbers stated in text

All 51 values were read in the document itself.

Of the 51 values, 4 are statements in words rather than numbers; they are marked * and left out of charts by default.

All 51 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).