Versions

One version is on record: the copy we retrieved on 26 Sep 2026. We know of no other.

  1. 26 Sep 2026

    Date retrieved

    Copy retrieved 26 Sep 2026

    Date from
    the date we retrieved it; the copy states no version date
    Copy
    Copy retrieved on 26 Sep 2026
    Values
    15 values recorded from this version

No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).

Revisions

No revisions are recorded for this document. We know of only one version of it.

Values

Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. 1 of the 15 has been blind-verified: a second reader found the same value without seeing ours.

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

M1 Evaluation awareness

1 value

M1 Evaluation awareness: values in the Gemini 2.5 Deep Think Model Card
ModelEvaluationConditionValueLocationChecked
Gemini 2.5 Deep ThinkSituational awareness challengesChallenges solvedNone statedFrontier Safety: deceptive alignmentUnverified

M3 Sabotage and sandbagging

1 value

M3 Sabotage and sandbagging: values in the Gemini 2.5 Deep Think Model Card
ModelEvaluationConditionValueLocationChecked
Gemini 2.5 Deep ThinkStealth challengesChallenges solvedNone statedFrontier Safety: deceptive alignmentUnverified

M4 Misalignment audits

1 value

M4 Misalignment audits: values in the Gemini 2.5 Deep Think Model Card
ModelEvaluationConditionValueLocationChecked
Gemini 2.5 Deep ThinkFrontier Safety Framework determination (misalignment)Instrumental Reasoning Level 1/2None statedFrontier Safety tableUnverified

M7 Harmful compliance and over-refusal

5 values

M7 Harmful compliance and over-refusal: values in the Gemini 2.5 Deep Think Model Card
ModelEvaluationConditionValueLocationChecked
Gemini 2.5 Deep ThinkImage to Text SafetyImage-to-text policy-violation rate deltavs Gemini 2.5 ProSafety evaluations tableUnverified
Gemini 2.5 Deep ThinkInstruction FollowingSafe instruction following deltavs Gemini 2.5 ProSafety evaluations tableUnverified
Gemini 2.5 Deep ThinkMultilingual SafetyMultilingual policy-violation rate deltavs Gemini 2.5 ProSafety evaluations tableUnverified
Gemini 2.5 Deep ThinkText to Text SafetyPolicy-violation rate deltavs Gemini 2.5 ProSafety evaluations tableUnverified
Gemini 2.5 Deep ThinkToneObjective tone of refusals deltavs Gemini 2.5 ProSafety evaluations tableVerified

M10 Dangerous capabilities and risk determinations

7 values

M10 Dangerous capabilities and risk determinations: values in the Gemini 2.5 Deep Think Model Card
ModelEvaluationConditionValueLocationChecked
Gemini 2.5 Deep ThinkFrontier Safety Framework determinationCBRN Uplift Level 1None statedFrontier Safety tableUnverified
Gemini 2.5 Deep ThinkFrontier Safety Framework determinationCyber Autonomy Level 1None statedFrontier Safety tableUnverified
Gemini 2.5 Deep ThinkFrontier Safety Framework determinationCyber Uplift Level 1None statedFrontier Safety tableUnverified
Gemini 2.5 Deep ThinkFrontier Safety Framework determinationML R&D Autonomy/Uplift Level 1None statedFrontier Safety tableUnverified
Gemini 2.5 Deep ThinkCyber autonomous offense suiteHard CTF challenges solvedNone statedFrontier Safety: cyberUnverified
Gemini 2.5 Deep ThinkCyber key skills benchmarkHard challenges solvedNone statedFrontier Safety: cyberUnverified
Gemini 2.5 Deep ThinkRE-BenchAverage normalised scoreNone statedFrontier Safety: ML R&DUnverified

Extraction coverage

What we read of this document, and where each value was read.

Our note First model to reach CBRN early-warning alert threshold

All 15 values were read in the document itself.

All 15 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).