Versions

One version is on record: the copy we retrieved on 26 Sep 2026. We know of no other.

  1. 26 Sep 2026

    Date retrieved

    Copy retrieved 26 Sep 2026

    Date from
    the date we retrieved it; the copy states no version date
    Copy
    Copy retrieved on 26 Sep 2026
    Values
    21 values recorded from this version

No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).

Revisions

No revisions are recorded for this document. We know of only one version of it.

Values

Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. None has been blind-verified yet.

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

M1 Evaluation awareness

1 value

M1 Evaluation awareness: values in the Gemini 2.5 Pro Model Card
ModelEvaluationConditionValueLocationChecked
Gemini 2.5 Pro PreviewPreviewSituational awareness challengesDisabled Tool success rateNone statedFrontier Safety: deceptive alignmentUnverified

M3 Sabotage and sandbagging

2 values

M3 Sabotage and sandbagging: values in the Gemini 2.5 Pro Model Card
ModelEvaluationConditionValueLocationChecked
Gemini 2.5 Pro PreviewPreviewStealth challengesDecision Sabotage success rateNone statedFrontier Safety: deceptive alignmentUnverified
Gemini 2.5 Pro PreviewPreviewStealth challengesOversight Pattern success rateNone statedFrontier Safety: deceptive alignmentUnverified

M4 Misalignment audits

2 values

M4 Misalignment audits: values in the Gemini 2.5 Pro Model Card
ModelEvaluationConditionValueLocationChecked
Gemini 2.5 ProFrontier Safety Framework determination (misalignment)Instrumental Reasoning Level 1None statedFrontier Safety tableUnverified
Gemini 2.5 ProFrontier Safety Framework determination (misalignment)Instrumental Reasoning Level 2None statedFrontier Safety tableUnverified

M7 Harmful compliance and over-refusal

10 values

M7 Harmful compliance and over-refusal: values in the Gemini 2.5 Pro Model Card
ModelEvaluationConditionValueLocationChecked
Gemini 2.5 ProImage to Text SafetyImage-to-text policy-violation rate deltavs Gemini 1.5 Pro 002Safety evaluations tableUnverified
Gemini 2.5 ProInstruction FollowingSafe instruction following deltavs Gemini 1.5 Pro 002Safety evaluations tableUnverified
Gemini 2.5 ProMultilingual SafetyMultilingual policy-violation rate deltavs Gemini 1.5 Pro 002Safety evaluations tableUnverified
Gemini 2.5 ProText to Text SafetyPolicy-violation rate deltavs Gemini 1.5 Pro 002Safety evaluations tableUnverified
Gemini 2.5 ProToneObjective tone of refusals deltavs Gemini 1.5 Pro 002Safety evaluations tableUnverified
Gemini 2.5 Pro Preview 05-06Preview 05-06Image to Text SafetyImage-to-text policy-violation rate deltavs Gemini 1.5 Pro 002Safety evaluations tableUnverified
Gemini 2.5 Pro Preview 05-06Preview 05-06Instruction FollowingSafe instruction following deltavs Gemini 1.5 Pro 002Safety evaluations tableUnverified
Gemini 2.5 Pro Preview 05-06Preview 05-06Multilingual SafetyMultilingual policy-violation rate deltavs Gemini 1.5 Pro 002Safety evaluations tableUnverified
Gemini 2.5 Pro Preview 05-06Preview 05-06Text to Text SafetyPolicy-violation rate deltavs Gemini 1.5 Pro 002Safety evaluations tableUnverified
Gemini 2.5 Pro Preview 05-06Preview 05-06ToneObjective tone of refusals deltavs Gemini 1.5 Pro 002Safety evaluations tableUnverified

M10 Dangerous capabilities and risk determinations

6 values

M10 Dangerous capabilities and risk determinations: values in the Gemini 2.5 Pro Model Card
ModelEvaluationConditionValueLocationChecked
Gemini 2.5 ProFrontier Safety Framework determinationCBRN Uplift Level 1None statedFrontier Safety tableUnverified
Gemini 2.5 ProFrontier Safety Framework determinationCyber Autonomy Level 1None statedFrontier Safety tableUnverified
Gemini 2.5 ProFrontier Safety Framework determinationCyber Uplift Level 1None statedFrontier Safety tableUnverified
Gemini 2.5 ProFrontier Safety Framework determinationML R&D Autonomy Level 1None statedFrontier Safety tableUnverified
Gemini 2.5 ProFrontier Safety Framework determinationML R&D Uplift Level 1None statedFrontier Safety tableUnverified
Gemini 2.5 ProCyber key skills benchmarkHard challenges solvedmodel version unclearFrontier Safety: cyberUnverified

Extraction coverage

What we read of this document, and where each value was read.

Our note Cyber figures differ from tech report (low confidence rows)

All 21 values were read in the document itself.

All 21 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).