Versions

One version is on record: the copy we retrieved on 26 Sep 2026. We know of no other.

  1. 26 Sep 2026

    Date retrieved

    Copy retrieved 26 Sep 2026

    Date from
    the date we retrieved it; the copy states no version date
    Copy
    Copy retrieved on 26 Sep 2026
    Values
    25 values recorded from this version

No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).

Revisions

No revisions are recorded for this document. We know of only one version of it.

Values

Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. 2 of the 25 have been blind-verified: a second reader found the same value without seeing ours.

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

M1 Evaluation awareness

1 value

M1 Evaluation awareness: values in the Gemini 2.5 Technical Report
ModelEvaluationConditionValueLocationChecked
Gemini 2.5 ProSituational awareness challengesChallenges solvedNone statedFrontier Safety: deceptive alignmentUnverified

M3 Sabotage and sandbagging

1 value

M3 Sabotage and sandbagging: values in the Gemini 2.5 Technical Report
ModelEvaluationConditionValueLocationChecked
Gemini 2.5 ProStealth challengesChallenges solvedNone statedFrontier Safety: deceptive alignmentUnverified

M5 Honesty and hallucination

2 values

M5 Honesty and hallucination: values in the Gemini 2.5 Technical Report
ModelEvaluationConditionValueLocationChecked
Gemini 2.5 ProFACTS GroundingFactuality scoreNone statedTable 3Unverified
Gemini 2.5 ProSimpleQAAccuracyNone statedTable 3Unverified

M7 Harmful compliance and over-refusal

7 values

M7 Harmful compliance and over-refusal: values in the Gemini 2.5 Technical Report
ModelEvaluationConditionValueLocationChecked
Gemini 2.5 FlashImage to Text SafetyImage-to-text policy-violation rate deltavs Gemini 1.5 Flash 002Table 7Unverified
Gemini 2.5 FlashInstruction FollowingSafe instruction following deltavs Gemini 1.5 Flash 002Table 7Verified
Gemini 2.5 FlashMultilingual SafetyMultilingual policy-violation rate deltavs Gemini 1.5 Flash 002Table 7Unverified
Gemini 2.5 FlashText to Text SafetyPolicy-violation rate deltavs Gemini 1.5 Flash 002Table 7Unverified
Gemini 2.5 FlashToneObjective tone of refusals deltavs Gemini 1.5 Flash 002Table 7Unverified
Gemini 2.5 FlashAutomated Red Teaming (helpfulness)Unhelpful response rateNone statedTable 8Unverified
Gemini 2.5 ProAutomated Red Teaming (helpfulness)Unhelpful response rateNone statedTable 8Unverified

M8 Jailbreak robustness

4 values

M8 Jailbreak robustness: values in the Gemini 2.5 Technical Report
ModelEvaluationConditionValueLocationChecked
Gemini 1.5 Pro 002002Automated Red Teaming (dangerous content)Dangerous content violation rateNone statedTable 8Unverified
Gemini 2.0 FlashAutomated Red Teaming (dangerous content)Dangerous content violation rateNone statedTable 8Unverified
Gemini 2.5 FlashAutomated Red Teaming (dangerous content)Dangerous content violation rateNone statedTable 8Verified
Gemini 2.5 ProAutomated Red Teaming (dangerous content)Dangerous content violation rateNone statedTable 8Unverified

M9 Prompt injection

6 values

M9 Prompt injection: values in the Gemini 2.5 Technical Report
ModelEvaluationConditionValueLocationChecked
Gemini 2.5 FlashIndirect prompt injection (adaptive attacks)Actor Critic ASRadaptive attackTable 9Unverified
Gemini 2.5 FlashIndirect prompt injection (adaptive attacks)Beam Search ASRadaptive attackTable 9Unverified
Gemini 2.5 FlashIndirect prompt injection (adaptive attacks)TAP ASRadaptive attackTable 9Unverified
Gemini 2.5 ProIndirect prompt injection (adaptive attacks)Actor Critic ASRadaptive attackTable 9Unverified
Gemini 2.5 ProIndirect prompt injection (adaptive attacks)Beam Search ASRadaptive attackTable 9Unverified
Gemini 2.5 ProIndirect prompt injection (adaptive attacks)TAP ASRadaptive attackTable 9Unverified

M10 Dangerous capabilities and risk determinations

4 values

M10 Dangerous capabilities and risk determinations: values in the Gemini 2.5 Technical Report
ModelEvaluationConditionValueLocationChecked
Gemini 2.5 ProCyber autonomous offense suiteHard CTF challenges solvedNone statedFrontier Safety: cyberUnverified
Gemini 2.5 ProCyber key skills benchmarkHard challenges solvedNone statedFrontier Safety: cyberUnverified
Gemini 2.5 ProRE-BenchBest score as % of expert solutionmax across tasksFrontier Safety: ML R&DUnverified
Gemini 2.5 ProFrontier Safety Framework determinationCyber Uplift Level 1None statedTable 10Unverified

Extraction coverage

What we read of this document, and where each value was read.

Our note Tables 7 and 9 headers garbled in extraction; column mapping inferred (medium confidence)

All 25 values were read in the document itself.

All 25 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).