Gemini 2.5 Technical Report
A technical report by Google DeepMind about Gemini 2.5 Pro and Gemini 2.5 Flash, published 17 Jun 2025. We recorded 25 values from it.
- Developer
- Google DeepMind
- Type
- Technical report
- Models covered
- Gemini 2.5 Pro, Gemini 2.5 Flash
- Published
- 17 Jun 2025
- Also at
- arXiv 2507.06261
- Archived copy
- No archived copy yet
- Changelog
- Not applicable
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Versions
One version is on record: the copy we retrieved on 26 Sep 2026. We know of no other.
26 Sep 2026
Date retrieved
Copy retrieved 26 Sep 2026
No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).
Revisions
No revisions are recorded for this document. We know of only one version of it.
Values
Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. 2 of the 25 have been blind-verified: a second reader found the same value without seeing ours.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M1 Evaluation awareness
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Gemini 2.5 Pro | Situational awareness challengesChallenges solved | None stated | Frontier Safety: deceptive alignment | Unverified |
M3 Sabotage and sandbagging
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Gemini 2.5 Pro | Stealth challengesChallenges solved | None stated | Frontier Safety: deceptive alignment | Unverified |
M5 Honesty and hallucination
2 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Gemini 2.5 Pro | FACTS GroundingFactuality score | None stated | Table 3 | Unverified | |
| Gemini 2.5 Pro | SimpleQAAccuracy | None stated | Table 3 | Unverified |
M7 Harmful compliance and over-refusal
7 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Gemini 2.5 Flash | Image to Text SafetyImage-to-text policy-violation rate delta | vs Gemini 1.5 Flash 002 | Table 7 | Unverified | |
| Gemini 2.5 Flash | Instruction FollowingSafe instruction following delta | vs Gemini 1.5 Flash 002 | Table 7 | Verified | |
| Gemini 2.5 Flash | Multilingual SafetyMultilingual policy-violation rate delta | vs Gemini 1.5 Flash 002 | Table 7 | Unverified | |
| Gemini 2.5 Flash | Text to Text SafetyPolicy-violation rate delta | vs Gemini 1.5 Flash 002 | Table 7 | Unverified | |
| Gemini 2.5 Flash | ToneObjective tone of refusals delta | vs Gemini 1.5 Flash 002 | Table 7 | Unverified | |
| Gemini 2.5 Flash | Automated Red Teaming (helpfulness)Unhelpful response rate | None stated | Table 8 | Unverified | |
| Gemini 2.5 Pro | Automated Red Teaming (helpfulness)Unhelpful response rate | None stated | Table 8 | Unverified |
M8 Jailbreak robustness
4 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Gemini 1.5 Pro 002002 | Automated Red Teaming (dangerous content)Dangerous content violation rate | None stated | Table 8 | Unverified | |
| Gemini 2.0 Flash | Automated Red Teaming (dangerous content)Dangerous content violation rate | None stated | Table 8 | Unverified | |
| Gemini 2.5 Flash | Automated Red Teaming (dangerous content)Dangerous content violation rate | None stated | Table 8 | Verified | |
| Gemini 2.5 Pro | Automated Red Teaming (dangerous content)Dangerous content violation rate | None stated | Table 8 | Unverified |
M9 Prompt injection
6 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Gemini 2.5 Flash | Indirect prompt injection (adaptive attacks)Actor Critic ASR | adaptive attack | Table 9 | Unverified | |
| Gemini 2.5 Flash | Indirect prompt injection (adaptive attacks)Beam Search ASR | adaptive attack | Table 9 | Unverified | |
| Gemini 2.5 Flash | Indirect prompt injection (adaptive attacks)TAP ASR | adaptive attack | Table 9 | Unverified | |
| Gemini 2.5 Pro | Indirect prompt injection (adaptive attacks)Actor Critic ASR | adaptive attack | Table 9 | Unverified | |
| Gemini 2.5 Pro | Indirect prompt injection (adaptive attacks)Beam Search ASR | adaptive attack | Table 9 | Unverified | |
| Gemini 2.5 Pro | Indirect prompt injection (adaptive attacks)TAP ASR | adaptive attack | Table 9 | Unverified |
M10 Dangerous capabilities and risk determinations
4 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Gemini 2.5 Pro | Cyber autonomous offense suiteHard CTF challenges solved | None stated | Frontier Safety: cyber | Unverified | |
| Gemini 2.5 Pro | Cyber key skills benchmarkHard challenges solved | None stated | Frontier Safety: cyber | Unverified | |
| Gemini 2.5 Pro | RE-BenchBest score as % of expert solution | max across tasks | Frontier Safety: ML R&D | Unverified | |
| Gemini 2.5 Pro | Frontier Safety Framework determinationCyber Uplift Level 1 | None stated | Table 10 | Unverified |
Extraction coverage
What we read of this document, and where each value was read.
Our note Tables 7 and 9 headers garbled in extraction; column mapping inferred (medium confidence)
All 25 values were read in the document itself.
All 25 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).