Gemini 2.5 Deep Think Model Card
A model card by Google DeepMind about Gemini 2.5 Deep Think, published 1 Aug 2025. We recorded 15 values from it.
- Developer
- Google DeepMind
- Type
- Model card
- Model covered
- Gemini 2.5 Deep Think
- Published
- 1 Aug 2025
- Archived copy
- No archived copy yet
- Changelog
- No changelog
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Versions
One version is on record: the copy we retrieved on 26 Sep 2026. We know of no other.
26 Sep 2026
Date retrieved
Copy retrieved 26 Sep 2026
No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).
Revisions
No revisions are recorded for this document. We know of only one version of it.
Values
Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. 1 of the 15 has been blind-verified: a second reader found the same value without seeing ours.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M1 Evaluation awareness
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Gemini 2.5 Deep Think | Situational awareness challengesChallenges solved | None stated | Frontier Safety: deceptive alignment | Unverified |
M3 Sabotage and sandbagging
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Gemini 2.5 Deep Think | Stealth challengesChallenges solved | None stated | Frontier Safety: deceptive alignment | Unverified |
M4 Misalignment audits
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Gemini 2.5 Deep Think | Frontier Safety Framework determination (misalignment)Instrumental Reasoning Level 1/2 | None stated | Frontier Safety table | Unverified |
M7 Harmful compliance and over-refusal
5 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Gemini 2.5 Deep Think | Image to Text SafetyImage-to-text policy-violation rate delta | vs Gemini 2.5 Pro | Safety evaluations table | Unverified | |
| Gemini 2.5 Deep Think | Instruction FollowingSafe instruction following delta | vs Gemini 2.5 Pro | Safety evaluations table | Unverified | |
| Gemini 2.5 Deep Think | Multilingual SafetyMultilingual policy-violation rate delta | vs Gemini 2.5 Pro | Safety evaluations table | Unverified | |
| Gemini 2.5 Deep Think | Text to Text SafetyPolicy-violation rate delta | vs Gemini 2.5 Pro | Safety evaluations table | Unverified | |
| Gemini 2.5 Deep Think | ToneObjective tone of refusals delta | vs Gemini 2.5 Pro | Safety evaluations table | Verified |
M10 Dangerous capabilities and risk determinations
7 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Gemini 2.5 Deep Think | Frontier Safety Framework determinationCBRN Uplift Level 1 | None stated | Frontier Safety table | Unverified | |
| Gemini 2.5 Deep Think | Frontier Safety Framework determinationCyber Autonomy Level 1 | None stated | Frontier Safety table | Unverified | |
| Gemini 2.5 Deep Think | Frontier Safety Framework determinationCyber Uplift Level 1 | None stated | Frontier Safety table | Unverified | |
| Gemini 2.5 Deep Think | Frontier Safety Framework determinationML R&D Autonomy/Uplift Level 1 | None stated | Frontier Safety table | Unverified | |
| Gemini 2.5 Deep Think | Cyber autonomous offense suiteHard CTF challenges solved | None stated | Frontier Safety: cyber | Unverified | |
| Gemini 2.5 Deep Think | Cyber key skills benchmarkHard challenges solved | None stated | Frontier Safety: cyber | Unverified | |
| Gemini 2.5 Deep Think | RE-BenchAverage normalised score | None stated | Frontier Safety: ML R&D | Unverified |
Extraction coverage
What we read of this document, and where each value was read.
Our note First model to reach CBRN early-warning alert threshold
All 15 values were read in the document itself.
All 15 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).