Gemini 2.5 Deep Think
15 values about Gemini 2.5 Deep Think from 1 document, in 5 metric families. In each family, values from its own document come first.
- Developer
- Google DeepMind
- Release date
- 1 Aug 2025The date of its first card in the dataset, the Gemini 2.5 Deep Think Model Card.
- Availability
- Public
- Its own document
- Gemini 2.5 Deep Think Model Card · 1 Aug 2025
- Also reported in
- No other document reports a value about it
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Values by metric family
Each value as the document printed it. Select a value for its source, its checks and its history.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M1 · Evaluation awareness 1 value
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGemini 2.5 Deep Think Model Card · 1 Aug 2025 | ||
| Situational awareness challengesChallenges solved | Gemini 2.5 Deep Think Model Card1 Aug 2025 | |
M3 · Sabotage and sandbagging 1 value
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGemini 2.5 Deep Think Model Card · 1 Aug 2025 | ||
| Stealth challengesChallenges solved | Gemini 2.5 Deep Think Model Card1 Aug 2025 | |
M4 · Misalignment audits 1 value
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGemini 2.5 Deep Think Model Card · 1 Aug 2025 | ||
| Frontier Safety Framework determination (misalignment)Instrumental Reasoning Level 1/2 | Gemini 2.5 Deep Think Model Card1 Aug 2025 | |
M7 · Harmful compliance and over-refusal 5 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGemini 2.5 Deep Think Model Card · 1 Aug 2025 | ||
| Image to Text SafetyImage-to-text policy-violation rate deltaCondition: vs Gemini 2.5 Pro | Gemini 2.5 Deep Think Model Card1 Aug 2025 | |
| Instruction FollowingSafe instruction following deltaCondition: vs Gemini 2.5 Pro | Gemini 2.5 Deep Think Model Card1 Aug 2025 | |
| Multilingual SafetyMultilingual policy-violation rate deltaCondition: vs Gemini 2.5 Pro | Gemini 2.5 Deep Think Model Card1 Aug 2025 | |
| Text to Text SafetyPolicy-violation rate deltaCondition: vs Gemini 2.5 Pro | Gemini 2.5 Deep Think Model Card1 Aug 2025 | |
| ToneObjective tone of refusals deltaCondition: vs Gemini 2.5 Pro | Gemini 2.5 Deep Think Model Card1 Aug 2025 | |
M10 · Dangerous capabilities and risk determinations 7 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentGemini 2.5 Deep Think Model Card · 1 Aug 2025 | ||
| Cyber autonomous offense suiteHard CTF challenges solved | Gemini 2.5 Deep Think Model Card1 Aug 2025 | |
| Cyber key skills benchmarkHard challenges solved | Gemini 2.5 Deep Think Model Card1 Aug 2025 | |
| Frontier Safety Framework determinationCBRN Uplift Level 1 | Gemini 2.5 Deep Think Model Card1 Aug 2025 | |
| Frontier Safety Framework determinationCyber Autonomy Level 1 | Gemini 2.5 Deep Think Model Card1 Aug 2025 | |
| Frontier Safety Framework determinationCyber Uplift Level 1 | Gemini 2.5 Deep Think Model Card1 Aug 2025 | |
| Frontier Safety Framework determinationML R&D Autonomy/Uplift Level 1 | Gemini 2.5 Deep Think Model Card1 Aug 2025 | |
| RE-BenchAverage normalised score | Gemini 2.5 Deep Think Model Card1 Aug 2025 | |
Risk determinations
The developer's formal decisions about Gemini 2.5 Deep Think under its framework, as printed. Levels from different frameworks do not map onto one another.
| Domain | Level as printed | Framework and document |
|---|---|---|
| Bio/chem | alert_threshold_reached code, not the printed wording Frontier Safety Framework Gemini 2.5 Deep Think Model Card · 1 Aug 2025 | Frontier Safety Framework Gemini 2.5 Deep Think Model Card · 1 Aug 2025 |
| Cyber | alert_threshold_reached code, not the printed wording Frontier Safety Framework Gemini 2.5 Deep Think Model Card · 1 Aug 2025 | Frontier Safety Framework Gemini 2.5 Deep Think Model Card · 1 Aug 2025 |
| Cyber | not_reached code, not the printed wording Frontier Safety Framework Gemini 2.5 Deep Think Model Card · 1 Aug 2025 | Frontier Safety Framework Gemini 2.5 Deep Think Model Card · 1 Aug 2025 |
| AI R&D / autonomy | not_reached code, not the printed wording Frontier Safety Framework Gemini 2.5 Deep Think Model Card · 1 Aug 2025 | Frontier Safety Framework Gemini 2.5 Deep Think Model Card · 1 Aug 2025 |
| Misalignment | not_reached code, not the printed wording Frontier Safety Framework Gemini 2.5 Deep Think Model Card · 1 Aug 2025 | Frontier Safety Framework Gemini 2.5 Deep Think Model Card · 1 Aug 2025 |
“Code, not the printed wording”: the version 0 file stored a code for this level rather than the words the document printed. The wording will be read again from the source.