Gemini 2.5 Pro
34 values about Gemini 2.5 Pro from 2 documents, in 8 metric families. In each family, values from its own documents come first.
- Developer
- Google DeepMind
- Release date
- 17 Jun 2025The date of its first document in the dataset, the Gemini 2.5 Technical Report.
- Availability
- Public
- Also printed as
- Gemini 2.5 Pro Preview · Gemini 2.5 Pro Preview 05-06
- Snapshots
- Preview 05-06 · PreviewValues about a snapshot are marked in the tables; they are not a separate model.
- Its own documents
- Gemini 2.5 Technical Report · 17 Jun 2025
- Gemini 2.5 Pro Model Card · 27 Jun 2025
- Also reported in
- No other document reports a value about it
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Values by metric family
Each value as the document printed it. Select a value for its source, its checks and its history.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M1 · Evaluation awareness 2 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documents | ||
| Situational awareness challengesChallenges solvedGemini 2.5 Technical Report17 Jun 2025 | Gemini 2.5 Technical Report17 Jun 2025 | |
| Situational awareness challengesDisabled Tool success rateSnapshot: PreviewGemini 2.5 Pro Model Card27 Jun 2025 | Gemini 2.5 Pro Model Card27 Jun 2025 | |
M3 · Sabotage and sandbagging 3 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documents | ||
| Stealth challengesChallenges solvedGemini 2.5 Technical Report17 Jun 2025 | Gemini 2.5 Technical Report17 Jun 2025 | |
| Stealth challengesDecision Sabotage success rateSnapshot: PreviewGemini 2.5 Pro Model Card27 Jun 2025 | Gemini 2.5 Pro Model Card27 Jun 2025 | |
| Stealth challengesOversight Pattern success rateSnapshot: PreviewGemini 2.5 Pro Model Card27 Jun 2025 | Gemini 2.5 Pro Model Card27 Jun 2025 | |
M4 · Misalignment audits 2 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentsGemini 2.5 Pro Model Card · 27 Jun 2025 | ||
| Frontier Safety Framework determination (misalignment)Instrumental Reasoning Level 1 | Gemini 2.5 Pro Model Card27 Jun 2025 | |
| Frontier Safety Framework determination (misalignment)Instrumental Reasoning Level 2 | Gemini 2.5 Pro Model Card27 Jun 2025 | |
M5 · Honesty and hallucination 2 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentsGemini 2.5 Technical Report · 17 Jun 2025 | ||
| FACTS GroundingFactuality score | Gemini 2.5 Technical Report17 Jun 2025 | |
| SimpleQAAccuracy | Gemini 2.5 Technical Report17 Jun 2025 | |
M7 · Harmful compliance and over-refusal 11 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documents | ||
| Automated Red Teaming (helpfulness)Unhelpful response rateGemini 2.5 Technical Report17 Jun 2025 | Gemini 2.5 Technical Report17 Jun 2025 | |
| Image to Text SafetyImage-to-text policy-violation rate deltaCondition: vs Gemini 1.5 Pro 002Gemini 2.5 Pro Model Card27 Jun 2025 | Gemini 2.5 Pro Model Card27 Jun 2025 | |
| Image to Text SafetyImage-to-text policy-violation rate deltaCondition: vs Gemini 1.5 Pro 002Snapshot: Preview 05-06Gemini 2.5 Pro Model Card27 Jun 2025 | Gemini 2.5 Pro Model Card27 Jun 2025 | |
| Instruction FollowingSafe instruction following deltaCondition: vs Gemini 1.5 Pro 002Gemini 2.5 Pro Model Card27 Jun 2025 | Gemini 2.5 Pro Model Card27 Jun 2025 | |
| Instruction FollowingSafe instruction following deltaCondition: vs Gemini 1.5 Pro 002Snapshot: Preview 05-06Gemini 2.5 Pro Model Card27 Jun 2025 | Gemini 2.5 Pro Model Card27 Jun 2025 | |
| Multilingual SafetyMultilingual policy-violation rate deltaCondition: vs Gemini 1.5 Pro 002Gemini 2.5 Pro Model Card27 Jun 2025 | Gemini 2.5 Pro Model Card27 Jun 2025 | |
| Multilingual SafetyMultilingual policy-violation rate deltaCondition: vs Gemini 1.5 Pro 002Snapshot: Preview 05-06Gemini 2.5 Pro Model Card27 Jun 2025 | Gemini 2.5 Pro Model Card27 Jun 2025 | |
| Text to Text SafetyPolicy-violation rate deltaCondition: vs Gemini 1.5 Pro 002Gemini 2.5 Pro Model Card27 Jun 2025 | Gemini 2.5 Pro Model Card27 Jun 2025 | |
| Text to Text SafetyPolicy-violation rate deltaCondition: vs Gemini 1.5 Pro 002Snapshot: Preview 05-06Gemini 2.5 Pro Model Card27 Jun 2025 | Gemini 2.5 Pro Model Card27 Jun 2025 | |
| ToneObjective tone of refusals deltaCondition: vs Gemini 1.5 Pro 002Gemini 2.5 Pro Model Card27 Jun 2025 | Gemini 2.5 Pro Model Card27 Jun 2025 | |
| ToneObjective tone of refusals deltaCondition: vs Gemini 1.5 Pro 002Snapshot: Preview 05-06Gemini 2.5 Pro Model Card27 Jun 2025 | Gemini 2.5 Pro Model Card27 Jun 2025 | |
M8 · Jailbreak robustness 1 value
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentsGemini 2.5 Technical Report · 17 Jun 2025 | ||
| Automated Red Teaming (dangerous content)Dangerous content violation rate | Gemini 2.5 Technical Report17 Jun 2025 | |
M9 · Prompt injection 3 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documentsGemini 2.5 Technical Report · 17 Jun 2025 | ||
| Indirect prompt injection (adaptive attacks)Actor Critic ASRCondition: adaptive attack | Gemini 2.5 Technical Report17 Jun 2025 | |
| Indirect prompt injection (adaptive attacks)Beam Search ASRCondition: adaptive attack | Gemini 2.5 Technical Report17 Jun 2025 | |
| Indirect prompt injection (adaptive attacks)TAP ASRCondition: adaptive attack | Gemini 2.5 Technical Report17 Jun 2025 | |
M10 · Dangerous capabilities and risk determinations 10 values
| Evaluation and metric | Value | Document |
|---|---|---|
| From its own documents | ||
| Cyber autonomous offense suiteHard CTF challenges solvedGemini 2.5 Technical Report17 Jun 2025 | Gemini 2.5 Technical Report17 Jun 2025 | |
| Cyber key skills benchmarkHard challenges solvedGemini 2.5 Technical Report17 Jun 2025 | Gemini 2.5 Technical Report17 Jun 2025 | |
| Frontier Safety Framework determinationCyber Uplift Level 1Gemini 2.5 Technical Report17 Jun 2025Restated later: compare with the later value | Gemini 2.5 Technical Report17 Jun 2025Restated later: compare with the later value | |
| RE-BenchBest score as % of expert solutionCondition: max across tasksGemini 2.5 Technical Report17 Jun 2025 | Gemini 2.5 Technical Report17 Jun 2025 | |
| Cyber key skills benchmarkHard challenges solvedCondition: model version unclearGemini 2.5 Pro Model Card27 Jun 2025 | Gemini 2.5 Pro Model Card27 Jun 2025 | |
| Frontier Safety Framework determinationCBRN Uplift Level 1Gemini 2.5 Pro Model Card27 Jun 2025 | Gemini 2.5 Pro Model Card27 Jun 2025 | |
| Frontier Safety Framework determinationCyber Autonomy Level 1Gemini 2.5 Pro Model Card27 Jun 2025 | Gemini 2.5 Pro Model Card27 Jun 2025 | |
| Frontier Safety Framework determinationCyber Uplift Level 1Gemini 2.5 Pro Model Card27 Jun 2025Restated: compare with the earlier value | Gemini 2.5 Pro Model Card27 Jun 2025Restated: compare with the earlier value | |
| Frontier Safety Framework determinationML R&D Autonomy Level 1Gemini 2.5 Pro Model Card27 Jun 2025 | Gemini 2.5 Pro Model Card27 Jun 2025 | |
| Frontier Safety Framework determinationML R&D Uplift Level 1Gemini 2.5 Pro Model Card27 Jun 2025 | Gemini 2.5 Pro Model Card27 Jun 2025 | |
Restated in later documents
A later document reported a value about Gemini 2.5 Pro again. Each pair is shown side by side: both values are kept with their own documents, and the later one does not replace the earlier.
| Earlier value | Later value |
|---|---|
| M10 · Dangerous capabilities and risk determinations | |
| Frontier Safety Framework determinationCyber Uplift Level 1 | |
| Gemini 2.5 Technical Report17 Jun 2025 · its own document | Gemini 2.5 Pro Model Card27 Jun 2025No reason stated |
Risk determinations
The developer's formal decisions about Gemini 2.5 Pro under its framework, as printed. Levels from different frameworks do not map onto one another.
| Domain | Level as printed | Framework and document |
|---|---|---|
| Cyber | alert_threshold_reached code, not the printed wording Frontier Safety Framework Gemini 2.5 Technical Report · 17 Jun 2025 | Frontier Safety Framework Gemini 2.5 Technical Report · 17 Jun 2025 |
| Bio/chem | not_reached code, not the printed wording Frontier Safety Framework Gemini 2.5 Pro Model Card · 27 Jun 2025 | Frontier Safety Framework Gemini 2.5 Pro Model Card · 27 Jun 2025 |
| Cyber | not_reached code, not the printed wording Frontier Safety Framework Gemini 2.5 Pro Model Card · 27 Jun 2025 | Frontier Safety Framework Gemini 2.5 Pro Model Card · 27 Jun 2025 |
| Cyber | alert_threshold_reached code, not the printed wording Frontier Safety Framework Gemini 2.5 Pro Model Card · 27 Jun 2025 | Frontier Safety Framework Gemini 2.5 Pro Model Card · 27 Jun 2025 |
| AI R&D / autonomy | not_reached code, not the printed wording Frontier Safety Framework Gemini 2.5 Pro Model Card · 27 Jun 2025 | Frontier Safety Framework Gemini 2.5 Pro Model Card · 27 Jun 2025 |
| AI R&D / autonomy | not_reached code, not the printed wording Frontier Safety Framework Gemini 2.5 Pro Model Card · 27 Jun 2025 | Frontier Safety Framework Gemini 2.5 Pro Model Card · 27 Jun 2025 |
| Misalignment | not_reached code, not the printed wording Frontier Safety Framework Gemini 2.5 Pro Model Card · 27 Jun 2025 | Frontier Safety Framework Gemini 2.5 Pro Model Card · 27 Jun 2025 |
| Misalignment | not_reached code, not the printed wording Frontier Safety Framework Gemini 2.5 Pro Model Card · 27 Jun 2025 | Frontier Safety Framework Gemini 2.5 Pro Model Card · 27 Jun 2025 |
“Code, not the printed wording”: the version 0 file stored a code for this level rather than the words the document printed. The wording will be read again from the source.