Gemini 3.1 Pro Model Card
A model card by Google DeepMind about Gemini 3.1 Pro, published 19 Feb 2026. We recorded 17 values from it.
- Developer
- Google DeepMind
- Type
- Model card
- Model covered
- Gemini 3.1 Pro
- Published
- 19 Feb 2026
- Archived copy
- No archived copy yet
- Changelog
- No changelog
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Versions
One version is on record: the copy we retrieved on 26 Sep 2026. We know of no other.
26 Sep 2026
Date retrieved
Copy retrieved 26 Sep 2026
No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).
Revisions
No revisions are recorded for this document. We know of only one version of it.
Values
Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. 1 of the 17 has been blind-verified: a second reader found the same value without seeing ours.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M1 Evaluation awareness
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Gemini 3.1 Pro | Situational awareness challengesSuccess rate on three hardest challenges | Deep Think | Frontier Safety: misalignment | Unverified |
M3 Sabotage and sandbagging
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Gemini 3.1 Pro | Stealth challengesComparison with Gemini 3 Pro | Deep Think | Frontier Safety: misalignment | Unverified |
M4 Misalignment audits
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Gemini 3.1 Pro | Frontier Safety Framework determination (misalignment)Instrumental reasoning | Deep Think | Frontier Safety table | Unverified |
M7 Harmful compliance and over-refusal
6 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Gemini 3.1 Pro | Child safety launch thresholdsLaunch thresholds met | None stated | Safety evaluations | Unverified | |
| Gemini 3.1 Pro | Image to Text SafetyImage-to-text policy-violation rate delta | vs Gemini 3 Pro | Safety evaluations table | Unverified | |
| Gemini 3.1 Pro | Multilingual SafetyMultilingual policy-violation rate delta | vs Gemini 3 Pro | Safety evaluations table | Unverified | |
| Gemini 3.1 Pro | Text to Text SafetyPolicy-violation rate delta | vs Gemini 3 Pro | Safety evaluations table | Verified | |
| Gemini 3.1 Pro | ToneObjective tone of refusals delta | vs Gemini 3 Pro | Safety evaluations table | Unverified | |
| Gemini 3.1 Pro | Unjustified refusalsUnjustified refusals delta | vs Gemini 3 Pro | Safety evaluations table | Unverified |
M10 Dangerous capabilities and risk determinations
8 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Gemini 3.1 Pro | Frontier Safety Framework determinationCBRN Uplift Level 1 | Deep Think | Frontier Safety table | Unverified | |
| Gemini 3.1 Pro | Frontier Safety Framework determinationCyber Uplift Level 1 | None stated | Frontier Safety table | Unverified | |
| Gemini 3.1 Pro | Frontier Safety Framework determinationHarmful manipulation | Deep Think | Frontier Safety table | Unverified | |
| Gemini 3.1 Pro | Frontier Safety Framework determinationML R&D | Deep Think | Frontier Safety table | Unverified | |
| Gemini 3 Pro | Harmful manipulation efficacy studyMax odds ratio vs non-AI baseline | None stated | Frontier Safety: manipulation | Unverified | |
| Gemini 3.1 Pro | Harmful manipulation efficacy studyMax odds ratio vs non-AI baseline | Deep Think | Frontier Safety: manipulation | Unverified | |
| Gemini 3 Pro | RE-BenchAverage normalised score | None stated | Frontier Safety: ML R&D | Unverified | |
| Gemini 3.1 Pro | RE-BenchAverage normalised score | Deep Think | Frontier Safety: ML R&D | Unverified |
Extraction coverage
What we read of this document, and where each value was read.
Our note We have not written a coverage note for this document yet.
All 17 values were read in the document itself.
Of the 17 values, 2 are statements in words rather than numbers; they are marked * and left out of charts by default.
All 17 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).