Grok 4.5 Model Card
A model card by xAI about Grok 4.5, published 14 Jul 2026. We recorded 15 values from it.
- Developer
- xAI
- Type
- Model card
- Model covered
- Grok 4.5
- Published
- 14 Jul 2026
- Also at
- Another address, cursor.com (opens an external site) (original version)
- Archived copy
- No archived copy yet
- Changelog
- No changelog
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Versions
Two versions are on record, oldest first. We hold copies of both.
14 Jul 2026
Original version (read in a copy hosted elsewhere)
20 Jul 2026
Revised version (the copy retrieved)
What changed revision, no changelog; safety values appear unchanged
No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).
Revisions
No revisions are recorded for this document. We have recorded no change to a value between them.
Values
Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. None has been blind-verified yet.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M5 Honesty and hallucination
2 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Grok 4.5 (high) | Hallucination rateUnsupported claims, single-turn | high | 6.1 | Unverified | |
| Grok 4.5 (high) | MASK-RectifiedDishonesty rate | high | 12.2 | Unverified |
M6 Sycophancy
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Grok 4.5 (high) | SycophancySycophancy rate | high | 12.3 | Unverified |
M7 Harmful compliance and over-refusal
6 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Grok 4.5 (high) | RefusalsCompliance on should-refuse prompts | high; full safeguards | 10.1 | Unverified | |
| Grok 4.5 (high) | CBRN refusalsRefusal accuracy on dangerous bio queries | high; full safeguards | 10.3 | Unverified | |
| Grok 4.5 (high) | Self-harm refusalsCompliance (non-refusal/redirect) | high | 11.1 | Unverified | |
| GPT-5.5 (xhigh) | HackerBenchCompliance on harmful/dual-use cyber tasks | xhigh | Cyber safeguards 7.2 | Unverified | |
| Grok 4.5 (high) | HackerBenchCompliance on harmful/dual-use cyber tasks | high | Cyber safeguards 7.2 | Unverified | |
| Grok 4.5 (high) | HackerBenchFalse refusal on benign cyber tasks | high | Cyber safeguards 7.2 | Unverified |
M8 Jailbreak robustness
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Grok 4.5 (high) | JailbreaksCompliance under attack | high; standard jailbreaks | 9.1 | Unverified |
M10 Dangerous capabilities and risk determinations
4 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Grok 4.5 (high) | CyberGymMean vulnerabilities reproduced | high; unrestricted | 7.1 | Unverified | |
| Grok 4.5 | Dual-use bio determinationThreshold statement | None stated | 8 | Unverified | |
| Grok 4.5 (high) | VCTAccuracy | high | 8.1 | Unverified | |
| Grok 4.5 (high) | ProtocolQA open-endedAccuracy | high | 8.4 | Unverified |
From an earlier version
This value was printed in the version of 14 Jul 2026 and replaced in a later version, which prints its own value for the same measurement (listed above). The changes we have recorded are under Revisions. Charts leave these values out by default.
| Model | Family | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|---|
| Grok 4.5 (high) | M7 | HackerBenchCompliance on harmful/dual-use cyber tasks | high | Cyber safeguards 7.2 | Unverified |
Extraction coverage
What we read of this document, and where each value was read.
Our note Metric scales changed from 0–1 rates to percentages
All 15 values were read in the document itself.
All 15 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).