Grok 4.6 Model Card
A model card by xAI about Grok 4.6, published 12 Aug 2026. We recorded 22 values from it.
- Developer
- xAI
- Type
- Model card
- Model covered
- Grok 4.6
- Published
- 12 Aug 2026
- Also at
- Another address, cursor.com (opens an external site) (original version)
- Archived copy
- No archived copy yet
- Changelog
- Partial changelog
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Revisions
Each change we have recorded between two versions: the old and the new text, whether the document's changelog explains it, how we know, and whether we have checked it. A revision is not in itself evidence of wrongdoing; most are corrections (Methodology §7).
16.7%6.9%Partly explained — changelog lists "corrected eval results" for this eval; no reason
Old value: New value:
How we know The old value was read in the earlier copy and the new one in the later copy. Blind check 26 Sep 2026 (both versions)
Checked Blind-verified
Version of 12 Aug 2026 (opens an external site) Version of 17 Aug 2026 (opens an external site)
3.8%1.90%Partly explained — listed as corrected; no reason
Old value: New value:
How we know The old value was read in the earlier copy and the new one in the later copy. Blind check 26 Sep 2026 (both versions)
Checked Blind-verified
Version of 12 Aug 2026 (opens an external site) Version of 17 Aug 2026 (opens an external site)
3.7%0.84%Partly explained — listed as corrected; no reason
Old value: New value:
How we know The old value was read in the earlier copy and the new one in the later copy. Blind check 26 Sep 2026 (both versions)
Checked Blind-verified
Version of 12 Aug 2026 (opens an external site) Version of 17 Aug 2026 (opens an external site)
no appreciable liftnoted in the biological domain but is limitedNot explained — The changelog does not mention this change.
Old value: New value:
How we know The old value was read in the earlier copy and the new one in the later copy. The FORTRESS-RN addition in the same revision is recorded separately.
Checked Not yet checked
Version of 12 Aug 2026 (opens an external site) Version of 17 Aug 2026 (opens an external site)
97.9%Not explained — Added without a changelog note.
New value:
How we know The value was read in the later copy; the earlier copy does not have it.
Checked Not yet checked
Version of 12 Aug 2026 (opens an external site) Version of 17 Aug 2026 (opens an external site)
Versions
Two versions are on record, oldest first. We hold copies of both.
12 Aug 2026
Original version (read in a copy hosted elsewhere)
17 Aug 2026
Revised version (the copy retrieved)
What changed 5 changes recorded as revisions. See them in redline
No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).
Values
Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. 9 of the 22 have been blind-verified: a second reader found the same value without seeing ours. 1 value was corrected after review.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M5 Honesty and hallucination
2 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Grok 4.6 (high) | Hallucination rateUnsupported claims, single-turn | high | 6.1 | Unverified | |
| Grok 4.6 (high) | MASK-RectifiedDishonesty rate | high | 12.1 | Verified |
M6 Sycophancy
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Grok 4.6 (high) | SycophancySycophancy rate | high | 12.2 | Unverified |
M7 Harmful compliance and over-refusal
7 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Grok 4.6 (high) | BioUseBenchSeverity-5 refusal rate | high | 8.3 | Unverified | |
| Grok 4.6 (high) | RefusalsCompliance on should-refuse prompts | high; full safeguards | 10.1 | Unverified | |
| Grok 4.6 (high) | FORTRESS-RNR/N refusal accuracy | high | 10.3 | Verified | |
| Grok 4.6 (high) | Self-harm refusalsCompliance (non-refusal/redirect) | high | 11.1 | Verified | |
| GPT-5.6 Sol (max) | HackerBenchCompliance on harmful/dual-use cyber tasks | max | Cyber 7.4 | Unverified | |
| Grok 4.6 (high) | HackerBenchCompliance on harmful/dual-use cyber tasks | high | Cyber 7.4 | Verified | |
| Grok 4.6 (high) | HackerBenchFalse refusal on benign cyber tasks | high | Cyber 7.4 | Verified |
M8 Jailbreak robustness
2 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Grok 4.6 (high) | JailbreaksCompliance under attack | high; standard jailbreaks | 9.1 | Unverified | |
| Grok 4.6 (high) | StrongREJECTCompliance under StrongREJECT attacks | high | 9.1 | Unverified |
M10 Dangerous capabilities and risk determinations
5 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Grok 4.6 (high) | CyberGymMean vulnerabilities reproduced | high; unrestricted | 7.1 | Unverified | |
| Grok 4.6 (high) | CVE-BenchReward | high; unrestricted | 7.2 | Unverified | |
| Grok 4.6 (high) | VCTAccuracy | high | 8.1 | Unverified | |
| Grok 4.6 (high) | WMDPCyber accuracy | high | 8.4 | Unverified | |
| Grok 4.6 | Bio capability liftLift vs Grok 4.5 | None stated | Dual-use summary | Corrected |
From an earlier version
These 5 values were printed in the version of 12 Aug 2026 and replaced in a later version, which prints its own value for each of the same measurements (listed above). The changes we have recorded are under Revisions. Charts leave these values out by default.
| Model | Family | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|---|
| Grok 4.6 (high) | M5 | MASK-RectifiedDishonesty rate | high | 12.1 | Verified | |
| Grok 4.6 (high) | M7 | Self-harm refusalsCompliance (non-refusal/redirect) | high | 11.1 | Verified | |
| Grok 4.6 (high) | M7 | HackerBenchCompliance on harmful/dual-use cyber tasks | high | Cyber 7.4 | Verified | |
| Grok 4.6 (high) | M7 | HackerBenchFalse refusal on benign cyber tasks | high | Cyber 7.4 | Unverified | |
| Grok 4.6 | M10 | Bio capability liftLift vs Grok 4.5 | None stated | Dual-use summary | Verified |
Extraction coverage
What we read of this document, and where each value was read.
Our note See also themidasproject.com/watchtower/xai-08172026 (opens an external site)
All 22 values were read in the document itself.
Of the 22 values, 1 was read off a figure and 2 are statements in words rather than numbers; they are marked * and left out of charts by default.
All 22 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).