Grok 4.20 System Card
A system card by xAI about Grok 4.20, published 7 Apr 2026. We recorded 18 values from it.
- Developer
- xAI
- Type
- System card
- Model covered
- Grok 4.20
- Published
- 7 Apr 2026
- Archived copy
- No archived copy yet
- Changelog
- No changelog
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Versions
One version is on record: the copy we retrieved on 26 Sep 2026. We know of no other.
26 Sep 2026
Date retrieved
Copy retrieved 26 Sep 2026
No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).
Revisions
No revisions are recorded for this document. We know of only one version of it.
Values
Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. 4 of the 18 have been blind-verified: a second reader found the same value without seeing ours.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M1 Evaluation awareness
2 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Grok 4 | Alignment auditVerbalized evaluation awareness rate | across all audits | Table 3 | Verified | |
| Grok 4.20 (single-agent) | Alignment auditVerbalized evaluation awareness rate | single-agent; across all audits | Table 3 | Verified |
M3 Sabotage and sandbagging
2 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Grok 4.20 (single-agent) | Alignment auditSubversive action rate vs user | single-agent; unethical-operator scenarios | Table 3 | Unverified | |
| Grok 4.20 (single-agent) | Alignment auditSubversive action rate vs xAI | single-agent | Table 3 | Unverified |
M5 Honesty and hallucination
2 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Grok 4.20 (single-agent) | HLE calibrationRMS calibration error | single-agent | Table 2 | Unverified | |
| Grok 4.20 (single-agent) | MASKDishonesty rate | single-agent | Table 2 | Unverified |
M6 Sycophancy
2 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Grok 4.20 (single-agent) | SycophancySycophancy rate | single-agent | Table 2 | Unverified | |
| Grok 4.20 (single-agent) | Alignment auditDelusion validation rate | single-agent | Table 3 | Unverified |
M7 Harmful compliance and over-refusal
4 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Grok 4.20 (single-agent) | AgentHarmAnswer rate on harmful agentic tasks | single-agent; no attack | Table 1 | Unverified | |
| Grok 4.1 | Alignment auditCooperation with misuse | chat | Table 3 | Verified | |
| Grok 4.20 (single-agent) | Alignment auditCooperation with misuse | single-agent; chat | Table 3 | Unverified | |
| Grok 4.20 (single-agent) | Alignment auditMisuse with system-prompt override | single-agent; system prompt override | Table 3 | Unverified |
M8 Jailbreak robustness
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Grok 4.20 (single-agent) | JailbreaksAnswer rate on should-refuse prompts | single-agent; user jailbreak | Table 1 | Unverified |
M9 Prompt injection
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Grok 4.20 (single-agent) | AgentDojoPrompt-injection attack success rate | single-agent | Table 1 | Unverified |
M10 Dangerous capabilities and risk determinations
4 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Grok 4.20 | Overall risk conclusion | with safeguards | Conclusion | Unverified | |
| Grok 4.20 (single-agent) | CloningScenariosAccuracy | without safeguards; single-agent | Table 4 | Unverified | |
| Grok 4.20 (single-agent) | CybenchUnguided success rate | without safeguards; single-agent | Table 4 | Verified | |
| Grok 4.20 (single-agent) | VCTAccuracy | without safeguards; single-agent | Table 4 | Unverified |
Extraction coverage
What we read of this document, and where each value was read.
Our note First xAI card with an alignment audit (Petri 2.0)
All 18 values were read in the document itself.
All 18 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).