Grok 4 Model Card
A model card by xAI about Grok 4, published 20 Aug 2025. We recorded 17 values from it.
- Developer
- xAI
- Type
- Model card
- Model covered
- Grok 4
- Published
- 20 Aug 2025
- Archived copy
- No archived copy yet
- Changelog
- No changelog
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Versions
Three versions are on record, oldest first. We hold a copy of one; the others are known only by their date.
20 Aug 2025
First published version
22 Aug 2025
Revision
What changed mentions of UK AISI removed (per Midas Project); no numeric changes documented
26 Sep 2026
Date retrieved
Copy retrieved 26 Sep 2026
No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).
Revisions
No revisions are recorded for this document. Without copies of the earlier versions, changes between them are not recorded value by value; what we know of each version is listed below.
Values
Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. None has been blind-verified yet.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M5 Honesty and hallucination
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Grok 4 (API) | MASKDishonesty rate | API | Table 2 | Unverified |
M6 Sycophancy
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Grok 4 (API) | SycophancySycophancy rate | API | Table 2 | Unverified |
M7 Harmful compliance and over-refusal
2 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Grok 4 (API) | AgentHarmAnswer rate on harmful agentic tasks | API; no attack | Table 1 | Unverified | |
| Grok 4 (API) | RefusalsAnswer rate on should-refuse prompts | system prompt with refusal policy; API | Table 1 | Unverified |
M8 Jailbreak robustness
3 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Grok 4 (API) | JailbreaksAnswer rate on should-refuse prompts | API; system prompt jailbreak | Table 1 | Unverified | |
| Grok 4 (API) | JailbreaksAnswer rate on should-refuse prompts | system prompt with refusal policy; API; user jailbreak | Table 1 | Unverified | |
| Grok 4 (Web) | JailbreaksAnswer rate on should-refuse prompts | grok.com; user jailbreak | Table 1 | Unverified |
M9 Prompt injection
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Grok 4 (API) | AgentDojoPrompt-injection attack success rate | with system prompt mitigation; API | Table 1 | Unverified |
M10 Dangerous capabilities and risk determinations
9 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Grok 4 | Overall risk conclusion | with safeguards | Conclusion | Unverified | |
| Grok 4 (API) | BioLP-BenchAccuracy | without safeguards; API | Table 3 | Unverified | |
| Grok 4 (API) | CybenchUnguided success rate | without safeguards; API | Table 3 | Unverified | |
| Grok 4 (API) | MakeMeSayPersuasion win rate | API | Table 3 | Unverified | |
| Grok 4 (API) | VCTAccuracy | without safeguards; API | Table 3 | Unverified | |
| Grok 4 (API) | WMDPBio accuracy | without safeguards; API | Table 3 | Unverified | |
| Grok 4 (API) | WMDPChem accuracy | without safeguards; API | Table 3 | Unverified | |
| Grok 4 (API) | WMDPCyber accuracy | without safeguards; API | Table 3 | Unverified | |
| Grok 4 (Web) | VCTAccuracy | grok.com | Table 3 | Unverified |
Extraction coverage
What we read of this document, and where each value was read.
Our note We have not written a coverage note for this document yet.
All 17 values were read in the document itself.
All 17 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).