Values by metric family

Each value as the document printed it. Select a value for its source, its checks and its history. Values from other documents are listed after the model’s own, each marked as the first report of that measure or as a restatement.

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

M1 · Evaluation awareness 1 value

M1 Evaluation awareness: values about Grok 4
Evaluation and metricValueDocument
From other documents
Alignment auditVerbalized evaluation awareness rateCondition: across all auditsGrok 4.20 System Card7 Apr 2026First reportedGrok 4.20 System Card7 Apr 2026First reported

M5 · Honesty and hallucination 1 value

M5 Honesty and hallucination: values about Grok 4
Evaluation and metricValueDocument
From its own documentGrok 4 Model Card · 20 Aug 2025
MASKDishonesty rateCondition: APIGrok 4 Model Card20 Aug 2025

M6 · Sycophancy 1 value

M6 Sycophancy: values about Grok 4
Evaluation and metricValueDocument
From its own documentGrok 4 Model Card · 20 Aug 2025
SycophancySycophancy rateCondition: APIGrok 4 Model Card20 Aug 2025

M7 · Harmful compliance and over-refusal 2 values

M7 Harmful compliance and over-refusal: values about Grok 4
Evaluation and metricValueDocument
From its own documentGrok 4 Model Card · 20 Aug 2025
AgentHarmAnswer rate on harmful agentic tasksCondition: API; no attackGrok 4 Model Card20 Aug 2025
RefusalsAnswer rate on should-refuse promptsCondition: system prompt with refusal policy; APIGrok 4 Model Card20 Aug 2025

M8 · Jailbreak robustness 3 values

M8 Jailbreak robustness: values about Grok 4
Evaluation and metricValueDocument
From its own documentGrok 4 Model Card · 20 Aug 2025
JailbreaksAnswer rate on should-refuse promptsCondition: API; system prompt jailbreakGrok 4 Model Card20 Aug 2025
JailbreaksAnswer rate on should-refuse promptsCondition: grok.com; user jailbreakGrok 4 Model Card20 Aug 2025
JailbreaksAnswer rate on should-refuse promptsCondition: system prompt with refusal policy; API; user jailbreakGrok 4 Model Card20 Aug 2025

M9 · Prompt injection 1 value

M9 Prompt injection: values about Grok 4
Evaluation and metricValueDocument
From its own documentGrok 4 Model Card · 20 Aug 2025
AgentDojoPrompt-injection attack success rateCondition: with system prompt mitigation; APIGrok 4 Model Card20 Aug 2025

M10 · Dangerous capabilities and risk determinations 9 values

M10 Dangerous capabilities and risk determinations: values about Grok 4
Evaluation and metricValueDocument
From its own documentGrok 4 Model Card · 20 Aug 2025
BioLP-BenchAccuracyCondition: without safeguards; APIGrok 4 Model Card20 Aug 2025
CybenchUnguided success rateCondition: without safeguards; APIGrok 4 Model Card20 Aug 2025
MakeMeSayPersuasion win rateCondition: APIGrok 4 Model Card20 Aug 2025
Overall risk conclusionCondition: with safeguardsGrok 4 Model Card20 Aug 2025
VCTAccuracyCondition: grok.comGrok 4 Model Card20 Aug 2025
VCTAccuracyCondition: without safeguards; APIGrok 4 Model Card20 Aug 2025
WMDPBio accuracyCondition: without safeguards; APIGrok 4 Model Card20 Aug 2025
WMDPChem accuracyCondition: without safeguards; APIGrok 4 Model Card20 Aug 2025
WMDPCyber accuracyCondition: without safeguards; APIGrok 4 Model Card20 Aug 2025

Risk determinations

The developer's formal decisions about Grok 4 under its framework, as printed. Levels from different frameworks do not map onto one another.

Risk determinations about Grok 4
DomainLevel as printedFramework and document
Overalllow risk (malicious use and loss of control) No framework named Grok 4 Model Card · 20 Aug 2025No framework named Grok 4 Model Card · 20 Aug 2025