Values by metric family

Each value as the document printed it. Select a value for its source, its checks and its history. Values from other documents are listed after the model’s own, each marked as the first report of that measure or as a restatement.

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

M5 · Honesty and hallucination 3 values

M5 Honesty and hallucination: values about Grok 4.6
Evaluation and metricValueDocument
From its own documentGrok 4.6 Model Card · 12 Aug 2026
Hallucination rateUnsupported claims, single-turnCondition: highGrok 4.6 Model Card12 Aug 2026
MASK-RectifiedDishonesty rateCondition: highEarlier versionGrok 4.6 Model Card12 Aug 2026Earlier version
MASK-RectifiedDishonesty rateCondition: highGrok 4.6 Model Card12 Aug 2026

M6 · Sycophancy 1 value

M6 Sycophancy: values about Grok 4.6
Evaluation and metricValueDocument
From its own documentGrok 4.6 Model Card · 12 Aug 2026
SycophancySycophancy rateCondition: highGrok 4.6 Model Card12 Aug 2026

M7 · Harmful compliance and over-refusal 10 values

M7 Harmful compliance and over-refusal: values about Grok 4.6
Evaluation and metricValueDocument
From its own documentGrok 4.6 Model Card · 12 Aug 2026
BioUseBenchSeverity-5 refusal rateCondition: highGrok 4.6 Model Card12 Aug 2026
FORTRESS-RNR/N refusal accuracyCondition: highGrok 4.6 Model Card12 Aug 2026
HackerBenchCompliance on harmful/dual-use cyber tasksCondition: highEarlier versionGrok 4.6 Model Card12 Aug 2026Earlier version
HackerBenchCompliance on harmful/dual-use cyber tasksCondition: highGrok 4.6 Model Card12 Aug 2026
HackerBenchFalse refusal on benign cyber tasksCondition: highEarlier versionGrok 4.6 Model Card12 Aug 2026Earlier version
HackerBenchFalse refusal on benign cyber tasksCondition: highGrok 4.6 Model Card12 Aug 2026
RefusalsCompliance on should-refuse promptsCondition: high; full safeguardsGrok 4.6 Model Card12 Aug 2026
Self-harm refusalsCompliance (non-refusal/redirect)Condition: highEarlier versionGrok 4.6 Model Card12 Aug 2026Earlier version
Self-harm refusalsCompliance (non-refusal/redirect)Condition: highGrok 4.6 Model Card12 Aug 2026
From other documents
HackerBenchCompliance on harmful/dual-use cyber tasksCondition: highGrok 4.7 Model Card21 Sep 2026First reportedGrok 4.7 Model Card21 Sep 2026First reported

M8 · Jailbreak robustness 2 values

M8 Jailbreak robustness: values about Grok 4.6
Evaluation and metricValueDocument
From its own documentGrok 4.6 Model Card · 12 Aug 2026
JailbreaksCompliance under attackCondition: high; standard jailbreaksGrok 4.6 Model Card12 Aug 2026
StrongREJECTCompliance under StrongREJECT attacksCondition: highGrok 4.6 Model Card12 Aug 2026

M10 · Dangerous capabilities and risk determinations 6 values

M10 Dangerous capabilities and risk determinations: values about Grok 4.6
Evaluation and metricValueDocument
From its own documentGrok 4.6 Model Card · 12 Aug 2026
Bio capability liftLift vs Grok 4.5Earlier versionGrok 4.6 Model Card12 Aug 2026Earlier version
Bio capability liftLift vs Grok 4.5Grok 4.6 Model Card12 Aug 2026
CVE-BenchRewardCondition: high; unrestrictedGrok 4.6 Model Card12 Aug 2026
CyberGymMean vulnerabilities reproducedCondition: high; unrestrictedGrok 4.6 Model Card12 Aug 2026
VCTAccuracyCondition: highGrok 4.6 Model Card12 Aug 2026
WMDPCyber accuracyCondition: highGrok 4.6 Model Card12 Aug 2026

Revisions

Changes made to a document after it was published that change a value about Grok 4.6.

  1. xAI · 12 Aug 2026 to 17 Aug 2026

    Grok 4.6 Model Card

    Bio capability lift statement relative to Grok 4.5, Grok 4.6

    M10 · Dangerous capabilities and risk determinations

    no appreciable liftnoted in the biological domain but is limited

    Not explained — The changelog does not mention this change.

    Old value: New value:

    How we know The old value was read in the earlier copy and the new one in the later copy. The FORTRESS-RN addition in the same revision is recorded separately.

    Checked Not yet checked

  2. xAI · 12 Aug 2026 to 17 Aug 2026

    Grok 4.6 Model Card

    FORTRESS-RN R/N refusal accuracy added, Grok 4.6 (high)

    M7 · Harmful compliance and over-refusal

    97.9%

    Not explained — Added without a changelog note.

    New value:

    How we know The value was read in the later copy; the earlier copy does not have it.

    Checked Not yet checked

  3. xAI · 12 Aug 2026 to 17 Aug 2026

    Grok 4.6 Model Card

    HackerBench v0.2 harmful/dual-use compliance, Grok 4.6 (high)

    M7 · Harmful compliance and over-refusal

    16.7%6.9%

    Partly explained — changelog lists "corrected eval results" for this eval; no reason

    Old value: New value:

    How we know The old value was read in the earlier copy and the new one in the later copy. Blind check 26 Sep 2026 (both versions)

    Checked Blind-verified

  4. xAI · 12 Aug 2026 to 17 Aug 2026

    Grok 4.6 Model Card

    MASK dishonesty, Grok 4.6 (high)

    M5 · Honesty and hallucination

    3.8%1.90%

    Partly explained — listed as corrected; no reason

    Old value: New value:

    How we know The old value was read in the earlier copy and the new one in the later copy. Blind check 26 Sep 2026 (both versions)

    Checked Blind-verified

  5. xAI · 12 Aug 2026 to 17 Aug 2026

    Grok 4.6 Model Card

    Self-harm compliance, Grok 4.6 (high)

    M7 · Harmful compliance and over-refusal

    3.7%0.84%

    Partly explained — listed as corrected; no reason

    Old value: New value:

    How we know The old value was read in the earlier copy and the new one in the later copy. Blind check 26 Sep 2026 (both versions)

    Checked Blind-verified