Revisions

Each change we have recorded between two versions: the old and the new text, whether the document's changelog explains it, how we know, and whether we have checked it. A revision is not in itself evidence of wrongdoing; most are corrections (Methodology §7).

  1. Value changed · 12 Aug 2026 to 17 Aug 2026

    HackerBench v0.2 harmful/dual-use compliance, Grok 4.6 (high)

    M7 · Harmful compliance and over-refusal

    16.7%6.9%

    Partly explained — changelog lists "corrected eval results" for this eval; no reason

    Old value: New value:

    How we know The old value was read in the earlier copy and the new one in the later copy. Blind check 26 Sep 2026 (both versions)

    Checked Blind-verified

  2. Value changed · 12 Aug 2026 to 17 Aug 2026

    MASK dishonesty, Grok 4.6 (high)

    M5 · Honesty and hallucination

    3.8%1.90%

    Partly explained — listed as corrected; no reason

    Old value: New value:

    How we know The old value was read in the earlier copy and the new one in the later copy. Blind check 26 Sep 2026 (both versions)

    Checked Blind-verified

  3. Value changed · 12 Aug 2026 to 17 Aug 2026

    Self-harm compliance, Grok 4.6 (high)

    M7 · Harmful compliance and over-refusal

    3.7%0.84%

    Partly explained — listed as corrected; no reason

    Old value: New value:

    How we know The old value was read in the earlier copy and the new one in the later copy. Blind check 26 Sep 2026 (both versions)

    Checked Blind-verified

  4. Wording changed · 12 Aug 2026 to 17 Aug 2026

    Bio capability lift statement relative to Grok 4.5, Grok 4.6

    M10 · Dangerous capabilities and risk determinations

    no appreciable liftnoted in the biological domain but is limited

    Not explained — The changelog does not mention this change.

    Old value: New value:

    How we know The old value was read in the earlier copy and the new one in the later copy. The FORTRESS-RN addition in the same revision is recorded separately.

    Checked Not yet checked

  5. Added · 12 Aug 2026 to 17 Aug 2026

    FORTRESS-RN R/N refusal accuracy added, Grok 4.6 (high)

    M7 · Harmful compliance and over-refusal

    97.9%

    Not explained — Added without a changelog note.

    New value:

    How we know The value was read in the later copy; the earlier copy does not have it.

    Checked Not yet checked

Versions

Two versions are on record, oldest first. We hold copies of both.

  1. 12 Aug 2026

    Original version (read in a copy hosted elsewhere)

    Date from
    our source registry
    Copy
    Earlier copy read on 26 Sep 2026
    Values
    5 values recorded from this version
  2. 17 Aug 2026

    Revised version (the copy retrieved)

    Date from
    the document's changelog
    Copy
    Copy retrieved on 26 Sep 2026
    Values
    17 values recorded from this version

    What changed 5 changes recorded as revisions. See them in redline

No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).

Values

Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. 9 of the 22 have been blind-verified: a second reader found the same value without seeing ours. 1 value was corrected after review.

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

M5 Honesty and hallucination

2 values

M5 Honesty and hallucination: values in the Grok 4.6 Model Card
ModelEvaluationConditionValueLocationChecked
Grok 4.6 (high)Hallucination rateUnsupported claims, single-turnhigh6.1Unverified
Grok 4.6 (high)MASK-RectifiedDishonesty ratehigh12.1Verified

M6 Sycophancy

1 value

M6 Sycophancy: values in the Grok 4.6 Model Card
ModelEvaluationConditionValueLocationChecked
Grok 4.6 (high)SycophancySycophancy ratehigh12.2Unverified

M7 Harmful compliance and over-refusal

7 values

M7 Harmful compliance and over-refusal: values in the Grok 4.6 Model Card
ModelEvaluationConditionValueLocationChecked
Grok 4.6 (high)BioUseBenchSeverity-5 refusal ratehigh8.3Unverified
Grok 4.6 (high)RefusalsCompliance on should-refuse promptshigh; full safeguards10.1Unverified
Grok 4.6 (high)FORTRESS-RNR/N refusal accuracyhigh10.3Verified
Grok 4.6 (high)Self-harm refusalsCompliance (non-refusal/redirect)high11.1Verified
GPT-5.6 Sol (max)HackerBenchCompliance on harmful/dual-use cyber tasksmaxCyber 7.4Unverified
Grok 4.6 (high)HackerBenchCompliance on harmful/dual-use cyber taskshighCyber 7.4Verified
Grok 4.6 (high)HackerBenchFalse refusal on benign cyber taskshighCyber 7.4Verified

M8 Jailbreak robustness

2 values

M8 Jailbreak robustness: values in the Grok 4.6 Model Card
ModelEvaluationConditionValueLocationChecked
Grok 4.6 (high)JailbreaksCompliance under attackhigh; standard jailbreaks9.1Unverified
Grok 4.6 (high)StrongREJECTCompliance under StrongREJECT attackshigh9.1Unverified

M10 Dangerous capabilities and risk determinations

5 values

M10 Dangerous capabilities and risk determinations: values in the Grok 4.6 Model Card
ModelEvaluationConditionValueLocationChecked
Grok 4.6 (high)CyberGymMean vulnerabilities reproducedhigh; unrestricted7.1Unverified
Grok 4.6 (high)CVE-BenchRewardhigh; unrestricted7.2Unverified
Grok 4.6 (high)VCTAccuracyhigh8.1Unverified
Grok 4.6 (high)WMDPCyber accuracyhigh8.4Unverified
Grok 4.6Bio capability liftLift vs Grok 4.5None statedDual-use summaryCorrected

From an earlier version

These 5 values were printed in the version of 12 Aug 2026 and replaced in a later version, which prints its own value for each of the same measurements (listed above). The changes we have recorded are under Revisions. Charts leave these values out by default.

Values from an earlier version of the Grok 4.6 Model Card
ModelFamilyEvaluationConditionValueLocationChecked
Grok 4.6 (high)M5MASK-RectifiedDishonesty ratehigh12.1Verified
Grok 4.6 (high)M7Self-harm refusalsCompliance (non-refusal/redirect)high11.1Verified
Grok 4.6 (high)M7HackerBenchCompliance on harmful/dual-use cyber taskshighCyber 7.4Verified
Grok 4.6 (high)M7HackerBenchFalse refusal on benign cyber taskshighCyber 7.4Unverified
Grok 4.6M10Bio capability liftLift vs Grok 4.5None statedDual-use summaryVerified

Extraction coverage

What we read of this document, and where each value was read.

Our note See also themidasproject.com/watchtower/xai-08172026 (opens an external site)

All 22 values were read in the document itself.

Of the 22 values, 1 was read off a figure and 2 are statements in words rather than numbers; they are marked * and left out of charts by default.

All 22 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).