Versions

Two versions are on record, oldest first. We hold copies of both.

  1. 14 Jul 2026

    Original version (read in a copy hosted elsewhere)

    Date from
    our source registry
    Copy
    Earlier copy read on 26 Sep 2026
    Values
    1 value recorded from this version
  2. 20 Jul 2026

    Revised version (the copy retrieved)

    Date from
    our source registry
    Copy
    Copy retrieved on 26 Sep 2026
    Values
    14 values recorded from this version

    What changed revision, no changelog; safety values appear unchanged

No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).

Revisions

No revisions are recorded for this document. We have recorded no change to a value between them.

Values

Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. None has been blind-verified yet.

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

M5 Honesty and hallucination

2 values

M5 Honesty and hallucination: values in the Grok 4.5 Model Card
ModelEvaluationConditionValueLocationChecked
Grok 4.5 (high)Hallucination rateUnsupported claims, single-turnhigh6.1Unverified
Grok 4.5 (high)MASK-RectifiedDishonesty ratehigh12.2Unverified

M6 Sycophancy

1 value

M6 Sycophancy: values in the Grok 4.5 Model Card
ModelEvaluationConditionValueLocationChecked
Grok 4.5 (high)SycophancySycophancy ratehigh12.3Unverified

M7 Harmful compliance and over-refusal

6 values

M7 Harmful compliance and over-refusal: values in the Grok 4.5 Model Card
ModelEvaluationConditionValueLocationChecked
Grok 4.5 (high)RefusalsCompliance on should-refuse promptshigh; full safeguards10.1Unverified
Grok 4.5 (high)CBRN refusalsRefusal accuracy on dangerous bio querieshigh; full safeguards10.3Unverified
Grok 4.5 (high)Self-harm refusalsCompliance (non-refusal/redirect)high11.1Unverified
GPT-5.5 (xhigh)HackerBenchCompliance on harmful/dual-use cyber tasksxhighCyber safeguards 7.2Unverified
Grok 4.5 (high)HackerBenchCompliance on harmful/dual-use cyber taskshighCyber safeguards 7.2Unverified
Grok 4.5 (high)HackerBenchFalse refusal on benign cyber taskshighCyber safeguards 7.2Unverified

M8 Jailbreak robustness

1 value

M8 Jailbreak robustness: values in the Grok 4.5 Model Card
ModelEvaluationConditionValueLocationChecked
Grok 4.5 (high)JailbreaksCompliance under attackhigh; standard jailbreaks9.1Unverified

M10 Dangerous capabilities and risk determinations

4 values

M10 Dangerous capabilities and risk determinations: values in the Grok 4.5 Model Card
ModelEvaluationConditionValueLocationChecked
Grok 4.5 (high)CyberGymMean vulnerabilities reproducedhigh; unrestricted7.1Unverified
Grok 4.5Dual-use bio determinationThreshold statementNone stated8Unverified
Grok 4.5 (high)VCTAccuracyhigh8.1Unverified
Grok 4.5 (high)ProtocolQA open-endedAccuracyhigh8.4Unverified

From an earlier version

This value was printed in the version of 14 Jul 2026 and replaced in a later version, which prints its own value for the same measurement (listed above). The changes we have recorded are under Revisions. Charts leave these values out by default.

Values from an earlier version of the Grok 4.5 Model Card
ModelFamilyEvaluationConditionValueLocationChecked
Grok 4.5 (high)M7HackerBenchCompliance on harmful/dual-use cyber taskshighCyber safeguards 7.2Unverified

Extraction coverage

What we read of this document, and where each value was read.

Our note Metric scales changed from 0–1 rates to percentages

All 15 values were read in the document itself.

All 15 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).