Versions

Three versions are on record, oldest first. We hold a copy of one; the others are known only by their date.

  1. 20 Aug 2025

    First published version

    Date from
    our source registry
    Copy
    Known to exist; no copy held
  2. 22 Aug 2025

    Revision

    Date from
    our source registry
    Copy
    Known to exist; no copy held

    What changed mentions of UK AISI removed (per Midas Project); no numeric changes documented

  3. 26 Sep 2026

    Date retrieved

    Copy retrieved 26 Sep 2026

    Date from
    the date we retrieved it; the copy states no version date
    Copy
    Copy retrieved on 26 Sep 2026
    Values
    17 values recorded from this version

No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).

Revisions

No revisions are recorded for this document. Without copies of the earlier versions, changes between them are not recorded value by value; what we know of each version is listed below.

Values

Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. None has been blind-verified yet.

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

M5 Honesty and hallucination

1 value

M5 Honesty and hallucination: values in the Grok 4 Model Card
ModelEvaluationConditionValueLocationChecked
Grok 4 (API)MASKDishonesty rateAPITable 2Unverified

M6 Sycophancy

1 value

M6 Sycophancy: values in the Grok 4 Model Card
ModelEvaluationConditionValueLocationChecked
Grok 4 (API)SycophancySycophancy rateAPITable 2Unverified

M7 Harmful compliance and over-refusal

2 values

M7 Harmful compliance and over-refusal: values in the Grok 4 Model Card
ModelEvaluationConditionValueLocationChecked
Grok 4 (API)AgentHarmAnswer rate on harmful agentic tasksAPI; no attackTable 1Unverified
Grok 4 (API)RefusalsAnswer rate on should-refuse promptssystem prompt with refusal policy; APITable 1Unverified

M8 Jailbreak robustness

3 values

M8 Jailbreak robustness: values in the Grok 4 Model Card
ModelEvaluationConditionValueLocationChecked
Grok 4 (API)JailbreaksAnswer rate on should-refuse promptsAPI; system prompt jailbreakTable 1Unverified
Grok 4 (API)JailbreaksAnswer rate on should-refuse promptssystem prompt with refusal policy; API; user jailbreakTable 1Unverified
Grok 4 (Web)JailbreaksAnswer rate on should-refuse promptsgrok.com; user jailbreakTable 1Unverified

M9 Prompt injection

1 value

M9 Prompt injection: values in the Grok 4 Model Card
ModelEvaluationConditionValueLocationChecked
Grok 4 (API)AgentDojoPrompt-injection attack success ratewith system prompt mitigation; APITable 1Unverified

M10 Dangerous capabilities and risk determinations

9 values

M10 Dangerous capabilities and risk determinations: values in the Grok 4 Model Card
ModelEvaluationConditionValueLocationChecked
Grok 4Overall risk conclusionwith safeguardsConclusionUnverified
Grok 4 (API)BioLP-BenchAccuracywithout safeguards; APITable 3Unverified
Grok 4 (API)CybenchUnguided success ratewithout safeguards; APITable 3Unverified
Grok 4 (API)MakeMeSayPersuasion win rateAPITable 3Unverified
Grok 4 (API)VCTAccuracywithout safeguards; APITable 3Unverified
Grok 4 (API)WMDPBio accuracywithout safeguards; APITable 3Unverified
Grok 4 (API)WMDPChem accuracywithout safeguards; APITable 3Unverified
Grok 4 (API)WMDPCyber accuracywithout safeguards; APITable 3Unverified
Grok 4 (Web)VCTAccuracygrok.comTable 3Unverified

Extraction coverage

What we read of this document, and where each value was read.

Our note We have not written a coverage note for this document yet.

All 17 values were read in the document itself.

All 17 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).