Versions

One version is on record: the copy we retrieved on 26 Sep 2026. We know of no other.

  1. 26 Sep 2026

    Date retrieved

    Copy retrieved 26 Sep 2026

    Date from
    the date we retrieved it; the copy states no version date
    Copy
    Copy retrieved on 26 Sep 2026
    Values
    39 values recorded from this version

No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).

Revisions

No revisions are recorded for this document. We know of only one version of it.

Values

Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. 3 of the 39 have been blind-verified: a second reader found the same value without seeing ours.

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

M1 Evaluation awareness

1 value

M1 Evaluation awareness: values in the Claude Haiku 4.5 System Card
ModelEvaluationConditionValueLocationChecked
Claude Haiku 4.5Automated behavioral auditShare of scenarios affected by verbalized eval awarenessNone statedSec 4.1.1Verified

M2 Reward hacking

9 values

M2 Reward hacking: values in the Claude Haiku 4.5 System Card
ModelEvaluationConditionValueLocationChecked
Claude Haiku 3.5Impossible tasksClassifier hack rateanti-hack promptTable 5.BUnverified
Claude Haiku 3.5Impossible tasksClassifier hack rateno anti-hack promptTable 5.BUnverified
Claude Haiku 3.5Reward-hack-prone coding tasksClassifier hack rateno anti-hack promptTable 5.BUnverified
Claude Haiku 3.5Reward-hack-prone coding tasksHidden-test hack rateno anti-hack promptTable 5.BUnverified
Claude Haiku 4.5Impossible tasksClassifier hack rateanti-hack promptTable 5.BUnverified
Claude Haiku 4.5Impossible tasksClassifier hack rateno anti-hack promptTable 5.BUnverified
Claude Haiku 4.5Reward-hack-prone coding tasksClassifier hack rateno anti-hack promptTable 5.BUnverified
Claude Haiku 4.5Reward-hack-prone coding tasksHidden-test hack rateno anti-hack promptTable 5.BUnverified
Claude Sonnet 4.5Impossible tasksClassifier hack rateno anti-hack promptTable 5.BUnverified

M7 Harmful compliance and over-refusal

17 values

M7 Harmful compliance and over-refusal: values in the Claude Haiku 4.5 System Card
ModelEvaluationConditionValueLocationChecked
Claude Haiku 3.5Single-turn violative requestsHarmless response rateoverall across thinking modesSec 2.1 tableUnverified
Claude Haiku 4.5Single-turn violative requestsHarmless response rateextended thinkingSec 2.1 tableUnverified
Claude Haiku 4.5Single-turn violative requestsHarmless response rateno extended thinkingSec 2.1 tableUnverified
Claude Sonnet 4.5Single-turn violative requestsHarmless response rateno extended thinkingSec 2.1 tableUnverified
Claude Haiku 3.5Single-turn benign requestsOver-refusal rateoverall across thinking modesSec 2.2 tableUnverified
Claude Haiku 4.5Single-turn benign requestsOver-refusal rateextended thinkingSec 2.2 tableVerified
Claude Haiku 4.5Single-turn benign requestsOver-refusal rateno extended thinkingSec 2.2 tableUnverified
Claude Haiku 3.5Malicious agentic codingSafety scorewithout safeguardsTable 3.1.1.AUnverified
Claude Haiku 4.5Malicious agentic codingSafety scorewithout safeguardsTable 3.1.1.AUnverified
Claude Sonnet 4.5Malicious agentic codingSafety scorewithout safeguardsTable 3.1.1.AUnverified
Claude Haiku 3.5Malicious Claude Code useRefusal rate, malicious requestswithout safeguardsTable 3.1.2.AUnverified
Claude Haiku 3.5Malicious Claude Code useSuccess rate, dual-use & benignwithout safeguardsTable 3.1.2.AUnverified
Claude Haiku 4.5Malicious Claude Code useRefusal rate, malicious requestswithout safeguardsTable 3.1.2.AUnverified
Claude Haiku 4.5Malicious Claude Code useSuccess rate, dual-use & benignwithout safeguardsTable 3.1.2.AUnverified
Claude Sonnet 4.5Malicious Claude Code useRefusal rate, malicious requestswithout safeguardsTable 3.1.2.AUnverified
Claude Sonnet 4.5Malicious Claude Code useSuccess rate, dual-use & benignwithout safeguardsTable 3.1.2.AVerified
Claude Haiku 4.5Malicious Claude Code useRefusal rate, malicious requestswith new mitigationsTable 3.1.2.BUnverified

M9 Prompt injection

6 values

M9 Prompt injection: values in the Claude Haiku 4.5 System Card
ModelEvaluationConditionValueLocationChecked
Claude Haiku 4.5Computer-use prompt injectionAttack prevention ratewith safeguardsTable 3.2.2.AUnverified
Claude Haiku 4.5Computer-use prompt injectionAttack prevention ratewithout safeguardsTable 3.2.2.AUnverified
Claude Haiku 3.5MCP prompt injectionAttack prevention ratewithout safeguardsTable 3.2.2.BUnverified
Claude Haiku 4.5MCP prompt injectionAttack prevention ratewithout safeguardsTable 3.2.2.BUnverified
Claude Haiku 3.5Tool-use prompt injectionAttack prevention ratewithout safeguardsTable 3.2.2.CUnverified
Claude Haiku 4.5Tool-use prompt injectionAttack prevention ratewithout safeguardsTable 3.2.2.CUnverified

M10 Dangerous capabilities and risk determinations

6 values

M10 Dangerous capabilities and risk determinations: values in the Claude Haiku 4.5 System Card
ModelEvaluationConditionValueLocationChecked
Claude Haiku 4.5RSP deployment standardASL standard deployed underNone statedSec 1 / Sec 6Unverified
Claude Haiku 4.5Long-form virology task 1ScoreNone statedSec 6.2.3.1Unverified
Claude Haiku 4.5SWE-bench Verified (hard subset)Problems solved, pass@1 avgof 45 problemsSec 6.3Unverified
Claude Sonnet 4SWE-bench Verified (hard subset)Problems solved, pass@1 avgof 42 problemsSec 6.3Unverified
Claude Haiku 4.5CybenchCTF challenges solvedof 32-challenge subsetSec 6.4Unverified
Claude Sonnet 4CybenchCTF challenges solvedof 32-challenge subsetSec 6.4Unverified

Extraction coverage

What we read of this document, and where each value was read.

Our note Read in full (~39 pages)

All 39 values were read in the document itself.

All 39 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).