Versions

Four versions are on record, oldest first. We hold a copy of one; the others are known only by their date.

  1. 29 Sep 2025

    First published version

    Date from
    our source registry
    Copy
    Known to exist; no copy held
  2. 10 Oct 2025

    Revision

    Date from
    the document's changelog
    Copy
    Known from the changelog; no copy held

    What changed footnote author fix

  3. 3 Dec 2025

    Revision

    Date from
    the document's changelog
    Copy
    Known from the changelog; no copy held

    What changed edit to §9.3.5

  4. 26 Sep 2026

    Date retrieved

    Copy retrieved 26 Sep 2026

    Date from
    the date we retrieved it; the copy states no version date
    Copy
    Copy retrieved on 26 Sep 2026
    Values
    56 values recorded from this version

No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).

Revisions

No revisions are recorded for this document. Without copies of the earlier versions, changes between them are not recorded value by value; what we know of each version is listed below.

Values

Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. 5 of the 56 have been blind-verified: a second reader found the same value without seeing ours. 1 value is disputed.

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

M1 Evaluation awareness

6 values

M1 Evaluation awareness: values in the Claude Sonnet 4.5 System Card
ModelEvaluationConditionValueLocationChecked
Claude Sonnet 4.5Automated behavioral auditShare of transcripts verbalizing eval awarenessNone statedSec 7.2Verified
Claude Opus 4.1Evaluation awarenessUnprompted verbalized eval awareness rateRun by UK AI Security InstituteNone statedSec 7.3.1Unverified
Claude Sonnet 4Evaluation awarenessUnprompted verbalized eval awareness rateRun by UK AI Security InstituteNone statedSec 7.3.1Unverified
Claude Sonnet 4.5Evaluation awarenessUnprompted verbalized eval awareness rateRun by UK AI Security InstituteNone statedSec 7.3.1Verified
Claude Opus 4.1Evaluation awarenessUnambiguous references to being evaluatedRun by Apollo ResearchNone statedSec 7.3.2Unverified
Claude Sonnet 4.5Evaluation awarenessUnambiguous references to being evaluatedRun by Apollo ResearchNone statedSec 7.3.2Verified

M2 Reward hacking

17 values

M2 Reward hacking: values in the Claude Sonnet 4.5 System Card
ModelEvaluationConditionValueLocationChecked
Claude Opus 4Impossible tasksClassifier hack rateanti-hack promptTable 6.1.AUnverified
Claude Opus 4Impossible tasksClassifier hack rateno anti-hack promptTable 6.1.AUnverified
Claude Opus 4Reward-hack-prone coding tasksClassifier hack rateno anti-hack promptTable 6.1.AUnverified
Claude Opus 4.1Impossible tasksClassifier hack rateanti-hack promptTable 6.1.AUnverified
Claude Opus 4.1Impossible tasksClassifier hack rateno anti-hack promptTable 6.1.AUnverified
Claude Opus 4.1Reward-hack-prone coding tasksClassifier hack rateno anti-hack promptTable 6.1.AVerified
Claude Sonnet 3.7Impossible tasksClassifier hack rateanti-hack promptTable 6.1.AUnverified
Claude Sonnet 3.7Impossible tasksClassifier hack rateno anti-hack promptTable 6.1.AUnverified
Claude Sonnet 3.7Reward-hack-prone coding tasksClassifier hack rateno anti-hack promptTable 6.1.AUnverified
Claude Sonnet 4Impossible tasksClassifier hack rateanti-hack promptTable 6.1.AUnverified
Claude Sonnet 4Impossible tasksClassifier hack rateno anti-hack promptTable 6.1.AUnverified
Claude Sonnet 4Reward-hack-prone coding tasksClassifier hack rateno anti-hack promptTable 6.1.AUnverified
Claude Sonnet 4.5Impossible tasksClassifier hack rateanti-hack promptTable 6.1.AUnverified
Claude Sonnet 4.5Impossible tasksClassifier hack rateno anti-hack promptTable 6.1.AUnverified
Claude Sonnet 4.5Reward-hack-prone coding tasksClassifier hack rateno anti-hack promptTable 6.1.AUnverified
Claude Sonnet 4.5Reward-hack-prone coding tasksHidden-test hack rateno anti-hack promptTable 6.1.AUnverified
Claude Sonnet 4.5Training distributionClassifier hack rateNone statedTable 6.1.AUnverified

M3 Sabotage and sandbagging

3 values

M3 Sabotage and sandbagging: values in the Claude Sonnet 4.5 System Card
ModelEvaluationConditionValueLocationChecked
Claude Opus 4.1SHADE-ArenaStealth success rateno extended thinkingSec 7.5.10.3Unverified
Claude Sonnet 4SHADE-ArenaStealth success rateno extended thinkingSec 7.5.10.3Unverified
Claude Sonnet 4.5SHADE-ArenaStealth success rateno extended thinkingSec 7.5.10.3Unverified

M4 Misalignment audits

1 value

M4 Misalignment audits: values in the Claude Sonnet 4.5 System Card
ModelEvaluationConditionValueLocationChecked
Claude Sonnet 4.5Automated behavioral auditImprovement on overall misaligned-behavior scorevs Sonnet 4Sec 7.1.3.1Unverified

M7 Harmful compliance and over-refusal

18 values

M7 Harmful compliance and over-refusal: values in the Claude Sonnet 4.5 System Card
ModelEvaluationConditionValueLocationChecked
Claude Opus 4Single-turn violative requestsHarmless response rateoverall across thinking modesTable 2.1.1.AUnverified
Claude Opus 4.1Single-turn violative requestsHarmless response rateoverall across thinking modesTable 2.1.1.AUnverified
Claude Sonnet 4Single-turn violative requestsHarmless response rateoverall across thinking modesTable 2.1.1.AUnverified
Claude Sonnet 4.5Single-turn violative requestsHarmless response rateextended thinkingTable 2.1.1.AUnverified
Claude Sonnet 4.5Single-turn violative requestsHarmless response rateno extended thinkingTable 2.1.1.AUnverified
Claude Sonnet 4.5Single-turn violative requestsHarmless response rateoverall across thinking modesTable 2.1.1.AVerified
Claude Opus 4Single-turn benign requestsOver-refusal rateoverall across thinking modesTable 2.1.2.AUnverified
Claude Opus 4.1Single-turn benign requestsOver-refusal rateoverall across thinking modesTable 2.1.2.AUnverified
Claude Sonnet 4Single-turn benign requestsOver-refusal rateoverall across thinking modesTable 2.1.2.AUnverified
Claude Sonnet 4.5Single-turn benign requestsOver-refusal rateoverall across thinking modesTable 2.1.2.AUnverified
Claude Sonnet 4Malicious agentic codingSafety scorewithout safeguardsTable 4.1.1.AUnverified
Claude Sonnet 4.5Malicious agentic codingSafety scorewithout safeguardsTable 4.1.1.AUnverified
Claude Sonnet 4Malicious Claude Code useRefusal rate, covert maliciouswithout safeguardsTable 4.1.AUnverified
Claude Sonnet 4Malicious Claude Code useRefusal rate, overt maliciouswithout safeguardsTable 4.1.AUnverified
Claude Sonnet 4.5Malicious Claude Code useRefusal rate, covert maliciouswithout safeguardsTable 4.1.AUnverified
Claude Sonnet 4.5Malicious Claude Code useRefusal rate, overt maliciouswithout safeguardsTable 4.1.AUnverified
Claude Sonnet 4.5Malicious Claude Code useSuccess rate, dual-use requestswithout safeguardsTable 4.1.AUnverified
Claude Sonnet 4.5Malicious Claude Code useRefusal rate, covert maliciouswith mitigationsTable 4.1.BUnverified

M9 Prompt injection

9 values

M9 Prompt injection: values in the Claude Sonnet 4.5 System Card
ModelEvaluationConditionValueLocationChecked
Claude Sonnet 4MCP prompt injectionAttack prevention ratewithout safeguardsTable 4.2.2.AUnverified
Claude Sonnet 4.5MCP prompt injectionAttack prevention ratewith safeguardsTable 4.2.2.AUnverified
Claude Sonnet 4.5MCP prompt injectionAttack prevention ratewithout safeguardsTable 4.2.2.AUnverified
Claude Sonnet 4Computer-use prompt injectionAttack prevention ratewithout safeguardsTable 4.2.3.AUnverified
Claude Sonnet 4.5Computer-use prompt injectionAttack prevention ratewith safeguardsTable 4.2.3.AUnverified
Claude Sonnet 4.5Computer-use prompt injectionAttack prevention ratewithout safeguardsTable 4.2.3.AUnverified
Claude Sonnet 4Tool-use prompt injectionAttack prevention ratewithout safeguardsTable 4.2.4.AUnverified
Claude Sonnet 4.5Tool-use prompt injectionAttack prevention ratewith safeguardsTable 4.2.4.AUnverified
Claude Sonnet 4.5Tool-use prompt injectionAttack prevention ratewithout safeguardsTable 4.2.4.AUnverified

M10 Dangerous capabilities and risk determinations

2 values

M10 Dangerous capabilities and risk determinations: values in the Claude Sonnet 4.5 System Card
ModelEvaluationConditionValueLocationChecked
Claude Sonnet 4.5RSP deployment standardASL standard deployed underNone statedSec 1.2Unverified
Claude Sonnet 4.5CybenchSuccess rate, pass@3037 of 40 challengesSec 5Disputed

Extraction coverage

What we read of this document, and where each value was read.

Our note Web reader stopped at §7.2; alignment numbers partly from secondary write-ups

Of the 56 values, 47 were read in the document itself and 9 in independent write-ups that quote it.

Values read somewhere other than the document stand in where the document's own section could not be read directly, and are flagged on every value (Methodology §2.2).

Of the 56 values, 1 was read off a figure; it is marked * and left out of charts by default.

All 56 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).