Claude Sonnet 4.5 System Card
A system card by Anthropic about Claude Sonnet 4.5, published 29 Sep 2025. We recorded 56 values from it.
- Developer
- Anthropic
- Type
- System card
- Model covered
- Claude Sonnet 4.5
- Published
- 29 Sep 2025
- Archived copy
- No archived copy yet
- Changelog
- Has a changelog
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Versions
Four versions are on record, oldest first. We hold a copy of one; the others are known only by their date.
29 Sep 2025
First published version
10 Oct 2025
Revision
What changed footnote author fix
3 Dec 2025
Revision
What changed edit to §9.3.5
26 Sep 2026
Date retrieved
Copy retrieved 26 Sep 2026
No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).
Revisions
No revisions are recorded for this document. Without copies of the earlier versions, changes between them are not recorded value by value; what we know of each version is listed below.
Values
Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. 5 of the 56 have been blind-verified: a second reader found the same value without seeing ours. 1 value is disputed.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M1 Evaluation awareness
6 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Sonnet 4.5 | Automated behavioral auditShare of transcripts verbalizing eval awareness | None stated | Sec 7.2 | Verified | |
| Claude Opus 4.1 | Evaluation awarenessUnprompted verbalized eval awareness rateRun by UK AI Security Institute | None stated | Sec 7.3.1 | Unverified | |
| Claude Sonnet 4 | Evaluation awarenessUnprompted verbalized eval awareness rateRun by UK AI Security Institute | None stated | Sec 7.3.1 | Unverified | |
| Claude Sonnet 4.5 | Evaluation awarenessUnprompted verbalized eval awareness rateRun by UK AI Security Institute | None stated | Sec 7.3.1 | Verified | |
| Claude Opus 4.1 | Evaluation awarenessUnambiguous references to being evaluatedRun by Apollo Research | None stated | Sec 7.3.2 | Unverified | |
| Claude Sonnet 4.5 | Evaluation awarenessUnambiguous references to being evaluatedRun by Apollo Research | None stated | Sec 7.3.2 | Verified |
M2 Reward hacking
17 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Opus 4 | Impossible tasksClassifier hack rate | anti-hack prompt | Table 6.1.A | Unverified | |
| Claude Opus 4 | Impossible tasksClassifier hack rate | no anti-hack prompt | Table 6.1.A | Unverified | |
| Claude Opus 4 | Reward-hack-prone coding tasksClassifier hack rate | no anti-hack prompt | Table 6.1.A | Unverified | |
| Claude Opus 4.1 | Impossible tasksClassifier hack rate | anti-hack prompt | Table 6.1.A | Unverified | |
| Claude Opus 4.1 | Impossible tasksClassifier hack rate | no anti-hack prompt | Table 6.1.A | Unverified | |
| Claude Opus 4.1 | Reward-hack-prone coding tasksClassifier hack rate | no anti-hack prompt | Table 6.1.A | Verified | |
| Claude Sonnet 3.7 | Impossible tasksClassifier hack rate | anti-hack prompt | Table 6.1.A | Unverified | |
| Claude Sonnet 3.7 | Impossible tasksClassifier hack rate | no anti-hack prompt | Table 6.1.A | Unverified | |
| Claude Sonnet 3.7 | Reward-hack-prone coding tasksClassifier hack rate | no anti-hack prompt | Table 6.1.A | Unverified | |
| Claude Sonnet 4 | Impossible tasksClassifier hack rate | anti-hack prompt | Table 6.1.A | Unverified | |
| Claude Sonnet 4 | Impossible tasksClassifier hack rate | no anti-hack prompt | Table 6.1.A | Unverified | |
| Claude Sonnet 4 | Reward-hack-prone coding tasksClassifier hack rate | no anti-hack prompt | Table 6.1.A | Unverified | |
| Claude Sonnet 4.5 | Impossible tasksClassifier hack rate | anti-hack prompt | Table 6.1.A | Unverified | |
| Claude Sonnet 4.5 | Impossible tasksClassifier hack rate | no anti-hack prompt | Table 6.1.A | Unverified | |
| Claude Sonnet 4.5 | Reward-hack-prone coding tasksClassifier hack rate | no anti-hack prompt | Table 6.1.A | Unverified | |
| Claude Sonnet 4.5 | Reward-hack-prone coding tasksHidden-test hack rate | no anti-hack prompt | Table 6.1.A | Unverified | |
| Claude Sonnet 4.5 | Training distributionClassifier hack rate | None stated | Table 6.1.A | Unverified |
M3 Sabotage and sandbagging
3 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Opus 4.1 | SHADE-ArenaStealth success rate | no extended thinking | Sec 7.5.10.3 | Unverified | |
| Claude Sonnet 4 | SHADE-ArenaStealth success rate | no extended thinking | Sec 7.5.10.3 | Unverified | |
| Claude Sonnet 4.5 | SHADE-ArenaStealth success rate | no extended thinking | Sec 7.5.10.3 | Unverified |
M4 Misalignment audits
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Sonnet 4.5 | Automated behavioral auditImprovement on overall misaligned-behavior score | vs Sonnet 4 | Sec 7.1.3.1 | Unverified |
M7 Harmful compliance and over-refusal
18 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Opus 4 | Single-turn violative requestsHarmless response rate | overall across thinking modes | Table 2.1.1.A | Unverified | |
| Claude Opus 4.1 | Single-turn violative requestsHarmless response rate | overall across thinking modes | Table 2.1.1.A | Unverified | |
| Claude Sonnet 4 | Single-turn violative requestsHarmless response rate | overall across thinking modes | Table 2.1.1.A | Unverified | |
| Claude Sonnet 4.5 | Single-turn violative requestsHarmless response rate | extended thinking | Table 2.1.1.A | Unverified | |
| Claude Sonnet 4.5 | Single-turn violative requestsHarmless response rate | no extended thinking | Table 2.1.1.A | Unverified | |
| Claude Sonnet 4.5 | Single-turn violative requestsHarmless response rate | overall across thinking modes | Table 2.1.1.A | Verified | |
| Claude Opus 4 | Single-turn benign requestsOver-refusal rate | overall across thinking modes | Table 2.1.2.A | Unverified | |
| Claude Opus 4.1 | Single-turn benign requestsOver-refusal rate | overall across thinking modes | Table 2.1.2.A | Unverified | |
| Claude Sonnet 4 | Single-turn benign requestsOver-refusal rate | overall across thinking modes | Table 2.1.2.A | Unverified | |
| Claude Sonnet 4.5 | Single-turn benign requestsOver-refusal rate | overall across thinking modes | Table 2.1.2.A | Unverified | |
| Claude Sonnet 4 | Malicious agentic codingSafety score | without safeguards | Table 4.1.1.A | Unverified | |
| Claude Sonnet 4.5 | Malicious agentic codingSafety score | without safeguards | Table 4.1.1.A | Unverified | |
| Claude Sonnet 4 | Malicious Claude Code useRefusal rate, covert malicious | without safeguards | Table 4.1.A | Unverified | |
| Claude Sonnet 4 | Malicious Claude Code useRefusal rate, overt malicious | without safeguards | Table 4.1.A | Unverified | |
| Claude Sonnet 4.5 | Malicious Claude Code useRefusal rate, covert malicious | without safeguards | Table 4.1.A | Unverified | |
| Claude Sonnet 4.5 | Malicious Claude Code useRefusal rate, overt malicious | without safeguards | Table 4.1.A | Unverified | |
| Claude Sonnet 4.5 | Malicious Claude Code useSuccess rate, dual-use requests | without safeguards | Table 4.1.A | Unverified | |
| Claude Sonnet 4.5 | Malicious Claude Code useRefusal rate, covert malicious | with mitigations | Table 4.1.B | Unverified |
M9 Prompt injection
9 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Sonnet 4 | MCP prompt injectionAttack prevention rate | without safeguards | Table 4.2.2.A | Unverified | |
| Claude Sonnet 4.5 | MCP prompt injectionAttack prevention rate | with safeguards | Table 4.2.2.A | Unverified | |
| Claude Sonnet 4.5 | MCP prompt injectionAttack prevention rate | without safeguards | Table 4.2.2.A | Unverified | |
| Claude Sonnet 4 | Computer-use prompt injectionAttack prevention rate | without safeguards | Table 4.2.3.A | Unverified | |
| Claude Sonnet 4.5 | Computer-use prompt injectionAttack prevention rate | with safeguards | Table 4.2.3.A | Unverified | |
| Claude Sonnet 4.5 | Computer-use prompt injectionAttack prevention rate | without safeguards | Table 4.2.3.A | Unverified | |
| Claude Sonnet 4 | Tool-use prompt injectionAttack prevention rate | without safeguards | Table 4.2.4.A | Unverified | |
| Claude Sonnet 4.5 | Tool-use prompt injectionAttack prevention rate | with safeguards | Table 4.2.4.A | Unverified | |
| Claude Sonnet 4.5 | Tool-use prompt injectionAttack prevention rate | without safeguards | Table 4.2.4.A | Unverified |
M10 Dangerous capabilities and risk determinations
2 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Sonnet 4.5 | RSP deployment standardASL standard deployed under | None stated | Sec 1.2 | Unverified | |
| Claude Sonnet 4.5 | CybenchSuccess rate, pass@30 | 37 of 40 challenges | Sec 5 | Disputed |
Extraction coverage
What we read of this document, and where each value was read.
Our note Web reader stopped at §7.2; alignment numbers partly from secondary write-ups
Of the 56 values, 47 were read in the document itself and 9 in independent write-ups that quote it.
Values read somewhere other than the document stand in where the document's own section could not be read directly, and are flagged on every value (Methodology §2.2).
Independent write-ups used
Of the 56 values, 1 was read off a figure; it is marked * and left out of charts by default.
All 56 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).