Claude Haiku 4.5 System Card
A system card by Anthropic about Claude Haiku 4.5, published 15 Oct 2025. We recorded 39 values from it.
- Developer
- Anthropic
- Type
- System card
- Model covered
- Claude Haiku 4.5
- Published
- 15 Oct 2025
- Archived copy
- No archived copy yet
- Changelog
- No changelog
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Versions
One version is on record: the copy we retrieved on 26 Sep 2026. We know of no other.
26 Sep 2026
Date retrieved
Copy retrieved 26 Sep 2026
No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).
Revisions
No revisions are recorded for this document. We know of only one version of it.
Values
Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. 3 of the 39 have been blind-verified: a second reader found the same value without seeing ours.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M1 Evaluation awareness
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Haiku 4.5 | Automated behavioral auditShare of scenarios affected by verbalized eval awareness | None stated | Sec 4.1.1 | Verified |
M2 Reward hacking
9 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Haiku 3.5 | Impossible tasksClassifier hack rate | anti-hack prompt | Table 5.B | Unverified | |
| Claude Haiku 3.5 | Impossible tasksClassifier hack rate | no anti-hack prompt | Table 5.B | Unverified | |
| Claude Haiku 3.5 | Reward-hack-prone coding tasksClassifier hack rate | no anti-hack prompt | Table 5.B | Unverified | |
| Claude Haiku 3.5 | Reward-hack-prone coding tasksHidden-test hack rate | no anti-hack prompt | Table 5.B | Unverified | |
| Claude Haiku 4.5 | Impossible tasksClassifier hack rate | anti-hack prompt | Table 5.B | Unverified | |
| Claude Haiku 4.5 | Impossible tasksClassifier hack rate | no anti-hack prompt | Table 5.B | Unverified | |
| Claude Haiku 4.5 | Reward-hack-prone coding tasksClassifier hack rate | no anti-hack prompt | Table 5.B | Unverified | |
| Claude Haiku 4.5 | Reward-hack-prone coding tasksHidden-test hack rate | no anti-hack prompt | Table 5.B | Unverified | |
| Claude Sonnet 4.5 | Impossible tasksClassifier hack rate | no anti-hack prompt | Table 5.B | Unverified |
M7 Harmful compliance and over-refusal
17 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Haiku 3.5 | Single-turn violative requestsHarmless response rate | overall across thinking modes | Sec 2.1 table | Unverified | |
| Claude Haiku 4.5 | Single-turn violative requestsHarmless response rate | extended thinking | Sec 2.1 table | Unverified | |
| Claude Haiku 4.5 | Single-turn violative requestsHarmless response rate | no extended thinking | Sec 2.1 table | Unverified | |
| Claude Sonnet 4.5 | Single-turn violative requestsHarmless response rate | no extended thinking | Sec 2.1 table | Unverified | |
| Claude Haiku 3.5 | Single-turn benign requestsOver-refusal rate | overall across thinking modes | Sec 2.2 table | Unverified | |
| Claude Haiku 4.5 | Single-turn benign requestsOver-refusal rate | extended thinking | Sec 2.2 table | Verified | |
| Claude Haiku 4.5 | Single-turn benign requestsOver-refusal rate | no extended thinking | Sec 2.2 table | Unverified | |
| Claude Haiku 3.5 | Malicious agentic codingSafety score | without safeguards | Table 3.1.1.A | Unverified | |
| Claude Haiku 4.5 | Malicious agentic codingSafety score | without safeguards | Table 3.1.1.A | Unverified | |
| Claude Sonnet 4.5 | Malicious agentic codingSafety score | without safeguards | Table 3.1.1.A | Unverified | |
| Claude Haiku 3.5 | Malicious Claude Code useRefusal rate, malicious requests | without safeguards | Table 3.1.2.A | Unverified | |
| Claude Haiku 3.5 | Malicious Claude Code useSuccess rate, dual-use & benign | without safeguards | Table 3.1.2.A | Unverified | |
| Claude Haiku 4.5 | Malicious Claude Code useRefusal rate, malicious requests | without safeguards | Table 3.1.2.A | Unverified | |
| Claude Haiku 4.5 | Malicious Claude Code useSuccess rate, dual-use & benign | without safeguards | Table 3.1.2.A | Unverified | |
| Claude Sonnet 4.5 | Malicious Claude Code useRefusal rate, malicious requests | without safeguards | Table 3.1.2.A | Unverified | |
| Claude Sonnet 4.5 | Malicious Claude Code useSuccess rate, dual-use & benign | without safeguards | Table 3.1.2.A | Verified | |
| Claude Haiku 4.5 | Malicious Claude Code useRefusal rate, malicious requests | with new mitigations | Table 3.1.2.B | Unverified |
M9 Prompt injection
6 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Haiku 4.5 | Computer-use prompt injectionAttack prevention rate | with safeguards | Table 3.2.2.A | Unverified | |
| Claude Haiku 4.5 | Computer-use prompt injectionAttack prevention rate | without safeguards | Table 3.2.2.A | Unverified | |
| Claude Haiku 3.5 | MCP prompt injectionAttack prevention rate | without safeguards | Table 3.2.2.B | Unverified | |
| Claude Haiku 4.5 | MCP prompt injectionAttack prevention rate | without safeguards | Table 3.2.2.B | Unverified | |
| Claude Haiku 3.5 | Tool-use prompt injectionAttack prevention rate | without safeguards | Table 3.2.2.C | Unverified | |
| Claude Haiku 4.5 | Tool-use prompt injectionAttack prevention rate | without safeguards | Table 3.2.2.C | Unverified |
M10 Dangerous capabilities and risk determinations
6 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Haiku 4.5 | RSP deployment standardASL standard deployed under | None stated | Sec 1 / Sec 6 | Unverified | |
| Claude Haiku 4.5 | Long-form virology task 1Score | None stated | Sec 6.2.3.1 | Unverified | |
| Claude Haiku 4.5 | SWE-bench Verified (hard subset)Problems solved, pass@1 avg | of 45 problems | Sec 6.3 | Unverified | |
| Claude Sonnet 4 | SWE-bench Verified (hard subset)Problems solved, pass@1 avg | of 42 problems | Sec 6.3 | Unverified | |
| Claude Haiku 4.5 | CybenchCTF challenges solved | of 32-challenge subset | Sec 6.4 | Unverified | |
| Claude Sonnet 4 | CybenchCTF challenges solved | of 32-challenge subset | Sec 6.4 | Unverified |
Extraction coverage
What we read of this document, and where each value was read.
Our note Read in full (~39 pages)
All 39 values were read in the document itself.
All 39 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).