Claude Opus 4.1 System Card (addendum)
A system card addendum by Anthropic about Claude Opus 4.1, published 5 Aug 2025. We recorded 34 values from it.
- Developer
- Anthropic
- Type
- System card addendum
- Model covered
- Claude Opus 4.1
- Published
- 5 Aug 2025
- Archived copy
- No archived copy yet
- Changelog
- Has a changelog
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Versions
Three versions are on record, oldest first. We hold a copy of one; the others are known only by their date.
5 Aug 2025
First published version
15 Sep 2025
Revision
What changed CBRN partner acknowledgments
26 Sep 2026
Date retrieved
Copy retrieved 26 Sep 2026
No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).
Revisions
No revisions are recorded for this document. Without copies of the earlier versions, changes between them are not recorded value by value; what we know of each version is listed below.
Values
Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. 1 of the 34 has been blind-verified: a second reader found the same value without seeing ours.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M2 Reward hacking
18 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Opus 4 | Impossible tasksClassifier hack rate | anti-hack prompt | Table 5.B | Unverified | |
| Claude Opus 4 | Impossible tasksClassifier hack rate | no anti-hack prompt | Table 5.B | Unverified | |
| Claude Opus 4 | Reward-hack-prone coding tasksClassifier hack rate | no anti-hack prompt | Table 5.B | Unverified | |
| Claude Opus 4 | Reward-hack-prone coding tasksHidden-test hack rate | no anti-hack prompt | Table 5.B | Unverified | |
| Claude Opus 4.1 | Impossible tasksClassifier hack rate | anti-hack prompt | Table 5.B | Unverified | |
| Claude Opus 4.1 | Impossible tasksClassifier hack rate | no anti-hack prompt | Table 5.B | Unverified | |
| Claude Opus 4.1 | Reward-hack-prone coding tasksClassifier hack rate | no anti-hack prompt | Table 5.B | Unverified | |
| Claude Opus 4.1 | Reward-hack-prone coding tasksHidden-test hack rate | no anti-hack prompt | Table 5.B | Unverified | |
| Claude Opus 4.1 | Training distributionClassifier hack rate | environment 1 | Table 5.B | Unverified | |
| Claude Opus 4.1 | Training distributionClassifier hack rate | environment 2 | Table 5.B | Unverified | |
| Claude Sonnet 3.7 | Impossible tasksClassifier hack rate | anti-hack prompt | Table 5.B | Unverified | |
| Claude Sonnet 3.7 | Impossible tasksClassifier hack rate | no anti-hack prompt | Table 5.B | Unverified | |
| Claude Sonnet 3.7 | Reward-hack-prone coding tasksClassifier hack rate | no anti-hack prompt | Table 5.B | Unverified | |
| Claude Sonnet 3.7 | Reward-hack-prone coding tasksHidden-test hack rate | no anti-hack prompt | Table 5.B | Unverified | |
| Claude Sonnet 4 | Impossible tasksClassifier hack rate | anti-hack prompt | Table 5.B | Unverified | |
| Claude Sonnet 4 | Impossible tasksClassifier hack rate | no anti-hack prompt | Table 5.B | Unverified | |
| Claude Sonnet 4 | Reward-hack-prone coding tasksClassifier hack rate | no anti-hack prompt | Table 5.B | Unverified | |
| Claude Sonnet 4 | Reward-hack-prone coding tasksHidden-test hack rate | no anti-hack prompt | Table 5.B | Unverified |
M4 Misalignment audits
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Opus 4.1 | Automated behavioral auditReduction in cooperation with egregious misuse | vs Opus 4; share of transcripts scoring >5/10 | Sec 4.1 / Fig 4.1.A | Unverified |
M7 Harmful compliance and over-refusal
6 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Opus 4 | Single-turn violative requestsHarmless response rate | overall across thinking modes | Table 2.1.A | Unverified | |
| Claude Opus 4.1 | Single-turn violative requestsHarmless response rate | extended thinking | Table 2.1.A | Unverified | |
| Claude Opus 4.1 | Single-turn violative requestsHarmless response rate | no extended thinking | Table 2.1.A | Unverified | |
| Claude Opus 4.1 | Single-turn violative requestsHarmless response rate | overall across thinking modes | Table 2.1.A | Unverified | |
| Claude Opus 4 | Single-turn benign requestsOver-refusal rate | overall across thinking modes | Table 2.1.B | Unverified | |
| Claude Opus 4.1 | Single-turn benign requestsOver-refusal rate | overall across thinking modes | Table 2.1.B | Unverified |
M10 Dangerous capabilities and risk determinations
9 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Opus 4.1 | RSP deployment standardASL standard deployed under | None stated | Sec 6 | Unverified | |
| Claude Opus 4 | Long-form virology task 1Score | None stated | Sec 6.3.1 | Unverified | |
| Claude Opus 4.1 | Long-form virology task 1Score | None stated | Sec 6.3.1 | Unverified | |
| Claude Opus 4 | AI R&D kernel taskBest speedup | None stated | Sec 6.4.1 | Verified | |
| Claude Opus 4 | SWE-bench Verified (hard subset)Problems solved, pass@1 avg | of 42 problems | Sec 6.4.1 | Unverified | |
| Claude Opus 4.1 | AI R&D kernel taskBest speedup | None stated | Sec 6.4.1 | Unverified | |
| Claude Opus 4.1 | SWE-bench Verified (hard subset)Problems solved, pass@1 avg | of 42 problems | Sec 6.4.1 | Unverified | |
| Claude Opus 4 | CybenchCTF challenges solved | of 35-challenge subset | Sec 6.5.1 | Unverified | |
| Claude Opus 4.1 | CybenchCTF challenges solved | of 35-challenge subset | Sec 6.5.1 | Unverified |
Extraction coverage
What we read of this document, and where each value was read.
Our note Read in full; carries corrected Opus 4 / Sonnet 4 reward-hacking numbers
All 34 values were read in the document itself.
All 34 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).