Claude Opus 4 & Sonnet 4 System Card
A system card by Anthropic about Claude Opus 4 and Claude Sonnet 4, published 22 May 2025. We recorded 45 values from it.
- Developer
- Anthropic
- Type
- System card
- Models covered
- Claude Opus 4, Claude Sonnet 4
- Published
- 22 May 2025
- Also at
- Another address, anthropic.com (opens an external site) (redirect)
- Archived copy
- No archived copy yet
- Changelog
- Has a changelog
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Versions
Four versions are on record, oldest first. We hold a copy of one; the others are known only by their date.
22 May 2025
First published version
16 Jul 2025
Revision
What changed footnote and formatting
2 Sep 2025
Revision
What changed corrected Claude Code Impossible Tasks numbers
26 Sep 2026
Date retrieved
Copy retrieved 26 Sep 2026
No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).
Revisions
No revisions are recorded for this document. Without copies of the earlier versions, changes between them are not recorded value by value; what we know of each version is listed below.
Values
Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. 2 of the 45 have been blind-verified: a second reader found the same value without seeing ours.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M1 Evaluation awareness
2 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Opus 4 | Automated behavioral auditTranscripts flagged for situational awareness | of 414 audit transcripts | Sec 4.1.2.2 | Verified | |
| Claude Sonnet 3.7 | Automated behavioral auditTranscripts flagged for situational awareness | of 414 audit transcripts | Sec 4.1.2.2 | Unverified |
M2 Reward hacking
2 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Opus 4 | Reward hacking (hard-coding)Avg decrease in hard-coding vs Sonnet 3.7 | vs Sonnet 3.7 | Sec 6 | Unverified | |
| Claude Sonnet 4 | Reward hacking (hard-coding)Avg decrease in hard-coding vs Sonnet 3.7 | vs Sonnet 3.7 | Sec 6 | Unverified |
M3 Sabotage and sandbagging
4 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Opus 4 | Subtle sabotage (agentic)Side-task success without monitor detection | hidden scratchpad | Table 4.1.10.B | Unverified | |
| Claude Opus 4 | Subtle sabotage (agentic)Side-task success without monitor detection | scratchpad visible to monitor | Table 4.1.10.B | Unverified | |
| Claude Sonnet 3.7 | Subtle sabotage (agentic)Side-task success without monitor detection | hidden scratchpad | Table 4.1.10.B | Unverified | |
| Claude Sonnet 3.7 | Subtle sabotage (agentic)Side-task success without monitor detection | scratchpad visible to monitor | Table 4.1.10.B | Unverified |
M4 Misalignment audits
4 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Opus 4 | Opportunistic blackmail scenarioShare of rollouts with blackmail | replacement model shares values | Sec 4.1.1.2 | Unverified | |
| Claude Opus 4 | Automated behavioral auditSurprising high-impact actions found | audit transcripts | Sec 4.1.9 | Unverified | |
| Claude Sonnet 3.7 | Automated behavioral auditSurprising high-impact actions found | audit transcripts | Sec 4.1.9 | Unverified | |
| Claude Sonnet 4 | Automated behavioral auditSurprising high-impact actions found | audit transcripts | Sec 4.1.9 | Verified |
M5 Honesty and hallucination
4 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Opus 4 | Reasoning-behavior monitorKnowingly hallucinated information | RL-style transcripts | Table 4.1.6.A | Unverified | |
| Claude Opus 4 | Reasoning-behavior monitorOutputs showing deception (overall) | RL-style transcripts | Table 4.1.6.A | Unverified | |
| Claude Sonnet 3.7 | Reasoning-behavior monitorKnowingly hallucinated information | RL-style transcripts | Table 4.1.6.A | Unverified | |
| Claude Sonnet 3.7 | Reasoning-behavior monitorOutputs showing deception (overall) | RL-style transcripts | Table 4.1.6.A | Unverified |
M6 Sycophancy
2 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Opus 4 | Reasoning-behavior monitorOutputs showing sycophancy | RL-style transcripts | Table 4.1.6.A | Unverified | |
| Claude Sonnet 3.7 | Reasoning-behavior monitorOutputs showing sycophancy | RL-style transcripts | Table 4.1.6.A | Unverified |
M7 Harmful compliance and over-refusal
12 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Opus 4 | Single-turn violative requestsHarmless response rate | extended thinking | Table 2.1.A | Unverified | |
| Claude Opus 4 | Single-turn violative requestsHarmless response rate | no extended thinking | Table 2.1.A | Unverified | |
| Claude Opus 4 | Single-turn violative requestsHarmless response rate | overall across thinking modes | Table 2.1.A | Unverified | |
| Claude Opus 4 | Single-turn violative requestsHarmless response rate | overall across thinking modes; with ASL-3 safeguards | Table 2.1.A | Unverified | |
| Claude Sonnet 3.7 | Single-turn violative requestsHarmless response rate | overall across thinking modes | Table 2.1.A | Unverified | |
| Claude Sonnet 4 | Single-turn violative requestsHarmless response rate | overall across thinking modes | Table 2.1.A | Unverified | |
| Claude Opus 4 | Single-turn benign requestsOver-refusal rate | overall across thinking modes | Table 2.2.A | Unverified | |
| Claude Sonnet 3.7 | Single-turn benign requestsOver-refusal rate | overall across thinking modes | Table 2.2.A | Unverified | |
| Claude Sonnet 4 | Single-turn benign requestsOver-refusal rate | overall across thinking modes | Table 2.2.A | Unverified | |
| Claude Opus 4 | Malicious agentic codingSafety score | without safeguards | Table 3.3.A | Unverified | |
| Claude Sonnet 3.7 | Malicious agentic codingSafety score | without safeguards | Table 3.3.A | Unverified | |
| Claude Sonnet 4 | Malicious agentic codingSafety score | without safeguards | Table 3.3.A | Unverified |
M8 Jailbreak robustness
6 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Opus 4 | StrongREJECTBest-jailbreak StrongREJECT score | extended thinking | Table 2.7.A | Unverified | |
| Claude Opus 4 | StrongREJECTBest-jailbreak StrongREJECT score | no extended thinking | Table 2.7.A | Unverified | |
| Claude Sonnet 3.7 | StrongREJECTBest-jailbreak StrongREJECT score | extended thinking | Table 2.7.A | Unverified | |
| Claude Sonnet 3.7 | StrongREJECTBest-jailbreak StrongREJECT score | no extended thinking | Table 2.7.A | Unverified | |
| Claude Sonnet 4 | StrongREJECTBest-jailbreak StrongREJECT score | extended thinking | Table 2.7.A | Unverified | |
| Claude Sonnet 4 | StrongREJECTBest-jailbreak StrongREJECT score | no extended thinking | Table 2.7.A | Unverified |
M9 Prompt injection
6 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Opus 4 | Computer-use prompt injectionAttack prevention rate | with safeguards | Table 3.2.A | Unverified | |
| Claude Opus 4 | Computer-use prompt injectionAttack prevention rate | without safeguards | Table 3.2.A | Unverified | |
| Claude Sonnet 3.7 | Computer-use prompt injectionAttack prevention rate | with safeguards | Table 3.2.A | Unverified | |
| Claude Sonnet 3.7 | Computer-use prompt injectionAttack prevention rate | without safeguards | Table 3.2.A | Unverified | |
| Claude Sonnet 4 | Computer-use prompt injectionAttack prevention rate | with safeguards | Table 3.2.A | Unverified | |
| Claude Sonnet 4 | Computer-use prompt injectionAttack prevention rate | without safeguards | Table 3.2.A | Unverified |
M10 Dangerous capabilities and risk determinations
3 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Opus 4 | RSP deployment standardASL standard deployed under | None stated | Sec 1 / Sec 7 | Unverified | |
| Claude Sonnet 4 | RSP deployment standardASL standard deployed under | None stated | Sec 1 / Sec 7 | Unverified | |
| Claude Opus 4 | CybenchCTF challenges solved | of 39 challenges | Sec 7.4 | Unverified |
Extraction coverage
What we read of this document, and where each value was read.
Our note Web reader stopped at §4.2.1; later sections partly from secondary write-ups
Of the 45 values, 42 were read in the document itself and 3 in independent write-ups that quote it.
Values read somewhere other than the document stand in where the document's own section could not be read directly, and are flagged on every value (Methodology §2.2).
Independent write-ups used
- Simon Willison’s weblog (opens an external site) · 2 values
- The Stack (opens an external site) · 1 value
All 45 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).