Claude Opus 4.6 System Card
A system card by Anthropic about Claude Opus 4.6, published Feb 2026. We recorded 24 values from it.
- Developer
- Anthropic
- Type
- System card
- Model covered
- Claude Opus 4.6
- Published
- Feb 2026 (month only; the day is not known)
- Also at
- Another address, www-cdn.anthropic.com (opens an external site) (earlier version)
- Archived copy
- No archived copy yet
- Changelog
- Has a changelog
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Versions
Five versions are on record, oldest first. We hold a copy of one; the others are known only by their date.
6 Feb 2026
Revision
10 Feb 2026
Revision
17 Feb 2026
Revision
6 Mar 2026
Revision
What changed (benchmark corrections, wording)
26 Sep 2026
Date retrieved
Copy retrieved 26 Sep 2026
No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).
Revisions
No revisions are recorded for this document. Without copies of the earlier versions, changes between them are not recorded value by value; what we know of each version is listed below.
Values
Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. None has been blind-verified yet.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
M3 Sabotage and sandbagging
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Opus 4.6 | Sandbagging transcript reviewExplicit sandbagging instances found | 1,000 eval transcripts | Sec 6 | Unverified |
M5 Honesty and hallucination
1 value
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Opus 4.6 | Honesty evaluationWin rate | full thinking | Sec 4 | Unverified |
M7 Harmful compliance and over-refusal
12 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Opus 4.6 | Malicious Claude Code useRefusal rate, malicious requests | with system prompt + FileRead reminder | Sec 5.1.2 | Unverified | |
| Claude Opus 4.6 | Malicious computer useRefusal rate | None stated | Sec 5.1.3 | Unverified | |
| Claude Haiku 4.5 | Single-turn violative requestsHarmless response rate | overall across thinking modes | Table 3.1.1.A | Unverified | |
| Claude Opus 4.5 | Single-turn violative requestsHarmless response rate | overall across thinking modes | Table 3.1.1.A | Unverified | |
| Claude Opus 4.6 | Single-turn violative requestsHarmless response rate | overall across thinking modes | Table 3.1.1.A | Unverified | |
| Claude Sonnet 4.5 | Single-turn violative requestsHarmless response rate | overall across thinking modes | Table 3.1.1.A | Unverified | |
| Claude Opus 4.5 | Single-turn benign requestsOver-refusal rate | overall across thinking modes | Table 3.1.2.A | Unverified | |
| Claude Opus 4.6 | Single-turn benign requestsOver-refusal rate | overall across thinking modes | Table 3.1.2.A | Unverified | |
| Claude Sonnet 4.5 | Single-turn benign requestsOver-refusal rate | overall across thinking modes | Table 3.1.2.A | Unverified | |
| Claude Opus 4.6 | Higher-difficulty violative requestsHarmless response rate | overall across thinking modes | Table 3.1.3.1.A | Unverified | |
| Claude Opus 4.6 | Higher-difficulty benign requestsOver-refusal rate | overall across thinking modes | Table 3.1.3.2.A | Unverified | |
| Claude Sonnet 4.5 | Higher-difficulty benign requestsOver-refusal rate | overall across thinking modes | Table 3.1.3.2.A | Unverified |
M9 Prompt injection
4 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Opus 4.6 | Agent Red Teaming (ART)Attack success rate within k=100 attemptsRun by Gray Swan | extended thinking | Sec 5.2.1 | Unverified | |
| Claude Opus 4.6 | Agent Red Teaming (ART)Attack success rate within k=100 attemptsRun by Gray Swan | no extended thinking | Sec 5.2.1 | Unverified | |
| Claude Opus 4.6 | Coding prompt injection (adaptive attacker)Attack success rate | all conditions | Sec 5.2.2.1 | Unverified | |
| Claude Opus 4.6 | Browser use prompt injectionAttack success rate per attacked session | None stated | Sec 5.2.2.3 | Unverified |
M10 Dangerous capabilities and risk determinations
6 values
| Model | Evaluation | Condition | Value | Location | Checked |
|---|---|---|---|---|---|
| Claude Opus 4.6 | RSP deployment standardASL standard deployed under | None stated | Sec 1.2.3 | Unverified | |
| Claude Opus 4.6 | AI R&D-4 thresholdThreshold crossed | None stated | Sec 1.2.4.1 | Unverified | |
| Claude Opus 4.6 | Internal AI R&D surveyRespondents saying model could automate entry-level researcher | of 16 respondents | Sec 1.2.4.1 | Unverified | |
| Claude Opus 4.6 | CBRN-4 thresholdThreshold crossed | None stated | Sec 1.2.4.2 | Unverified | |
| Claude Opus 4.6 | CybenchSuccess rate, pass@30 | None stated | Sec 1.2.4.3 | Unverified | |
| Claude Opus 4.6 | CyberGymTargeted vuln reproduction rate | None stated | Sec 1.2.4.3 | Unverified |
Extraction coverage
What we read of this document, and where each value was read.
Our note Web reader stopped ~p.58; sabotage figures live in a separate Sabotage Risk Report (not extracted)
Of the 24 values, 16 were read in the document itself and 8 in independent write-ups that quote it.
Values read somewhere other than the document stand in where the document's own section could not be read directly, and are flagged on every value (Methodology §2.2).
Independent write-ups used
Of the 24 values, 1 is a statement in words rather than a number; it is marked * and left out of charts by default.
All 24 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).