Versions

Five versions are on record, oldest first. We hold a copy of one; the others are known only by their date.

  1. 6 Feb 2026

    Revision

    Date from
    the document's changelog
    Copy
    Known from the changelog; no copy held
  2. 10 Feb 2026

    Revision

    Date from
    the document's changelog
    Copy
    Known from the changelog; no copy held
  3. 17 Feb 2026

    Revision

    Date from
    the document's changelog
    Copy
    Known from the changelog; no copy held
  4. 6 Mar 2026

    Revision

    Date from
    the document's changelog
    Copy
    Known from the changelog; no copy held

    What changed (benchmark corrections, wording)

  5. 26 Sep 2026

    Date retrieved

    Copy retrieved 26 Sep 2026

    Date from
    the date we retrieved it; the copy states no version date
    Copy
    Copy retrieved on 26 Sep 2026
    Values
    24 values recorded from this version

No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).

Revisions

No revisions are recorded for this document. Without copies of the earlier versions, changes between them are not recorded value by value; what we know of each version is listed below.

Values

Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. None has been blind-verified yet.

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

M3 Sabotage and sandbagging

1 value

M3 Sabotage and sandbagging: values in the Claude Opus 4.6 System Card
ModelEvaluationConditionValueLocationChecked
Claude Opus 4.6Sandbagging transcript reviewExplicit sandbagging instances found1,000 eval transcriptsSec 6Unverified

M5 Honesty and hallucination

1 value

M5 Honesty and hallucination: values in the Claude Opus 4.6 System Card
ModelEvaluationConditionValueLocationChecked
Claude Opus 4.6Honesty evaluationWin ratefull thinkingSec 4Unverified

M7 Harmful compliance and over-refusal

12 values

M7 Harmful compliance and over-refusal: values in the Claude Opus 4.6 System Card
ModelEvaluationConditionValueLocationChecked
Claude Opus 4.6Malicious Claude Code useRefusal rate, malicious requestswith system prompt + FileRead reminderSec 5.1.2Unverified
Claude Opus 4.6Malicious computer useRefusal rateNone statedSec 5.1.3Unverified
Claude Haiku 4.5Single-turn violative requestsHarmless response rateoverall across thinking modesTable 3.1.1.AUnverified
Claude Opus 4.5Single-turn violative requestsHarmless response rateoverall across thinking modesTable 3.1.1.AUnverified
Claude Opus 4.6Single-turn violative requestsHarmless response rateoverall across thinking modesTable 3.1.1.AUnverified
Claude Sonnet 4.5Single-turn violative requestsHarmless response rateoverall across thinking modesTable 3.1.1.AUnverified
Claude Opus 4.5Single-turn benign requestsOver-refusal rateoverall across thinking modesTable 3.1.2.AUnverified
Claude Opus 4.6Single-turn benign requestsOver-refusal rateoverall across thinking modesTable 3.1.2.AUnverified
Claude Sonnet 4.5Single-turn benign requestsOver-refusal rateoverall across thinking modesTable 3.1.2.AUnverified
Claude Opus 4.6Higher-difficulty violative requestsHarmless response rateoverall across thinking modesTable 3.1.3.1.AUnverified
Claude Opus 4.6Higher-difficulty benign requestsOver-refusal rateoverall across thinking modesTable 3.1.3.2.AUnverified
Claude Sonnet 4.5Higher-difficulty benign requestsOver-refusal rateoverall across thinking modesTable 3.1.3.2.AUnverified

M9 Prompt injection

4 values

M9 Prompt injection: values in the Claude Opus 4.6 System Card
ModelEvaluationConditionValueLocationChecked
Claude Opus 4.6Agent Red Teaming (ART)Attack success rate within k=100 attemptsRun by Gray Swanextended thinkingSec 5.2.1Unverified
Claude Opus 4.6Agent Red Teaming (ART)Attack success rate within k=100 attemptsRun by Gray Swanno extended thinkingSec 5.2.1Unverified
Claude Opus 4.6Coding prompt injection (adaptive attacker)Attack success rateall conditionsSec 5.2.2.1Unverified
Claude Opus 4.6Browser use prompt injectionAttack success rate per attacked sessionNone statedSec 5.2.2.3Unverified

M10 Dangerous capabilities and risk determinations

6 values

M10 Dangerous capabilities and risk determinations: values in the Claude Opus 4.6 System Card
ModelEvaluationConditionValueLocationChecked
Claude Opus 4.6RSP deployment standardASL standard deployed underNone statedSec 1.2.3Unverified
Claude Opus 4.6AI R&D-4 thresholdThreshold crossedNone statedSec 1.2.4.1Unverified
Claude Opus 4.6Internal AI R&D surveyRespondents saying model could automate entry-level researcherof 16 respondentsSec 1.2.4.1Unverified
Claude Opus 4.6CBRN-4 thresholdThreshold crossedNone statedSec 1.2.4.2Unverified
Claude Opus 4.6CybenchSuccess rate, pass@30None statedSec 1.2.4.3Unverified
Claude Opus 4.6CyberGymTargeted vuln reproduction rateNone statedSec 1.2.4.3Unverified

Extraction coverage

What we read of this document, and where each value was read.

Our note Web reader stopped ~p.58; sabotage figures live in a separate Sabotage Risk Report (not extracted)

Of the 24 values, 16 were read in the document itself and 8 in independent write-ups that quote it.

Values read somewhere other than the document stand in where the document's own section could not be read directly, and are flagged on every value (Methodology §2.2).

Of the 24 values, 1 is a statement in words rather than a number; it is marked * and left out of charts by default.

All 24 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).