Dataset releases

  1. 27 Sep 2026

    Dataset v0.1

    Follows methodology version 0.1 · the current release

    The first release. The version 0 feasibility dataset, one flat file, was moved into the related tables described in the data dictionary: documents and their versions, evaluations and their comparability groups, values, checks, revisions, restatements and risk determinations. Values keep the text their documents print. Every judgment the move needed was recorded with its reason, and a random sample of those judgments was decided again independently.

    It holds 1,645 values from 70 documents, covering 87 models, plus two named only in values reported for several models at once (Llama 4 Maverick and Llama 4 Scout). A sample of 96 values was checked blind against the sources, and 95 matched. Checks of values chosen because the site shows them are counted apart (12 values).

    Independent re-decision of the recorded judgments: 261 of 306 agreed (85.3%). Of the disagreements, 32 kept the first decision and 13 adopted the independent one, each with a recorded reason.

Methodology versions

  1. 27 Sep 2026

    Methodology 0.1

    First published methodology, covering the version 0 feasibility dataset.

Corrections to values

When a check finds that we recorded a value differently from how its document prints it, we correct the record, list the correction here, and keep what we had recorded in the value's history (Methodology §11). Each entry strikes through what we had recorded and underlines what the document prints. Select a value for its source and its checks.

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

  1. 27 Sep 2026

    Claude Opus 4.8 · Verbalized grader speculation

    Share of training episodes (RL training), in the Claude Opus 4.8 System Card, Sec 6.2.2

    Recorded now

    0.1%roughly 0.1%

    Corrected

    Found by a blind check of the document on 27 Sep 2026. Changed: the qualifier (about, less than …).

    Both blind checks of the card read "roughly 0.1%". The record, taken from a secondary write-up, had no qualifier; the card prints an approximation.

  2. 27 Sep 2026

    Claude Opus 4.8 · Long-form virology task 2

    Score, in the Claude Opus 4.8 System Card, Table 2.2.3.B

    Recorded now

    0.90.90

    Corrected

    Found by a blind check of the document on 27 Sep 2026. Changed: the value as printed.

    Both blind checks read "0.90" in the table, in the prose and in the change log. The record had dropped the printed trailing zero.

  3. 27 Sep 2026

    Grok 4.6 · Bio capability lift

    Lift vs Grok 4.5, in the Grok 4.6 Model Card, Dual-use summary

    Recorded now

    lift noted but limitednoted in the biological domain but is limited

    Corrected

    Found by a blind check of the document on 27 Sep 2026. Changed: the value as printed and the value in words.

    Both blind checks read "noted in the biological domain but is limited". The recorded "lift noted but limited" was a paraphrase, not the printed wording.

    The revision that records this change of wording was corrected with it (the new wording). See it in the revisions ledger

  4. 27 Sep 2026

    How a revision of the GPT-6 Astra System Card is described

    A change recorded in the GPT-6 Astra System Card, confirmed from the document's own changelog.

    How we describe the change

    Alignment section: "alignment faking" renamed "oversight gaming"; limitations expandedAlignment section clarified and its limitations expanded; section on verbalized metagaming and oversight gaming renamed and revised

    Corrected

    Changed: the description of the change, the note on how we know and the summary of the version.

    Both checks found two change-log entries dated 9 Sep 2026: one on the Alignment section (how the evaluations were built; limitations expanded), one on naming and substance in the section on verbalized metagaming and oversight gaming. Neither names "alignment faking", which our description said was renamed. §8.7 says "oversight gaming" replaces "undermining evaluation validity", a term from OpenAI's previous system card. The revision is confirmed; its description is corrected to what the change log says.

    See the revision in the ledger

Changes to the site

  1. 28 Sep 2026

    Phase 2, stage 3: the pages that complete the core site

    The home page and the first finding; the methodology as published; the source registry; downloads of every table with a data dictionary, checksums and suggested citations; this changelog; and the pages about the project, privacy, corrections and search.

  2. 28 Sep 2026

    Phase 2, stage 2: overview and detail pages

    A page for every metric family, model, developer and document, each opening with a summary and ending with where to go next, and the revisions ledger, which lists every recorded change to a document in redline with whether its changelog explains it.

  3. 28 Sep 2026

    Data explorer: changes after the first review

    Date controls moved under the chart; plainer words where a line joins documents on the assumption that the test did not change; controls that fit a phone; a Filters button that says how many values the filters leave; evaluations named with who ran them.

  4. 27 Sep 2026

    Phase 2, stage 1: the data explorer

    Every value in the dataset, one evaluation at a time, with lines only within one definition of a test. Any value opens its source, its checks and its history; the view is kept in the address; charts export as images and CSV.

  5. 27 Sep 2026

    Phase 1: the data

    The version 0 feasibility dataset moved into the tables the site is built from, each with a schema and checked on every build. Every judgment the move needed is recorded with its reason, and a random sample was decided a second time, independently.

  6. 26 Sep 2026

    Phase 0: the design system

    Type, colour, light and dark themes, the chart kit, tables, value chips and the details drawer that shows where each value comes from, collected in a component gallery that uses real rows from the data.

An RSS feed of this page is planned for Phase 3.