Project
Changelog
What changed in the dataset, the methodology and the site, newest first. Corrections to values are listed here, and what we had recorded before stays in each value's history.
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Dataset releases
27 Sep 2026
Dataset v0.1
The first release. The version 0 feasibility dataset, one flat file, was moved into the related tables described in the data dictionary: documents and their versions, evaluations and their comparability groups, values, checks, revisions, restatements and risk determinations. Values keep the text their documents print. Every judgment the move needed was recorded with its reason, and a random sample of those judgments was decided again independently.
It holds 1,645 values from 70 documents, covering 87 models, plus two named only in values reported for several models at once (Llama 4 Maverick and Llama 4 Scout). A sample of 96 values was checked blind against the sources, and 95 matched. Checks of values chosen because the site shows them are counted apart (12 values).
Independent re-decision of the recorded judgments: 261 of 306 agreed (85.3%). Of the disagreements, 32 kept the first decision and 13 adopted the independent one, each with a recorded reason.
Methodology versions
27 Sep 2026
Methodology 0.1
First published methodology, covering the version 0 feasibility dataset.
Corrections to values
When a check finds that we recorded a value differently from how its document prints it, we correct the record, list the correction here, and keep what we had recorded in the value's history (Methodology §11). Each entry strikes through what we had recorded and underlines what the document prints. Select a value for its source and its checks.
KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.
27 Sep 2026
Claude Opus 4.8 · Verbalized grader speculation
Share of training episodes (RL training), in the Claude Opus 4.8 System Card, Sec 6.2.2
Recorded now
0.1%roughly 0.1%Corrected
Found by a blind check of the document on 27 Sep 2026. Changed: the qualifier (about, less than …).
Both blind checks of the card read "roughly 0.1%". The record, taken from a secondary write-up, had no qualifier; the card prints an approximation.
27 Sep 2026
Claude Opus 4.8 · Long-form virology task 2
Score, in the Claude Opus 4.8 System Card, Table 2.2.3.B
Recorded now
0.90.90Corrected
Found by a blind check of the document on 27 Sep 2026. Changed: the value as printed.
Both blind checks read "0.90" in the table, in the prose and in the change log. The record had dropped the printed trailing zero.
27 Sep 2026
Grok 4.6 · Bio capability lift
Lift vs Grok 4.5, in the Grok 4.6 Model Card, Dual-use summary
Recorded now
lift noted but limitednoted in the biological domain but is limitedCorrected
Found by a blind check of the document on 27 Sep 2026. Changed: the value as printed and the value in words.
Both blind checks read "noted in the biological domain but is limited". The recorded "lift noted but limited" was a paraphrase, not the printed wording.
The revision that records this change of wording was corrected with it (the new wording). See it in the revisions ledger
27 Sep 2026
How a revision of the GPT-6 Astra System Card is described
A change recorded in the GPT-6 Astra System Card, confirmed from the document's own changelog.
How we describe the change
Alignment section: "alignment faking" renamed "oversight gaming"; limitations expandedAlignment section clarified and its limitations expanded; section on verbalized metagaming and oversight gaming renamed and revisedCorrected
Changed: the description of the change, the note on how we know and the summary of the version.
Both checks found two change-log entries dated 9 Sep 2026: one on the Alignment section (how the evaluations were built; limitations expanded), one on naming and substance in the section on verbalized metagaming and oversight gaming. Neither names "alignment faking", which our description said was renamed. §8.7 says "oversight gaming" replaces "undermining evaluation validity", a term from OpenAI's previous system card. The revision is confirmed; its description is corrected to what the change log says.
Changes to the site
28 Sep 2026
Phase 2, stage 3: the pages that complete the core site
The home page and the first finding; the methodology as published; the source registry; downloads of every table with a data dictionary, checksums and suggested citations; this changelog; and the pages about the project, privacy, corrections and search.
28 Sep 2026
Phase 2, stage 2: overview and detail pages
A page for every metric family, model, developer and document, each opening with a summary and ending with where to go next, and the revisions ledger, which lists every recorded change to a document in redline with whether its changelog explains it.
28 Sep 2026
Data explorer: changes after the first review
Date controls moved under the chart; plainer words where a line joins documents on the assumption that the test did not change; controls that fit a phone; a Filters button that says how many values the filters leave; evaluations named with who ran them.
27 Sep 2026
Phase 2, stage 1: the data explorer
Every value in the dataset, one evaluation at a time, with lines only within one definition of a test. Any value opens its source, its checks and its history; the view is kept in the address; charts export as images and CSV.
27 Sep 2026
Phase 1: the data
The version 0 feasibility dataset moved into the tables the site is built from, each with a schema and checked on every build. Every judgment the move needed is recorded with its reason, and a random sample was decided a second time, independently.
26 Sep 2026
Phase 0: the design system
Type, colour, light and dark themes, the chart kit, tables, value chips and the details drawer that shows where each value comes from, collected in a component gallery that uses real rows from the data.
An RSS feed of this page is planned for Phase 3.