Data
Download the data
Every table of the dataset as CSV, in one zip or as one JSON file, with a data dictionary, checksums and suggested citations. Free to reuse under CC BY 4.0.
Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog
Current release
- Dataset
- v0.1
- Released
- 27 Sep 2026
- Methodology
- Version 0.1
Dataset v0.1 holds 1,645 values read in 70 documents by 8 developers. They cover 87 models, plus two named only in values reported for several models at once (Llama 4 Maverick and Llama 4 Scout), and 303 evaluations in 12 metric families. Each value is recorded as the document prints it, with where it was found.
Before release, a sample of 96 values was looked up again by a check that could not see the recorded value (Methodology §4.2), and 95 matched. A further 12 values were checked because the site shows them individually; those are counted apart, because values picked for a reason are not a sample. Values corrected after a check are listed in the changelog.
Judgment calls in preparing the data, such as which names refer to one model or which values form one series, are recorded with a reason. A random sample of them was decided a second time by an independent pass that could not see the first answer: 261 of 306 agreed (85.3%), and each disagreement was settled with a recorded reason.
Or every table in one JSON file (2.5 MB), or each table as CSV.
Files and checksums
Each CSV file is a table exactly as the dataset keeps it: UTF-8, comma-separated, one header row, list fields separated by semicolons. Each checksum is computed from the file as served.
| File | What it holds | Contents | Size | SHA-256 |
|---|---|---|---|---|
| The whole dataset | ||||
| safety-card-ledger-v0.1.zip Everything, zipped SHA-256 089fcf1e097a017becee2dcf46ac826768f4908de67983abd58aa70087b6049e | Everything, zippedEvery table as CSV with its schema, the source registry, a read-me, the licence and CITATION.cff. | 34 files | 176 KB | 089fcf1e097a017becee2dcf46ac826768f4908de67983abd58aa70087b6049e |
| safety-card-ledger-v0.1.json Every table in one JSON file SHA-256 236c93b0d758b63bbb60f384945abacf37ba772b2c5efd8579328ea667ee9035 | Every table in one JSON fileThe same tables, typed from the schemas, with the release, methodology version and licence in a header. | 16 tables | 2.5 MB | 236c93b0d758b63bbb60f384945abacf37ba772b2c5efd8579328ea667ee9035 |
| One table per file (CSV) | ||||
| labs.csv Developers SHA-256 b58fdab6ce5fa4afda1a2f6c8a0a99dbfdafb6be66584209cbe72c564696001e | DevelopersOne row per AI developer whose documents are in the dataset.Fields of labs.csv | 8 rows | 567 bytes | b58fdab6ce5fa4afda1a2f6c8a0a99dbfdafb6be66584209cbe72c564696001e |
| models.csv Models SHA-256 6a7862f630196d52a0f9152262c63bd993dbfcf3cfcccd762ab9ae9761da3269 | ModelsOne row per model. Snapshots (pre-release checkpoints, dated updates, previews) are not separate models; they are recorded on each measurement as snapshot_label.Fields of models.csv | 89 rows | 16 KB | 6a7862f630196d52a0f9152262c63bd993dbfcf3cfcccd762ab9ae9761da3269 |
| documents.csv Documents SHA-256 34da1e9891383442d93a319bccebf4d9e549f546fd69c98699cb2667feda9c44 | DocumentsOne row per developer document (system card, model card, framework report, technical report, launch or policy post). Values are stored against the document they were read in.Fields of documents.csv | 70 rows | 21 KB | 34da1e9891383442d93a319bccebf4d9e549f546fd69c98699cb2667feda9c44 |
| document_versions.csv Document versions SHA-256 ac95932485001f5a98f9794f618546bb1e1f95ae9f520e8dcdc194cf00e6ffc5 | Document versionsOne row per known state of a document: the copy we retrieved, earlier copies we read, and states known only from a changelog or the publication date.Fields of document_versions.csv | 126 rows | 35 KB | ac95932485001f5a98f9794f618546bb1e1f95ae9f520e8dcdc194cf00e6ffc5 |
| secondary_sources.csv Secondary sources SHA-256 1f7502710e25bae17f597fb9ead5db5ec48501c66fe179f8110107738f5525e5 | Secondary sourcesIndependent write-ups and developer summary pages that some values were read from, because the card section itself could not be read in version 0.Fields of secondary_sources.csv | 25 rows | 9.4 KB | 1f7502710e25bae17f597fb9ead5db5ec48501c66fe179f8110107738f5525e5 |
| evaluators.csv Evaluators SHA-256 8db7dcbc2e8cac129829abad258ec03d93a4ddb19fccb81a7d0bbaf4678ff1d0 | EvaluatorsWho ran an evaluation: the developer itself, a third party, or a government body.Fields of evaluators.csv | 14 rows | 492 bytes | 8db7dcbc2e8cac129829abad258ec03d93a4ddb19fccb81a7d0bbaf4678ff1d0 |
| evals.csv Evaluations SHA-256 8449fcc93763aedbe96074ca2256c81646e1b7a7e54968990d67a59c7924f03e | EvaluationsOne row per evaluation: a named test run by one evaluator. An evaluation belongs to one metric family.Fields of evals.csv | 303 rows | 84 KB | 8449fcc93763aedbe96074ca2256c81646e1b7a7e54968990d67a59c7924f03e |
| comparability_groups.csv Comparability groups SHA-256 a7695c2cfbeef247db8b78db175d9de5d0ad39868194a105c31c8d2acf343196 | Comparability groupsValues in one group come from the same test definition, reported by one developer, and may be joined in a series. A new group starts when there is a sign the test changed.Fields of comparability_groups.csv | 349 rows | 111 KB | a7695c2cfbeef247db8b78db175d9de5d0ad39868194a105c31c8d2acf343196 |
| measurements.csv Measurements SHA-256 fbdd429f29724199caa4295a54df5269d86ea31afd0e56475ef73b02b6c72300 | MeasurementsOne row per value a developer printed (or, for a few version 0 rows, stated only in a figure or in words). value_printed is the text as recorded; value_num exists only to position a point on a chart.Fields of measurements.csv | 1,645 rows | 641 KB | fbdd429f29724199caa4295a54df5269d86ea31afd0e56475ef73b02b6c72300 |
| revisions.csv Revisions SHA-256 0afc8e9fd9f437309dacd66b7aa276d1134c7b2fa2d6a684783460fb29a2f496 | RevisionsOne row per change between two versions of a document.Fields of revisions.csv | 15 rows | 5.8 KB | 0afc8e9fd9f437309dacd66b7aa276d1134c7b2fa2d6a684783460fb29a2f496 |
| verifications.csv Verifications SHA-256 89df534ab5b690ee197df2a68795caf4c9fb49d27aebe9ed2057517331d40079 | VerificationsOne row per check of a measurement or a revision.Fields of verifications.csv | 112 rows | 25 KB | 89df534ab5b690ee197df2a68795caf4c9fb49d27aebe9ed2057517331d40079 |
| restatement_links.csv Restatements SHA-256 45e1cac82065455dd919483c13f47f09d2812da8dea4ee15bf8a1e6221687524 | RestatementsLinks between a model's value in its own (or first-reporting) document and the value a later document gives for the same model, evaluation and metric.Fields of restatement_links.csv | 138 rows | 35 KB | 45e1cac82065455dd919483c13f47f09d2812da8dea4ee15bf8a1e6221687524 |
| determinations.csv Risk determinations SHA-256 7fbbfb8eb841311b4e469dc783ae35fd0e206fc61261150e97af3f20b0089d45 | Risk determinationsFormal decisions under a developer's safety framework: a level, threshold or risk conclusion, as printed, one row per domain. Levels are never mapped between frameworks.Fields of determinations.csv | 142 rows | 33 KB | 7fbbfb8eb841311b4e469dc783ae35fd0e206fc61261150e97af3f20b0089d45 |
| releases.csv Releases SHA-256 9c1f57b14b18584a7465fe4614f2efd67d6648e9d8bd8552d3fc6d1f0b75bb4a | ReleasesOne row per data release.Fields of releases.csv | 1 row | 399 bytes | 9c1f57b14b18584a7465fe4614f2efd67d6648e9d8bd8552d3fc6d1f0b75bb4a |
| id_map.csv ID map SHA-256 f752b0fae719089d56656fa240638395af02f7a0d200f8ae805ec88976f13f18 | ID mapPermanent map from version 0 rows to measurement IDs. Re-running the migration reuses it and only appends.Fields of id_map.csv | 1,645 rows | 25 KB | f752b0fae719089d56656fa240638395af02f7a0d200f8ae805ec88976f13f18 |
| sources.csv Source registry SHA-256 c4be115b618171d24877bd8764f72cd00480f922e51cf68eeb98ad2cc54fb38f | Source registryEvery document in scope and the other sources we track, with links, dates, known revisions and extraction notes (Methodology §2.1). It comes with the version 0 feasibility dataset and has no schema file.Fields of sources.csv | 133 rows | 40 KB | c4be115b618171d24877bd8764f72cd00480f922e51cf68eeb98ad2cc54fb38f |
| Checksums | ||||
| SHA256SUMS Checksums SHA-256 8be3cf66dc6fc96d2a2a4b1f327809edec0d29668a2ff364347adbf03251ca63 | ChecksumsThe SHA-256 checksum of every file above, in the format sha256sum reads (sha256sum -c SHA256SUMS). | 18 checksums | 1.5 KB | 8be3cf66dc6fc96d2a2a4b1f327809edec0d29668a2ff364347adbf03251ca63 |
To check what you downloaded, save SHA256SUMS in the same folder and run the first line there (Linux, or Git Bash on Windows) or the second (macOS, which also lists the files you did not download as missing):
sha256sum -c SHA256SUMS --ignore-missing
shasum -a 256 -c SHA256SUMSFiles stay at addresses that name their release, so a link to a file keeps leading to the same bytes.
Data dictionary
Every table and every field, from the schemas that check the data on every build (Frictionless Table Schema; the zip includes them). Open a table to see its fields.
Developers
One row per AI developer whose documents are in the dataset.
Each row is identified by lab_id
lab_idtext · required · uniquePermanent identifier: a short lowercase slug.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$Example
google-deepmindnametext · requiredName as used on the site.
Example
Google DeepMindshort_nametextShorter name for tight spaces such as chart legends.
websitetextThe developer's main website, when the source registry shows it.
Pattern
^https?://\S+$docs_index_urltextThe developer page listing its system or model cards, monitored for new and revised documents.
Pattern
^https?://\S+$frameworkslist of text, separated by “;”Names of the safety frameworks under which the developer records formal risk determinations, as they appear in its documents.
aliaseslist of text, separated by “;”Other names the developer appears under in sources.
Example
Zhipu AI (Z.ai)notestextFree-text notes.
Models
One row per model. Snapshots (pre-release checkpoints, dated updates, previews) are not separate models; they are recorded on each measurement as snapshot_label.
Each row is identified by model_id
model_idtext · required · uniquePermanent identifier: the model name as a slug, without a developer prefix.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$Example
claude-opus-4-5lab_idtext · requiredThe developer of the model (not necessarily the developer whose document reports a value about it).
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$Refers to
labs.lab_idnametext · requiredName as the developer writes it.
Example
Claude Opus 4.5aliaseslist of text, separated by “;”Other names the model appears under in documents, as printed.
Example
GPT-5 (thinking)familytextThe developer's product line, used to group models on the site.
Example
Claude Opusrelease_datetextRelease date. Defaults to the published date of the earliest document whose subject is the model (SPEC §5.3). Empty when no document in the dataset is about this model.
Pattern
^\d{4}-\d{2}(-\d{2})?$release_date_sourcetextWhere the release date comes from.
One of
documentannouncementotheravailabilitytextHow the model was made available, as the documents describe it.
One of
publiclimitedinternalopen_weightsnotestextFree-text notes.
Documents
One row per developer document (system card, model card, framework report, technical report, launch or policy post). Values are stored against the document they were read in.
Each row is identified by doc_id
doc_idtext · required · uniquePermanent identifier: the source_id from the source registry.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$Example
anthropic-claude-opus-4-5-system-cardlab_idtext · requiredThe developer that published the document.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$Refers to
labs.lab_idtitletext · requiredTitle as published.
doc_typetext · requiredKind of document.
One of
system_cardaddendummodel_cardfsf_reportsafety_reporttech_reportlaunch_postpolicy_postlab_summary_pagesubject_model_idslist of text, separated by “;”The models the document is about (its subject), as opposed to models it only compares against. Decided in review.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$Each item refers to
models.model_idpublished_datetext · requiredDate the document was first published.
Pattern
^\d{4}-\d{2}(-\d{2})?$date_precisiontext · requiredWhether published_date is known to the day or only to the month.
One of
daymonthcanonical_urltextThe developer's own URL for the document. Empty when no primary copy has been located.
Pattern
^https?://\S+$alt_urlslist of text, separated by “;”Other URLs for the same document, with a short note in brackets where the registry gives one.
has_changelogtext · requiredWhether the document carries a changelog of its revisions.
One of
yespartialnounknownnot_applicableextraction_coverage_notetextWhich parts of the document were read, and known gaps.
notestextFree-text notes.
Document versions
One row per known state of a document: the copy we retrieved, earlier copies we read, and states known only from a changelog or the publication date.
Each row is identified by version_id
Download document_versions.csv
version_idtext · required · uniquePermanent identifier: {doc_id}--{yyyy-mm-dd}, with -b appended if two versions share a date. The date is the version date when known to the day, otherwise the retrieval date.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*--\d{4}-\d{2}-\d{2}(-b)?$Example
xai-grok-4-6-model-card--2026-08-17doc_idtext · requiredThe document.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$Refers to
documents.doc_idversion_labeltextShort description of this state in our words.
version_datetextDate of this state, to the day or to the month.
Pattern
^\d{4}-\d{2}(-\d{2})?$date_sourcetext · requiredHow version_date is known. changelog: the document's changelog. header: the document's own header. retrieval: the date we retrieved it. report: our source registry or a third-party report.
One of
changelogheaderretrievalreportstatustext · requiredWhat we hold for this state. retrieved: a copy we retrieved. archived: an earlier copy we read elsewhere (for example a copy the developer hosted with a partner). known_from_changelog: a state known only from the document's own changelog. known_not_retrieved: a state known from the publication date, a document header or a third-party report, of which we hold no copy.
One of
retrievedarchivedknown_from_changelogknown_not_retrievedurltextWhere this state was read, when it was read.
Pattern
^https?://\S+$sha256textSHA-256 of the retrieved file (from dataset v0.2).
Pattern
^[0-9a-f]{64}$retrieved_attextDate the copy was retrieved.
Pattern
^\d{4}-\d{2}-\d{2}$archive_urltextAn archived snapshot of this state.
Pattern
^https?://\S+$changelog_summarytextWhat changed in this state, in our words, from the developer's changelog or registry notes.
parsed_pagestextPages parsed from this copy (from dataset v0.2).
notestextFree-text notes.
Secondary sources
Independent write-ups and developer summary pages that some values were read from, because the card section itself could not be read in version 0.
Each row is identified by secondary_id
Download secondary_sources.csv
secondary_idtext · required · uniqueIdentifier: the source_id from the source registry.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$titletext · requiredTitle or site name as in the registry.
author_or_sitetextAuthor or publishing site.
urltext · requiredURL read.
Pattern
^https?://\S+$published_datetextPublication date, when the URL or registry states it.
Pattern
^\d{4}-\d{2}(-\d{2})?$covers_doc_idslist of text, separated by “;”Documents whose values were read from this source.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$Each item refers to
documents.doc_idnotestextFree-text notes.
Evaluators
Who ran an evaluation: the developer itself, a third party, or a government body.
Each row is identified by evaluator_id
evaluator_idtext · required · uniquePermanent identifier: a short slug.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$Example
apollonametext · requiredName as used on the site.
Example
Apollo Researchtypetext · requiredKind of evaluator.
One of
developerthird_partygovernmenturltextThe evaluator's website, when the source registry shows it.
Pattern
^https?://\S+$
Evaluations
One row per evaluation: a named test run by one evaluator. An evaluation belongs to one metric family.
Each row is identified by eval_id
eval_idtext · required · uniquePermanent identifier: {evaluator-or-lab}-{eval-slug}.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$Example
openai-destructive-action-avoidancenametext · requiredName as the documents print it.
familytext · requiredMetric family code (METHODOLOGY §1.1).
One of
M1M2M3M4M5M6M7M8M9M10M11M12evaluator_idtext · requiredWho ran it.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$Refers to
evaluators.evaluator_idtest_typetext · requiredKind of test. Evaluation-awareness evaluations (M1) use the six kinds in Finding 1; other families use the general list.
One of
developer_audit_verbalizedexternal_apolloexternal_uk_aisideveloper_other_setupsdeployment_real_or_simulatedwhite_box_or_trainingstatic_benchmarkadaptive_attackerred_teambug_bountyautomated_auditscenario_evalcapability_benchmarkuplift_studyproduction_trafficdeployment_simulationtraining_monitoringinternal_monitoringsurveyframework_determinationqualitative_assessmentdescriptiontextWhat the evaluation measures, in our words.
notestextFree-text notes.
Comparability groups
Values in one group come from the same test definition, reported by one developer, and may be joined in a series. A new group starts when there is a sign the test changed.
Each row is identified by group_id
Download comparability_groups.csv
group_idtext · required · uniquePermanent identifier: {eval_id}-v{n} or a descriptive suffix.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$Example
openai-destructive-action-avoidance-v1eval_idtext · requiredThe evaluation.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$Refers to
evals.eval_idlabeltext · requiredShort label shown in the explorer.
definition_notestextWhat defines this version of the test, and what changed from the previous group, in our words.
join_basistext · requiredstated: the developer's documents say the test is unchanged, or all values come from one document. inferred: values from several documents were joined because nothing indicated a change; flagged to readers.
One of
statedinferredfirst_doc_idtextEarliest document with a value in this group.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$Refers to
documents.doc_idlast_doc_idtextLatest document with a value in this group.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$Refers to
documents.doc_idsupersedes_group_idtextThe group this one replaces, when the test changed.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$Refers to
comparability_groups.group_idnotestextFree-text notes.
Measurements
One row per value a developer printed (or, for a few version 0 rows, stated only in a figure or in words). value_printed is the text as recorded; value_num exists only to position a point on a chart.
Each row is identified by measurement_id
measurement_idtext · required · uniquePermanent identifier: m- and six digits.
Pattern
^m-\d{6}$Example
m-000812legacy_row_idintegerrow_id in the version 0 file, for rows that came from it.
Range 0 or more
split_indexinteger0, or 1, 2 … when one version 0 row became several measurements.
Range 0 or more
version_idtext · requiredThe document version the value was read in (or, for secondary rows, the version the value is about).
Pattern
^[a-z0-9]+(-[a-z0-9]+)*--\d{4}-\d{2}-\d{2}(-b)?$Refers to
document_versions.version_idsubject_kindtext · requiredWhat the value is about: a model, a deployed system with safeguards, several models at once, or a baseline.
One of
modelsystemmultibaselinemodel_idtextThe model, when subject_kind is model or system.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$Refers to
models.model_idsubject_model_idslist of text, separated by “;”The models covered, when subject_kind is multi.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$Each item refers to
models.model_idsubject_labeltext · requiredThe subject exactly as the version 0 row names it.
snapshot_labeltextA snapshot of the model (pre-release checkpoint, dated update, preview), as printed.
is_subject_modeltrue or false · requiredtrue when the model is a subject of the document; false when it appears only as a comparison.
eval_idtext · requiredThe evaluation.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$Refers to
evals.eval_idgroup_idtext · requiredThe comparability group.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$Refers to
comparability_groups.group_idmetric_labeltext · requiredShort name of the metric within the evaluation.
metric_definitiontext · requiredThe metric as described in version 0 (the extractor's wording).
condition_efforttextReasoning effort or budget, as printed.
condition_modetextMode, such as extended thinking or non-reasoning.
condition_safeguardstextWhether safeguards or mitigations were on.
condition_surfacetextDeployment surface, such as API or web.
condition_othertextAny other measurement condition (attempts, prompt set, scope).
context_notetextContext printed with the value that is not a condition (for example a human baseline).
value_printedtext · requiredThe value exactly as recorded in version 0, which keeps the printed form (including trailing zeros). This is what the site displays.
value_numnumberA number parsed from value_printed, only to position a point. Never displayed. For printed ranges it is the stored midpoint.
value_texttextThe value as words, for categorical values.
value_qualifiertextHow the printed value is bounded or approximated.
One of
eqltlegtgeapproxrangerange_lownumberLower end of a printed range.
range_highnumberUpper end of a printed range.
unittext · requiredUnit, from the controlled vocabulary.
One of
percentratepp_changerelative_change_percentcountscoremultiplierminutesusdodds_ratiocorrelationcategoricalunit_as_printedtext · requiredThe unit as recorded in version 0.
denominatorintegerDenominator for counts out of a fixed number.
Range 1 or more
scaletextScale of a score, as stated.
ci_lownumberLower bound of a printed confidence interval.
ci_highnumberUpper bound of a printed confidence interval.
nintegerSample size, when printed with the value.
Range 0 or more
directiontext · requiredWhether a higher value is better or worse, as the developer frames it.
One of
higher_worsehigher_bettercategoricalneutralvalue_statustext · requiredprinted: a number or statement printed in the text or a table. figure_read: read off a chart. qualitative: a statement in words rather than a value.
One of
printedfigure_readqualitativelocationtext · requiredSection, table or page where the value appears.
source_typetext · requiredWhere the value was read: the developer's document, a developer summary page, or an independent secondary source.
One of
primarysecondarylab_summaryvia_urltextThe URL actually read, when it is not the document version's own URL (secondary sources, summary pages, HTML section pages).
Pattern
^https?://\S+$confidencetext · requiredConfidence grade (METHODOLOGY §4.1).
One of
highmediumlowverification_statustext · requiredVerification status (METHODOLOGY §4.2).
One of
unverifiedblind_verifieddisputedcorrectedextraction_methodtext · requiredHow the value was extracted.
One of
v0_web_readerpdf_text_v1manualsuperseded_bytextThe measurement that replaces this one, if any. IDs are never removed.
Pattern
^m-\d{6}$Refers to
measurements.measurement_idnotestextFree-text notes: the version 0 extraction note, then any migration note.
Revisions
One row per change between two versions of a document.
Each row is identified by revision_id
revision_idtext · required · uniquePermanent identifier: rev- and a descriptive slug.
Pattern
^rev-[a-z0-9]+(-[a-z0-9]+)*$doc_idtext · requiredThe document.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$Refers to
documents.doc_idfrom_version_idtext · requiredThe earlier version.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*--\d{4}-\d{2}-\d{2}(-b)?$Refers to
document_versions.version_idto_version_idtext · requiredThe later version.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*--\d{4}-\d{2}-\d{2}(-b)?$Refers to
document_versions.version_idmeasurement_id_oldtextThe measurement in the earlier version, when recorded.
Pattern
^m-\d{6}$Refers to
measurements.measurement_idmeasurement_id_newtextThe measurement in the later version, when recorded.
Pattern
^m-\d{6}$Refers to
measurements.measurement_idchange_typetext · requiredKind of change.
One of
value_changedaddedremovedwordingdetermination_changedunknownwhattext · requiredWhat changed, in our words.
old_texttextThe earlier text or value, as printed.
new_texttextThe later text or value, as printed.
explainedtext · requiredWhether the document's changelog explains the change.
One of
yespartlynono_changelogexplanation_summarytextThe developer's explanation, in our words.
verification_statustext · requiredHow the revision was confirmed.
One of
unverifiedblind_verifiedchangelog_confirmeddisputedcorrectednotestextFree-text notes.
Verifications
One row per check of a measurement or a revision.
Each row is identified by verification_id
verification_idtext · required · uniqueIdentifier: v-, then the measurement or revision ID, the method and the date.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$measurement_idtextThe measurement checked.
Pattern
^m-\d{6}$Refers to
measurements.measurement_idrevision_idtextThe revision checked.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$Refers to
revisions.revision_idmethodtext · requiredHow it was checked.
One of
blind_recheckfull_text_matchchangelog_readmanualselectiontext · requiredWhy it was chosen for checking.
One of
featuredrandomnew_rowstargetedchecked_sourcetext · requiredWhich kind of source was re-read. A check against the same secondary source is weaker than a check against the card.
One of
primarysecondarylab_summaryfound_valuetextThe value the check found, as it wrote it.
resulttext · requiredOutcome.
One of
matchmismatchno_printed_valuenot_foundchecked_attext · requiredDate of the check.
Pattern
^\d{4}-\d{2}-\d{2}$checked_bytextWho or what ran the check.
notestextThe checker's note, in its words.
Restatements
Links between a model's value in its own (or first-reporting) document and the value a later document gives for the same model, evaluation and metric.
Each row is identified by link_id
Download restatement_links.csv
link_idtext · required · uniqueIdentifier: rs-, then both measurement IDs.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$own_measurement_idtext · requiredThe own-document or first-reported value.
Pattern
^m-\d{6}$Refers to
measurements.measurement_idlater_measurement_idtext · requiredThe later document's value.
Pattern
^m-\d{6}$Refers to
measurements.measurement_idstatustext · requiredWhether the link has been confirmed.
One of
confirmedcandidaterejectedstated_reasontextThe reason the later document gives for a difference, in our words, or the condition difference.
notestextFree-text notes, including who confirmed the link and why.
Risk determinations
Formal decisions under a developer's safety framework: a level, threshold or risk conclusion, as printed, one row per domain. Levels are never mapped between frameworks.
Each row is identified by determination_id
determination_idtext · required · uniqueIdentifier: d-, the measurement ID and the domain (with a suffix when one value covers two thresholds in a domain).
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$measurement_idtext · requiredThe categorical measurement the determination comes from.
Pattern
^m-\d{6}$Refers to
measurements.measurement_idmodel_idtextThe model.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*$Refers to
models.model_idversion_idtext · requiredThe document version.
Pattern
^[a-z0-9]+(-[a-z0-9]+)*--\d{4}-\d{2}-\d{2}(-b)?$Refers to
document_versions.version_idframeworktextThe framework's name, as the developer writes it. Empty when the document, as recorded, names no framework for the conclusion; the notes say so.
framework_versiontextThe framework version, when stated.
domaintext · requiredRisk domain.
One of
overallbio_chemcyberai_rnd_autonomymanipulationmisalignmentotherlevel_as_printedtext · requiredThe level or conclusion. Where version 0 stored a code rather than the printed wording, the code is kept and flagged in notes until the wording is recovered from the source.
ordinal_within_frameworkintegerPosition of the level within its own framework, for ordering only; never compared across frameworks.
notestextFree-text notes.
Releases
One row per data release.
Each row is identified by release
releasetext · required · uniqueDataset version.
Pattern
^v\d+\.\d+$datetext · requiredRelease date.
Pattern
^\d{4}-\d{2}-\d{2}$methodology_versiontext · requiredMethodology version the release follows.
n_measurementsinteger · requiredMeasurements in the release.
Range 0 or more
n_documentsinteger · requiredDocuments with at least one measurement.
Range 0 or more
n_modelsinteger · requiredModels with at least one measurement.
Range 0 or more
blind_check_sampleintegerValues blind-checked for this release.
Range 0 or more
blind_check_matchesintegerBlind checks that matched.
Range 0 or more
notestextFree-text notes, including the agreement rate of the independent re-decision of review decisions.
ID map
Permanent map from version 0 rows to measurement IDs. Re-running the migration reuses it and only appends.
Each row is identified by these together legacy_row_idsplit_index
legacy_row_idinteger · requiredrow_id in the version 0 file.
Range 0 or more
split_indexinteger · required0, or 1, 2 … when one row became several measurements.
Range 0 or more
measurement_idtext · required · uniqueThe measurement ID assigned.
Pattern
^m-\d{6}$
Source registry
The source registry supplied with the version 0 feasibility dataset (Methodology §2.1). It has no schema file; its columns are listed as they appear in its header row.
source_idcategorylabtitledoc_typemodels_coveredpublishedurlalt_urlsrevisions_knownchangelogextraction_notesv0_rowsv0_card_labellast_checked
Licence
The data and the text of this site are published under the Creative Commons Attribution 4.0 licence (CC BY 4.0 (opens an external site)). You may copy, share and adapt them for any purpose, including commercial use.
Attribution means naming Safety Card Ledger and the dataset version you used, linking to the licence, and saying whether you changed anything; the citations below do the first part.
The values were published by the developers named in the data, in documents the data links to. Those documents are theirs: we link to them and do not re-host them.
Cite the dataset
Please cite the version you used, so a reader can find the same data.
APA
Safety Card Ledger. (2026). Safety Card Ledger (Version 0.1) [Data set]. https://safetycardledger.example/download
BibTeX
@misc{safetycardledger_v0_1,
author = {{Safety Card Ledger}},
title = {Safety Card Ledger},
year = {2026},
month = sep,
version = {0.1},
howpublished = {\url{https://safetycardledger.example/download}},
note = {Dataset v0.1, methodology 0.1. CC BY 4.0},
}CITATION.cff gives the same citation in the Citation File Format, which reference managers and code hosts read. The zip includes a copy.
Previous releases
None yet: v0.1 is the first release. Every release keeps its files and their checksums on this page, at addresses that name its version, so a citation of an earlier version still leads to the data it used.
API
A static JSON API is planned for Phase 3. These files will be published at these addresses, alongside one file per entity; none of them exists yet.
- /api/v0/measurements.json
- /api/v0/models.json
- /api/v0/documents.json
- /api/v0/document_versions.json
- /api/v0/evals.json
- /api/v0/revisions.json
- /api/v0/restatements.json
- /api/v0/determinations.json
- /api/v0/releases.json