What this is

Developers of frontier AI models publish safety results about their own models in system cards, model cards and reports: how often a model seems to notice it is being tested, how it behaves when it could cheat or undermine a task, how it responds to harmful requests and attempts to get round its safeguards, what dangerous capabilities it shows, and whether it crosses a risk threshold the developer has set. The results are printed in prose and tables across many documents. This project records them as data.

The dataset holds 1,645 values from 70 documents by 8 developers. Each value is kept exactly as the document prints it, with the document, version and section it came from, how it was measured, and how it has been checked. Where a later document reports an earlier model again with a different value, the two are linked; where a developer edits a document after publishing it, the change is recorded.

It is meant for anyone who needs to know what a card said, whether it changed, and how a number was measured: researchers, journalists, staff of AI safety institutes and regulators, projects that track developers, and developers' own staff checking how their results are recorded.

What it is not

Values are recorded as published by each developer. Safety Card Ledger does not rate the safety of models.

  • Not a rating or a ranking. The site does not score models, rank developers or say whether a model is safe.
  • Not a comparison across developers. Developers measure the same property with different tests, prompts and graders, so their numbers are never put on one scale or joined into one line (Methodology §5).
  • Not our own evaluations. Every value is one a developer printed, or one an evaluator the developer quotes. Evaluations we run ourselves are planned for later, and will be labelled and versioned separately (Methodology §9).
  • Not an incident tracker or a forum. It records published evaluation results, not incidents, and has no accounts or comments.

How it is built

Values are read from developers' documents, each listed on the Sources page with a link to the original. The first version of the dataset was extracted with the help of AI research agents and then checked: a sample of values was looked up again by a separate pass that could not see the recorded value, and the match rate is published with each release. Judgments about how rows fit together, such as which names refer to the same model or test and which values form one series, are recorded with their reasons. Every release has a version number and names the methodology it follows; the Methodology page sets all of this out.

Who runs it

To be written by the owner before launchWho runs the site

Who runs Safety Card Ledger: a person, an organisation, or neither named (SPEC §15, “Named on About”). Until this is decided, citations name the project itself.

Independence and funding

To be written by the owner before launchIndependence and funding statement

Required before launch (SPEC §11.7): who pays for the work, whether any AI developer has funded, reviewed or been consulted on it, and how conflicts of interest are handled.

Contact

Email corrections@safetycardledger.example (a placeholder: no one reads this address yet). To report a value that does not match its source, see Corrections.

Licence and citation

The data and the text are published under CC BY 4.0 (opens an external site): you may reuse them with credit. Please cite the dataset version you used; suggested citations are on the Download page.

Acknowledgements

To be written by the owner before launchAcknowledgements

Whom the owner wishes to thank: people who advised on, reviewed or checked the work, and any organisation that supported it.