Safety Card Ledger · Methodology, version 0.1 · effective 27 Sep 2026 · https://safetycardledger.example/methodology/v0-1
This page explains how Safety Card Ledger collects, records, checks and presents the safety evaluation results that AI developers publish about their own models. It is versioned. When the method changes, the version number goes up, the change is described in the changelog at the end, and every data release states which methodology version it follows.
1. What we track
Link to section 1We record quantitative and categorical safety results that developers publish about their models: the numbers in system cards, model cards, safety framework reports, technical reports and official launch material.
We do not score models, rank them, or say whether a model is safe. Every value on this site is a value a developer (or an evaluator the developer quotes) printed, recorded with where it came from and how it was measured.
1.1 Metric families
Link to section 1.1Each value is assigned to one of twelve families. The families group results that address the same underlying property. They do not make the results comparable with each other; see section 5.
| Code | Family | What it covers |
|---|---|---|
| M1 | Evaluation awareness | How often a model shows signs of recognizing it is being tested, whether stated in its reasoning or detected by other means |
| M2 | Reward hacking | Exploiting flaws in tests or graders, special-casing tests, and behavior on deliberately impossible tasks |
| M3 | Sabotage and sandbagging | Covertly completing a hidden side task, undermining assigned work, or deliberately underperforming |
| M4 | Misalignment audits | Automated behavioral audits, agentic misalignment scenarios, covert-action and scheming rates, destructive actions |
| M5 | Honesty and hallucination | Deception, fabrication, false claims of task completion, hallucination rates |
| M6 | Sycophancy | Telling users what they want to hear at the expense of accuracy |
| M7 | Harmful compliance and over-refusal | Responses to disallowed requests, and refusals of benign ones |
| M8 | Jailbreak robustness | Resistance to attempts to bypass safeguards |
| M9 | Prompt injection | Resistance to instructions planted in content an agent reads |
| M10 | Dangerous capabilities and risk determinations | Capability evaluations in biology, chemistry, cyber and AI research, and the developer’s threshold or risk-level decisions |
| M11 | Self-preservation | Shutdown resistance, self-exfiltration, resource or power seeking |
| M12 | Chain-of-thought monitorability | Whether a model’s written reasoning is legible and faithful enough to monitor |
A value that fits two families is assigned to the one closest to what the developer says the test measures, and the other is noted.
2. Sources
Link to section 22.1 Inclusion
Link to section 2.1A document is in scope if it is published by a model developer and reports safety results about one or more of its models. This covers:
- system cards, model cards and their addenda;
- safety framework reports (for example, frontier safety framework or preparedness reports);
- technical reports and papers with safety sections;
- official launch posts, when they contain results not found elsewhere;
- developer summary pages that restate card results (recorded as a separate source type).
Results produced by third-party evaluators (for example Apollo Research or the UK AI Security Institute) are included when a developer’s document reports them, and are attributed to the named evaluator.
The full list of documents, with links, dates, known revisions and extraction notes, is published as the source registry (sources.csv) and on the Sources page.
2.2 Source types
Link to section 2.2Every value records the kind of source it was read from:
| Source type | Meaning | How it is shown |
|---|---|---|
| Primary | Read in the developer’s own document | Solid marker |
| Developer summary page | Read on a developer page that restates card results | Solid marker, flagged in details |
| Secondary | Read in an independent write-up that quotes the card, because the card section could not be read directly | Hollow marker, flagged in details |
Secondary values are a stopgap. They are replaced with primary values as soon as the relevant section of the card is extracted.
2.3 What we leave out, and what we flag
Link to section 2.3- Values shown only in charts or figures are not extracted from new documents. Reading values off a chart introduces error we cannot check. If a developer prints only a chart, we record that the result exists but not the number.
- Capability benchmarks with no safety purpose (general coding, math, knowledge), unless the developer reports them as part of a dangerous-capability assessment.
- Qualitative statements (“sycophancy was reduced”) are not treated as values, unless they are a formal determination such as a risk level.
The version 0 dataset predates these rules in a few places: it contains a small number of values read from charts and some qualitative statements. They are kept for transparency, marked with a value status of figure read or qualitative, excluded from charts and counts by default, and will be replaced or removed as full-text extraction reaches them.
3. Extraction
Link to section 3For each document we record every in-scope value as one row with:
- the model the value is about, and whether that model is the document’s subject or a comparison model;
- the evaluation’s name as printed, its version if stated, and the measurement conditions (for example reasoning effort, safeguards on or off, number of attempts, language scope);
- the value exactly as printed, its unit, and whether higher is better or worse;
- the section, table or page where it appears;
- the source link and source type;
- a confidence grade (section 4).
Values are stored as printed. We keep the printed text of every value (including trailing zeros, such as “1.90%”), and a number parsed from it for charts. We do not convert a developer’s “violation rate” into a “safe rate”, rescale units, or combine sub-scores. Where a developer prints a bound (“less than 0.01%”), an approximation (“about 9%”) or a range (“1.5–2%”), we record that qualifier and the printed bounds; charts show the qualifier in tooltips and never present an approximate value as exact.
Version 0 of the dataset was extracted with the help of AI research agents reading documents through a web reader, followed by the checks in section 4. That reader stopped partway through some long PDFs (around pages 50–60), so later sections of several cards were not read directly. Those gaps are listed per document in the source registry. From dataset version 0.2, documents are downloaded and parsed in full.
4. Confidence and verification
Link to section 44.1 Confidence grades
Link to section 4.1| Grade | Meaning |
|---|---|
| High | We read the value directly from a table or sentence, and the model, condition and unit are unambiguous |
| Medium | We read the value directly, but the table layout, column mapping or condition needed interpretation, or it came from a secondary source that quotes it plainly |
| Low | Our reading is approximate (for example, from a chart) or the source paraphrases; kept for completeness and excluded from featured charts |
A value that the source itself prints as approximate (“about 9%”) is recorded with an “approximately” qualifier. That does not by itself lower confidence, because our reading of the printed text is exact.
4.2 Verification
Link to section 4.2Each value carries a verification status:
| Status | Meaning |
|---|---|
| Unverified | Extracted once |
| Blind-verified | Looked up again by a separate pass that could not see the recorded value, and matched |
| Disputed | A re-check found a different value or no printed value; under review |
| Corrected | Changed after review; the previous value is kept in the history |
| Confirmed from changelog | Revisions only: the change is confirmed by the developer’s own changelog entry, read by us. This does not verify the values on either side of the change, which carry their own status |
Each check also records which source was re-read (the developer’s document, a developer summary page, or a secondary source), because a check against the same secondary source is weaker than a check against the card. Revisions recorded from a document’s changelog are verified by reading the changelog itself.
Every value shown individually (named in text, labelled in a chart, used in a revision or restatement, or highlighted by default) must be blind-verified before publication. Aggregates such as counts and coverage shares are computed from all values. A random sample of every data release is also blind-verified, and the match rate is published with the release.
For version 0, 96 values were blind-checked: 54 because the initial findings relied on them and 42 chosen at random. 95 matched the recorded value. The other had been read off a chart rather than printed, and is marked as disputed.
5. Comparability
Link to section 5Safety results from different documents often cannot be compared directly, even when they share a name. The site is built around this.
- Across developers. Different developers use different tests, prompt sets, graders and definitions for the same property. We do not rank developers against each other on reported numbers. Cross-developer comparisons are only made with evaluations we run ourselves under one protocol (planned; see section 9).
- Across documents from the same developer. A developer may change a test’s prompt set, grader or definition between documents. Each evaluation is recorded with a comparability group: values in the same group are plotted as one series. We keep one developer’s values for one evaluation in one group unless there is a sign the test changed: the developer says so, or the published definitions differ in substance. When the group changes, charts show a break and a note, and do not draw a line across it. Many documents do not say whether a test is unchanged. Where we connect values from such documents because nothing suggests the test changed, we say so: a note under the chart, and “test assumed unchanged” in the tooltip and details of each value on that line.
- Across conditions. Values measured under different conditions (reasoning effort, safeguards, number of attempts) are kept as separate series.
Deciding how rows fit together (which names refer to the same model or the same evaluation, which values form one series, which statements are formal risk determinations) involves judgment. We record every such decision with its reason. For version 0.1 these decisions were made with the help of an AI assistant. A random sample was decided a second time by an independent pass that could not see the first answer, and the rate at which the two agreed is published with each data release. If you disagree with how rows were grouped, report it as a correction (section 11).
6. Restatements
Link to section 6Developers often re-report older models as comparisons in a new document, sometimes with a different value than the older model’s own document gave. We call this a restatement.
- Every value is stored against the document it came from. Values for the same model from different documents are never merged or averaged.
- By default, a model’s value in a series is the value from its own document (a document whose subject is that model). If a model has more than one own document, we use the earliest; if they disagree, both are shown and the difference is flagged.
- If no own document reports the value, we use the value first reported in a later document, and say so wherever it appears.
- Restated values are shown on request and linked to the original. Differences in stated conditions (for example, a re-run) do not stop two values from being linked; they are shown as the reason.
- When a later document explains why a value changed (a new prompt set, a re-run, a corrected configuration), the explanation is recorded.
7. Revisions
Link to section 7Developers sometimes edit a document after publishing it.
- Each document is stored as a series of versions, each with its date and link. Versions known only from a changelog are recorded as such. From dataset version 0.2, each retrieved version also carries a file hash and an archived snapshot.
- When a new version appears, its values are extracted and compared with the previous version. Every changed, added or removed value is recorded as a revision, with the old value, the new value, the date, and whether the document’s changelog explains the change.
- We report revisions factually. A revision is not in itself evidence of wrongdoing; most are corrections.
- Where a developer publishes no changelog, we say so, and describe what changed based on the two versions.
8. Risk determinations
Link to section 8Developers make formal decisions about whether a model crosses a risk threshold, under frameworks with different names and levels (for example Anthropic’s AI Safety Levels and capability thresholds, OpenAI’s Preparedness Framework levels, Google DeepMind’s Critical Capability Levels, xAI’s and Meta’s frameworks). We record each determination as printed, with the framework name and version it was made under. We do not map levels between frameworks.
9. Our own evaluations (planned)
Link to section 9Reported numbers cannot answer cross-developer questions, and several developers of open-weight models publish few or no safety results. We plan to run a fixed set of public safety evaluations on open-weight models under one published protocol. Those results will be clearly labeled as ours, versioned separately, and never mixed with developer-reported values.
10. Updates and releases
Link to section 10- Once the refresh pipeline is running, sources are checked for new and revised documents at least weekly. Until then, checks are manual and dated.
- Each data release has a version number, a date, the methodology version it follows, row counts, the blind-check match rate, and a list of changes.
- Every page states the date its data was last updated.
11. Corrections
Link to section 11If you believe a value is wrong, report it through the correction link on any value or page, with the source location. We re-check the value against the source, publish the outcome in the changelog, and keep the previous value in the history. Developers are welcome to flag errors in how we recorded their results.
12. Known limitations
Link to section 12- Version 0 relies partly on a web reader that did not reach the later sections of some long PDFs. About 6% of version 0 values come from secondary write-ups or a developer summary page instead of the card itself. Both are flagged per value.
- Version 0 contains a few values read from charts and some qualitative statements (see 2.3). They are flagged and excluded from charts by default.
- Extraction and verification in version 0 used the same kind of reader, so a systematic misreading of a table layout could repeat in both.
- Values shown only in figures are missing by design.
- The source registry covers the major U.S. developers well, and other developers only partly.
13. Licence and citation
Link to section 13The dataset is published under Creative Commons Attribution 4.0 (CC BY 4.0). Please cite the dataset version you used; a suggested citation is on the Download page.
Changelog
Link to the Changelog section| Version | Date | Changes |
|---|---|---|
| 0.1 | 2026-09-27 | First published methodology, covering the version 0 feasibility dataset |