Tracking the safety results AI developers publish about their own models — every number with its source, its method, and its history.

Data as of 26 Sep 2026 · Dataset v0.1 · Methodology v0.1 · Changelog

1,645 values · 70 documents · 8 developers

Recent revisions

The latest changes developers made to their documents after publishing them, where a value is recorded on both sides and both have been checked. The ledger holds all 15 revisions, including changes of wording and documents marked updated with no record of what changed.

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

  1. Grok 4.6 Model Card

    xAI · 3 of its 5 changes

    1. HackerBench v0.2 harmful/dual-use compliance, Grok 4.6 (high)

      16.7%6.9%

      Partly explained — changelog lists "corrected eval results" for this eval; no reason

      Old value New value In the ledger: HackerBench v0.2 harmful/dual-use compliance, Grok 4.6 (high)

    2. MASK dishonesty, Grok 4.6 (high)

      3.8%1.90%

      Partly explained — listed as corrected; no reason

      Old value New value In the ledger: MASK dishonesty, Grok 4.6 (high)

    3. Self-harm compliance, Grok 4.6 (high)

      3.7%0.84%

      Partly explained — listed as corrected; no reason

      Old value New value In the ledger: Self-harm compliance, Grok 4.6 (high)

All 15 revisions, newest first

Findings

What the data shows, one chart each, with the values behind it and how they were chosen and checked.

  1. Finding 1 ·

    33 reported numbers on evaluation awareness come from six kinds of test

    How often a model shows signs of recognising a test is measured in six different ways, and the same model gets different numbers from different kinds of test.

Not written yet: planned for Phase 3

  • Finding 2Later documents report earlier models again, sometimes with different values (not written yet)
  • Finding 3Documents change after release, with and without a changelog (not written yet)
  • Finding 4Within one developer and one test, values form a series until the test changes (not written yet)
  • Finding 5Who reports what (not written yet)

All findings

Coverage at a glance

Which metric families each developer’s documents report with a printed number. This describes what developers publish, not the safety of their models.

Metric families reported, by developer

Share of each developer’s documents with at least one printed number in each metric family. Developers are ordered by number of documents (in brackets), then by name. A cell opens the data explorer at that family and developer.

Share of each developer's documents in the version 0 dataset with at least one numeric value in each metric family.
DeveloperM1M2M3M4M5M6M7M8M9M10M11M12
OpenAI (22)18%5%18%45%41%5%86%45%55%59%0%18%
Google DeepMind (16)38%0%31%0%13%0%88%13%6%38%0%6%
Anthropic (15)60%53%53%53%60%13%80%13%67%87%7%20%
xAI (8)13%0%13%0%100%88%100%88%63%100%0%0%
Meta (5)40%20%40%40%40%40%40%40%60%60%0%0%
Moonshot AI (2)0%0%0%0%0%0%50%50%50%50%0%0%
DeepSeek (1)0%0%0%0%0%0%0%0%0%0%0%0%
Zhipu AI (1)0%0%0%0%0%0%100%0%0%0%0%0%

Share of documents0%1-24%25-49%50-74%75-100%

M1
Evaluation awareness
M2
Reward hacking
M3
Sabotage and sandbagging
M4
Misalignment audits
M5
Honesty and hallucination
M6
Sycophancy
M7
Harmful compliance and over-refusal
M8
Jailbreak robustness
M9
Prompt injection
M10
Dangerous capabilities and risk determinations
M11
Self-preservation
M12
Chain-of-thought monitorability

A document counts for a family when it prints at least one number in it. Values read off a figure, stated in words, or given as a category such as a risk level do not count. Some documents were not read to the end in version 0, so read the shares as a lower bound.

Source: Safety Card Ledger v0.1 · Data from developer system cards · CC BY 4.0

Recently added or revised documents

Ordered by the latest date we know for each document: its publication date, or the date of a later version when the document’s changelog or header gives one. The day we retrieved a copy does not count.

  1. Claude Opus 5.5 System Card

    Anthropic · System card · Published; no later version recorded

  2. GPT-6 Astra System Card

    OpenAI · System card · Revised 22 Sep 2026; first published 3 Sep 2026

  3. Grok 4.7 Model Card

    xAI · Model card · Published; no later version recorded

  4. Our framework for reporting model misalignment

    OpenAI · Policy post · Published; no later version recorded

  5. ChatGPT Images 2.5 System Card

    OpenAI · System card · Published; no later version recorded

All 70 documents

About

Safety Card Ledger records the safety evaluation results that AI developers publish about their own models: each number as the document prints it, where it was read, how it was checked, and how it changed when a document was revised. Values are recorded as published by each developer. Safety Card Ledger does not rate the safety of models. Developers are never ranked on the numbers they report.

About the project