The sources at a glance

Every value in the dataset was read in one of the sources below. Of the 1,645 values, 1,541 were read in developers’ own documents: 70 documents from eight developers, each with its own page here. Where a document’s section could not be read, 14 values were read on a developer summary page and 90 in 24 independent write-ups that quote the document.

Values read on a summary page or in a write-up are flagged wherever they appear, and are to be replaced with the document’s own values once its full text is read.

The registry also lists the 5 developer indexes we check for new documents, 21 documents we know of but have not read yet, and 12 related projects.

Sources by kind: how many of each, and how many values were read in them
Kind of sourceSourcesValues read there
Developer documents701,541
Developer summary pages114
Secondary write-ups2490
Developer indexes we monitor5none
Documents not yet read21none
Related projects12none
All sources1331,645

The table counts each value once, where it was read. A document's entry below counts every value recorded about it, including any read on a summary page or in a write-up.

Every source

By kind, as the source registry records them. Search by title or developer, or narrow the list by kind and developer.

Showing all 133 sources.

Developer documents

70 sources

System cards, model cards, reports and posts, by developer and newest first. Each title opens the document's page here, with its versions, every value recorded from it and what we read of it. No version has an archived snapshot of our own yet; from dataset v0.2 each retrieved version carries one (Methodology §7).

Anthropic

DeepSeek

Google DeepMind

Meta

Moonshot AI

OpenAI

xAI

Zhipu AI

Developer summary pages

1 source

A developer's own page restating results from its cards. We read values here only where a card's section was out of reach; each is marked “Developer summary page” wherever it appears (Methodology §2.2).

Secondary write-ups

24 sources

Independent articles and posts that quote a document, listed by the document they are about. We read values here only where the document's own section could not be read; each is marked “Secondary source” and drawn hollow in charts, and is to be replaced with the document's own value once its full text is read (Methodology §2.2).

Developer indexes we monitor

5 sources

Pages where developers list their documents. We check them for new and revised documents; until the refresh pipeline runs, the checks are by hand and dated (Methodology §10).

Documents not yet read

21 sources

Documents we know of but have not read, with the reason. None of their results are in the dataset yet; they are queued for the full-text pass that comes with the refresh pipeline.

  • Qwen3 technical report; Qwen3.6-27B model card

    Alibaba

    A first reading found no safety results given as numbers.

  • August 2026 Risk Report

    Anthropic

    The Claude Fable 5.1 & Claude Mythos 5.1 System Card refers to it; not read yet.

  • Claude Opus 4.6 Sabotage Risk Report

    Anthropic

    Holds the sabotage results that the Claude Opus 4.6 System Card leaves to a separate report.

  • DeepSeek-V4 technical report and model card

    DeepSeek

    A first reading found no safety results given as numbers.

  • Gemini 2.5 Computer Use model card

    Google DeepMind

    A first reading found no safety results given as numbers.

  • Gemini 2.5 Flash-Lite model card

    Google DeepMind

    Reports safety results as changes from an earlier model; not extracted yet.

  • Gemini Omni Flash model card

    Google DeepMind · 27 Aug 2026

    A first reading found no table of safety results.

  • Kimi K2.5 model card

    Moonshot AI

    A first reading found no safety results given as numbers.

  • ChatGPT agent system card

    OpenAI · 17 Jul 2025

    Not collected for the first version of the dataset.

  • codex-1 addendum

    OpenAI · 16 May 2025

    Not collected for the first version of the dataset.

  • Deep research system card

    OpenAI · 25 Feb 2025

    Not collected for the first version of the dataset.

  • GPT-4.5 system card

    OpenAI · 27 Feb 2025

    Not collected for the first version of the dataset.

  • GPT-5 sensitive-conversations addendum

    OpenAI · 27 Oct 2025

    Not collected for the first version of the dataset.

  • gpt-oss-safeguard report

    OpenAI · 29 Oct 2025

    Not collected for the first version of the dataset.

  • o3-mini system card

    OpenAI · 31 Jan 2025

    Not collected for the first version of the dataset.

  • Operator system card

    OpenAI · 23 Jan 2025

    Not collected for the first version of the dataset.

  • Sora 2 system card

    OpenAI · 30 Sep 2025

    Not collected for the first version of the dataset.

  • GLM-5 / GLM-5.3 reports

    Zhipu AI

    A first reading found no safety results given as numbers.

  • Anthropic “Agentic Misalignment in Summer 2026” cross-lab study

    Third-party evaluation

    A study of several developers’ models, published by one developer; a candidate for inclusion. alignment.anthropic.com (opens an external site)

  • SaferAI evaluation of GLM-5.2

    Third-party evaluation

    An evaluation by an independent organisation rather than by the developer; a candidate for inclusion.

  • UK AISI and US CAISI evaluations of open-weight models (Kimi K3, DeepSeek V4 Pro)

    Third-party evaluation

    Evaluations run by government institutes rather than by the developers; how to record them is not decided yet.