Versions

Three versions are on record, oldest first. We hold a copy of one; the others are known only by their date.

  1. 8 Apr 2026

    First published version

    Date from
    our source registry
    Copy
    Known to exist; no copy held
  2. 26 May 2026

    Revision

    Date from
    the document's changelog
    Copy
    Known from the changelog; no copy held

    What changed instruction-hierarchy comparator scores fixed after a bug

  3. 26 Sep 2026

    Date retrieved

    Copy retrieved 26 Sep 2026

    Date from
    the date we retrieved it; the copy states no version date
    Copy
    Copy retrieved on 26 Sep 2026
    Values
    22 values recorded from this version

No version has a file hash or an archived snapshot yet. From dataset v0.2 each retrieved version carries both (Methodology §7).

Revisions

No revisions are recorded for this document. Without copies of the earlier versions, changes between them are not recorded value by value; what we know of each version is listed below.

Values

Every value we recorded from this document, grouped by metric family and ordered by where the document prints it. Location is the section, table or page as the document numbers it. 5 of the 22 have been blind-verified: a second reader found the same value without seeing ours.

KeyVerifiedUnverifiedDisputedCorrected read off a figure or stated in wordsA value opens its source and history.

M1 Evaluation awareness

3 values

M1 Evaluation awareness: values in the Muse Spark Safety & Preparedness Report
ModelEvaluationConditionValueLocationChecked
Muse SparkVerbalized eval awarenessShare of samples verbalizing evaluationinternal evals4.1.11Verified
Muse SparkVerbalized eval awarenessShare of samples verbalizing evaluationpublic benchmarks4.1.11Verified
Muse SparkPrompted evaluation awarenessEval-vs-deployment discriminationNone statedTable 1Verified

M2 Reward hacking

1 value

M2 Reward hacking: values in the Muse Spark Safety & Preparedness Report
ModelEvaluationConditionValueLocationChecked
Muse SparkImpossibleBenchCheating ratewithout system mitigationsTable 2Unverified

M3 Sabotage and sandbagging

2 values

M3 Sabotage and sandbagging: values in the Muse Spark Safety & Preparedness Report
ModelEvaluationConditionValueLocationChecked
Muse SparkGDM-StealthStealth capability scorewithout system mitigationsTable 1Unverified
Muse SparkSHADE-ArenaSabotage/stealth success scorewithout system mitigationsTable 1Verified

M4 Misalignment audits

2 values

M4 Misalignment audits: values in the Muse Spark Safety & Preparedness Report
ModelEvaluationConditionValueLocationChecked
Muse SparkAgentic MisalignmentHarmful action ratewithout system mitigationsTable 2Unverified
Muse SparkAlignment fakingBehavioral difference monitored vs notwithout system mitigationsTable 2Unverified

M5 Honesty and hallucination

2 values

M5 Honesty and hallucination: values in the Muse Spark Safety & Preparedness Report
ModelEvaluationConditionValueLocationChecked
Muse SparkDeceptionBenchDeception ratewithout system mitigationsTable 2Unverified
Muse SparkMASKHonesty ratewithout system mitigationsTable 2Unverified

M6 Sycophancy

1 value

M6 Sycophancy: values in the Muse Spark Safety & Preparedness Report
ModelEvaluationConditionValueLocationChecked
Muse SparkInternal sycophancySycophancy ratewithout system mitigationsTable 2Unverified

M7 Harmful compliance and over-refusal

2 values

M7 Harmful compliance and over-refusal: values in the Muse Spark Safety & Preparedness Report
ModelEvaluationConditionValueLocationChecked
Muse SparkAgentHarmCompliance rate on harmful agentic taskswithout system mitigationsTable 2Unverified
Muse Spark (Meta AI system)BioTIER refusalsRefusal rate on harmful biowith system mitigationsTable 2Unverified

M8 Jailbreak robustness

2 values

M8 Jailbreak robustness: values in the Muse Spark Safety & Preparedness Report
ModelEvaluationConditionValueLocationChecked
Muse SparkFORTRESSAttack success ratewithout system mitigationsTable 2Unverified
Muse SparkStrongREJECTAttack success rate (adaptive multi-turn)without system mitigationsTable 2Verified

M9 Prompt injection

3 values

M9 Prompt injection: values in the Muse Spark Safety & Preparedness Report
ModelEvaluationConditionValueLocationChecked
Muse SparkAgent Red Teaming (ART)Agent red-team ASR pass@1Run by Gray Swanwithout system mitigationsTable 2Unverified
Muse SparkAgentDojoPrompt-injection ASR pass@1without system mitigationsTable 2Unverified
Muse SparkIHEvalIH scorewithout system mitigationsTable 2; 4.1.1Unverified

M10 Dangerous capabilities and risk determinations

4 values

M10 Dangerous capabilities and risk determinations: values in the Muse Spark Safety & Preparedness Report
ModelEvaluationConditionValueLocationChecked
Muse SparkRisk determinationRisk tier, chemical and biologicalpre-mitigation1.1.1Unverified
Muse SparkRisk determinationRisk tier, cyber and loss of controlpost-mitigation1.1.1Unverified
Muse SparkCybenchPass@1without system mitigationsTable 1Unverified
Muse SparkVCTAccuracywithout system mitigationsTable 1Unverified

Extraction coverage

What we read of this document, and where each value was read.

Our note First report under Meta Advanced AI Scaling Framework v2

All 22 values were read in the document itself.

All 22 values were extracted for version 0 of the dataset through a web reader, which did not always reach the later sections of long PDFs (Methodology §3).