Redwood Research Blog
Avsnitt

“SOTA alignment assessments don’t strongly update us against misalignment” by Alexa Pan

Dela

Anthropic concluded in the April Mythos Preview alignment risk update that the model “does not possess any unknown propensities that would increase alignment risk.” The report argues that if Mythos Preview were coherently misaligned[1], it likely would have been detected by the assessment (following Anthropic, I will call this “reliability of the assessment”[3]).

While I agree with the report on the above bottom-line conclusions (substantially on priors), I think there are gaps in its argument which weaken the current assessment and might invalidate future assessments. In particular, the report often uses weak evidence to justify reliability.

  1. The report gives fairly weak experimental evidence for Mythos Preview having insufficient capabilities to evade monitoring. The model is plausibly often eval-aware and underelicited in the relevant capability evaluations. So, it might silently sandbag if coherently misaligned, or unintentionally underperform if otherwise misaligned.

    1. This limitation is important: one could argue that lack of covert capabilities for sophisticated sabotage (a subset of the capabilities I discuss here) is the single most load bearing argument in alignment risk reports.

    2. Authors of the report could have made calibrated guesses about Mythos Preview's covert capabilities, especially for covert sabotage, based on other [...]

---

Outline:

(03:16) How reliability fits into the overall safety argument

(05:22) Reliability claims by AI companies

(05:59) Reliability claims by external evaluators

(06:37) Alignment assessments are less reliable than developers claim

(07:21) 1: Measuring capabilities to covertly undermine alignment assessments

(10:15) Issues with evaluation awareness

(13:39) Issues with underestimating covert capabilities

(16:45) Issues with sandbagging rule-out

(19:20) 2: Stress-testing alignment assessments with auditing games

(20:26) An auditing failure with Mythos

(22:21) AuditBench results

(24:17) 3: Conditioning on misalignment should make us think that certain covert capabilities are better than expected

(26:20) Bottom line on the strength of current alignment assessments

(29:07) Conclusion

(29:44) Appendix:

(29:47) Why I focus on motive / alignment assessments in alignment risk reports

(30:59) Auditability vs. Trustedness

(33:27) More reliability claims by developers and third party evaluators

(33:43) Mythos Alignment Risk Update

(35:01) Opus 4.6 Sabotage Risk Report

(35:46) GPT 5.5 System card

(36:49) Muse Spark system card

(37:36) Mythos Alignment Risk Update, safety arguments against sandbagging

(38:40) UK AISI evaluations for Opus 4.7

(40:01) Past auditing games by Anthropic

(42:24) Anti-auditing capability measurements

(43:51) Conditioning on coherent misalignment updates us on certain covert capabilities

The original text contained 92 footnotes which were omitted from this narration.

---

First published:
July 31st, 2026

Source:
https://blog.redwoodresearch.org/p/sota-alignment-assessments-dont-strongly

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Flowchart showing argument structure forText about alignment assessment findings onText about AI modelText excerpt discussing AI models' misaligned goals and power-seeking motives.Text about alignment auditing exercises with highlighted passages.Mathematical equations defining likelihood ratio for detection probabilities.Document excerpt titledText excerpt discussing model alignment with highlighted sentence.Text listing three claims about misaligned goals in Claude Opus 4.6.Document section on GPT-5.5 misalignment risks and internal deployment.Text section titledText passage discussing sandbagging in AI model evaluations, highlighted.Text about Mythos Preview evaluation, with highlighted sandbagging statement.Highlighted text discussing evaluation awareness limiting research results interpretation.Table listing exercises and auditor types across five examples.Diagram linking AI behaviors to training, assessment, and control risks.

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Podden och tillhörande omslagsbild på den här sidan tillhör Redwood Research. Innehållet i podden är skapat av Redwood Research och inte av, eller tillsammans med, Poddtoppen.