Anthropic Opens Models to Australia as Evaluators Face Their Own Weak Spots
Wednesday, October 7, 2026 · 9 min

Anthropic says Australia's AI Safety Institute will soon run independent evaluations of its frontier models. Meanwhile METR shows an agent could have rewritten the Inspect transcripts reviewers rely on, and California is setting rules for the evaluators themselves.
Listen
Show notes
Anthropic says Australia's AI Safety Institute will soon run independent evaluations of its frontier models. Meanwhile METR shows an agent could have rewritten the Inspect transcripts reviewers rely on, and California is setting rules for the evaluators themselves.
In this episode
- Anthropic opening its models up to Australian AI Safety Institute — Forbes Australia
Anthropic opening its models up to Australian AI Safety Institute # Anthropic will let Australia stress test dangerous new models By Daniel Van Boom Published on October 6, 2026 ##### The AI giant told a parliamentary inquiry that it’s finalising an arrangement for Australia’s AI Safety Institute to evaluate its latest and greatest (and most threatening) models. Dario Amodei, CEO of…
- Research Note: Filtering Subversion-Relevant Information From Pretraining Data Is Feasible — LessWrong — LessWrong
# Research Note: Filtering Subversion-Relevant Information From Pretraining Data Is Feasible * By [Kyle O’Brien](/users/kyle-o-brien), [Spencer Kitts](/users/spencer-kitts-1), [Cam](/users/cam-tice), [Alek Westover](/users/alek-westover) * 2026-10-05 18:26:12Z * 45 points * Tag: [AI Control](/w/ai-control) * Tag: [Alignment Pretraining](/w/alignment-pretraining) * Tag: [AI](/w/ai) * Frontpage *…
- AI systems could cover up misbehavior — METR
Recent AI misalignment incidents have shown AI systems capably pursuing goals their human supervisors would not approve of, like hacking other companies. Fortunately, current AIs still seem relatively bad at concealing their misbehavior from human reviewers: these recent misalignment incidents have left significant amounts of evidence in reasoning traces, logs, and other telemetry. This has made…
- Newest Slate of AI Laws Keeps California at the Forefront of US Regulation — Latham & Watkins
Newest Slate of AI Laws Keeps California at the Forefront of US Regulation # Newest Slate of AI Laws Keeps California at the Forefront of US Regulation October 5, 2026 During the 2026 legislative session, California enacted more than two dozen AI laws regulating transparency, chatbot safety, automated decision-making, and other domains. ## Key points - Expanded AI transparency…
- ‘Evaluators’ are supposed to keep AI from killing us all. No pressure | LAist — LAist
‘Evaluators’ are supposed to keep AI from killing us all. No pressure | LAist By Khari Johnson | CalMatters Published October 6, 2026 at 9:37 AM PDT # ‘Evaluators’ are supposed to keep AI from killing us all. No pressure By Khari Johnson | CalMatters Published October 6, 2026 at 9:37 AM PDT Artificial intelligence evaluators are increasingly called upon to audit powerful new…