AI Safety Daily

Anthropic Opens Models to Australia as Evaluators Face Their Own Weak Spots

Wednesday, October 7, 2026 · 9 min

AI Safety Daily cover art

Anthropic says Australia's AI Safety Institute will soon run independent evaluations of its frontier models. Meanwhile METR shows an agent could have rewritten the Inspect transcripts reviewers rely on, and California is setting rules for the evaluators themselves.

Listen

Listen to the audio episode

Read the episode transcript

Show notes

Anthropic says Australia's AI Safety Institute will soon run independent evaluations of its frontier models. Meanwhile METR shows an agent could have rewritten the Inspect transcripts reviewers rely on, and California is setting rules for the evaluators themselves.

In this episode

  1. Anthropic opening its models up to Australian AI Safety Institute — Forbes Australia

    Anthropic opening its models up to Australian AI Safety Institute # Anthropic will let Australia stress test dangerous new models By Daniel Van Boom Published on October 6, 2026 ##### The AI giant told a parliamentary inquiry that it’s finalising an arrangement for Australia’s AI Safety Institute to evaluate its latest and greatest (and most threatening) models. Dario Amodei, CEO of…

  2. Research Note: Filtering Subversion-Relevant Information From Pretraining Data Is Feasible — LessWrong — LessWrong

    # Research Note: Filtering Subversion-Relevant Information From Pretraining Data Is Feasible * By [Kyle O’Brien](/users/kyle-o-brien), [Spencer Kitts](/users/spencer-kitts-1), [Cam](/users/cam-tice), [Alek Westover](/users/alek-westover) * 2026-10-05 18:26:12Z * 45 points * Tag: [AI Control](/w/ai-control) * Tag: [Alignment Pretraining](/w/alignment-pretraining) * Tag: [AI](/w/ai) * Frontpage *…

  3. AI systems could cover up misbehavior — METR

    Recent AI misalignment incidents have shown AI systems capably pursuing goals their human supervisors would not approve of, like hacking other companies. Fortunately, current AIs still seem relatively bad at concealing their misbehavior from human reviewers: these recent misalignment incidents have left significant amounts of evidence in reasoning traces, logs, and other telemetry. This has made…

  4. Newest Slate of AI Laws Keeps California at the Forefront of US Regulation — Latham & Watkins

    Newest Slate of AI Laws Keeps California at the Forefront of US Regulation # Newest Slate of AI Laws Keeps California at the Forefront of US Regulation October 5, 2026 During the 2026 legislative session, California enacted more than two dozen AI laws regulating transparency, chatbot safety, automated decision-making, and other domains. ## Key points - Expanded AI transparency…

  5. ‘Evaluators’ are supposed to keep AI from killing us all. No pressure | LAist — LAist

    ‘Evaluators’ are supposed to keep AI from killing us all. No pressure | LAist By Khari Johnson | CalMatters Published October 6, 2026 at 9:37 AM PDT # ‘Evaluators’ are supposed to keep AI from killing us all. No pressure By Khari Johnson | CalMatters Published October 6, 2026 at 9:37 AM PDT Artificial intelligence evaluators are increasingly called upon to audit powerful new…