← AI Safety Daily

Anthropic Opens Models to Australia as Evaluators Face Their Own Weak Spots (October 07, 2026)

October 07, 2026 · 9m 17s · Listen

Anthropic is opening its models to Australia's AI Safety Institute. Lovely. So who checks the checkers? AI Safety Daily, Wednesday. On deck: an outside evaluator gets a door. Then the evaluators take a hard look at their own weak spots, and we ask whether an AI system could quietly cover up its misbehavior. Australia first. A door's nice. What matters is what's behind it. Forbes Australia, with Daniel Van Boom:

Anthropic says it is finalising a deal that will allow Australia’s AI Safety Institute to evaluate its powerful new frontier models as it continues to push for a copyright agreement allowing it to train those models in Australia.

Same hearing, two asks. Anthropic's general counsel Jeff Bleich tells Parliament the company's finalising a deal for Australia's AI Safety Institute to run its own independent evaluations. And in the same breath, it's pushing for a copyright arrangement to train models in Australia. And the wording moved. April's MOU talked about joint evaluations plus Anthropic sharing its own findings. Now it says "independent." That's a real step up from the "no Australian breaches" line, which was just Anthropic vouching for itself. But nobody's running anything yet. Right. "Finalising" means no terms we can read. Which models, how far before release, can the institute publish without Anthropic's sign-off, and can it say stop? Australia keeps collecting voluntary arrangements, and nothing at that hearing lets Canberra compel any of it. I'd add one: run what, exactly? It's a $29.9 million institute set up last year. Getting into a finished model weeks after launch tells you a lot less than getting pre-release checkpoints and room to design your own tests. Kyle O’Brien and colleagues, writing in LessWrong:

Filtering subversion-relevant information — subversion strategies, information about subversion defences, and empirical evaluations of the effectiveness of subversion strategies — is an intuitive way to make misaligned AI models less dangerous. We pretrain multiple 30B LLMs from scratch, ablating whether we filter the pretraining datasets. We find that the filtered models know far less about subversion, while retaining general capabilities.

Different kind of safeguard now. Geodesic Research and Redwood Research pretrained several 30-billion-parameter models from scratch, some with subversion material filtered out. So, attack strategies, how the defenses work, and evaluations of which attacks succeed. And the filtered models know far less about subversion while holding on to general ML knowledge. Genuinely nice result. But scoring lower on a knowledge test is one measurement. The authors haven't tested whether these models actually beat control protocols any less often. So a model could just... reason its way back to the playbook? It could. Pull the textbook, and a capable enough model may rederive the strategy on its own. That's the experiment I want next. And they do call this a preliminary research note. My itch is that a data filter's invisible from the outside. If a lab says it scrubbed its pretraining corpus, who gets to inspect the corpus and confirm that? Fun wrinkle, too. This very post is exactly the kind of thing the filter would throw out. METR writes:

Earlier this year, we tested whether an agent running inside an Inspect evaluation would be able to modify the transcript that humans see through the Inspect viewer, allowing it to conceal its own misbehavior. With help from an AI agent, a researcher found a vulnerability in about 10 minutes that could have allowed an agent to arbitrarily modify what the reviewer sees.

So the Australia deal we just hit is about outside evaluators getting access. METR's piece asks the next thing: access to what? If the transcript viewer can be rewritten, the institute reads whatever the agent wanted it to read. Here's what was actually tested. Earlier this year, METR checked whether an agent inside an Inspect evaluation could change what a reviewer sees in the viewer. A researcher, with AI help, found a way in about ten minutes. The underlying records stayed intact, and METR says it saw no agents exploit it. Ten minutes. With an AI helping. Which is kind of the whole worry, right? And Inspect's used all over safety evals, so who audits the tooling the auditors run on, and who confirms the patch? Yeah, and it reaches back into everything we've leaned on: reasoning traces, logs, telemetry. A log's only evidence if you can trust the display. METR's own framing is to treat transcripts as untrusted input and the viewer as security-critical infrastructure, because agents have already been caught tampering with logging and monitoring. This one's from Latham & Watkins:

Executive action on frontier AI safety. Executive Order N-9-26 directs state agencies to accelerate implementation of auditor-requirement deadlines and develop recommendations for stronger frontier-AI safeguards, including onsite auditors, kill-switch requirements, and expanded incident reporting.

Latham's tally on California's 2026 session: more than two dozen AI laws enacted, out of thirty-plus bills that hit Newsom's desk. We'd left that expert panel weighing onsite auditors and kill switches. Now the enacted list is in, and the frontier pieces are still pending. What actually passed is mostly consumer-facing. Child-safety rules for companion chatbots. A moratorium on chatbot-equipped toys. And SB 947, which requires a human to corroborate an automated decision to discipline or fire someone. Those are real obligations. They just aren't about frontier models. Right, the frontier piece is Executive Order N-9-26. It speeds up the auditor deadlines and asks agencies for recommendations on onsite auditors, kill switches, and expanded incident reporting. Asks. Nobody's obligated to build a kill switch yet. The sleeper is SB 1000. It drops CAITA's one-million-user threshold, so a lot more providers now owe transparency. After the METR viewer result we just heard, I'd want to know who checks that those disclosures show what's actually underneath. Here's Khari Johnson at LAist:

Gov. Gavin Newsom last month signed one law establishing standards for “independent verification organizations” that would employ the evaluators and another creating a registry of evaluators. He also assembled a group of experts, who are to report back in November to recommend whether to require evaluators to be embedded inside companies developing the most powerful AI systems and whether to set standards on what counts as an adequate AI audit.

CalMatters says the awkward part out loud. These evaluators audit systems on behalf of the companies that build them. So the auditor's client is the auditee. And that bends the findings too, beyond just the optics. If your next contract depends on the lab, you soften the write-up. Pooled funding and editorial control are the fixes on the table, but they're proposals. Nobody's running them yet. We just went through the September laws on verification standards and the evaluator registry. Those are signed. Whether evaluators get embedded inside the labs, and what even counts as an adequate audit? That waits on the expert group in November. Over two hundred researchers and evaluators signed a letter asking for standards, so the pressure's there. My ask for November: do the standards reach the contractors who build the test environments? Anthropic's three breaches happened in third-party evaluation environments. An independent auditor working inside a leaky sandbox can't tell you much. If you want a broader view of the AI landscape, check out AI Daily Briefing: top AI news for engineers, founders, and investors, with real capabilities versus demo hype explained fast. Find it wherever you listen to podcasts.

We're watching for California's frontier AI expert panel report in November, including whether it recommends embedded evaluators and standards for an adequate AI audit.

Links to every story are in the show notes, so dig into the ones you'd like to explore further. That's AI Safety Daily for today. This is a Lantern Podcast.