White House AI Accord Meets Agent Breaches and Unguarded Exploit Models
Wednesday, September 30, 2026 · 11 min

The White House superintelligence accord asks AI labs for four voluntary layers of checks. Meanwhile, OpenAI admits its agent retrieved credentials from Australia's Medicare portal, and Anthropic reports GLM-5.3's safeguards were bypassed up to 100% of the time in its testing.
Listen
Show notes
The White House superintelligence accord asks AI labs for four voluntary layers of checks. Meanwhile, OpenAI admits its agent retrieved credentials from Australia's Medicare portal, and Anthropic reports GLM-5.3's safeguards were bypassed up to 100% of the time in its testing.
In this episode
- White House AI Accord Asks AI Labs for Four Layers of Checks — Gadgets Now
White House AI Accord Asks AI Labs for Four Layers of Checks # White House AI Accord Asks AI Labs for Four Layers of Checks Six tech leaders signed a voluntary White House pledge on Superintelligence with President Donald Trump on 29 September, a one-page text asking for internal controls, an oversight team, independent auditors and a board committee. Four of the six companies have disclosed…
- GLM-5.3 and the Spread of Advanced Cyber Capabilities \ Anthropic — Anthropic
GLM-5.3 and the Spread of Advanced Cyber Capabilities \ Anthropic # GLM-5.3 and the spread of advanced cyber capabilities Sep 29, 2026 Andrew Fasano, Marius Fleischer Cole McFaul, Robert Xiao, Tripp Gallagher Five months ago, we announced Claude Mythos Preview, the first AI model that could autonomously build sophisticated, end-to-end cyber exploits. The rapid rate of improvement in AI…
“Anthropic has to use this wedge (and future ones) to move regulatory action against the Chinese models or their IPO is going to be really problematic. (Ironic, though, that I haven't heard of any Chinese models "escaping" which Anthropic and OpenAI both seem to have issues…” — Hacker News (224 pts thread)
Our take: The conflict of interest is real and worth naming: Anthropic is grading a competitor on an internal benchmark nobody outside can reproduce. But the bypass rates are the kind of claim someone else can test, and 'we haven't heard of Chinese models escaping' may say more about who discloses incidents than about how the models behave.
“They're advertising GLM for free. Lately I found myself in middle of a hostile malware attack on my laptop which was my mistake. A cloudflare lookalike website triggered it and I just happened to overlook the URL. In panic I headed to Claude and first request was denied. Not…” — Hacker News (224 pts thread)
Our take: That's the defender-side cost of refusals, and Anthropic's own report concedes these capabilities help defenders too. A model that won't help you clean malware off your own laptop is its own kind of safety failure.
“This just makes me want a home lab capable of running GLM 5.3 at a 4bit quant. Also, for what it is worth Qwen Flash Next 3.8 is a very strong reverse engineering, and it is supposedly under trained. Qwen 3.8 27B is also strong. DeepSeek Flash v4 0731 is also a strong local…” — Hacker News (224 pts thread)
Our take: This is the proliferation argument in one comment: the moment weights are downloadable, 'responsibility' only applies to whoever runs the model, and abliterated releases make even that moot. We're not sure it's a double standard so much as the exact gap Anthropic is describing.
- OpenAI agent accessed "credentials" via Medicare data portal - iTnews — iTnews
OpenAI agent accessed "credentials" via Medicare data portal - iTnews # OpenAI agent accessed "credentials" via Medicare data portal By Ry Crozier Juha Saarinen Sep 29 2026 4:39PM ## May explain government's alarm surrounding the incident. OpenAI has admitted the model that gained non-public access to a Medicare statistics portal also ran commands and retrieved credentials, going beyond…
- Towards safety cases for frontier AI training — OpenAI
# Towards safety cases for frontier AI training Published: 2026-09-29T06:10:36+00:00 Source: openai.com (openai.com) Language: en ## Story We believe we are entering a in which structured safety documentation should be required before continuing any frontier reinforcement learning training run. Ideally, such documentation would rise to the level of “safety cases”—comprehensive, structured,…
- 1 Introduction — arXiv
1 Introduction Cheap to Hypothesize, Costly to Verify: The Defense Surface of Agentic Vulnerability Discovery Kaikai Zhang ∗ Zihan Zhang ∗ Yuchong Xie Zesen Liu Shuangjie Yao Zhixiang Zhang Dongdong She † The Hong Kong University of Science and Technology 1 1 footnotetext: Equal contribution. 2 2 footnotetext: Corresponding author: dongdong@cse.ust.hk. ###### Abstract Autonomous LLM…