AI Safety Moves From Benchmarks to Courts and Covert Agents
Thursday, September 24, 2026 · 10 min

OpenAI faces a British Columbia lawsuit over alleged ignored ChatGPT risk flags, while new security work shows Claude-assisted hacking, MCP tool hijacking, and covert agent collusion pressuring AI safety governance beyond benchmark scores.
Listen
Show notes
OpenAI faces a British Columbia lawsuit over alleged ignored ChatGPT risk flags, while new security work shows Claude-assisted hacking, MCP tool hijacking, and covert agent collusion pressuring AI safety governance beyond benchmark scores.
In this episode
- British Columbia Is Suing OpenAI Over a School Shooting—and Testing the Management-Responsibility Doctrine in Court – Forkast — Forkast
British Columbia Is Suing OpenAI Over a School Shooting—and Testing the Management-Responsibility Doctrine in Court – Forkast # British Columbia Is Suing OpenAI Over a School Shooting—and Testing the Management-Responsibility Doctrine in Court OpenAI's safety team flagged the shooter's ChatGPT conversations and recommended contacting police. Leadership overruled them. Now a foreign government…
- Researchers used Claude to hack OpenAI | Malwarebytes — Malwarebytes
Researchers used Claude to hack OpenAI | Malwarebytes # Researchers used Claude to hack OpenAI by Danny Bradbury | September 22, 2026 We’ve heard of OpenAI’s AI agents running amok and hacking other companies. Now, a cybersecurity company has turned the tables on the ChatGPT operator by using AI to help hack OpenAI itself. The hack, which also exposed a bug affecting dozens of other major…
- A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem — arXiv
A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem # A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem Laizhen Li Affiliation: Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences Affiliation: University of Chinese Academy of Sciences Xuan Wang Affiliation: Nanyang Technological University Peicheng Zhao Affiliation: Shenzhen Institutes of Advanced Technology,…
- AI Agents Teamed Up to Cheat at Blackjack. Their Collusion Is Getting Harder to Spot | WIRED — WIRED
# AI Agents Teamed Up to Cheat at Blackjack. Their Collusion Is Getting Harder to Spot | WIRED Author: Will Knight Published: 2026-09-23T14:30:00-04:00 Source: wired.com (wired.com) Language: en ## Story AI Agents Teamed Up to Cheat at Blackjack. Their Collusion Is Getting Harder to Spot | WIRED Will Knight Sep 23, 2026 2:30 PM # AI Agents Teamed Up to Cheat at Blackjack. Their Collusion Is…
- WorkspaceBench: Evaluating Interpretability Methods for the Global Workspace — AI Alignment Forum
TL;DR We introduce WorkspaceBench, a set of evaluations for how well an activation-to-text tool can read the contents of the “global workspace” of a model, i.e. the intermediate variables during a forward pass. The benchmark comprises 3,356 questions across 27 eval families, spanning topics in safety, logical reasoning, and multihop computation, with a subset for single-token-output tools. A…