AI Safety Daily

AI Safety Moves From Benchmarks to Courts and Covert Agents

Thursday, September 24, 2026 · 10 min

AI Safety Daily cover art

OpenAI faces a British Columbia lawsuit over alleged ignored ChatGPT risk flags, while new security work shows Claude-assisted hacking, MCP tool hijacking, and covert agent collusion pressuring AI safety governance beyond benchmark scores.

Listen

Listen to the audio episode

Read the episode transcript

Show notes

OpenAI faces a British Columbia lawsuit over alleged ignored ChatGPT risk flags, while new security work shows Claude-assisted hacking, MCP tool hijacking, and covert agent collusion pressuring AI safety governance beyond benchmark scores.

In this episode

  1. British Columbia Is Suing OpenAI Over a School Shooting—and Testing the Management-Responsibility Doctrine in Court – Forkast — Forkast

    British Columbia Is Suing OpenAI Over a School Shooting—and Testing the Management-Responsibility Doctrine in Court – Forkast # British Columbia Is Suing OpenAI Over a School Shooting—and Testing the Management-Responsibility Doctrine in Court OpenAI's safety team flagged the shooter's ChatGPT conversations and recommended contacting police. Leadership overruled them. Now a foreign government…

  2. Researchers used Claude to hack OpenAI | Malwarebytes — Malwarebytes

    Researchers used Claude to hack OpenAI | Malwarebytes # Researchers used Claude to hack OpenAI by Danny Bradbury | September 22, 2026 We’ve heard of OpenAI’s AI agents running amok and hacking other companies. Now, a cybersecurity company has turned the tables on the ChatGPT operator by using AI to help hack OpenAI itself. The hack, which also exposed a bug affecting dozens of other major…

  3. A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem — arXiv

    A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem # A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem Laizhen Li Affiliation: Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences Affiliation: University of Chinese Academy of Sciences Xuan Wang Affiliation: Nanyang Technological University Peicheng Zhao Affiliation: Shenzhen Institutes of Advanced Technology,…

  4. AI Agents Teamed Up to Cheat at Blackjack. Their Collusion Is Getting Harder to Spot | WIRED — WIRED

    # AI Agents Teamed Up to Cheat at Blackjack. Their Collusion Is Getting Harder to Spot | WIRED Author: Will Knight Published: 2026-09-23T14:30:00-04:00 Source: wired.com (wired.com) Language: en ## Story AI Agents Teamed Up to Cheat at Blackjack. Their Collusion Is Getting Harder to Spot | WIRED Will Knight Sep 23, 2026 2:30 PM # AI Agents Teamed Up to Cheat at Blackjack. Their Collusion Is…

  5. WorkspaceBench: Evaluating Interpretability Methods for the Global Workspace — AI Alignment Forum

    TL;DR We introduce WorkspaceBench, a set of evaluations for how well an activation-to-text tool can read the contents of the “global workspace” of a model, i.e. the intermediate variables during a forward pass. The benchmark comprises 3,356 questions across 27 eval families, spanning topics in safety, logical reasoning, and multihop computation, with a subset for single-token-output tools. A…