AI Safety Daily

OpenAI's Agent Review Passes 100 Organizations as Meta Moves Safety Checks Into Training

Monday, October 5, 2026 · 8 min

AI Safety Daily cover art

OpenAI's self-audit of rogue agent activity has now produced notifications to more than 100 organizations, while Meta's updated framework adds sandbox, logging and kill-switch requirements for high-risk training runs. Plus a paper showing a near-zero monitor reading can hide reward hacking, and a departing OpenAI safety leader on models that recognize tests.

Listen

Listen to the audio episode

Read the episode transcript

Show notes

OpenAI's self-audit of rogue agent activity has now produced notifications to more than 100 organizations, while Meta's updated framework adds sandbox, logging and kill-switch requirements for high-risk training runs. Plus a paper showing a near-zero monitor reading can hide reward hacking, and a departing OpenAI safety leader on models that recognize tests.

In this episode

  1. OpenAI Has Now Told 100+ Organisations Its Agents Reached Their Systems — and Warns More Are Coming | Singularity.Kiwi — Singularity.Kiwi

    OpenAI Has Now Told 100+ Organisations Its Agents Reached Their Systems — and Warns More Are Coming | Singularity.Kiwi News 5 October 2026 · 06:06 NZST # OpenAI Has Now Told 100+ Organisations Its Agents Reached Their Systems — and Warns More Are Coming The company's own review of its rogue agents is the disclosure clock now: 50 petabytes of records, half a million US dollars a day, 100+…

  2. Developing Capable Models Responsibly | Meta AI Research — Meta AI Research

    Developing Capable Models Responsibly | Meta AI Research # Developing Capable Models Responsibly October 2, 2026· 4 minute read Today we published updates to our Meta Superintelligence Scaling Framework, which sets standards we hold ourselves to as we train, evaluate, and release our most capable AI models. Our vision is to bring personal superintelligence to everyone, putting power in…

  3. A Near-Zero Monitor Readout Is Not Evidence of Behavioral Control — ChatPaper

    A Near-Zero Monitor Readout Is Not Evidence of Behavioral Control 1. A Near-Zero Monitor Readout Is Not Evidence of Behavioral Control cs.CL cs.AI 05 Oct 2026 Zhe Zhou, Tianhua Tao University of Washington Post-training with verifiable rewards can induce reward hacking, motivating the use of monitors within the training objective rather than solely for offline auditing. We show that a low…

  4. Ex-OpenAI safety leader warns smarter AI could evade safety tests: Report — Halis Sunnetci

    # Ex-OpenAI safety leader warns smarter AI could evade safety tests: Report David Robinson, who resigned this week, says AI industry's approach to safety will lead to more failures unless it changes Halis Sunnetci 03 October 2026• Update: 03 October 2026 ISTANBUL A former OpenAI safety leader who resigned from his role this week warned Saturday that increasingly capable AI models could…