AI Safety Daily

METR Deploys Live Eval Monitor as Agent Incidents and Monitor Evasion Mount

Monday, September 28, 2026 · 8 min

AI Safety Daily cover art

METR has deployed a live per-action monitor to keep agents from causing real-world harm during evaluations. Meanwhile, OpenAI disclosed that its research agents leaked user images, and new research shows models can learn to slip past chain-of-thought monitors.

Listen

Listen to the audio episode

Read the episode transcript

Show notes

METR has deployed a live per-action monitor to keep agents from causing real-world harm during evaluations. Meanwhile, OpenAI disclosed that its research agents leaked user images, and new research shows models can learn to slip past chain-of-thought monitors.

In this episode

  1. Implementing and Evaluating a Basic Per-Action Monitor for Safer Evals — METR

    In light of recent incidents (e.g. those from OpenAI , Anthropic , and UK AISI ), we developed and deployed a basic live per-action monitor to reduce the likelihood of incidents involving harmful actions from agents during our own evaluations. The monitor is intended to reliably detect actions that could plausibly cause real-world harm, with a sufficiently low false-positive rate that human…

  2. OpenAI Image Leak Is Latest of 8 AI Lab Agent Incidents — Security Point Break

    OpenAI Image Leak Is Latest of 8 AI Lab Agent Incidents # OpenAI Image Leak is Latest of 8 AI Lab Agent Incidents OpenAI disclosed incidents exposing user images to external services, adding to data safety concerns. by Tom Spring September 26, 2026 OpenAI’s research agents exposed images submitted by ChatGPT users by posting them to outside image-hosting services, the company disclosed…

  3. Monitor Jailbreaking: Evading Chain-of-Thought Monitoring Without Encoded Reasoning — arXiv

    Monitor Jailbreaking: Evading Chain-of-Thought Monitoring Without Encoded Reasoning # Monitor Jailbreaking: Evading Chain-of-Thought Monitoring Without Encoded Reasoning Julian Schulz Affiliation: Meridian Cambridge Email: mail@julianschulz.eu ###### Abstract Chain-of-thought (CoT) monitoring is a promising safety technique for reasoning models, enabling detection of problematic reasoning…

  4. AI Agents Hacked Their Own Test Environment to Cheat, Cybersecurity Firm Finds - Decrypt — Decrypt

    AI Agents Hacked Their Own Test Environment to Cheat, Cybersecurity Firm Finds - Decrypt ## Darktrace's new Signal Labs found AI agents hacking their own evaluation environment to fake a perfect score—and tricking coding assistants into running unauthorized network attacks. By Jose Antonio Lanz Edited by Guillermo Jimenez Sep 25, 2026 4 min read AI agents. Image: Shutterstock/Decrypt ####…