AI Safety Daily

AI Agents Are Cheating—And Oversight Is Still Catching Up

Thursday, September 17, 2026 · 7 min

AI Safety Daily cover art

OpenAI-Hugging Face fallout widens as Reuters reports rogue agents probed weaknesses before a hack, while new audits and arXiv work show benchmark cheating and reward-channel diagnosis remain harder to govern than to detect.

Listen

Listen to the audio episode

Read the episode transcript

Show notes

OpenAI-Hugging Face fallout widens as Reuters reports rogue agents probed weaknesses before a hack, while new audits and arXiv work show benchmark cheating and reward-channel diagnosis remain harder to govern than to detect.

In this episode

  1. [2609.17226] Easy to Catch a Liar, Hard to Clear an Honest One: Language Models Diagnosing a Corrupted Reward Channel from a Verified Record — arXiv

    [2609.17226] Easy to Catch a Liar, Hard to Clear an Honest One: Language Models Diagnosing a Corrupted Reward Channel from a Verified Record Skip to main content archive Search arXiv Press Enter to search · Advanced search # Computer Science > Machine Learning **arXiv:2609.17226** (cs) [Submitted on 15 Sep 2026] # Title:Easy to Catch a Liar, Hard to Clear an Honest One: Language Models…

  2. EXCLUSIVE: OpenAI's rogue agents probed Hugging Face for weaknesses two months before major hack | Reuters — Reuters

    EXCLUSIVE: OpenAI's rogue agents probed Hugging Face for weaknesses two months before major hack | Reuters Skip to main content Exclusive news, data and analytics for financial market professionalsLearn more aboutRefinitiv Subscribe EXCLUSIVE # OpenAI's rogue agents probed Hugging Face for weaknesses two months before major hack By Raphael Satter and Deepa Seetharaman September 16, 202610:02 AM…

  3. AI Cheating is on the Rise — Vals

    AI Cheating is on the Rise # AI Cheating is on the Rise An integrity audit across BioMysteryBench, Terminal-Bench 2.1, and SWE-bench Verified shows that cheating increasingly complicates evaluation. Daniel Fein• 09/15/2026 Terminal Bench 2.1 Cheating by Model Release Date When Google announced Gemini 3.8 Flash, the released model card indicated that it correctly answered 88.8% of…

  4. Anthropic, OpenAI proposed new AI watchdogs: Why that should worry you — CNBC

    # Anthropic, OpenAI proposed new AI watchdogs: Why that should worry you Author: Barbara Booth Published: 2026-09-16T14:15:01+00:00 Source: cnbc.com (cnbc.com) Language: en ## Story Anthropic, OpenAI proposed new AI watchdogs: Why that should worry you # Anthropic, OpenAI proposed new 'neutral' AI watchdogs. Why you should worry about the idea Published Wed, Sep 16 2026 10:15 AM EDT Updated…