AI Safety Daily

Reward-Hacking Monitors Meet the Verification Gap

Friday, September 18, 2026 · 9 min

AI Safety Daily cover art

GoodfireAI says activation probes can catch reward hacking in real time, while ImpossibleRubrics and alignment-drift tests show how reward criteria become attack surfaces. The governance side adds Proof-of-Control and a reminder that auditors cannot replace basic security.

Listen

Listen to the audio episode

Read the episode transcript

Show notes

GoodfireAI says activation probes can catch reward hacking in real time, while ImpossibleRubrics and alignment-drift tests show how reward criteria become attack surfaces. The governance side adds Proof-of-Control and a reminder that auditors cannot replace basic security.

In this episode

  1. Thread By @GoodfireAI - Models know when they’re reward... — Unrollnow

    Thread By @GoodfireAI - Models know when they’re reward... 7 Tweets 32 minutes ago Published Sep 18 • 32 minutes ago • 7 tweets • Read on X Models know when they’re reward hacking. But they still do it a ton - in 50-96% of rollouts we studied! We built activation monitors that detect the behavior behind the Hugging Face hack in real time. This can help us stop hacks now - and train future…

  2. ImpossibleRubrics Tests the Weaknesses of Generated Rubrics - CCTest — CCTest

    ImpossibleRubrics Tests the Weaknesses of Generated Rubrics - CCTest # ImpossibleRubrics: When Reward Criteria Become an Attack Surface Sep 16, 2026 3 min read The benchmark contains 169 impossible tasks, each paired with a verifiable oracle certificate, plus 48 answerable controls. Researchers generate rubrics downstream and then search for answers that violate the certificate while still…

  3. AI agents reportedly reward hack more often after doing so on a similar task · Digg — Digg

    # AI agents reportedly reward hack more often after doing so on a similar task · Digg Published: 2026-09-17T16:50:50+00:00 Source: digg.com (digg.com) Language: en ## Story # AI agents reportedly reward hack more often after doing so on a similar task Researchers report reward-hacking rates rising from 10% to 64% in one GPT-5.5 example when a similar earlier task involved reward hacking. Both…

  4. AI labs want in-house auditors — but maybe they should shut the front door first | TechCrunch — TechCrunch

    AI labs want in-house auditors — but maybe they should shut the front door first | TechCrunch Image Credits: IR_Stone / Getty Images # AI labs want in-house auditors — but maybe they should shut the front door first Tim Fernholz 11:25 AM PDT · September 16, 2026 Last weekend, after one of his researchers resigned over fears that AI could lead to human extinction, Anthropic CEO Dario Amodei…

  5. Advanced AI Society joins the Linux Foundation, launches open verification ecosystem as Congress moves on agent security — The Next Web

    Advanced AI Society joins the Linux Foundation, launches open verification ecosystem as Congress moves on agent security # Advanced AI Society joins the Linux Foundation, launches open verification ecosystem as Congress moves on agent security September 17, 2026 - 8:25 pm Credit: Advanced AI Society New York, NY, September 17th, 2026, TechnologyWire As bipartisan legislation demands…