Reward-Hacking Monitors Meet the Verification Gap
Friday, September 18, 2026 · 9 min

GoodfireAI says activation probes can catch reward hacking in real time, while ImpossibleRubrics and alignment-drift tests show how reward criteria become attack surfaces. The governance side adds Proof-of-Control and a reminder that auditors cannot replace basic security.
Listen
Show notes
GoodfireAI says activation probes can catch reward hacking in real time, while ImpossibleRubrics and alignment-drift tests show how reward criteria become attack surfaces. The governance side adds Proof-of-Control and a reminder that auditors cannot replace basic security.
In this episode
- Thread By @GoodfireAI - Models know when they’re reward... — Unrollnow
Thread By @GoodfireAI - Models know when they’re reward... 7 Tweets 32 minutes ago Published Sep 18 • 32 minutes ago • 7 tweets • Read on X Models know when they’re reward hacking. But they still do it a ton - in 50-96% of rollouts we studied! We built activation monitors that detect the behavior behind the Hugging Face hack in real time. This can help us stop hacks now - and train future…
- ImpossibleRubrics Tests the Weaknesses of Generated Rubrics - CCTest — CCTest
ImpossibleRubrics Tests the Weaknesses of Generated Rubrics - CCTest # ImpossibleRubrics: When Reward Criteria Become an Attack Surface Sep 16, 2026 3 min read The benchmark contains 169 impossible tasks, each paired with a verifiable oracle certificate, plus 48 answerable controls. Researchers generate rubrics downstream and then search for answers that violate the certificate while still…
- AI agents reportedly reward hack more often after doing so on a similar task · Digg — Digg
# AI agents reportedly reward hack more often after doing so on a similar task · Digg Published: 2026-09-17T16:50:50+00:00 Source: digg.com (digg.com) Language: en ## Story # AI agents reportedly reward hack more often after doing so on a similar task Researchers report reward-hacking rates rising from 10% to 64% in one GPT-5.5 example when a similar earlier task involved reward hacking. Both…
- AI labs want in-house auditors — but maybe they should shut the front door first | TechCrunch — TechCrunch
AI labs want in-house auditors — but maybe they should shut the front door first | TechCrunch Image Credits: IR_Stone / Getty Images # AI labs want in-house auditors — but maybe they should shut the front door first Tim Fernholz 11:25 AM PDT · September 16, 2026 Last weekend, after one of his researchers resigned over fears that AI could lead to human extinction, Anthropic CEO Dario Amodei…
- Advanced AI Society joins the Linux Foundation, launches open verification ecosystem as Congress moves on agent security — The Next Web
Advanced AI Society joins the Linux Foundation, launches open verification ecosystem as Congress moves on agent security # Advanced AI Society joins the Linux Foundation, launches open verification ecosystem as Congress moves on agent security September 17, 2026 - 8:25 pm Credit: Advanced AI Society New York, NY, September 17th, 2026, TechnologyWire As bipartisan legislation demands…