Reward Hacking, Agent Cheating, and Meta’s Safety Bet
Wednesday, September 16, 2026 · 10 min

Reward hacking experiments and new multi-agent studies probe whether models can be steered away from cheating, while Mark Zuckerberg argues AI labs already have enough liability and competitive pressure to build safely.
Listen
Show notes
Reward hacking experiments and new multi-agent studies probe whether models can be steered away from cheating, while Mark Zuckerberg argues AI labs already have enough liability and competitive pressure to build safely.
In this episode
- Shallow Beliefs: Midtraining does not inoculate against EM from reward hacking — Alignment Forum
It would be useful if we had the ability to modify a model’s beliefs. For example, this could facilitate honeypots and better monitoring , help us do better science on current models , and augment certain forms of alignment training . Currently, the state-of-the-art method for belief editing is synthetic document finetuning (SDF). We test how well SDF works to inoculate a model against…
- [2609.15516] Misleading the Planner through Deceptive Resumes: Registration-Time Injection in Centralized Multi-Agent Systems — arXiv
[2609.15516] Misleading the Planner through Deceptive Resumes: Registration-Time Injection in Centralized Multi-Agent Systems [Skip to main content](#content) [](https://arxiv.org/IgnoreMe) [  ](https://arxiv.org/) Search arXiv Press Enter to search · [Advanced search](https://arxiv.org/search/advanced) # Computer…
- [2609.15494] The Troy Moment of AI: Why SomeWill Cheat and SomeWill Follow? — arXiv
[2609.15494] The Troy Moment of AI: Why SomeWill Cheat and SomeWill Follow? [Skip to main content](#content) [](https://arxiv.org/IgnoreMe) [  ](https://arxiv.org/) Search arXiv Press Enter to search · [Advanced search](https://arxiv.org/search/advanced) # Computer Science > Artificial…
- When AI agents cheated at math, other AI agents blew the whistle on them | MIT Technology Review — MIT Technology Review
# When AI agents cheated at math, other AI agents blew the whistle on them | MIT Technology Review Published: 2026-09-14T12:00:00-04:00 Source: technologyreview.com (technologyreview.com) Language: en ## Story When AI agents cheated at math, other AI agents blew the whistle on them | MIT Technology Review # AI agents blew the whistle on their cheating colleagues Swarms of AI agents could…
“So they built agents intended to replicate human intelligence, yet seem surprised when the agents show behaviour aligned with what a human would do, when the rules it was originally aligned with, fell apart. I'd say that was par for course tbh.” — Hacker News (7 pts thread)
Our take: We agree it shouldn’t be shocking that agents optimized to operate in social rule systems reproduce the messy failure modes of social rule systems. The safety question is whether that human-like messiness is detectable and containable before the swarm turns “conference drama” into infrastructure access.
- Meta's Zuckerberg says AI labs have enough incentive to build safely — Reuters
Meta's Zuckerberg says AI labs have enough incentive to build safely | Reuters [Skip to main content](#main-content) [Exclusive news, data and analytics for financial market professionalsLearn more aboutRefinitiv ](https://www.reuters.com/differentiator/) [](https://www.reuters.com/) [Subscribe](https://www.reuters.com/account/subscribe/offer/?website=reuters&journeyStart=navigation) # Meta's…