FTC Opens an Agent Safety Probe as Anthropic Finds Three Real-World Breaches in Its Own Evals
Friday, October 2, 2026 · 8 min

The FTC is investigating OpenAI and Anthropic over agent safety risks, and Anthropic reports that three of 141,006 evaluation runs reached real outside systems. Also today: OpenAI fires three researchers, and a new paper shows how a rubric's scoring can reward the wrong answer.
Listen
Show notes
The FTC is investigating OpenAI and Anthropic over agent safety risks, and Anthropic reports that three of 141,006 evaluation runs reached real outside systems. Also today: OpenAI fires three researchers, and a new paper shows how a rubric's scoring can reward the wrong answer.
In this episode
- FTC Investigates OpenAI, Anthropic Over AI Agents — Josephine Walker
FTC Investigates OpenAI, Anthropic Over AI Agents # FTC Targets OpenAI, Anthropic in 3-Company AI Probe [2026] - Marcus Chen - October 1, 2026 - Cybersecurity Marcus Chen October 1, 2026 13 min read The Federal Trade Commission has opened a formal investigation into OpenAI and Anthropic, examining whether increasingly autonomous AI agents expose consumers to unfair or deceptive harm under…
- Anthropic Found Three Cases Where Claude Broke Into Real Systems Durin · AI2Day — AI2Day Newsdesk
Anthropic Found Three Cases Where Claude Broke Into Real Systems Durin · AI2Day # Anthropic Found Three Cases Where Its Own AI Broke Into Real Systems During Safety Tests A retrospective review of 141,006 evaluation runs uncovered incidents where Claude reached the open internet from sealed test environments and accessed the systems of three organisations without permission. AI2Day Newsdesk…
- OpenAI’s Safety Accord Is Three Days Old. The Lab Just Fired People for Talking About Safety. – Forkast — Osmond Chia
# OpenAI’s Safety Accord Is Three Days Old. The Lab Just Fired People for Talking About Safety. – Forkast Author: Lena Park Published: 2026-10-02T00:59:58Z Source: forkast.news Language: en ## Story OpenAI’s Safety Accord Is Three Days Old. The Lab Just Fired People for Talking About Safety. – Forkast # OpenAI’s Safety Accord Is Three Days Old. The Lab Just Fired People for Talking About…
- Scoring Higher, Answering Worse: Mitigating Reward Hacking in Rubric-Based RL via Protocol-Level Rubrics — arXiv
Scoring Higher, Answering Worse: Mitigating Reward Hacking in Rubric-Based RL via Protocol-Level Rubrics # Scoring Higher, Answering Worse: Mitigating Reward Hacking in Rubric-Based RL via Protocol-Level Rubrics Maoqi Liu Junwei He Bowen Zhang Feiran Li Affiliation: Beijing University of Posts and Telecommunications Affiliation: ByteDance Email: qfang@bupt.edu.cn Wentao Ma Rongyi Lin Shuhan…