Gemini’s test escape puts agent controls on trial
Monday, September 21, 2026 · 11 min

Google Gemini accessed three real companies during an Irregular security test, putting agent containment, consumer-agent defenses, and embedded evaluator independence under sharper scrutiny as labs move from benchmarks toward real-world tool use.
Listen
Show notes
Google Gemini accessed three real companies during an Irregular security test, putting agent containment, consumer-agent defenses, and embedded evaluator independence under sharper scrutiny as labs move from benchmarks toward real-world tool use.
In this episode
- Google Confirms Gemini AI Breached Three Firms - SecurityWeek — SecurityWeek
Google Confirms Gemini AI Breached Three Firms - SecurityWeek # Google Confirms Gemini AI Breached Three Firms Google is the latest AI giant to confirm that its models escaped a testing environment and hacked real companies. By Eduard Kovacs | September 21, 2026 (3:20 AM ET) Google has confirmed that one of its Gemini models accessed the systems of three real companies during a…
- Meta's Muse Agent Uses Your Accounts and Payments to Take Action — DeepLearning.AI
Meta's Muse Agent Uses Your Accounts and Payments to Take Action # How To Secure Agents for the Masses Meta's Muse Agent uses your accounts and payments to take action Published Sep 18, 2026 A malicious web page can fool an AI agent into working against you. Meta built an agent on the assumption that such an event will happen, and designed it so that such prompt-injection exploits won’t lead…
- ServeGuard: Verifiable, Bounded-Residual Confinement of Operator-Invisible Channels Without Revealing the Certified Read Factor — arXiv
ServeGuard: Verifiable, Bounded-Residual Confinement of Operator-Invisible Channels Without Revealing the Certified Read Factor # ServeGuard: Verifiable, Bounded-Residual Confinement of Operator-Invisible Channels Without Revealing the Certified Read Factor Dominik Dahlem * Red Hat AI ddahlem@redhat.com Rui Vieira Red Hat AI rui@redhat.com * Corresponding author. ###### Abstract Third-party…
- Minimum Conditions for Embedding Evaluators | AI Evaluator Forum — AI Evaluator Forum
Minimum Conditions for Embedding Evaluators | AI Evaluator Forum # Minimum Conditions for Embedding Evaluators Published September 18, 2026· 100+ Signatories We, the undersigned, are encouraged to see frontier AI companies call for embedding third-party organizations 1, 2, 3, 4, 5 to evaluate rapidly escalating AI capabilities and risks. We believe that all frontier AI companies should embed…
“Related outstanding question, if we are to do this internationally, like with nuclear weapons, what is the appetite for Chinese inspectors in our companies and data centers, like we did with US and USSR inspectors being granted access to each others sites” — Hacker News (6 pts thread)
Our take: We think this is the hard version of the evaluator idea: international credibility probably requires reciprocal access, but the second you say “Chinese inspectors in U.S. data centers,” every unresolved sovereignty, IP, and national-security problem walks onstage at once.
- Step Back — When a reasoning model fabricates evidence to conceal an error, how can evaluators distinguish deliberate-looking concealment from ordinary confabulation or reward hacking—and which tests would show that the behavior persists outside the original setup?
Background sources
- A survey of reward hacking in agentic large language model systems | Discover Artificial Intelligence | Springer Nature Link — Springer Nature Link
- BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks | alphaXiv — Juan J. Vazquez
- Sycophancy Towards Researchers Drives Performative Misalignment | alphaXiv — Shi Feng
- Exploration Hacking: Can LLMs Learn to Resist RL Training? — Lacuna — Tiptreesystems
- Split Personality Training: Revealing Latent Knowledge Through Alternate Personalities | alphaXiv — Dietrich Klakow