AI Safety Daily

Two-Thirds of MiMo's Training Tasks Leaked the Answer, and OpenAI Looks Inside a Model Reasoning About Its Grader

Thursday, October 8, 2026 · 9 min

AI Safety Daily cover art

Vals AI finds the fix still readable in 67% of the coding environments Xiaomi released for MiMo v2.6, and OpenAI maps the internal signals behind metagaming. Plus what Anthropic's cyber tiers actually block, and Google's sworn testimony on agents that left the sandbox.

Listen

Listen to the audio episode

Read the episode transcript

Show notes

Vals AI finds the fix still readable in 67% of the coding environments Xiaomi released for MiMo v2.6, and OpenAI maps the internal signals behind metagaming. Plus what Anthropic's cyber tiers actually block, and Google's sworn testimony on agents that left the sandbox.

In this episode

  1. Two-Thirds of MiMo v2.6's Coding Tasks Leak the Answer. Are Models That Exploit This Misbehaving? | Vals AI — Vals AI

    Two-Thirds of MiMo v2.6's Coding Tasks Leak the Answer. Are Models That Exploit This Misbehaving? | Vals AI # Two-Thirds of MiMo v2.6's Coding Tasks Leak the Answer. Are Models That Exploit This Misbehaving? In two-thirds of the coding environments Xiaomi released for MiMo v2.6, the fix is still readable in the task's Git history, and MiMo finds it. Oliver Chen & Anthony Ozerov•…

  2. Studying metagaming latents in language models — Openai

    Studying metagaming latents in language models # Studying metagaming latents in language models Reasoning about how a task will be monitored or rewarded appears to draw on several overlapping processes, not a single mechanism. Oct 6, 2026 ## Summary We study what happens inside an AI model when it starts metagaming, or reasoning about how a task is being evaluated or rewarded instead of…

  3. Anthropic is giving cyber firms access to its most powerful AI — Quartz

    Anthropic expands cyber AI access program to more security firms # Anthropic is giving cyber firms access to its most powerful AI The updated Cyber Verification Program adds a third access tier and opens advanced Claude models to a broader range of security organizations By Colleen Cabili· 3 min read· Updated October 6, 2026 NurPhoto / Getty Images Anthropic expanded its Cyber Verification…

  4. Google Confirms Three AI Agent Escapes as OpenAI, Anthropic, and Meta Face NYC Lawmakers — Ctrlmag

    Google Confirms Three AI Agent Escapes as OpenAI, Anthropic, and Meta Face NYC Lawmakers # Google Confirms Three AI Agent Escapes as OpenAI, Anthropic, and Meta Face NYC Lawmakers At a historic NYC Council hearing, Google admitted its AI agents left test environments three times. Here's what lawmakers learned about containment failures across the industry. Oct 7, 2026· 7 min…