AI Safety Daily

AI Control Warnings Move From Incidents to Institutions

Tuesday, September 22, 2026 · 10 min

AI Safety Daily cover art

OpenAI-Hugging Face incident anchors a new UN scientific brief on agent misalignment and loss of control, while embedded-evaluator partnerships, OVERT runtime-evidence standards, and a reported RubyGems disclosure gap test whether oversight can get deeper access, stronger logs, and real consequences.

Listen

Listen to the audio episode

Read the episode transcript

Show notes

OpenAI-Hugging Face incident anchors a new UN scientific brief on agent misalignment and loss of control, while embedded-evaluator partnerships, OVERT runtime-evidence standards, and a reported RubyGems disclosure gap test whether oversight can get deeper access, stronger logs, and real consequences.

In this episode

  1. Thematic Brief on AI Agents, Misalignment and the Risk of Losing Human Control | Independent International Scientific Panel on AI — Independent International Scientific Panel on AI

    # Thematic Brief on AI Agents, Misalignment and the Risk of Losing Human Control | Independent International Scientific Panel on AI Published: 2026-09-21T10:20:02+00:00 Source: un.org (un.org) Language: en ## Story Thematic Brief on AI Agents, Misalignment and the Risk of Losing Human Control | Independent International Scientific Panel on AI Independent International Scientific Panel on…

  2. Partnering with Accenture on embedded evaluation - Extrapolator AI — Extrapolator AI

    Partnering with Accenture on embedded evaluation - Extrapolator AI # Partnering with Accenture on embedded evaluation On 18 September 2026, Anthropic announced a five-year partnership with Accenture's Faculty division to function as an embedded evaluator — a structural reconfiguration of how frontier model safety is assessed that moves the evaluative function inside the lab rather than…

  3. Glacis to place OVERT, its open standard for verifying AI safeguards, under CHAI and AIGovOps Foundation stewardship — PR Newswire

    Glacis to place OVERT, its open standard for verifying AI safeguards, under CHAI and AIGovOps Foundation stewardship Glacis Technologies, Inc. Sep 21, 2026, 09:01 ET Agreement provides for shared stewardship of the open, royalty-free standard for AI operational evidence, with CHAI contributing healthcare expertise and Glacis serving as non-voting editor SEATTLE, Sept. 21, 2026 /PRNewswire/…

  4. RubyGems Supply Chain Breach Was Never Reported to Brussels Under EU AI Act Rules — TechTimes

    # RubyGems Supply Chain Breach Was Never Reported to Brussels Under EU AI Act Rules Author: Mark Rutherford Published: Sep 20 2026, 2:43 AM EDT Published: 2026-09-20T02:43:22-04:00 Source: techtimes.com (techtimes.com) Language: en ## Story RubyGems Supply Chain Breach Was Never Reported to Brussels Under EU AI Act Rules # RubyGems Supply Chain Breach Was Never Reported to Brussels…

  5. Step Back — If frontier labs bring outside safety evaluators inside, what would make those evaluations genuinely independent rather than outsourced red-teaming—who controls model access, test design, logs, and publication, and what findings could actually force a deployment change?

    Background sources