AI Safety’s Next Fight: Codes, Evaluators, and Proof
Tuesday, September 15, 2026 · 10 min

Microsoft AI opened consultation on a Humanist AI Code of Conduct, while frontier-lab governance debates shift from pledges to verifiable evidence: embedded evaluators, EU-style codes, reward-hacking benchmarks, and assurance claims all face the same question—what proof actually constrains deployment?
Listen
Show notes
Microsoft AI opened consultation on a Humanist AI Code of Conduct, while frontier-lab governance debates shift from pledges to verifiable evidence: embedded evaluators, EU-style codes, reward-hacking benchmarks, and assurance claims all face the same question—what proof actually constrains deployment?
In this episode
- Humanist AI Code of Conduct — Microsoft AI
Humanist AI Code of Conduct | Microsoft AI Preface This document outlines the intended behavior and values of MAI models, the models developed by Microsoft AI. It summarizes our approach to training and operating them, following a set of design principles we call Humanist AI. This document, and our approach more generally, is still under development so we are not using it to train our…
- What on earth is an 'embedded evaluator'? One of the most important jobs in AI, according to frontier lab chiefs. — Business Insider Africa
# What on earth is an 'embedded evaluator'? One of the most important jobs in AI, according to frontier lab chiefs. Published: 2026-09-14T12:53:58+01:00 Source: africa.businessinsider.com (africa.businessinsider.com) Language: en ## Story The embedded evaluators' job is to check whether the company "is actually following the training, deployment, operational, and safeguards practices they…
- Consider how your global governance proposal is different from the EU Code of Practice - LessWrong 2.0 viewer — GreaterWrong
Consider how your global governance proposal is different from the EU Code of Practice - LessWrong 2.0 viewer # Consider how your global governance proposal is different from the EU Code of Practice David Matolcsi 13 Sep 2026 16:15 UTC (As an employee of the European AI Office, it’s important for me to emphasize this point: The views and opinions of the author expressed herein are personal and…
- AI Agent Evaluation and Reward Hacking | EM360Tech — EM360Tech
AI Agent Evaluation and Reward Hacking | EM360Tech AI 14 September 2026 # Your AI Agent Can Game the Test You Use to Prove It Works We look at why AI agent evaluation can mistake metric success for real task success, and what enterprises need to trust the results. We test software because we want evidence. Does it work? Can it complete the task? Does it behave the way we expected? If the…
- Access Is Not Yet Verifiability: Toward a Claim-Preserving Evidence Contract for AI Assurance — Hugging Face
# Access Is Not Yet Verifiability: Toward a Claim-Preserving Evidence Contract for AI Assurance Published September 13, 2026 Ali Toygar Abak Standards can preserve records without guaranteeing that the next dashboard, adapter or evaluation report preserves the limits of the claims those records support. Consider a hypothetical assurance pipeline. A runtime records a denied tool call. A…