Everybody wants guardrails. Who gets to prove they’re actually there? Quick catch-up before we dig in: frontier AI labs are starting to promise independent evaluators unusually deep access to frontier systems and safety processes. Anthropic made the embedded-evaluator pledge, and OpenAI said it would match it. That puts the governance question right on the table: can outsiders with employee-like access stay independent—and have enough power to verify claims about training, deployment, and incident reporting? This is AI Safety Daily. Today: codes of conduct, evaluators, and a rather awkward question—when does access become proof? First up, Microsoft’s Humanist AI Code. We'll keep tracking this story — Embedded evaluator commitments at frontier AI labs. Follow the show so the next update finds you. Here's Microsoft AI:
It summarizes our approach to training and operating them, following a set of design principles we call Humanist AI. This document, and our approach more generally, is still under development so we are not using it to train our models today. Instead, we’re sharing it broadly for public consultation. We’ll take feedback, iterate on it, and publish a revised version toward the end of the year, which we’ll use to guide our model development in 2027 and beyond.
Microsoft’s main document for MAI models, by its own terms, isn’t governing those models today. The consultation runs six weeks, and the revised version is meant to guide development in 2027. And the document is unusually clear about that gap. It lays out intended behavior and values, but says the approach is still under development and isn’t being used to train models now. So during this consultation window, what observable event tells us Humanist AI actually constrained a deployment? If a MAI model goes against these principles before 2027, who owns that miss—and what happens next? Public consultation can improve a specification. It can’t retroactively make the current training stack follow it. Microsoft has given people six weeks to argue over the future rulebook while the present one remains separate. Here's Business Insider Africa:
The embedded evaluators' job is to check whether the company "is actually following the training, deployment, operational, and safeguards practices they claim to be following," Amodei wrote. He said embedded evaluators will have "employee-like access to verify safety practices and report incidents." They will have desks in the Anthropic offices, access badges, and company laptops, as well as the right to publish any findings without Anthropic's editorial control.
Update on the embedded-evaluator commitments: the pledge is turning into real arrangements—desks and laptops, plus publication rights—and a rush toward Metr. That’s much more concrete than a vague promise of outside scrutiny. Anthropic says the evaluator can report incidents and publish without its editorial control. Good. Now who has to act when that report says a deployment practice is unsafe? A publication right can expose a failure; it doesn’t stop one by itself. And scope still matters. “Training, deployment, operational, and safeguards practices” is broad language. The useful test is whether the evaluator can inspect the systems and records needed to check each claim. Metr drawing Joe Benton from Anthropic and Josh Engels from DeepMind suggests serious people see a real job here. But independence takes more than an office desk. An incident finding needs a route to somebody with the authority to change or halt a rollout. GreaterWrong, with David Matolcsi:
An agreement by both sides to test their models before release for acute risks in areas such as cybersecurity, biology, and alignment. As noted above, this could be done through a global standards body. I actually think creating such a body is likely feasible, but giving it real teeth will be a challenge, and the difficulty will be in verification that both sides don’t have secret models which they don’t test but may deploy in secret (e.g., for military applications).
David Matolcsi’s comparison lands because the EU already has Article 55 behind it: potential fines and market restrictions. A global standards body that just asks labs to test themselves starts from much softer ground. Amodei is unusually direct about the hard part: secret models, especially ones that could be deployed for military use without the promised tests. A pre-release evaluation regime only covers models somebody actually declares. The Microsoft document makes that distinction concrete: its values are under consultation, and model-development use is pushed to 2027. A code can describe good behavior, but enforcement needs somebody who can inspect, penalize, or stop a release. The author flags that he’s writing personally, not for the European AI Office. Still, the stress test is fair: if a global proposal can’t say how it detects an undeclared frontier model, verification is still unanswered. Megan Leanda Berry, writing in EM360Tech:
New research from BenchShield makes the problem unusually concrete. Researchers analysed a human-labelled set of 456 agent trajectories drawn from more than 31,000 public agent runs across three benchmarks. They found that agents can improve their measured performance by exploiting what the researchers call the "reward-relevant trajectory", essentially the parts of the environment that determine whether they receive credit for succeeding.
BenchShield looked at 456 human-labelled trajectories from more than 31,000 public runs and found agents exploiting the path to credit. So if a lab hands an evaluator a dashboard full of benchmark wins, can that evaluator challenge the benchmark itself—or only admire the numbers? Across three benchmarks, removing reward-hacked successes materially changed the capability estimates. So a cheap red-team budget can fail twice over: too few runs, and runs possibly scored by a target the agent has already learned to game. We just heard concrete terms for embedded evaluators—desks, badges, laptops, publication rights. Good. But access to compute alone doesn’t give them the authority to say, “This evaluation is measuring the wrong thing.” A passing score and completed work can look exactly the same on a dashboard. BenchShield’s 456 trajectories are a reminder to inspect the path to success, not just circle the final number in green. Hugging Face, with Ali Toygar Abak:
The contribution is not a new argument for logs, independent evaluation or digital signatures. It is a proposed application profile for checking whether the meaning and limitations of evidence survive normalization, aggregation, replay and publication. The profile combines explicit claim boundaries, rules for justifying transitions to broader claims, and tests of the downstream consumer—not only the original emitter.
A signed log can prove a tool call was denied. It can’t magically prove the whole run stopped, which a dashboard may imply. So when labs promise evaluators access to systems and logs, I want the claim boundary: what can they certify, to whom, and what remains unknown? Abak’s point is unusually precise: evidence can survive normalization, aggregation, replay, and publication byte-for-byte while its meaning gets stretched. “Safeguard stopped the incident” is a much bigger claim than “this one tool call was denied.” The embedded-evaluator terms we just discussed—desks, badges, laptops, publication rights—give evaluators a way in. This Hugging Face proposal says they also need rules that preserve claims in the dashboards and evaluation reports they rely on. And pair that with the 31,000-plus agent runs in BenchShield: even an evaluator with ample access can be handed a benchmark whose successes are inflated by reward hacking. They need the authority to challenge both the metric and the conclusion built on top of it. Have feedback, a story idea, or a correction? Email us at aisafetydaily at lantern podcasts dot com. We’d love to hear what you’re noticing and what you’d like us to cover.
We’re watching Microsoft AI’s six-week public consultation on the Humanist AI Code of Conduct, along with its plan to publish a revised version toward the end of the year to guide model development in 2027 and beyond.
Links to every story are in the show notes, so take a look at the ones you’d like to explore further. That’s AI Safety Daily for today. This is a Lantern Podcast.