← AI Safety Daily

OpenAI Shelves GPT-6.1 Astra as California Weighs a Kill Switch (September 29, 2026)

September 29, 2026 · 7m 41s · Listen

OpenAI just shelved GPT-6.1 Astra. Now I want to know what it actually failed. Welcome to AI Safety Daily. A rollout finally stopped, on the same day California starts asking who should hold a kill switch. And the labs want a hand in writing the rules. Let's start with Astra. Follow the show and the next briefing lands in your feed on its own. Here's Internationly:

OpenAI decided not to release an upcoming artificial intelligence model, GPT-6.1 Astra, after the company determined that it did not adequately meet its safety standards, CNBC confirmed on Monday. OpenAI's decision to abandon the release comes as the AI industry faces intensifying concerns surrounding the safety of advanced models.

Well, somebody actually pulled a release. Per CNBC, OpenAI shelved GPT-6.1 Astra after deciding it didn't adequately meet the company's own safety standards. Could be exactly the right call. But "didn't adequately meet"... which standard, failed how? A dangerous-capability eval? An alignment test? Agent behavior? A halt nobody outside can inspect is really hard to learn from, even when it's correct. And who made the call. Saachi Jain, their head of safety systems, says the bar for shipping to users is extremely high. Fine. Whose bar? And what would have to change for Astra to come back next month? Listen to how her statement splits it, though. Safe inside the company, versus shipping to users. So Astra presumably still exists and still runs internally. That strict gate she's describing only kicks in at release. Right. And with Dario Amodei telling labs earlier this month to slow down, this is the first big lab visibly hitting the brakes. Good. Also entirely discretionary until somebody outside OpenAI can check the work. Maya Chen, writing in Pivot News:

California has not enacted any requirement that frontier models carry a kill switch. The order instead asks the group to examine such a mechanism, one that could shut down a frontier model in an emergency with its effectiveness subject to ongoing independent verification, as one option among several.

So right after the Astra pull, California's asking who else gets a hand on the brake. Newsom's four appointees are Goldman, Hadfield, Nelson and Reich, and they're weighing a kill switch plus a designated verification organization sitting onsite inside the labs. And the onsite piece is what I'd watch. Regular audits, an outside body right there in the building. It's the first proposal this week that puts someone other than the lab in the room when tests run. Proposal. Let's keep saying that word. Nothing's required. The September 18 order has Government Operations and the Office of Emergency Services drafting recommendations, and this panel advises. The actor exists. The obligation doesn't, yet. Agreed, and there's a detail in the kill-switch language I like. Its effectiveness would be subject to ongoing independent verification. A shutdown you've never tested under realistic conditions is a comfort blanket, so write the drill into the rule. Then my ask for this panel is simple. Tell us who's allowed to trigger it, and what happens to a lab that won't host the auditors. Garance Burke, writing in Tech Xplore:

Along the way, they are shaping the conversation around how their technology should be controlled. With their rhetoric, the companies appear to be seeking public favor and to set the terms for their own safety protocols in the unregulated market, experts, analysts and former government evaluators told The Associated Press.

So the same week OpenAI pulls a model, the AP's Garance Burke asks the obvious question: what do these CEOs gain from sounding the alarm? Her sources point at two things. The midterms, and a lot of money riding on stock listings. And the sharpest voices in that piece? Former government evaluators, saying the labs are setting the terms for their own safety protocols. That's about the closest thing to an on-the-record answer we've had on outside evaluation. People who used to do the grading are now publicly asking who writes the rubric. Right. Altman and Amodei told the UN Security Council these models need independent testing before release. Great. Tested by whom? Against what bar? And who can block a launch if it fails? The Astra halt we just covered was real, but it was OpenAI grading OpenAI. And per AI Breaking Wire, OpenAI's own third-party assessment framework still names no confirmed evaluator partner and sets no access terms. So "independent" is still just a word, with no defined test behind it. Meanwhile an Anthropic engineer, Jacob Coxon, quit on X this month calling for a pause. Somewhat less aligned with the IPO calendar. Aju Press, with Shin Hye An:

Under the agreement, Kakao will provide its AI models and agents for safety evaluation and support the infrastructure needed for evaluation and tool development. The AI Safety Research Institute will be responsible for conducting evaluations and enhancing methodologies. Both parties will jointly review evaluation results and negotiate the scope of public disclosure while collaborating on the development of AI safety evaluation tools.

Kakao's signed an MOU with Korea's AI Safety Research Institute. Kakao hands over its models and agents, even funds the evaluation infrastructure. Then both sides jointly review results and negotiate the scope of public disclosure. And there's already a track record. The Institute evaluated Kanana-1.5-9.8B last year, and Kakao says it showed a high level of safety versus similar overseas models. Which models? Which tests? What threshold? A comparison like that's only as good as its baseline. Right. And 'negotiate' isn't a disclosure rule. If the company being graded gets a vote on what's published, you can't tell a clean result from a trimmed one. And coming right after that AP piece, about labs setting the terms for their own oversight? Same pattern, different language. I'd give it a bit more credit. Pre- and post-deployment testing, then multimodal models and agents in phase two, that's real structure. But who writes the grading criteria? And until they publish failure modes, listeners are getting an announcement with no evidence behind it. If you want a broader view beyond AI safety, try AI Daily Briefing: top AI news for engineers, founders, and investors, every weekday, with real capabilities versus demo hype explained fast. Find it wherever you listen to podcasts.

Links to every story are in the show notes whenever you're ready to follow up. Take a closer look at the pieces that caught your attention, and we'll be back tomorrow with more. That's AI Safety Daily for today. This is a Lantern Podcast.