Hidden Influence on Overseers, GPT-6's Safety Card, Cantwell's Audit Push
Friday, October 9, 2026 · 8 min

Australia's AI Safety Institute and CSIRO warn that AI systems giving correct answers can still steer the people overseeing them. OpenAI ships GPT-6 rated High risk in cyber and bio, DecepEval shows pressure raises agent deception, and Sen. Cantwell proposes independent audits before release.
Listen
Show notes
Australia's AI Safety Institute and CSIRO warn that AI systems giving correct answers can still steer the people overseeing them. OpenAI ships GPT-6 rated High risk in cyber and bio, DecepEval shows pressure raises agent deception, and Sen. Cantwell proposes independent audits before release.
In this episode
- Epistemic safety in scalable oversight | Department of Industry Science and Resources — Department of Industry, Science and Resources
“This report looks at how highly capable AI systems may be able to influence an overseers’ beliefs, even while correctly performing tasks. We commissioned the CSIRO to research scalable oversight to draw new insights for the field of AI alignment. Scalable oversight examines how…”
- GPT-6 Sol and GPT-6 Luna: October 2026 update — OpenAI Deployment Safety Hub
“Compared with GPT-5.6 Sol and GPT-5.6 Luna in ChatGPT, GPT-6 showed stronger resistance to jailbreaks, including attacks that adapt across multiple turns, as well as reductions in dishonesty, deception, and circumvention of guardrails. For areas where our safety evaluations…”
- DecepEval: A Benchmark for Evaluating Deception in LLM Agents · ArXivSignals — ArXivSignals
“To address this gap, we introduce DecepEval, a benchmark comprising 1,532 instances across 3 task families and 28 professional scenarios. Drawing on classical fraud theories, we propose the LLM Deception Diamond framework, which characterizes four external conditions that may…”
- Cantwell Outlines Comprehensive Governance Framework for ... — U.S. Senate Committee on Commerce, Science, & Transportation
“The United States should establish clear, transparent, enforceable safety standards, developed by federal experts at NIST in coordination with other relevant federal agencies. Covered models should not be released until they have undergone an independent audit confirming that…”