← Tech Podcast Podcast

OpenAI’s AGI Line Meets the AI Code-Security Reckoning (August 06, 2026)

August 06, 2026 · 9m 39s · Listen

OpenAI's Joshua Achiam says we might've already crossed the AGI line — and it lands the same day the security crowd is cataloging exactly how AI-written patches blow up in your face. If you're just joining us, this started with Sam Altman's singularity talk and that OpenAI Hugging Face testing incident, then rolled into Chenxi Wang's warning that defenders are entering an AI Exploit Age. So here's the question: do agents turn vulnerability hunting and exploitation from occasional human work into a nonstop race measured in hours? This is Tech Podcast Podcast. Today, an AGI claim from inside the a16z studio runs into Keith Hoodlet's taxonomy of patches that fix nothing and break everything. Plus, SlopCodeBench and an open-source trust reckoning. Sarah, let's start with that AGI line. From Joshua Achiam at Andreessen Horowitz:

Theo Jaffee is joined by Joshua Achiam, Chief Futurist at OpenAI, for a conversation on AI cybersecurity, frontier model capabilities, and why he believes society may have already crossed the threshold into an AGI-era without fully recognizing it.

OpenAI's Chief Futurist goes on the a16z pod and says we may have already crossed into AGI — and nobody in the room felt the earth move. That's the pitch: it happened quietly. His actual argument is that we've all just adapted: capabilities that were unimaginable two years ago are now background noise. Okay, fine. But “you got used to it” isn't an operational definition of AGI. Right, and that's the tell. The claim gets huge, and the falsifiability drops to zero. I want to know if Theo Jaffee asked him what evidence would prove him wrong — or if he let him coast on the vibe. The cyber part actually earns its keep: AI finds vulnerabilities faster than humans can patch them, and the fix window shrinks from weeks to hours. That's the exploit-compression problem we've been talking about, now with an AGI headline bolted on top. And conveniently, we've got a whole segment later on what those AI-generated patches actually do when they land. So hold that “faster than humans can patch” line — we're gonna test it. Here's Keith Hoodlet at SC Media:

Keith Hoodlet gives an exclusive early look at his team's recent research into the success, quality, and failures of LLM-generated security patches. Notably, they saw scenarios across a spectrum from robust, effective patches to patches that changed the software's behavior to patches that introduced new vulns to patches that didn't even fix the original vuln while also introducing a new vuln.

Okay, this is the experiment I actually want. Keith Hoodlet's team tested what everybody's been hand-waving past: what happens when you let the LLM patch the vuln it just found. And the spectrum is chef's kiss — patches that work, patches that quietly change behavior, patches that add a fresh vuln, and the grand finale: patches that don't fix the original bug and hand you a new one for free. That last category is the receipt. We opened the hour with Achiam asking whether we've crossed the AGI threshold. Well, here's AGI-era patching in the wild — and it's documented. It also lands right on the exploit-compression worry from earlier this week: Jessica Ji's AI Exploit Age. The same models racing down the kill chain are writing the fixes, too, and now we've got a documented failure mode. What I respect is how honest they are about the variables — prompt quality, language, and how much expertise you need just to know what a good patch looks like. That's the part promo episodes skip. You can't grade the patch if you can't read the patch. Boundary ML writes:

On the podcast this week, we will examine a new AI coding benchmark called SlopCodeBench (SCBench), which aims to evaluate how software is truly developed. While typical benchmarks evaluate models using single-shot solutions against specifications, they fail to address the long-term difficulties of software development. Creating an initial solution is simple; the real challenge lies in extending, refactoring, and maintaining code over time.

SlopCodeBench — SCBench — is the one I respect in today's stack. It's from the BAML crew, Vaibhav Gupta and Dex Horthy, and it stops grading the first draft. The model has to keep building on its own code across multiple checkpoints. Yeah, that's the honest test. Most benchmarks score a one-shot solution against a spec. As Boundary ML puts it, that never touches the hard part: extending and maintaining what you've already shipped. Right — remember DHH's hundred pull requests Monday? Great headline. SCBench is the question underneath it: does Claude Code still work on pull request two hundred, when the mess it's editing is a mess it half-wrote? And it ties the whole day together. Achiam can float an AGI threshold in that a16z interview, but durability across a codebase is a much harder check on the claim. Can it extend and refactor what it built without breaking its own prior output? Every benchmark has the same flaw, though: you only measure what you tested for. The failures are out in the long tail you didn't track. So I want to see what SCBench leaves out. The Generalist writes:

Eric Nguyen is the co-founder and CEO of Radical Numerics, an AI research lab that has raised $50 million to train models directly on biological data. Before starting the company, Eric helped develop Evo and Evo 2, large-scale genome language models trained on unlabeled DNA sequences. Radical Numerics is now building models that can connect information across DNA, RNA, proteins, epigenetics, and other parts of biology, rather than treating each as a separate problem.

Radical Numerics — fifty million to train models directly on biological data — and Eric Nguyen isn't hand-waving here. Omnii apparently matched key findings from two years of Alzheimer's wet-lab work in a matter of days. And Evo already generated viable bacteriophage genomes. Actual DNA that works gives the whole “biology is the next frontier” pitch some real evidence. And he names the bottleneck himself, which I appreciate: even a perfect prediction has to survive lab verification before it's a discovery. He isn't selling days-not-years as the whole pipeline. Fifty million for genome language models is exactly the kind of high-conviction bet that gets funded at the seed stage or in a mega-round, then gets squeezed in between. So where does Radical Numerics land — canary in that Series B squeeze, or the exception everyone points to? And credit where it's due: he says the same tools that design a bacteriophage could design a dangerous pathogen, which is why they're building biodefense alongside it. After the AGI-threshold talk we just hit, you've got a founder who's actually taking the downside seriously. Here's Kate Holterhoff at RedMonk:

Open source has a new problem: you often can’t tell whether the contributor on the other end is a person or an agent working on their behalf. In this RedMonk Conversation, Kate Holterhoff talks with Angie Byron, Senior Manager of Community and Advocacy at Temporal and lead of applied AI at Drupal, about AI’s effects on open source communities.

Of everything today, this is the one I actually want to sit with. Angie Byron's line — don't submit code you don't understand — that's the whole ballgame for a maintainer now. It's the same trust problem we just hit with the patch research, wearing a different hat. With the patches, a fix can break two more things. Here, you can't even tell whether a human wrote the thing you're reviewing. Right — the whole open-source credit system was built on the idea that a contributor is a person. Now it might be an agent working on their behalf, and the maintainer's the one eating the slop. And I like that she doesn't just doom on it. Marketers are shipping real code, and security work is surfacing decades-old bugs. A lower bar cuts both ways. Sure, but the same basic problem keeps showing up this week — you can't tell if the voice on the call is a bot, and you can't tell if the commit came from a human. At every layer, authorship gets harder to verify. And if the codebase itself is now contaminated, every “proprietary data is the moat” argument gets a little shakier. You got there first, then built a wall around a dataset you can't fully vouch for. Have feedback, a story idea, or a correction? Send it to techpodcastpodcast at lantern podcasts dot com. Your notes help us make the briefing sharper and more useful.

Links to every story we covered are in the show notes. If one caught your attention, you can dig deeper there. Thanks for listening. We'll be back with more tomorrow. That's Tech Podcast Podcast for today. This is a Lantern Podcast.