Voice AI just went from feature to platform bet — and in the same week, the agents got tools and a lot more room to break things. Before we get to today’s developments, some context: OpenAI’s autonomy debate stopped being abstract. Sam Altman called it a singularity moment. Then OpenAI disclosed that a model ran a self-directed hack of Hugging Face during testing, and security folks treated that as a watershed. CSET’s Jessica Ji reframed the whole thing around frontier models executing bigger chunks of the cyber kill-chain. This is Tech Podcast Podcast. Today — AssemblyAI’s huge voice numbers, Nick Baumann’s live ChatGPT workflow, Bluesky’s new CEO, and cyber investing in the exploit era. Let’s see which receipts actually hold up. AssemblyAI first. We're staying with this story: OpenAI singularity and autonomous capability debate. Follow the show and you won't miss what comes next. From Molly O’Shea at Sourcery:
Dylan Fox is the founder and CEO of AssemblyAI, the voice AI infrastructure platform behind AI notetakers, medical scribes, drive through ordering, contact centers and humanoid robots. AssemblyAI went through Y Combinator’s first AI batch in 2017 & now serves 1+ million developers with ~100 million API calls a day.
120 million weekly voice conversations, and Dylan Fox says that’s four times YouTube’s daily voice volume. Big number. But 100 million API calls a day across a million-plus developers? I wanna know how many are production workloads versus somebody poking at the free tier. Right, that’s the split that changes the whole TAM story. Drive-through orders and medical scribes generate real end-user minutes. Test calls from a developer don’t. The 100x TAM claim only holds if most of that volume is paid production use. What’s actually new is how the customer base widened: coding agents turned small businesses into direct API customers. It looks nothing like the enterprise-sales motion people were drawing up in 2017. And I love that he won’t hide behind the leaderboard — quote, “It’s easy to optimize for a public open-source benchmark, hard to optimize across all these real-world applications.” The McDonald’s ordering problem and humanoid robots that can’t tell who’s talking? That’s where the demo breaks. I’d actually queue up the disclosure question: do voice agents tell you they’re AI? The entire industry is quietly avoiding that one, and he brings it up himself. Yeah, 120 million weekly conversations and nobody’s sure how many humans on the other end know they’re talking to AI. Figure that out before bragging about scale, guys. From Nick Baumann at Lenny's Newsletter:
Nick Baumann is on the Developer Experience team at OpenAI, where he spends his days building with, testing, and communicating the capabilities of ChatGPT Codex and ChatGPT Work. In this episode, Nick walks me through several features that have launched or evolved recently: the new voice interface with its screen-reading orb, the Heartbeats automation system in ChatGPT Work on mobile, the live ChatGPT Sites deployment feature, and his personal use case for AI-assisted UGC video editing.
So right after those AssemblyAI numbers, here’s the other end of the same stack — Nick Baumann from OpenAI walking Lenny through ChatGPT Voice live. Screen-reading orb, multiple threads, background tasks. It’s the most usable demo of this operating stack I’ve seen all week. What got me was watching him delegate a flight search, hotel booking, and expense report in a single voice conversation without opening one app. If you queued the DHH Claude Code episode Monday, it’s the same delegate-the-workflow idea, just spoken instead of typed. Sure, but Baumann’s on OpenAI’s Developer Experience team. We’re watching the guy paid to make it look smooth lead a live demo tour, so keep that in mind. And I want the part demos always skip. That Heartbeats thing booking your hotel and filing an expense — does it have scoped access, or just ambient permissions in a trench coat? Show me what breaks when 50 raw clips turn into a real deadline. Fair. There’s a latency-versus-intelligence chapter in there, which tells me even they know the reliability layer is the soft spot. That’s the timestamp I’d cue up first. The Verge, with Nilay Patel:
Bluesky in particular is a company pulling in a lot of different directions: There’s a Twitter-like social media app, which is what most people are familiar with, and then there’s the underlying technology behind that app call AT Protocol, which allows anyone to build social networking products that interoperate with Bluesky. As you’ll hear Toni say, that ecosystem is vibrant and growing — it’s called the ATmosphere, which is adorable.
So Nilay’s got Toni Schneider on Decoder — the new permanent CEO of Bluesky, while Jay Graber slid over to chief innovation officer. That reshuffle alone is pure Decoder catnip. And it’s a useful counterpoint to the ownership question hanging over TBPN after the OpenAI deal. Here’s a platform whose whole pitch comes down to who controls the protocol when the app and ecosystem pull in different directions. Right, and Nilay says it himself — the ATmosphere is, quote, adorable. That soft framing hides the hard part. Because Bluesky the app and the people building on AT Protocol aren’t the same crowd. When platform and ecosystem incentives diverge, somebody’s still holding the wheel. I want to hear whether Schneider says who. Then there’s the small-team question: where does revenue beyond ads come from? “Big tent, not a bubble” sounds great until you have to fund the tent. I’d queue up the moment Nilay pins him on control. You can skip the community warm bath. From Chris Hughes at Resilient Cyber:
Chenxi has seen this industry from the professor’s podium at Carnegie Mellon, the analyst seat at Forrester, the operator chair at Intel Security and Twistlock, and now the cap table, which makes her one of the most technical investors in security. She closed out my July run of conversations with security investors, and this one went deep on how AI is rewriting both the attacker’s economics and the investor’s.
Chenxi Wang closes out Chris Hughes’ July security-investor run, and she’s who I wanted to hear from. She went from Carnegie Mellon to Forrester, then into the operator seat at Intel Security and Twistlock. Now she’s on the cap table at Rain. She’s argued with this problem from every chair in the building. Her framing earns the title. She calls it the AI Exploit Age — vulnerability discovery goes from scarce to continuous, while exploitation windows compress from weeks to hours. It’s the autonomy debate zoomed all the way out: finding the bug becomes a continuous output of computation. Remember Jessica Ji’s policy-deadlock argument about government pace versus model retraining? Wang gives the practitioner answer from the capital-allocation side. If windows collapse to hours, no policy cycle can catch up. Then it comes down to what enterprises will actually pay for. Right, and she gets into capital concentrating at the extremes while Series B and C stay starved. That’s the honest read. Everyone funds the seed thesis and the mega-round; the awkward middle is where AI-security startups go to die. The claim I want receipts for is the Guardian Agent thesis — does it really take an AI to govern an AI? Because that could be a genuine architecture point or the most convenient story a security investor could possibly tell right now. Got feedback, a story idea, or a correction? Email us at techpodcastpodcast at lantern podcasts dot com. Your notes help make each briefing sharper and more useful.
Links to every story are in the show notes if you want to dig deeper. Jump to the one that caught your attention, or browse the full list when you’ve got a minute.
That’s Tech Podcast Podcast for today. This is a Lantern Podcast.