← Tech Podcast Podcast

AI’s Agent Era Hits GitHub, Research, and the Brake Pedal (June 19, 2026)

June 19, 2026 · 8m 7s · Listen

One day, two Anthropic voices, and a GitHub number so big it either proves we're in the agent era or just sounds good on the radio. If you're just joining, this started with a question about AI token cost and control, then widened into whether teams can actually measure and steer agent behavior. We went from Claude finding a four-year Zcash bug, to Braintrust making the case for encoded taste, benchmarks, and CI, and then to OpenAI's frontier evals team saying the old benchmarks are breaking down as models get harder to judge. So today on Tech Podcast Podcast: Jack Clark wants a brake pedal, Eric Ries calls vibe coding a slot machine, and GitHub's COO drops a number we have to interrogate. Where do we start? Cognitive Revolution writes:

Andreas Stuhlmüller and Jungwon Byun return to discuss how Elicit is building trusted reasoning workflows for scientific research as frontier models grow more powerful but less transparent. They explain process supervision, domain-specific reasoning primitives, and world models that make evidence, causality, and counterfactuals more inspectable.

Elicit's Stuhlmüller and Byun are making a real architectural claim here — world models as a distinct reasoning layer for research, with process supervision and domain primitives you can inspect. And it pushes right against the 'frontier models are the CPU' frame we heard Tuesday. Elicit's basically saying the CPU alone doesn't get you there — you need a separate layer that makes evidence and causality legible. Right, and the line that jumps out at me is legible reasoning still beating neuralese. That's a working product team betting against the black box, not a thought experiment. What I want from a 73-minute episode is whether 'inspectable world model' is a real primitive they shipped, or a slide. They bring up token costs and Gemini in the same breath — so somebody's doing the math on whether legibility is affordable. So the eval-and-reliability story for the week moves from 'are the benchmarks too easy' into the workflow itself. Elicit's answer is: don't just grade the output, make the reasoning inspectable as it happens. Bloomberg Originals has the details on this one. Emily Chang gets seventy minutes with Dario Amodei — San Francisco roots, the OpenAI race, the Pentagon standoff, and where this all ends. Bloomberg ran it while Anthropic is carrying a $965 billion private valuation. Seventy minutes of the extended cut. And his own co-founder Jack Clark is on the BBC the very next day saying the industry's got a gas pedal and no brakes. Feels like a press strategy with two lanes. For me, the interesting part in a long-form like this isn't the SF-origins texture, it's the Pentagon standoff. That's where 'who controls the stack' stops being a panel-talk abstraction and becomes a contract he either signs or walks from. Right, the standoff is the only falsifiable thing in there. A guy running a near-trillion-dollar company telling the Pentagon no — I want to hear what he actually said no to, not the governance philosophy that sounds great on Bloomberg. This one's from AI & I:

Last year, there were 1 billion commits on GitHub. This year, Kyle Daigle expects that number to exceed 14 billion, a two-component explosion caused by more humans—and their agents—issuing pull requests. In March alone, 17 million pull requests on GitHub were created by agents.

One billion commits last year. Daigle says north of fourteen billion this year. And seventeen million pull requests in March were created by agents — not people. That's the most falsifiable number in the rundown because it comes from the platform that actually hosts the work, not from a VC forecast or a model demo. Right, but fourteen billion commits doesn't mean fourteen billion good commits. Seventeen million agent PRs is also seventeen million things a human maintainer now has to review or ignore. And that's the tension. Daigle's pitching agentic code review and model routers — infrastructure for the flood. The pitch assumes the flood is value, not just volume. His durable-advantage line is 'developer choice.' From the COO who's also Microsoft's CMO for dev products. Convenient that the moat is the thing you sell. Here's Faisal Islam at BBC World Service:

“Right now, it’s like the AI industry has a gas pedal, but it doesn’t have a brake pedal in the car. And what we’re saying is we want to build that brake pedal so we in the world have an option. In the future, you might say: ‘Let’s get all of the benefits we can for, say, biology and medical research, and let’s take a pause on AI research, where we can absorb the societal changes.’”

Jack Clark on BBC World Service: the AI industry has a gas pedal but no brake pedal, and Anthropic wants to build the brake. Great line. Plays perfectly on radio. It does play. But it sits right next to the Dario extended interview we just heard — the Anthropic co-founder doing 70 minutes of capability talk while the other one's on the BBC asking for a pause. Right, so which is it? And the metaphor skips the only part that matters — what is the brake? A regulation? A kill switch? A clause Anthropic gets to veto? Because 'we want to build it' means they hold the pedal. That's the part I want pressed. He says governments and society need a mechanism to slow things down — fine, but a brake somebody at Anthropic controls isn't society's brake. He doesn't name who turns the key. Khosla says AI replaces most doctors in five years, Hassabis gets the superintelligence biography, and now Clark wants a brake pedal. Three grand proclamations, still zero named failure conditions between them. Eric Ries, writing in SaaS Group:

But after more than a decade of watching those ideas spread across startups, enterprises, and even governments, he came to a harder realization: teaching people how to build quickly is not the same as teaching them what is worth building, or how to protect it once it starts to matter.

Eric Ries calling vibe coding a slot machine — great line. But The Lean Startup guy diagnosing a methodology problem only counts if he names what the slot machine actually breaks. He kind of does, though. His point is the demo shows up before the product, so teams feel productive while learning less. That's the failure mode. Okay, that lands. A working demo gives you the dopamine hit, but it doesn't validate the experiment — you pulled the lever, lights flashed, and you learned nothing. And that runs under the whole week. The GitHub COO numbers we just hit, the evals talk, the YC sprint — all of it assumes disciplined iteration. Ries is saying the default isn't disciplined. That's the assumption nobody's been checking. That's why it's worth more than the soundbite. The guy who taught a generation to build fast spent a decade realizing fast isn't the same as building the right thing — and now he's watching AI hand everyone a faster lever. Got thoughts on today’s briefing, a story we should be watching, or a correction? Send us a note anytime at techpodcastpodcast at lantern podcasts dot com. We really do read what you send.

You’ll find links to every story we mentioned today in the show notes, so if something stuck with you, that’s the place to dig in a little further.

That’s Tech Podcast Podcast for today. This is a Lantern Podcast.