OpenAI just bought the package manager and linter every Python developer leans on — so I'm less worried about uv dying than about what OpenAI thinks owning that toolchain gets them. If you're just joining, this week's been about operating discipline more than agent demos. Box's Aaron Levie was pushing CIOs toward agentic AI, Claude's Zcash bug find raised trust questions around automated audits, and Braintrust's Ankur Goyal made the case that teams need encoded taste, evals, and CI to make coding agents actually reliable. This is Tech Podcast Podcast. Today — Charlie Marsh on whether uv is toast, a frontier eval lead saying the old tests are too easy, and a YC Demo Day stacked with founders. Astral first. Charlie Marsh laid out his answer on Talk Python 552. Recorded June 2nd, published the 17th — which means he had two weeks to figure out exactly what he wanted to say about getting acquired. Two weeks to polish, and the framing they led with was, 'wait, is uv toast?' You don't lead with that if you're totally relaxed about the integration. Here's what I want from him — did OpenAI buy the tooling, the team, or the ecosystem trust? Because those three don't survive a merger at the same rate. And it's the second time this week. Evan You drew a hard line around what Cloudflare wasn't getting from Vite. Now we've got two cases — does Marsh draw the same line, or does he give away more? Two open-source tools acquired by infrastructure players in one week. At this point it feels like a board with names on it. Who's next? The part I keep coming back to: uv and Ruff sit under Python AI development everywhere. The access fight isn't just compute pricing anymore; it's who controls the toolchain. Different layer, same lock-in. Speaking of which — Tejal Patwardhan on the OpenAI Podcast. She runs frontier evals there, and she's saying the old benchmarks are getting too easy. Okay, this is the escalation. All week, we've had pundits saying benchmarks are saturating. Now the person whose literal job is benchmarks is saying it. So name one. Which test did you retire? If she names a specific benchmark they walked away from, that's the thing worth holding onto. Otherwise it's vibes with a title attached. That's the testable version of what Goyal was circling Monday with evals-as-PRD. He was selling a product; she's the institutional weight behind the claim. Then the YC Demo Day lightning round on TBPN — seventeen-plus founders, all application layer, under three hours. Harj Taggar's on the panel. Harj asks the sharp follow-up, which is the fun part — with that many founders rotating through, I'm listening for who drops a real production number, and who just pitches the category and bounces. It also tests that whole 'commoditize the infra, value moves up' thesis. If that migration's actually happening, this is the most concentrated sample we'll get of it in real time. TBPN also opens with SpaceX ripping at 01:05 — which lines up against that secondaries-and-tenders plumbing we were chewing on Monday. Curious if there's an actual new number in it or just a headline. TBPN writes:
Andrew Lee, co-founder of Firebase (acquired by Google in 2014), discusses his latest venture, Tasklet, an AI agent platform that integrates with various work tools to automate workflows. He highlights Tasklet's rapid growth, achieving a $7 million run rate, and its ability to dynamically generate integrations using AI, enabling connections to both public and internal APIs.
YC Demo Day is a two-hour-fifty lightning round, with dozens of founders rotating through. I'm watching for the production numbers — who actually has one, and who just pitches the category. And we got one. Andrew Lee at Tasklet: a seven-million-dollar run rate, a twenty-million-dollar round, and an AI agent platform that generates integrations on the fly. Finally, an actual number instead of a vibe. Lee's the Firebase guy — built developer infrastructure, sold it to Google in 2014. So when he says Tasklet's positioned against the major AI labs, he's not bluffing about what it takes to build that stack. Right, but 'competitive positioning against major AI labs' is exactly the phrase I'd make him defend. Dynamically generating integrations to internal APIs is the interesting part — that I'd queue. Then there's Eden Robotics — a wheeled semi-humanoid claiming over 80% of industrial tasks, a twenty-hour battery, and labor-as-a-service pricing by the hour of operation. That pricing model is the part I'd argue with. Per-hour robot labor is at least falsifiable — you can put it next to a wage. Eighty percent of industrial tasks, though? Show me the eighty percent. Talk Python To Me Podcast writes:
OpenAI just acquired Astral, the company behind uv, Ruff, and ty. And if your first thought was "wait, is uv toast?", you are not alone. But here's the twist Charlie Marsh shared with me: he thinks they may ship more open source at OpenAI than they ever did at Astral.
OpenAI just bought Astral — uv, Ruff, ty — the Rust-fast tooling half of Python now ships on. And Charlie Marsh's pitch to Talk Python is that OpenAI might ship more open source than Astral did on its own. That's the line everyone's going to clip, and it's exactly the one I'd hold him to. uv and Ruff matter because they're under half the Python toolchains out there now. And notice the dates — recorded June 2nd, published the 17th. Charlie had two weeks to sand that answer down. So what did OpenAI actually buy: the tooling, the team, or the ecosystem trust? My read: they bought standard-setting. Own the package manager and the linter, and you quietly own how AI-generated Python gets written and shipped. The open-source promise is the part that has to survive contact with that incentive. Right. 'More open source' is easy to say while OpenAI controls the roadmap. I want one falsifiable thing — does uv stay Apache-licensed with outside maintainers who can say no? From GoLoud:
Drawing on his experience building Lattice from startup to multi-billion-dollar company, Altman explains how founders should think about customer requests, when to pivot, and why some of the hardest decisions come from balancing conviction with market feedback. He shares lessons from the early days of Lattice, including finding product-market fit, building a sales motion, hiring the first employees, and navigating the tradeoffs that emerge as companies scale.
Jack Altman on Speedrun talking product-market fit. Twenty-nine minutes on customer requests and when to pivot. I've heard this episode before — I just don't know who recorded it. The reason I wouldn't skip it: Altman built Lattice to a multi-billion-dollar outcome — full-cycle founder, not pundit. I want to know whether that arc gets him past 'listen to your customers.' Right, and Lattice has been through actual public scrutiny over its own AI product calls. If he talks about that — the conviction-versus-feedback decision when it goes wrong — that's the operating detail. If it's just the clean version, skip it. The pivot section is where I'd start. Anybody can say 'balance conviction with market feedback.' I want the specific moment a Lattice customer asked for something and he said no — and what it cost. This one's from OpenAI:
The old tests are getting too easy. Tejal Patwardhan leads OpenAI’s frontier evals team, which is finding new ways to measure and forecast progress as models become more capable. She and host Andrew Mayne discuss why evals matter for research, how benchmarks can break or get gamed, and what models need to be judged on next.
This is the concrete version of Monday's Braintrust conversation — evals as the modern PRD. Now you've got Tejal Patwardhan, who actually leads frontier evals at OpenAI, saying the old tests are getting too easy. And that's the escalation that matters. Monday it was a startup selling eval CI. Today it's the person inside the lab whose whole job is measuring this, telling you the rulers are breaking. The chapter I want is 11:20 — why old benchmarks stopped working. If she names a specific benchmark they retired, that's the falsifiable piece. Otherwise it's just a vibe about saturation. Right, 'getting gamed or getting too easy' is the whole framing — but which one? Tell me the test you killed and what replaced it. That's where the receipt is. The Blog of Author Tim Ferriss, with Tim Ferriss:
Sebastian Mallaby (@scmallaby) is the Paul A. Volcker senior fellow for international economics at the Council on Foreign Relations, a two-time Pulitzer Prize finalist, and the author of six books, including More Money Than God, The Power Law, The Man Who Knew, and The World’s Banker. His latest book is The Infinity Machine: Demis Hassabis, DeepMind, and the Quest for Superintelligence.
Okay, so it's Tim Ferriss versus a Council on Foreign Relations fellow who just wrote a Demis Hassabis biography — The Infinity Machine. Mallaby talked to a hundred-plus AI insiders for this thing. And he comes in with real reporter weight: The Power Law on venture capital, More Money Than God on hedge funds. He translates closed worlds for outsiders. Which is exactly why 'The Religion of AI' in the title makes me twitch. A serious reporter reaching for superintelligence-as-faith framing — is that his read, or the marketing department's? The question I'd hold him to: with a hundred sources inside DeepMind, what did he learn that you can't get from Hassabis's own podcast tour? A biographer's job is the stuff the founder won't say on mic. Right — give me one operating detail about how DeepMind actually decided to bet on AlphaFold over the safer thing. If the chapter just says 'the race to superintelligence,' that's marketing copy. If Tech Podcast Podcast is part of your daily routine, consider subscribing wherever you're listening. And if you have a moment, leave a quick review — it really helps other people find the show.
You'll find links to every story we covered today in the show notes, so if something stuck with you, that's the place to dig in a little more.
Thanks for listening. That's Tech Podcast Podcast for today. This is a Lantern Podcast.