← Tech Podcast Podcast

AI Agents Hit the Real Bottlenecks: Power, Evals, Access (June 16, 2026)

June 16, 2026 · 11m 5s · Listen

Turns out tokens and talent aren't the only things throttling AI agents. The power grid, your eval suite, and who can even buy in are starting to matter. If you're just joining us: mega-startups are staying private longer, so more of their early compounding happens before public investors can touch them. SpaceX is the biggest test case — launch economics, Starlink, Starship reuse, even orbital compute all feed the question of what you'd actually be buying at IPO. This is Tech Podcast Podcast. Today: the Perplexity CEO defending the export controls he profited under, a researcher declaring prompt engineering dead, and the firm that quietly built a twenty-billion-dollar SpaceX stake. Let's find the numbers. Okay, 137 Ventures — over one percent of SpaceX, roughly twenty billion dollars. After three days of guys nodding along that the future is electric, here's somebody who actually built the position. Right, and I want to know whether Justin Fishner-Wolfson explains the actual secondary-market plumbing — how you accumulate that in a company almost nobody can touch — or whether we just get a conviction story with a big number on it. It's the same illiquid private market that's leaving European funds stuck, unable to exit or raise — and 137 built a fifteen-billion-dollar platform inside it. What did they do differently? That's the part I want. And it lands right on the IPO math. Instead of speculating about when SpaceX goes public, we finally have a real equity stake to reason from. Perplexity tripled revenue to over five hundred million ARR this year — and Aravind Srinivas is out there arguing export controls helped China. So does the revenue number make him credible, or does it just mean he's been winning while the exact policy he's defending was in place? The mechanism he names matters. He's claiming Chinese labs got better because they were forced off US chips. I want to hear that chain, because he's making a causal claim instead of just doing think-tank vibes. And he says power is the bottleneck now, not chips. If compute's gated by watts instead of dollars, that quietly shrinks the whole Google-cuts-token-prices-eighty-percent threat everyone's been waving around. Both can be true at once, though — cost and access on Monday, power today. That's an update worth pressing, not a contradiction to wave away. Then Jim Fan comes in with: prompt engineering becomes irrelevant. That's a direct shot at an entire consultant cottage industry — and there's no falsifiable number under it. It's only useful if he names what replaces it as the scarce skill. His Imbue framing is about scaling data for embodied agents — same 'no Common Crawl for the physical world' problem a16z raised Tuesday, now with a named researcher and a domain. Pair it with Khosla saying AI replaces most doctors in five years — two giant proclamations, same news cycle, not one named failure condition between them. That's the pattern. Braintrust's Ankur Goyal gives the other half of this: evals as the modern PRD. Put that next to the governance question: who's accountable when an agent fails in production. And if Fan's models are inspecting factory robots in the real world, the eval is the safety layer. Same question Levie got Thursday, just with a robot arm attached. Here's Imbue:

“The reason prompt engineering will not be relevant forever is because RLHF – why prompt engineering even exists in the first place – is because these systems are misaligned with what humans want, so we have to kind of coerce the model to give us what we want by typing out very unnatural sentences, and to essentially trick the model into solving the task.”

Jim Fan, an NVIDIA research scientist, says prompt engineering isn't a real job. Bold thing to put in print when there's an entire consulting cottage industry built on exactly that. But there's a mechanism under it — he ties it to RLHF. The claim is that as models get tuned on human feedback, the careful phrasing stops mattering because the model meets you halfway. Okay, that's at least a falsifiable version. What I want is the part he skips: if prompt engineering dies, what's the scarce skill that replaces it? Because the usable takeaway is the replacement, not the obituary. Right, and his whole Imbue framing points there — embodied agents, MineDojo, the Minecraft benchmark. The bottleneck moves from how you ask to how you scale the data the agent learns from. There's no Common Crawl for physical-world tasks. Here's Podcast Alpha:

The industry earns over $200 billion per year in exports and employs millions of Indian engineers doing work that spans software development, business process outsourcing, and IT support. Vinod argues AI agents can perform most of that cognitive and process work at a fraction of the cost. The displacement is not a question of if, only when.

Khosla says AI replaces most doctors in five years. Same news cycle, Jim Fan says prompt engineering is dead — and you know what neither of those grand claims comes with? A falsifiable number. Except here, Khosla actually gives one — his son's company layering domain-specific AI on GPT-5, near-zero triage error versus 20 to 30 percent for general LLMs. That's the rare claim with a stat bolted on. The India number is where I'd pressure-test him. He says the $200 billion IT services industry is just gone — flat, no hedge — employing millions of engineers. And the mechanism is real enough: agents doing BPO and IT support at a fraction of the cost. But 'pivot to AI deployment at scale' is a huge assumption for an industry that size. Here's Aravind Srinivas at TechFounder:

Aravind Srinivas is the Founder and CEO of Perplexity, one of the fastest-growing AI companies in the world. Since the start of the year, Perplexity has tripled revenue to well over $500M in ARR. Aravind has raised over $1BN for the company with reported valuations reaching $20BN.

Perplexity tripled revenue to north of $500 million in ARR since January — and the same guy putting up those numbers is out here arguing U.S. export controls actually helped China. That's a contrarian geopolitical call from someone with real skin in the game, rather than a think-tank op-ed. Right, but I wanna know the mechanism. Srinivas saying export controls helped China — what's the actual claim? That cutting them off US chips forced them to get efficient? Because 'constraint breeds innovation' is the oldest hand-wave in the book unless he names what specifically got better. And the other live wire is the bottleneck flip. He's saying it's power now — watts, not tokens, not chips. That's a direct update on the cost-as-access framing we ran Monday. And that one actually has teeth. If compute's gated by the grid instead of dollars, then the whole 'Google cuts token prices 80% and wins overnight' thesis shrinks — you can't price-lever your way past a power constraint. That's a counter with some weight behind it. Claire Vo, writing in Lenny's Newsletter:

We get into how coding agents can take on deeply technical architecture and infrastructure work that no single human engineer could tackle before, and then we demystify evals so you can use them to make your AI products better without touching the implementation.

Ankur Goyal's line from this Lenny episode is the one I'd actually pull: evals are the modern PRD. Braintrust runs the evals layer for Notion, Stripe, Vercel — so when he says that, it's an operator describing how they ship, not a slogan. And right after Jim Fan tells us prompt engineering's dead, here's the guy saying the new scarce skill is writing the eval that encodes your taste. That's the falsifiable version of Fan's claim — somebody actually building the thing that replaces it. The concrete detail is the Codex bit — he's running week-long benchmark experiments across database indexes and column-store formats. Agents that benchmark tirelessly, so there's no excuse to skip the rigor. Which is the answer to the accountability question, right? When an agent ships something broken in production, the eval is the thing that catches it. 'Evals are the PRD' is also 'evals are the receipt.' Justin Fishner-Wolfson, writing in Sourcery:

Justin Fishner-Wolfson is Co-Founder & Managing Partner of 137 Ventures, the firm that turned a contrarian read on private markets into a $15B platform & one of the largest SpaceX positions in venture. His firm now owns more than 1% of SpaceX, a stake worth roughly $20B at the company’s $1.77T listing valuation.

Okay, this is the one I've been waiting for. After a week of people just agreeing the future is electric, here's the firm that actually built a 1%-plus position in SpaceX — roughly twenty billion dollars — and Justin Fishner-Wolfson walks through how. And the access gap finally has a cap-table face. He's been buying since SpaceX was a one-billion-dollar company — roughly two dozen times, all secondaries and tenders, never sold a share. That's the operating detail: he names the mechanism — secondaries and tenders, the two-hundred-forty-billion-dollar secondary market. It's how a single firm touches a company most people can't. And the part I'd press him on — he says partial liquidity sharpens founder focus instead of dulling it. That's a real claim about why companies stay private longer, and it's the underwriting thesis underneath the whole $20B stake. If Tech Podcast Podcast is part of your routine, take a second to subscribe and leave a review wherever you're listening. It really helps other people find the show, and it helps us keep bringing you these daily updates.

We'll be watching when SpaceX's post-listing lockup lifts — that's the next checkpoint for whether 137 holds, trims, or recycles its more-than-1% stake.

Links to every story we covered today are in the show notes, so if one of them is still on your mind, you can head there and read further. That's Tech Podcast Podcast for today. This is a Lantern Podcast.