← Tech Podcast Podcast

Edge AI, AI Spend Cuts, and the New VC Liquidity Machine (July 06, 2026)

July 06, 2026 · 9m 16s · Listen

A founder wants the model living on your phone, a crypto giant just cut its AI bill in half, and Goldman's parking twenty-two billion in venture. Monday's busy. If you're just joining: AI teams have realized inference isn't a rounding error. Token pricing turns every prompt, every agent loop, every chatty tool call into an operating expense. The push to cut spend started with the boring mechanics — shorter prompts, cheaper models, tools that force fewer tokens — not some grand debate about AI's future. This is Tech Podcast Podcast, and after four days of scale-is-everything takes, today gives us a device-native AI pitch and a 50% budget cut on the same rundown. Something's cracking. Let's start with Liquid AI's Ramin Hasani and where the model actually lives. Cognitiverevolution writes:

If the dominant story of modern AI is "scale is all you need," this episode is the loyal opposition's most technically grounded rebuttal. Ramin Hasani, co-founder and CEO of Liquid AI, joins Nathan for ninety minutes on what it actually takes to pack the maximum amount of intelligence into the smallest possible package — and why the answer to that question keeps turning out to be different from the answer to "how do you build the most capable model in the world."

Okay, ninety minutes with Ramin Hasani, and it starts with a worm's 300-neuron nervous system and winds up in a Mercedes dashboard. That's the range. But here's what I actually care about — he's the loyal opposition to 'scale is all you need.' The bet is that the smallest package wins, not the biggest cloud model. Right, and the technical claim is specific: gated convolutions and selective attention over raw scale. He's making an architecture argument here, not a vibes argument. If the model lives on the device, the whole inference-cost story we've been chewing on all week changes shape. That's the part I want pinned down. If it's device-native, who's paying the inference bill? Nobody, ideally — and that's exactly why Liquid's pitch sits so awkwardly next to every scale benchmark we've heard. What makes this more than a smaller-model story is the path from bio-inspired MIT research to hardware-aware architecture search. He's saying the answer to 'most intelligence per gram' is different from the answer to 'most capable model on earth.' That distinction is the whole episode. And it's a real test of that leading-indicator idea — is edge AI inventing its own path, or just following cloud AI with a lag? Hasani's whole bet is that it's a different path. I want to hear if Nathan makes him defend that or lets him coast. Harry Stebbings picks up the cost-pressure side of this on The Twenty Minute VC. Okay, the headline is 'Dario declares war on open-source,' but the line that actually matters is tucked right up top: Coinbase cut AI spend by fifty percent. That's the operator data point we've been circling all week. A company deep in AI, halving the bill — that's a real corporate number, not surplus-compute theory. And I want to know if Stebbings actually asks why, or just uses it as debate fuel for the Anthropic segment. Because 'AI token bubble bursting' and 'Anthropic warns open-source will destroy the business model' — those two agenda items are quietly at war with each other. Right — if customers are pulling back on inference costs, that's exactly the pressure that makes Liquid AI's device-native pitch look a lot more practical. Anthropic's framing is basically: open-source could destroy the commercial AI business model. That's a strange thing to warn about if your product is supposedly this insurmountable moat. And it turns the open-versus-closed fight into a margin fight. Somebody's economics are under threat, and they're saying the quiet part out loud. From Hans Swildens at The Split:

Hans Swildens started Industry Ventures in 1999, and we recently sat down to record his first podcast since getting acquired by Goldman Sachs in 2026.Hans has been buying venture secondaries longer than almost anyone, and the combined business is one of the largest VC portfolios in the world: 525 firms, 1,600 funds, and over $22B in capital commitments.

So this is the number sitting under the whole week. Swildens sold Industry Ventures to Goldman — 525 firms, 1,600 funds, over 22 billion in commitments — and the line that matters is secondaries going from a 250 million market to 150 billion in 25 years. Right, and he built the thing by buying fund stakes at 99% discounts during the Dot Com collapse. That's the whole game — he's a distressed buyer who's now the exit. And that reframes the IPO chatter from the 20VC episode we just hit. When Swildens says VCs are manufacturing their own exits, this is the mechanism — asset managers buying venture firms because the liquidity has to come from somewhere. The one I keep chewing on is the seed line — most seed funds without a real angle have five years left. Coming from a guy who buys the wreckage, that sounds like him telling you which funds become his inventory. And he passed on a multi-billion-dollar data center deal, which — given every AI-spend story this week — makes me want his actual reasoning, not just the war story. This one's from Security Conversations:

Three Buddy Problem – Episode 104: We discuss the return of Anthropic’s Fable 5 from export-control suspension with guardrails so aggressive that spelling “exploit” gets you downgraded. Plus, a debate on AI frontier labs killing businesses at scale, and OpenAI offering equity to the US government.

Buried on page nine of the Scattered Spider indictment — Microsoft's GDID. A persistent Windows device fingerprint nobody had ever documented publicly, and it's what tied a teenager to the breach. Yeah, the mechanism is the real story here: a silent identifier baked into Windows that survives whatever OPSEC the kid thought he was running. And it shows up in a court filing, not a Microsoft disclosure. That's how we learn it exists — from prosecutors describing how they caught someone. Guerrero-Saade and Raiu are the right people to sit with that, too. If GDID is persistent across reinstalls, every APT-tracking assumption about clean machines just shifted. I want that conversation more than the Fable 5 guardrail comedy. Okay, but the guardrail comedy is real — Anthropic's Fable 5 comes back from its export timeout and now spelling 'exploit' gets you downgraded. Try doing malware analysis with a model that flinches at the word. Here's Paperdive:

Tell an AI coding agent "careful, this is production" and, measurably, almost nothing changes — agents acted 65.5% of the time on throwaway surfaces and 64% on production-like ones. A new benchmark of over two thousand prompts finds agents respond to what's missing from an instruction, not to how much damage a command could do, and that refusal is nearly extinct.

Okay, this is the number I've been waiting for all week. 65.5% action rate on throwaway surfaces, 64% on production-like ones. You tell the agent 'careful, this is production' and it changes its behavior by a point and a half. And that's a falsifiable finding — over two thousand prompts in an arXiv paper from Ji, Zhang, and Xu. They measured it: saying 'be careful' does functionally nothing. Which lands right on top of every 'human in the loop' recipe we heard on the outer-loop pitch a couple days back. Turns out the human saying words is the loop, and the words don't matter. The lever that does work is almost boring — name the exact resource, skip the warning. The agent responds to what's missing from the instruction, not to how much damage the command could do. And refusal is basically extinct. No config in the study refused more than 2.5% of the time. Even at maximum ambiguity, the most cautious system still just went ahead in 36% of runs. If Tech Podcast Podcast is part of your daily routine, take a second to subscribe and leave a quick review wherever you're listening. It helps other people find us, and it really does make a difference.

You'll find links to everything we covered today in the show notes if you want to dig into the pieces that caught your ear. That's Tech Podcast Podcast for today. This is a Lantern Podcast.