AI’s appetite for training data is turning into a very expensive habit. Somebody, apparently, is finally getting paid. This is AI Daily Briefing. Today: the data-economy number behind the infrastructure splurge, plus two science results and an agent benchmark that actually has to make it through eight hours. Start with Micro1—the headline is impressive, but the number underneath may explain the whole buildout. Tap follow so the next episode finds you. TechCrunch, with Marina Temkin:
One of these fast-growing businesses is Micro1, a four-year-old startup that expanded its gross annual run rate from $100 million to $500 million over the past eight months, according to a person familiar with the company. Like its peers that hire domain experts such as doctors, lawyers, and scientists on a contract basis, Micro1 retains roughly 60% to 70% of that figure, putting its net annual run rate between $150 million and $200 million.
Micro1 going from $100 million to $500 million in gross run rate in eight months is the first clean read on demand in this infrastructure boom. The appetite isn't stopping at GPU leases; it's reaching doctors, lawyers, and scientists labeling the material labs want. And $500 million gross becomes roughly $150 million to $200 million net once 60% to 70% goes back out to contractors. Healthy business. But the number you can actually operate on is a lot smaller than the headline. Still, that net run rate is real evidence that somebody can fund this stack with customer demand, not just infrastructure financing. Micro1 says automated video descriptions and reusable off-the-shelf datasets can push gross margins to 80% or 90%—which raises the obvious question: how carefully will buyers police where that same data gets sold? Exactly. If labs decide their training data is a moat, some of this work moves in-house fast. Micro1's margin expansion depends on synthetic and repeat-sale data staying valuable after the customer has seen it once. Here's Lixue Cheng at Nature Communications:
Here, we bring this paradigm to fruition with Orbformer, a transferable wavefunction model pretrained on 22,000 equilibrium and dissociating structures that can be fine-tuned on unseen molecules reaching an accuracy–cost ratio rivalling classical multireference methods. On established benchmarks as well as more challenging bond dissociations and Diels–Alder reactions, Orbformer is the only method that consistently converges to chemical accuracy (1 kcal/mol).
Orbformer is the kind of benchmark claim I’ll take seriously: 1 kilocalorie per mole on bond dissociation and Diels–Alder reactions, consistently. Pretrain on 22,000 structures, fine-tune on an unfamiliar molecule, and stop paying full multireference-compute price every time. And Microsoft Research put the methodology in Nature Communications. There’s a technical report, established benchmarks, and a precise accuracy target—1 kcal/mol—not a cinematic molecule animation and a promise that it’s “transformative.” If that accuracy-to-cost ratio holds outside this paper’s test set, chemistry teams will care a lot less about the model’s mystique than about getting classical multireference-quality answers cheaper. That makes a real case for deployment. We just talked about demand spreading through the training-data supply chain. Here’s the other side of the infrastructure story: the biggest balance sheets are funding compute buildings and, increasingly, the scientific tools that determine what those buildings get used for. University of Waterloo, with Benjamin Schneider:
We evaluate ASH on two complementary environments demanding multi-hour planning: Pokémon Emerald, a turn-based RPG, and The Legend of Zelda: The Minish Cap, a real time action-adventure game. In both games, behavioral cloning, retrieval-augmented and zero-shot foundation-model baselines plateau, while ASH sustains progression across our 8-hour evaluation. ASH reaches an average of 11.2/12 milestones in Pokémon Emerald and 9.9/12 in Legend of Zelda, while the strongest baseline gets stuck in both environments at an average of 6.5/12 and 6.0/12 milestones, respectively.
Eight hours in Pokémon Emerald and Minish Cap is a much better agent test than the usual little three-step routine. ASH gets 11.2 of 12 Pokémon milestones; the best baseline stalls at 6.5. That gap’s hard to ignore. And Benjamin Schneider at Waterloo is doing this from unlabeled, noisy internet video—no hand-built reward and no expert action labels. A result this specific from a single-researcher thesis deserves attention, even if it isn’t a major lab announcement. I still want the production translation. Pokémon is controlled: you’ve got a finite world and finite controller actions, and all 12 milestones are countable. Does that 11.2 survive the step-seven failure that wrecks a ten-step workflow with messy company data? Unknown—but at least this test gives us a real place to start asking. And there’s an actual 14-megabyte technical document attached, with an eight-hour evaluation and baseline comparisons. After the Orbformer paper we just covered, this is the standard: show the methodology and numbers so we can interrogate the claim. Ning-Zheng Li; Zi-Yu Li; Qiang Shi; Qing-Yu Liu; Sheng-Gui He, writing in Nature Communications:
Based on ≈ 93,000 global-minimum structures obtained via automated first-principles calculations, our model—the cluster-transformer-encoder network—enables reliable predictions of atomization energies (≈ 40 meV per atom accuracy) for 9.13 million metal cluster compositions, covering 30 d-block metals and 4 chemically relevant ligand elements (C, N, O, and S).
Ninety-three thousand automated first-principles calculations to map 9.13 million cluster compositions—that’s the compute appetite people miss when they reduce the whole buildout to chatbot inference. Materials science is pulling on the grid too. And this is a sensible use of ML: spend the expensive quantum-compute budget on 93,000 global-minimum structures, then use a composition model for the first pass across 30 d-block metals. Forty meV per atom is a number a chemist can argue with, which I respect. Nature Communications has the methodology attached, plus a stated accuracy and clear scope: small clusters with C, N, O, and S ligands. Same pattern we saw with Orbformer—science claims with enough detail to actually interrogate. Don’t turn 9.13 million predictions into 9.13 million validated materials. This is first-pass triage. But if it reliably narrows a search that used to eat months of first-principles compute, the cost curve changes long before anyone puts a new catalyst in a reactor. If you follow AI’s impact on business and technology, you may also enjoy The Data Center Daily, a daily briefing on AI compute, hyperscaler capex, the power grid, semiconductor supply, and energy markets. Find it wherever you listen to podcasts.
Links to every story are in the show notes, if you want to dig into anything you heard. Thanks for listening, and we’ll be back tomorrow. That’s AI Daily Briefing for today. This is a Lantern Podcast.