GPUs, TPUs, & The Economics of AI Explained | Gavin Baker Interview
GPUs, TPUs, & The Economics of AI Explained | Gavin Baker Interview
Summary
- Gemini 3’s real significance: pre-training scaling laws held — stated “unequivocally,” which matters because “no one on planet Earth knows how or why scaling laws for pre-training work.” Baker’s contrarian frame: given the 200,000-Hopper coherence ceiling and Blackwell’s brutal transition, “there really should have been no progress in ‘24 and ‘25” — “reasoning kind of saved AI,” bridging an 18-month gap (ARC-AGI: 0→8% in four years, then 8%→95% in three months). The scaling laws are multiplicative, so “the Blackwell models are going to be amazing.”
- Google’s low-cost-token advantage is temporary. Gemini 3 was trained on 2024-25-era TPU v6/v7 — “F-4 Phantoms” next to Blackwell’s F-35. Once GB300s (drop-in compatible with GB200 racks) shift to inference, vertically integrated Blackwell users become the low-cost producers, and Google’s rational strategy of “sucking the economic oxygen out of the AI ecosystem” at negative-30% margins gets painful: “it might start to impact their stock.” When Reuben lands, “the gap is going to expand significantly.”
- The ASIC field likely narrows to TPU and Trainium. Broadcom takes 50-55% gross margin on the TPU back end — ~$15B of a ~$30B 2027 program against ~$5B of divisional opex — so in-housing is “absolutely inevitable” (MediaTek is the warning shot). It takes three generations to make a good chip, and Nvidia’s answer to every in-house ASIC is annual cadence: “you cannot keep up with us.”
- Only four labs matter — OpenAI, Gemini, Anthropic, xAI — and the gap is compounding. Reasoning restarted the data flywheel (user feedback as verifiable reward), internal checkpoints train the next model, and Meta’s failure shows how hard it is: Zuckerberg “was as wrong as it was possible to be.” China’s refusal of Blackwell will blow the gap out — DeepSeek admitted the compute shortfall in its v3.2 paper — and China realizes “whoopsy daisy, we do need the Blackwells” around late ‘26, by which point rare-earth leverage is solved.
- The board repositions around token cost: OpenAI is a high-cost producer paying margins for compute — “you go from $1.4 trillion rough vibes to code red pretty fast” — while Anthropic is “burning dramatically less cash than OpenAI and growing faster” and its $5B Nvidia deal gives Jensen three fighters against Google.
- ROI is “empirically, factually, unambiguously” positive — audited ROIC at the big GPU buyers is up, and Q3 was the first quarter non-tech Fortune 500s quantified uplift: C.H. Robinson quoting truckloads in seconds at 100% of inbound requests (vs 15-45 minutes at 60%), stock +~20%. The risk he watches: a Blackwell ROI “air gap” while the chips do training, since “there’s no ROI on training — the ROI comes from inference.”
- Application SaaS is repeating brick-and-mortar’s e-commerce mistake: clinging to 80% gross margins while AI natives run agents at sub-35-40% — “you are guaranteeing that you will not succeed at AI.” It’s “a life-or-death decision that essentially everyone except Microsoft is failing”; the activist playbook (show low-margin AI revenue) is open to Salesforce, ServiceNow, HubSpot, GitLab, Atlassian.
- Bear cases and bubbles: edge AI — a free, pruned Gemini 5 or Grok 4.1 on-phone at 30-60 tokens/second, “clearly Apple’s strategy” — is “by far the most plausible and scariest bear case.” The rolling bubble has moved from EVs and meme stocks to nuclear and quantum, where no public vehicle is a real leader. Data centers in space are his “most important thing in the next 3-4 years” — “in every way… superior to data centers on earth.”
Deep dive
1. Gemini 3 proved scaling holds — and reasoning “saved AI” while Blackwell was late
- Baker’s first rule for processing any release: use it yourself, and pay. He’s amazed how many “famous and August investors” reach definitive conclusions from the free tier — “the free tier is like you’re dealing with a 10-year-old” and extrapolating to the adult; the $200/month tiers (Gemini Ultra, SuperGrok) are “a fully-fledged 30, 35-year-old.” The rest of the process: AI happens on X, follow the 500-1,000 people on Earth who truly understand it (“everything Andrej Karpathy writes, you have to read it three times — minimum”), and listen whenever anyone from the four labs that matter — OpenAI, Gemini, Anthropic, xAI — goes on a podcast.
- Gemini 3 stated “unequivocally” that pre-training scaling laws are intact — important because it’s not a law but an empirical observation measured precisely without being understood. His analogy: we’re like the ancient Egyptians and the sun — able to align the pyramids perfectly to the equinoxes with zero grasp of orbital mechanics. “It’s really important every time we get a confirmation.”
- The part he thinks public investors misunderstand: “there really should have been no progress in ‘24 and ‘25.” You can’t get more than ~200,000 Hoppers coherent, and Blackwell was “by far the most complex product transition we’ve ever gone through in technology” — air to liquid cooling, racks from ~1,000 to ~3,000 lbs, 30kW to 130kW (“imagine if to get a new iPhone you had to change all the outlets in your house… and reinforce the floor”). “Reasoning kind of saved AI” — RLVR and test-time compute bridged the ~18-month gap: ARC-AGI went 0→8% over four years, then 8%→95% in three months after the first OpenAI reasoning model.
- The kicker: scaling laws are multiplicative, and Gemini 3 was trained on 2024-25-era TPU v6/v7 — in his fighter-jet taxonomy, Hopper is the P-51 Mustang, those TPUs are F-4 Phantoms, Blackwell is an F-35. Apply the two new scaling laws to Blackwell-trained base models and “the Blackwell models are going to be amazing.”
2. Blackwell flips the low-cost-token game — and Google’s calculus with it
- AI is “the first time in my career as a tech investor that being the low-cost producer has ever mattered” — Apple, Microsoft and Nvidia aren’t worth trillions for being cheap. Google, as the low-cost producer of tokens, has been “sucking the economic oxygen out of the AI ecosystem” — running AI at a negative-30% margin is “by far the rational decision” when your competitors need outside funding and you don’t.
- The first Blackwell-trained models arrive early 2026, and the first will come from xAI: per Jensen — on the record — “no one builds data centers faster than Elon,” and a new chip needs 6-9 months of tuning just to outperform the previous generation. xAI effectively debugs Blackwell for everyone else.
- Then the flip: the GB300 is drop-in compatible with GB200 racks — “no new power walls” — and whoever runs GB300s, especially vertically integrated, becomes the low-cost token producer. Once Google loses that title, negative margins turn painful — “it might start to impact their stock” — and when Reuben arrives, “the gap is going to expand significantly versus TPUs” and all other ASICs.
- The board repositions on cost. OpenAI pays a margin to others for compute — “maybe the people who run their compute are not the best at running GPUs” — hence Stargate, and hence “you go from $1.4 trillion rough vibes to code red pretty fast.” Anthropic is “burning dramatically less cash than OpenAI and growing faster,” and its $5B Nvidia deal — Dario understanding Blackwell and Reuben versus TPU — takes Nvidia “from two fighters to three” against Google. The OpenRouter tell: xAI processed 1.35T tokens versus Google’s ~800-900B and Anthropic’s ~700B.
3. The ASIC shakeout: Broadcom’s 50-55% margin is the crack in the story
- Google does the TPU front end; Broadcom does the back end and manages TSMC at 50-55% gross margins. On a ~$30B 2027 TPU program, that’s ~$15B to Broadcom — whose entire semiconductor-division opex is ~$5B. “Google can go to every person who works in Broadcom semi, double their comp, and make an extra $5 billion.” Bringing in MediaTek, at much lower Taiwanese-ASIC margins, is “the first shot across the bow”; a really good SerDes is foundational, and he discusses its value in roughly the $10B-$25B-a-year range.
- “It takes at least three generations to make a good chip.” TPU wasn’t “even vaguely competitive” until v3/v4. Amazon has “the best ASIC team at any semiconductor company” (Graviton, Nitro/SuperNIC) and still needed until Trainium 3 for “okay.” He’ll “be surprised if there are a lot of ASICs other than Trainium and TPU” — and both will eventually run on customer-owned tooling: “no matter what the companies say… the economics make it absolutely inevitable.”
- Meanwhile Lisa Su and Jensen’s answer to every customer ASIC is annual cadence — “we’re just going to accelerate… you cannot keep up with us.” His riff on the aspirants: “Oh wow, you made your own accelerator. What’s the NIC going to be? The scale-up switch? The optics?… Oh s*, I made this tiny little chip. I thought this was easy.”
- Nvidia can’t run that cadence alone — a Blackwell rack has thousands of parts and Nvidia makes maybe 200-300 — which is why the semiconductor-VC renaissance (average founder ~50 years old, ignited single-handedly by Nvidia’s market cap) is foundational: the whole ecosystem has to accelerate together. “My little firm maybe has done more semiconductor deals in the last seven years than the top 10 VCs combined.”
4. Reasoning restarted the flywheel — Meta’s failure shows how hard it is
- He used to quote Eric Vishria [likely] — foundation models without unique data and internet-scale distribution are the fastest-depreciating assets in history — but “reasoning fundamentally changed that in a really profound way”: when many users consistently like or dislike an answer, that’s a verifiable reward you can feed back into the model. The Bezos flywheel that made Netflix, Amazon, Meta and Google increasing-returns businesses “has started to spin” at the frontier labs. “It’s early… but you can see it beginning to spin.”
- The empirical test that it’s hard: Zuckerberg said in January that Meta would have the best AI at some point in 2025 — “I don’t know if he’s in the top hundred… he was as wrong as it was possible to be.” Microsoft failed too (Inflection), Amazon’s Nova isn’t top-20 (Adept). Why: wild variation in running GPUs — “if you have 30% uptime on that cluster and you’re competing with somebody who has 90% uptime, you’re not even competing” — plus “taste,” the intuition for which experiments deserve 50,000 GPUs for days. His retail analogy: run a thousand stores clean, well-lit, stocked, “staffed by friendly employees who are not stealing from you,” and you’re a $20-30B company — roughly 15 companies ever did it.
- The compounding layer: every top lab runs a more advanced internal checkpoint and uses it to train the next model — without one, “it’s getting really hard to catch up.” Which is why “Chinese open source is a gift from God to Meta”: the bootstrap checkpoint.
- And China is squandering that card: forcing its labs onto domestic chips just as Blackwell lands. DeepSeek’s v3.2 paper conceded — in “very politically correct, still a little bit risky” language — that it lacks the compute to match American labs. The gap blows out, China realizes “whoopsy daisy, we do need the Blackwells” — probably late ‘26 — by which point rare earths get solved “way faster than anyone thinks” (DARPA enzyme-refining programs, deposits in friendly countries; “they’re obviously not that rare, they’re just misnamed”). If Blackwell does return to China, “Chinese open source will be back” — which he’d count as good.
5. ROI has been “empirically, factually, unambiguously” positive — and Q3 was the tell
- His irritation with the debate: the largest GPU buyers are public companies with audited quarterly financials, and their ROIC is higher than before the spending ramp — partly opex savings, largely moving ad-recommender systems from CPUs to GPUs, which accelerated revenue growth. “But so what? The ROI has been there.” Inside every big internet company the revenue owners fight researchers for chips: “It’s a very linear equation. If you give me more GPUs, I will drive more revenue.” And with Blackwell — “for sure with Reuben” — economics will finally dominate the prisoner’s dilemma that has driven spending (a dilemma fed partly by a religious belief in ASI: “almost all of them want to live forever”).
- The Q3 milestone: the first quarter non-tech Fortune 500s gave quantitative AI uplift. C.H. Robinson, the freight broker, now quotes price and availability on 100% of inbound requests in seconds — versus 60% of requests in 15-45 minutes before — and the stock went up ~20% on earnings.
- The adoption clock rhymes with cloud: every startup ran on AWS by the first re:Invent in 2013; the Fortune 500 standardized ~5 years later. VCs are more bullish than public investors because they see revenue-per-employee gains in AI-native companies directly (ICONIQ’s charts, a16z David George’s “model busters”) — and the young founders “get more polished faster… because they’re talking to the AI.”
- His worry was a Blackwell ROI air gap: capex unimaginably high while the chips do training, and “there’s no ROI on training… the ROI comes from inference” — Meta already printed a declining-ROIC quarter that hurt the stock. C.H. Robinson-type quarters suggest the gap is navigable. Side call: VC-run AI holding companies won’t out-execute buyout firms — “you’re just not going to beat private equity at their game” — but PE will apply AI systematically.
6. From intelligence to usefulness — and edge AI, the bear case that scares him
- The event path: GB300 (and probably more the MI450 than the MI355) collapses per-token cost, letting models think much longer. Gemini 3 booked him a restaurant reservation — “the first time it’s done something for me” — and reservation → hotel → flight → Uber means “all of a sudden you got an assistant.” At tech-forward big companies “50%-plus of customer support is already done by AI” — a $400B industry — and AI excels at persuasion, i.e., sales: of a company’s three functions (make, sell, support), AI may be good at two by late ‘26.
- The organizing principle, via Karpathy: with software, anything you can specify you can automate; “with AI, anything you can verify, you can automate.” Do the books reconcile? Did the sale convert? “That’s just like AlphaGo… most important functions are important because they can be verified.”
- Progress is going invisible to non-experts — at the paid tier you must probe true-expertise questions (PCIe vs Ethernet for scale-up networking) to see model differences, though “these new models are quite a bit better” at his charity fantasy-football lineups. So “we need to shift from getting more intelligent to more useful” — then usefulness must hand off to scientific breakthroughs. Usefulness means doing things consistently while holding all your context (his trip-planning test: east-facing balcony for Huberman-style morning sun, Starlink on the plane, the whole family’s preferences) — METR-style task length has to keep extending.
- The scariest bear case, other than scaling laws breaking: edge AI. In ~3 years a bigger, bulkier phone runs a pruned Gemini 5 or Grok 4.1 at 30-60 tokens/second — free. “This is clearly Apple’s strategy”: privacy-safe distributor of AI, calling the “god models in the cloud” only when needed. If ~115 IQ on-device is good enough, “I think that’s a bear case” — by far the most plausible and scariest bear case.
7. Data centers in space: “superior in every way”
- His flat claim: “the most important thing that’s going to happen in the world in the next 3 to 4 years is data centers in space” — with “really profound implications for everyone building a power plant or a data center on planet Earth.” From first principles a data center is power, cooling, and chips. In orbit the satellite sits in sun 24 hours a day at 30% higher intensity — 6x the irradiance — and needs no battery, killing a giant cost line. Cooling, a majority of rack mass on Earth, is free: “you just put a radiator on the dark side of the satellite… as close to absolute zero as you can get.”
- Networking improves too: “the only thing faster than a laser going through a fiber optic cable is a laser going through absolute vacuum” — a faster, more coherent cluster than on Earth. Training in space takes longer “just because it’s so big,” but for inference, Starlink’s demonstrated direct-to-cell means phone → satellite → answer, skipping the whole metro-fiber round trip. The friction is launch: “we need a lot of those Starships.”
- The convergence: Elon said yesterday Tesla, SpaceX and xAI are converging — xAI as Optimus’s intelligence module, Tesla Vision as its perception, SpaceX’s orbital data centers powering it all — “each one is kind of creating competitive advantage for the other.”
8. Governors, gluts, and the rolling bubble now in nuclear and quantum
- Will the iron law of gluts-follow-shortages apply? AI differs from software: every use consumes compute. Every lab could absorb 10x more (Mark Chen is on record) — the $200 tier gets better, the free tier becomes today’s $200 tier. Google monetizing AI Mode with ads “will give everyone else permission” to put ads in free tiers — plus agent commissions (“here are your three vacations, would you like me to book one?”).
- Two natural governors are holding: power, and TSMC’s caution. TSMC — “the guys who met with Sam Altman and laughed and said he’s a podcast bro” — is “in the process of making a mistake” by under-expanding, which will eventually fill Intel’s empty fabs (Lip-Bu Tan [likely] is reaping Patrick Gelsinger’s [likely] strategy; firing him was “shameful”). A third governor to watch: the first true DRAM cycle since the late ’90s, when a wafer was valued “like a 5-carat diamond” — prices rising “by X’s instead of percentages… that’s a whole different game.” His verdict: governors are good — “smoother and longer is good.”
- While watts bind, “the price you pay for compute is irrelevant” — 3-5x more tokens per watt is literally 3-5x more revenue (his example: a $50B data center pumping out $25B of revenue versus a $35B ASIC data center pumping out $8B), so the best products win irrespective of price with “crazy pricing power.” Nuclear can’t be built fast enough — “one ant” that could simply be relocated can delay a plant — so the answers are natural gas and solar (hence Abilene; Caterpillar just said it would increase capacity by 75% over the next few years).
- The rolling-bubble sequence since 2020 — non-Tesla EV startups (down 99%), meme stocks — has moved to nuclear/fusion/SMRs and quantum: transformative themes where “none of the public ways you can invest… are likely to succeed or have any real fundamental support.” The real quantum leaders in his view: Google, IBM, and Honeywell’s quantum unit; quantum supremacy just means some calculations classical computers can’t do — “it doesn’t mean that quantum takes over the world.” His eerie meta-observation: “for the last two years, whatever AI needs to keep growing and advancing, it gets” — nuclear opinion flipped overnight, space opens just as terrestrial power binds — Kevin Kelly’s technium made real.
9. SaaS is repeating brick-and-mortar’s e-commerce mistake
- The framing: application SaaS companies are making exactly the error brick-and-mortar retailers made — they saw customer demand for e-commerce but hated its margin structure, so they didn’t invest; now Amazon’s North American retail margins exceed many mass-market retailers’. “If there’s a fundamentally transformative new technology that customers are demanding, it’s always a mistake not to embrace it.”
- The mechanics: SaaS is write-once, distribute-cheap at 70-90% gross margins; AI recomputes the answer every time, so a good AI company runs ~40% gross margins — yet generates cash earlier than SaaS ever did “not because they have high gross margins, but because they have very few human employees.” An agent priced to protect 80% margins “is never going to succeed… If you are trying to preserve an 80% gross margin structure, you are guaranteeing that you will not succeed at AI. Absolute guarantee.”
- Investors have already proven they’ll tolerate this: the cloud transition — Adobe’s revenues and margins imploded moving off on-premise; Microsoft was a tough stock early on, then bought GitHub as the Copilot distribution channel and built “a giant business” at much lower gross margins. That’s why “this is a life-or-death decision, and essentially everyone except Microsoft is failing it — their platforms are burning,” a deliberate echo of the [likely Nokia] memo.
- The playbook he thinks an activist — “or constructive activist” — should force on Salesforce, ServiceNow, HubSpot, GitLab, Atlassian: report AI revenues and AI gross margins separately (“you know it’s real AI because it’s low gross margins”), even run them at zero for a while, backed by a cash-generative core the venture competitors lack. The alternative is grim: “right now another agent, made by someone else, is accessing your systems… pulling the data into their system, and you will eventually be turned off.”
10. “Investing is the search for hidden truths”
- Asked to pitch the career to Patrick’s kids: “investing is the search for truth — and if you find truth first, and you’re right about it being a truth, that’s how you generate alpha… you’re searching for hidden truths.” His formula: the most thorough knowledge possible of history intersected with the most accurate read of current events, to form a differential opinion in “the greatest game of skill and chance imaginable” — which stock is mispriced in the Perry Mutual system. It started with a sophomore investment-bank internship mailing research reports; by day three he was devouring them, then Peter Lynch, Buffett’s shareholder letters twice, and self-taught accounting — abandoning a plan-of-record to be a ski bum, river guide, and novelist.
- The confession that lands: picked last for every sports team, “a small fortune on private skiing lessons” without getting good, never beat a park chess player — “this is the only thing I’ve been vaguely competitive at. I’d love to be good at something else. I’m just not.” And cleaning toilets at Alta “permanently impacted how I treat other people.”