Pioneers Insight Method Research Author
Groq Founder, Jonathan Ross: OpenAI & Anthropic Will Build Their Own Chips & Will NVIDIA Hit $10TRN
Back to Episodes

Groq Founder, Jonathan Ross: OpenAI & Anthropic Will Build Their Own Chips & Will NVIDIA Hit $10TRN

Summary

  • Ross’s headline call: “I personally would be surprised if in 5 years Nvidia wasn’t worth 10 trillion” — but the sharper version is that Nvidia keeps a majority of revenue on a minority of units, “maybe 51% of the revenue and 10% of the chips sold,” as the 35-36 customers accounting for 90%-99% of total spend gain the power to buy on merit rather than brand. And the question he’d rather you ask: “Will Groq be worth 10 trillion in 5 years? Possible.”
  • The central compute wager: “If OpenAI were given twice the inference compute they have today, if Anthropic was given twice… within one month their revenue would almost double.” Anthropic’s one of its biggest complaints is rate limits; OpenAI throttles chat speed and loses engagement. Two weeks ago a customer asked Groq for 5x its total capacity — no one could serve it.
  • Everyone — OpenAI, Anthropic, every hyperscaler — will build their own chips, not necessarily to beat Nvidia: the prize is “control over your own destiny” against Nvidia’s monopsony on HBM (Nvidia could fab 50M GPU die a year but ships ~5.5M GPUs). For new chip startups, though, “it’s too late — that ship has already sailed”: 3 years design-to-silicon if perfect, and only 14% of first tape-outs work.
  • On the bubble question: ask what the smart money is doing — all doubling down, because at a Goldman Abu Dhabi event zero of 50+ managers of $10B+ AUM were sure AI couldn’t do their job in a decade. “Of course they’re going to be spending like drunken sailors” — and AI breaks SaaS math anyway, since spending more compute per query directly improves the product.
  • Misconception busted: Chinese models are ~10x more expensive to run than GPT OSS — they were optimized for cheap training, and captive-market pricing confused price with cost. China can win its home game (150 nuclear reactors, subsidies), but the US has a 2-3 year window in the “away game” for allied nations that can’t build power plants.
  • “The countries that control compute will control AI, and you cannot have compute without energy.” Europe competes “through legislation” instead of building; if it doesn’t act at speed, “Europe’s economy is going to be a tourist economy” — Norway, with wind power at 5x its hydro power, could provide as much energy as the US, and Japan (2nm fab, $65B for AI, reactors restarting) shows the tempo required.
  • The labor counternarrative: AI means massive labor shortages, not unemployment — deflationary pressure across everything, people opting out of work earlier, and jobs uncontemplatable today. “We’re going to be able to add more labor to the economy by producing more compute… That has never happened in the history of the economy before.”
  • Landscape calls: the enduring duel is OpenAI vs Google (Anthropic “does something different” — coding); OpenAI at $500B and Anthropic at $180B are both “highly undervalued” because labs become “the Mag 9, the Mag 11, the Mag 20”; Microsoft resets short-term on the OpenAI relationship; CUDA lock-in “is true for training, but not true for inference.”

Deep dive

1. Don’t ask “is it a bubble” — ask what the smart money is doing

  • Ross’s reframe: if a question keeps not getting an answer, ask a different one — what are Google, Microsoft, Amazon, and nations doing? All doubling down, spending more with every announcement. His best evidence: Microsoft deployed GPUs in one quarter and then withheld them from Azure “because they made more money using them themselves than renting them out.”
  • His market analogy: the early days of oil drilling — “a lot of dry holes and a couple of gushers,” with ~35-36 companies responsible for 99% of revenue, or at least the token spend. Lumpiness means instinct still beats science, “and right now is the best time for investors” — once it becomes predictable, the good investors make less.
  • The spend isn’t purely economic. At Goldman’s inaugural Abu Dhabi event he asked 50+ people managing $10B+ AUM who was 100% convinced AI won’t do their job in ten years: “No hands went up. I’m like, great. That’s how the hyperscalers feel. So of course they’re going to be spending like drunken sailors, because the alternative is that they’re completely locked out of their business.”
  • Stebbings pushes: returns must eventually materialize, Mag 7 or not. Ross concedes (“that’s correct”) but insists value is landing now: a customer asked for a feature, he “was prompt engineering the engineers,” and it shipped to production four hours later with no human-written code. Six months out it could happen before the meeting ends — “a qualitative difference,” winning deals competitors can’t.

2. The insatiable-compute thesis: double the inference, double the revenue

  • The episode’s central wager: “If OpenAI were given twice the inference compute that they have today, if Anthropic was given twice… within one month from now their revenue would almost double.” Mechanism: Anthropic’s one of the biggest complaints is rate limits; OpenAI regulates its chat service by running it slower, sacrificing engagement.
  • Groq’s own funnel proves the constraint: 100% of customers arrive asking for speed, and none keep asking once they see the shortage — “the real value prop is can you provide more compute capacity.” Two weeks ago a customer wanted 5x Groq’s total capacity; neither Groq nor any hyperscaler could take them. (Harry’s instinct — “should I just buy the shit out of CoreWeave?” — gets the response that CoreWeave is a great company but has a finite GPU allocation like everyone else.)
  • Ten years of data-center forecasting has erred one way: everyone builds too little, raises projections above their most optimistic case, and still builds too little. Compute is “the easiest knob” — algorithms rarely improve, data is hard to get, but compute gets better every year and a big enough check reliably buys it. “And yet we still underestimate how much we need.”
  • The inference surprise he admits to: “What I never expected was that AI was going to be based on language… I thought it was going to be more like AlphaGo.” Language made AI trivial to use — “I expected AI to come sooner and grow slower. It came later and it’s growing faster than I ever imagined.” 10% of the world’s population is described as a “GBT” weekly active user.

3. Speed compounds into brand — the latency-tolerant future is “100% wrong”

  • To the claim that users will happily fire off prompts and walk away: “100% wrong.” His evidence is CPG ranked by margin — smoking tobacco, then chewing tobacco, then soft drinks, down to water — where margin tracks the speed at which the ingredient acts on you. The dopamine cycle builds brand affinity; Google and Facebook were built on it, with “every 100 milliseconds of speedup” worth “about an 8% conversion rate.”
  • When early viewers of Groq’s speed demos asked why output needed to be faster than you can read: “Why does a web page need to load faster than you can read?” People, he says, are consistently bad at predicting what drives engagement — a lesson from the early internet.

4. Everyone will build chips — for destiny, not necessarily superiority

  • The TPU is seen as proof hyperscaler silicon works, but it was one of roughly three concurrent chip efforts at Google and the only one that beat GPUs; Dojo just got cancelled. Building a chip to take on Nvidia “is a little bit like saying Google Search is pretty nice, let’s go replicate it. It’s insane.”
  • Still, he has “no doubt” OpenAI, Anthropic, and every hyperscaler will build chips — but read the motivation. His Google lab-tour story: 10,000 AMD servers built, then the chips popped off and thrown in a trash can — pre-ordained, because the point was extracting a discount on the Intel chips Google actually bought.
  • The binding constraint is Nvidia’s monopsony on HBM: the GPU die uses phone-chip processes and Nvidia could build 50 million a year, but HBM and interposer capacity cap it at ~5.5M GPUs this year. When a hyperscaler asks for a million GPUs, gets refused, and threatens to build its own — “all of a sudden those GPUs are found.” Your own chip buys “control over your own destiny,” even if it costs more.
  • Sarah Hooker’s “hardware lottery” is the trap for entrants: models are designed for incumbent hardware — attention works well on GPUs, so a better architecture “is not going to run well,” so it isn’t better. Incumbents can plan two years out; a challenger needs a faster loop.

5. The $100B “infinite money loop” and Nvidia at $10 trillion

  • On Nvidia’s $100B into OpenAI: “It’s not round-tripping if actual productive outcomes are occurring” — at least 40% flows out to suppliers building infrastructure. Harry: so a partial loop, 60% back to Nvidia plus a couple-hundred-billion stock bump. Ross: economically, “why not do that all day long” — the premium is justified only if lock-in holds, “and I would actually say with Nvidia that’s probably true,” because there simply isn’t enough compute in the world.
  • The five-year map: Nvidia keeps over 50% of revenue on perhaps 10% of chips sold. Brand pays — “no one’s going to get fired for buying from Nvidia” — but the 35-36 customers accounting for 90%-99% of total spend gain the power to choose on business merit, so other chips get used.
  • Over/under $10T in five years: “I personally would be surprised if in 5 years Nvidia wasn’t worth 10 trillion. The question you should ask is will Groq be worth 10 trillion in 5 years? Possible.” Groq’s claim: with fewer supply-chain constraints, “we can produce nearly unlimited quantities” of the scarcest asset in the market.
  • Quickfire on Nvidia’s biggest misconception: “That Nvidia’s software is a moat.” CUDA lock-in is “true for training, but not true for inference” — Groq has 2.2M signed-up developers against CUDA’s claimed 6M.

6. Chip economics: two hurdles to clear — and the ship has sailed for new entrants

  • A chip’s life has two calculations: deployment must beat capex, staying in service only has to beat opex. The industry-wide bet is that new chips won’t push old-chip value below operating cost — which is why near-5-year-old H100s still rent profitably: “You would never deploy an H100 today, but they’re still profitable to run.” Only compute scarcity keeps that true.
  • Groq refuses the long bet: Ross says people think about amortization over a longer period than he would; he says “5 to 6 years,” then says “a little bit less” after Harry says “which would be like 3 years.” Groq upgrades chips roughly annually, because at five years old chips fall below their electricity and data-center cost. Long-contract holders face a third calculation — is breaking the contract cheaper than running at a loss? What happens then? “I can’t tell you, because we’re trying to avoid that situation.”
  • Founding Groq today? “I wouldn’t do chips. That ship has already sailed.” Perfect execution means three years from design to production, and only 14% of first tape-outs work — Ross has done three chips, all working first-silicon, and even scheduled a respin for V2 that turned out unnecessary “to our shock.” Groq is now on a one-year cadence (V2→V3→V4) versus Nvidia’s pipelined 3-4 years per chip.
  • The actual moat is the supply chain: unlike the HBM-constrained GPU supply chain, “you write us a check for a million LPUs and the first of those LPUs starts showing up 6 months later” — versus two years for GPUs. Pitching a hyperscaler’s head of infrastructure, speed and cost got polite nods; the six-month supply chain “was the only thing he cared about.”

7. Busting the China misconception — and the home game vs the away game

  • “Let’s just start busting every misconception.” Chinese models are not cheaper to run — they’re “about 10x as expensive” as the GPT OSS model. They were optimized to be cheap to train, and captive-market pricing let providers charge for scarcity: “people were confusing the cost with the price.” The intelligence squeezed into OSS shows “the US still has a training advantage,” built on chip access. Ross admits “even I was snookered a little bit at first.”
  • China can win its home game — 150 nuclear reactors planned, subsidies covering inefficient chips. The away game is different: an ally with 100 megawatts of grid can’t just build a reactor, so the more energy-efficient chip wins. “For the next 2 to 3 years, the United States has a clear advantage in that away game” — if it moves fast enough to bring allies in.
  • The open-source play: he’d predicted OpenAI’s open-source release on brand strength alone (“they could probably use Llama 2… and people would still use it”). Anthropic “should be open sourcing their previous generation” — because prompt compatibility is the new software compatibility: people adopted OpenAI’s OSS model over Chinese ones because their prompts were reusable, and low-cost users graduate to the premium model later.
  • GPT-5’s efficiency focus isn’t a retreat from scaling: to win India you need 99 rupees a month (~$1.13), against an alternative of no AI at all.

8. “Countries that control compute will control AI” — Europe’s tourist-economy warning

  • The line he wants everyone to leave with: “The countries that control compute will control AI. And you cannot have compute without energy.” And it’s not nuclear-or-nothing: Norway has ~80% wind utilization, and with wind power at 5x its hydro power, “Norway itself could provide as much energy as the United States” — consistently, from one country.
  • Europe’s failure mode is embracing “mistakes of omission” and competing “through legislation” — data-residency rules instead of megawatts — while “missing out is more expensive than fumbling something” in a growth economy. The permitting rot is American too: a nuclear-company board member told him they spend three times as much on permitting as on the power plant. Why no nuclear? “Fear.” (Trump, on this axis, is “definitely help.”)
  • Japan is his tempo benchmark: slow to decide, fast to move — a 2nm fab already producing wafers (yield not production-grade yet), $65B allocated to AI, reactors coming back online. “When Japan is going to turn their nuclear reactors back on, Europe needs to listen to that.”
  • If Europe doesn’t move at speed: “Europe’s economy is going to be a tourist economy. People are going to come here to see the quaint old buildings and that’s going to be it.” Model sovereignty won’t save it — “you could have a model that is 10 times smarter than OpenAI’s, and if [OpenAI] has 10 times the compute, OpenAI’s model is going to be better.” He partners with Mistral and likes it; the fix is compute so likely Mistral can compete. Interim hack: Saudi “data embassies” — sovereign data oversight riding their 3-4GW build-out.

9. AI breaks SaaS economics — and creates labor shortages, not unemployment

  • AI ≠ SaaS because quality isn’t fixed at ship time: “I can improve the quality of my product by running two instances of the prompt and picking the better answer.” That’s why token-as-a-service bills nearly match revenue, and why OpenAI just announced compute-heavy products for limited users at higher prices — “we want to see what happens when we give more compute to the AI.”
  • On margins, his heterodoxy: high margin exists only to buffer volatility, and “your margin is my opportunity.” He wants Groq’s margin “as low as I possibly can make it while keeping my business stable” — because “trust pays interest” and Jevan’s paradox does the rest: “If we produce 10x the compute, we will have 10x the sales.”
  • The macro chain: “The most valuable thing in the economy is labor. And now we’re going to be able to add more labor to the economy by producing more compute and better AI. That has never happened in the history of the economy before.” Three consequences: massive deflationary pressure (robot-farmed, genetically engineered coffee, cheaper housing), people opting out — fewer hours, earlier retirement — and jobs uncontemplatable today (agriculture went 98%→2% of the US workforce; “influencers wouldn’t have made sense 100 years ago”). Net: “massive labor shortages,” not mass unemployment.
  • On S&P-near-7,000 toppiness: separate the weighing machine from the popularity contest. He’s never bought a Bitcoin — “I can’t play in the popularity contest.” The tell that AI is weighing-machine value: PE firms are all over Groq for cheap compute to change portfolio-company bottom lines. A downturn is unpredictable in principle (“if a prediction affects the prediction, you cannot predict it”); his overheating test is whether the economy impedes company success — and the one distortion he names is every good engineer raising “10, 20, 100 million, a billion” to start yet another lab instead of joining one. “Please stop doing that.”

10. The landscape: OpenAI vs Google, both labs “highly undervalued,” differentiate or die

  • The enduring duel is OpenAI and Google — “Anthropic does something different”: coding. Google is his most impressive turnaround (an engineer-driven culture is “a systemic advantage”) and Gemini a success on adoption numbers, though consumer integration is “practically unusable” in Gmail — darts worth taking, Chrome-from-Google-TV style, but with 10% of the world’s population described as a “GBT” weekly active user, he concedes “Google may be too late.”
  • OpenAI at $500B or Anthropic at $180B? “I’d want to invest in both… They’re both undervalued. Highly undervalued.” The mistake is treating it as a finite market: labs are expanding it with R&D and will simply join the index — “the Mag 9, the Mag 11, the Mag 20.” Among incumbents: Microsoft resets short-term on the OpenAI relationship but keeps the deployed compute (“compute is like gold”); Amazon “doesn’t have AI DNA” but has compute; Meta and Google always had the DNA.
  • On Elon: xAI can probably pull it off, but differently — it has a coding model with no coding distribution, and Anthropic’s brilliance was refusing to do everything. “If you do not differentiate, you die.”
  • Groq’s own line in the sand: it will not create its own models, so customers can trust building on it — “I could be making a huge mistake on that call.” It raised $750M at almost $7B (planned $300M, 4x oversubscribed) with positive hardware margins. Internally, engineers must use AI but choose their tools — Sourcegraph, then Anthropic, now Codex, “next month it’ll probably be Sourcegraph again.” His closer, the Galileo analogy: LLMs are “the telescope of the mind” — they make us feel small now, but “in a hundred years, we’re going to realize that intelligence is more vast than we could ever have imagined and we’re going to think that’s beautiful.”