Pioneers Insight Method Research Author
Back to Pioneers
Jonathan Ross
Founders 6 Curated Dialogues

Jonathan Ross

Groq · Founder & CEO

Frontier Insights

Core Thesis: AI economics will be overwhelmingly dominated by inference (up to 95% of compute), where token deflation massively expands consumption. As models commoditize with zero switching costs, enduring value concentrates in infrastructure, distribution brands, and apps.

Strategy: Groq positions its HBM-free LPU architecture against Nvidia as an energy- and cost-efficient inference layer (5x cheaper, 1/3 power). Backed by sovereign capital like Aramco, Ross is aggressively scaling deployment to capture 50% of global inference capacity by 2027.

Risks: Gridlock from power and datacenter overbuilding, brutal custom-silicon lead times, geopolitical export leakage via cloud rental, and post-capex economic sustainability.

Key Views & Dialogues

Why Asking the Right Questions Is the Most Important Skill in the AI Age

  • 🗓️ Date2026-07-05 | 🎙️ Show:David Senra

Groq’s rumored $20B NVIDIA partnership moved from first call to money in the bank in about three weeks, after Jensen Huang saw its GPU-LPU integration and proposed making it available to all NVIDIA customers. The architecture treats GPUs and LPUs as complements for different bottlenecks, while faster inference can deepen model search and compound across agentic AI, leaving compute scarcity and whether capital still creates advantage worth monitoring.

View Dialogue Notes & Key Takeaways
  • The rumored $20B NVIDIA partnership went from first call to money in the bank in about three weeks. Jonathan Ross went to Jensen Huang asking to buy roughly 100,000 GPUs to deploy a hybrid GPU-LPU system Groq had already built; Jensen instead “thought maybe it would be better to make this available to all of their customers.” Ross corrects the desperation narrative: Groq wasn’t cash-strapped this time; the value of the agreement was only a little over 2x the prior valuation, and Groq could have raised at the amount of the licensing.

  • The technical thesis is that LPUs and GPUs are complements, not substitutes. “18-wheelers or vans for last-mile delivery—which one would you pick? The answer is both”: compute-constrained matrix multiplies go to the GPU, memory-throughput-constrained ones to the LPU, because “there is no one perfect architecture” and combining them “defeats the bottlenecks.” What many people get wrong is splitting prefill from generation across hardware—the wrong cut, since “the generation of tokens is the hard part.”

  • Speed makes models smarter, not just faster—Ross’s proof is AlphaGo on the TPU he created at Google. Ross gave approximate Elo figures of about 3,200 for AlphaGo on GPUs and about 3,550 for Lee Sedol, while warning he might have a leading digit wrong; the TPU result was described in the exchange as roughly 3,900 or 2,900. The same model found Move 37, a 1-in-10,000 move, only after the hardware enabled deeper search. “Being able to think faster makes you think smarter”—amplified by agentic AI-to-AI traffic, where speed compounds exponentially.

  • Fast inference was out of favor as recently as three or four years ago—including inside Groq itself. People were leaving while saying it added no value, customers asked “why do I need an LLM to be faster than I can read,” and Ross twice let his team talk him out of LLM opportunities—including a direct call from GitHub’s CEO wanting chips for code completion. What finally worked was letting people try it: a viral X video of an LLM running on Groq did what years of first-principles argument could not.

  • Capital is no longer the winning bet in AI, as Senra frames Ross’s VC story through the Keynesian beauty contest. VC herding was once rational because the most-funded startup could appear advantaged, but “for the first time in history, startups are not starved for cash… putting more money in is not an advantage. But people are still acting as if” it is. The kicker: typical West Coast VCs passed on Groq, while East Coast crossover funds invested in what Senra called NVIDIA’s biggest deal “by almost 3x.”

  • The defining AI-age skill is asking questions, not answering them. “Success in the information age was about being able to answer questions; success in the AI age will be about being able to ask the right questions.” Everyone shifts from ICs to “leaders of AI,” and Ross argues that school curricula should center on real community problems that students decompose into questions for AI.

  • Code’s marginal cost “is approaching zero,” opening software creation to many more founders. Ross’s EA now builds travel apps without knowing how to code; the literacy analogy says software creation is becoming accessible to people who previously lacked technical ability but may have good taste. He expects individual founders without large teams to create valuable companies.

  • Ross’s current manufactured discontent is the world’s compute scarcity. “If it takes us an extra year to cure cancer because we don’t have enough compute, that’s my fault.” The backstory: Groq was once three weeks from running out of money and preserved the team through salary-for-equity “Groq bonds,” in which 80% of employees participated and about half went down to the statutory minimum.

  • 🔗 Original source & video: Why Asking the Right Questions Is the Most Important Skill in the AI Age

Listen to full conversation →


The Inference Revolution: Groq, Nvidia and the Future of AI

  • 🗓️ Date2026-05-18 | 🎙️ Show:Sohn Conference Foundation

Memory’s pricing power may be self-limiting: Deep Seek’s V4 reportedly compressed likely KV cache by 90%, while the proposed response is to build more memory fabs as bottlenecks attract solutions. The unresolved intelligence debate matters for AI economics, with agentic systems and self-training potentially favoring the smartest model despite humanly invisible differences.

View Dialogue Notes & Key Takeaways
  • At the Irish Open conference, the Groq founder/CEO and Google TPU inventor — introduced as Nvidia’s chief software architect, having “been in the news a little bit lately” — sat down with hedge-fund manager John. His core framework: there is no permanent bottleneck in AI — “every time a bottleneck gets big enough, people solve it,” so components can charge a premium while they’re a problem but not a huge one; as they become bigger problems, people start solving them.

  • The direct application to today’s tightest trade: memory’s pricing power is self-limiting. Memory was “the most commoditized segment of the semiconductor supply chain” and is discussed through a Giffen-good analogy rather than a luxury good — you can raise the price of rice only until people switch to corn. Deep Seek’s V4 release was cited as an example of algorithmic efficiency, reportedly compressing what was likely KV cache by 90%: “it’s the tall poppy. As soon as it gets too tall, it gets chopped down.”

  • He rejects the Jevons-paradox defense of memory scarcity on opportunity-cost grounds: engineers working on that problem represented an opportunity cost — “if they had enough memory, they would have worked on something else.” His prescription is blunt: “start building more memory fabs.”

  • The pair’s long-standing disagreement is the thesis-relevant one: John argues intelligence has diminishing returns above PhD level and that open-source models are roughly 6 months behind closed-weight model labs, so the open-weight landscape catches up to the closed-weight landscape if his diminishing-returns argument holds. Jonathan’s rebuttal: “there’s no way to satiate the appetite for intelligence” — unsolved problems (cancer, aging), insufficient compute, and human competition mean even imperceptible model gaps show up “in my returns.”

  • Agentic AI hardens the moat for the smartest model: “AI likes to use AI” and will recognize and use the smarter model even when humans can’t tell the difference. Supporting anecdote: LLM-screening recruiters prefer resumes written by their own model — so write “one resume with Claude Opus 47 and one with ChatGPT.”

  • His case against a capability plateau underwrites the build-out: models now generate, prune, and retrain on their own data, “improving at a pretty linear rate,” and AI’s intuition already exceeds ours — Waymo’s daily data, he said, was starting to approach a human lifetime of driving data, though he didn’t know whether it had reached that amount. Bonus definition worth keeping: intelligence is a stationary capability—the ability to make a prediction or influence an outcome; sentience is “your rate of improvement in your intelligence” — “a property of a civilization,” and an accelerating feedback loop in society.

  • 🔗 Original source & video: The Inference Revolution: Groq, Nvidia and the Future of AI

Listen to full conversation →


Groq founder and TPU creator Jonathan Ross GPU ♥ LPU Everything You Wanted to Know Nvidia GTC 2026

  • 🗓️ Date2026-03-26 | 🎙️ Show:NoRush Invest

Nvidia-Groq’s LPX rack is already in production for Q3 availability, splitting inference between LPU FFN layers and GPU attention layers to improve utilization, latency, and throughput per megawatt. Ross expects fast-tier pricing to remain super-linear because speed increases iteration capacity, while inference revenue scales with users; energy supply, networking, and whether customers can support ultra-tier token spending remain the watchpoints.

View Dialogue Notes & Key Takeaways
  • The Nvidia-Groq integration is real, broad, and shipping: Ross confirms the LPX rack is “already in production” with Q3 availability — “probably one of the fastest ramps of a semiconductor in history” (legal inserted “probably”). The deal went from Groq COO Sunny Madra reaching out to Jensen about NVLink access to a deal done in weeks, with Ross working at Nvidia full-time from December 25th.

  • The product splits the LLM decode workload: FFN layers run on LPUs, attention layers on GPUs — Ross’s analogy is a logistics network of 18-wheelers (GPU long-haul throughput) and delivery vans (LPU low-latency). Result: utilization rises on both chips and the Pareto curve “bends up” at high speed, reaching “thousands of tokens per second” that are “otherwise impossible to get.”

  • Ross expects fast-tier pricing to stay super-linear: Anthropic and likely OpenAI already charge more-than-linear for fast tiers and Ross expects that to continue — his coal-vs-oil frame: oil costs ~7x per BTU and people pay it because “you cannot fly a plane with coal.” Enterprises should give ultra-tier tokens to their best engineers; today’s “$10K/month” spenders “might be on the low end,” heading toward engineers using millions of dollars of tokens per year.

  • “Speed is intelligence” — and can beat model quality: internally Groq found Qwen-32B, “a good model not a great model,” solved every formal-reasoning math problem faster and cheaper than Anthropic’s Opus by simply iterating more — fast feedback loops beat fewer, expensive iterations.

  • Inference drives the revenue cycle: “training scales with the number of researchers you have and inference scales with the number of users you have. Revenue comes from users, not researchers” — a virtuous cycle where inference leadership drives customer revenue drives more training hardware purchases.

  • Energy limits inference scale: “The world right now doesn’t have enough inference compute. It doesn’t have enough energy to power enough inference compute” — the hybrid rack’s pitch is more tokens per megawatt, framed around Jensen’s gigawatt-datacenter monetization question.

  • Culture read on NVDA from the inside: Nvidia, with over 40,000 people, moves “just as fast, no slower” than 450-person Groq — “no real bureaucracy” — with a pointed on-stage dig at Google’s pace.

  • 🔗 Original source & video: Groq founder and TPU creator Jonathan Ross GPU ♥ LPU Everything You Wanted to Know Nvidia GTC 2026

Listen to full conversation →


Groq Founder, Jonathan Ross: OpenAI & Anthropic Will Build Their Own Chips & Will NVIDIA Hit $10TRN

  • 🗓️ Date2025-09-29 | 🎙️ Show:20VC

Jonathan Ross says OpenAI, Anthropic and every hyperscaler will build chips for control over their own destiny, while Nvidia could be worth $10 trillion in five years. Doubling inference compute could nearly double their revenue within one month as rate limits suppress engagement, making HBM, energy, Groq’s six-month supply chain and the United States’ 2-3 year away-game window key signals.

View Dialogue Notes & Key Takeaways
  • Ross’s headline call: “I personally would be surprised if in 5 years Nvidia wasn’t worth 10 trillion” — but the sharper version is that Nvidia keeps a majority of revenue on a minority of units, “maybe 51% of the revenue and 10% of the chips sold,” as the 35-36 customers accounting for 90%-99% of total spend gain the power to buy on merit rather than brand. And the question he’d rather you ask: “Will Groq be worth 10 trillion in 5 years? Possible.”

  • The central compute wager: “If OpenAI were given twice the inference compute they have today, if Anthropic was given twice… within one month their revenue would almost double.” Anthropic’s one of its biggest complaints is rate limits; OpenAI throttles chat speed and loses engagement. Two weeks ago a customer asked Groq for 5x its total capacity — no one could serve it.

  • Everyone — OpenAI, Anthropic, every hyperscaler — will build their own chips, not necessarily to beat Nvidia: the prize is “control over your own destiny” against Nvidia’s monopsony on HBM (Nvidia could fab 50M GPU die a year but ships ~5.5M GPUs). For new chip startups, though, “it’s too late — that ship has already sailed”: 3 years design-to-silicon if perfect, and only 14% of first tape-outs work.

  • On the bubble question: ask what the smart money is doing — all doubling down, because at a Goldman Abu Dhabi event zero of 50+ managers of $10B+ AUM were sure AI couldn’t do their job in a decade. “Of course they’re going to be spending like drunken sailors” — and AI breaks SaaS math anyway, since spending more compute per query directly improves the product.

  • Misconception busted: Chinese models are ~10x more expensive to run than GPT OSS — they were optimized for cheap training, and captive-market pricing confused price with cost. China can win its home game (150 nuclear reactors, subsidies), but the US has a 2-3 year window in the “away game” for allied nations that can’t build power plants.

  • “The countries that control compute will control AI, and you cannot have compute without energy.” Europe competes “through legislation” instead of building; if it doesn’t act at speed, “Europe’s economy is going to be a tourist economy” — Norway, with wind power at 5x its hydro power, could provide as much energy as the US, and Japan (2nm fab, $65B for AI, reactors restarting) shows the tempo required.

  • The labor counternarrative: AI means massive labor shortages, not unemployment — deflationary pressure across everything, people opting out of work earlier, and jobs uncontemplatable today. “We’re going to be able to add more labor to the economy by producing more compute… That has never happened in the history of the economy before.”

  • Landscape calls: the enduring duel is OpenAI vs Google (Anthropic “does something different” — coding); OpenAI at $500B and Anthropic at $180B are both “highly undervalued” because labs become “the Mag 9, the Mag 11, the Mag 20”; Microsoft resets short-term on the OpenAI relationship; CUDA lock-in “is true for training, but not true for inference.”

  • 🔗 Original source & video: Groq Founder, Jonathan Ross: OpenAI & Anthropic Will Build Their Own Chips & Will NVIDIA Hit $10TRN

Listen to full conversation →


Jonathan Ross, Founder & CEO @ Groq: NVIDIA vs Groq - The Future of Training vs Inference | E1260

  • 🗓️ Date2025-02-17 | 🎙️ Show:20VC

Groq’s differentiated inference wedge avoids scarce HBM, supporting claimed 5x lower cost and one-third the energy per token while Nvidia concentrates on training. Aramco-backed capex and $1.5B in revenue—not fundraising—support a target of at least half of global inference compute by end-2027, but a power bottleneck in three to four years and Nvidia’s uncertain trajectory remain risks.

View Dialogue Notes & Key Takeaways
  • Groq’s headline number is being misread: “we did not raise 1.5 billion — that’s revenue,” which Jonathan Ross sizes at “about 30% of the revenue of OpenAI.” The Aramco/Saudi deal puts up the capex, Groq deploys and repays to an agreed IRR before the split flips — “we are not limited by capital anymore.” The stated goal: at least half the world’s AI inference compute by end of 2027, run on the creed that “when you are growing faster than exponential there is no amount of profit that you can make that matters.”

  • Ross refuses the Nvidia-rivalry frame and inverts it: training is “a solved problem” he happily cedes — he tells his own customers to “buy every single GPU you can get your hands on” — while Groq takes low-margin (~20% upfront), high-volume inference off Nvidia’s 70-80%-margin hands. Claimed economics: more than 5x lower cost, one-third the energy per token, and “just the memory alone in the latest GPUs costs more than our fully loaded capex per chip deployed.”

  • The structural wedge: Nvidia is a monopsony for HBM — only SK Hynix, Samsung and Micron make it, and it’s the hard-to-ramp part of a GPU. Groq’s LPUs skip HBM entirely and run on the same silicon process as phone chips, so “we effectively have almost no limit on how much we can scale up”: 640 chips in production at the start of 2024, 40,000 by year-end, 2M+ targeted this year — needing nearly all of its fab’s capacity.

  • His biggest macro worry is a power crunch in three to four years: ~20GW of potential power against ~15GW of existing data centers worldwide, overbuilding now; when burned builders retrench and chip counts keep doubling every 18-24 months toward 120-240GW of chip-driven demand, “that power will become a hard bottleneck.” Meanwhile “fake data centers” — real-estate-focused builders with no generators or water — inflate apparent supply.

  • On scaling laws: the apparent asymptote is an artifact of assuming all data is equal quality — an LLM generating and pruning its own synthetic data climbs beyond that apparent asymptote, compounding with test-time reasoning into “geometrically increasing improvement.” Compute is only a “soft bottleneck,” DeepSeek was an algorithmic tweak, and Nvidia’s 15% post-DeepSeek drop was “the popularity contest side of the market… nothing to do with the weighing machine.”

  • Bubble verdict, both halves intact: “I can guarantee you that a huge amount of money will be incinerated, but I also bet that in total more money will be made than will be put in.” The novelty this cycle is a Keynesian beauty contest “gone completely amok” — multiple rivals each holding billions, so capital no longer anoints winners: “the people who have the best products are actually going to be the winners because everyone can be capitalized.”

  • The positioning sermon, earned over seven years without product-market fit: “your job is not to follow the wave, your job is to get positioned for the wave.” Groq survived on “Groq bonds” — 80% of employees traded salary for equity, half down to the statutory minimum — and his map of the next waves names four era-defining companies: whoever solves hallucination, agentic sub-goals, “invention,” and decision proxying.

  • The China tell isn’t Blackwell access — cloud rentals and Malaysia/Singapore “wink wink” GPU deployments may make physical access less decisive — but censorship: if China isn’t permissive of “more open, truthful models,” it’s inherently disadvantaged. And on Nvidia’s next decade he’s genuinely split: “I wouldn’t be surprised if they were 3x bigger; I also wouldn’t be surprised if they stayed around the same.”

  • 🔗 Original source & video: Jonathan Ross, Founder & CEO @ Groq: NVIDIA vs Groq - The Future of Training vs Inference | E1260

Listen to full conversation →


Jonathan Ross: DeepSeek Special - How Should OpenAI and the US Government Respond | E1253

  • 🗓️ Date2025-01-30 | 🎙️ Show:20VC

DeepSeek’s real breakthrough was fully automated verifiable-reward RL, while its $6M training figure excludes substantial distillation or scraping costs. Ross argues LLMs are now commoditized with no switching cost, pushing OpenAI toward open source while inference—not training—could reach 95% of compute spending. That expands the Nvidia opportunity through Jevons paradox, but CCP data access and cloud-based GPU workarounds leave export controls unresolved.

View Dialogue Notes & Key Takeaways
  • Ross’s headline verdict: DeepSeek is “Sputnik 2.0” — the NASA-space-pen-versus-Russian-pencil story “just happened again” — but the $6M figure is marketing. “It is true that they spent about $6 million or whatever on the training — they spent a lot more distilling or scraping the OpenAI model,” and since OpenAI reportedly loses money on every API token, it was “effectively subsidizing, accidentally, the training of this model.” The genuine innovation was fully automated verifiable-reward RL, no humans in the loop.

  • Models are now nakedly commoditized — “if there was any doubt before, that doubt’s over” — and LLMs have “no switching cost whatsoever,” which kills the cloud analogy. Ross’s move if he were Sam Altman: gear up to open-source OpenAI’s models in response — “open always wins, always,” Linux proved it — because “it’s pretty clear you’re going to lose that, so you might as well try and win all the users and the love,” then fall back on OpenAI’s real Seven-Powers moat: brand.

  • Stargate’s $500B is “not enough spending,” not too much: Google spent 10–20x more on inference than training in Ross’s TPU days, Harry thinks Jenson said half of Nvidia’s revenue is already inference, and Ross thinks inference could reach 95% — “you don’t train to become a cardiovascular surgeon and then perform for 5% of your life.” Test-time compute compounds it: one DeepSeek answer burned 18,000 intermediate tokens.

  • The tradeable call: Harry “just bought a shitload of Nvidia” on the 16% dump — “the most screaming buy of the century” — and Ross agrees on the weighing-machine view: Nvidia is “actually more valuable thanks to DeepSeek, not less.” Jevons paradox: compute cost drops ~1,000x a decade and consumption rises ~100,000x, so spend rises 100x. Training is the high-margin “mainframe” niche; inference is the larger market, and Groq taking low-margin volume is “probably the best thing that’s ever happened for Nvidia stock.”

  • The CCP risk is the data, and it’s about to get worse: “right now the CCP is probably going to be taking the safeties off the weapons… now we want the data” — Harry puts it at 100% that Beijing treats DeepSeek as another TikTok. Meanwhile export controls are theater: “you can literally log in, swipe a credit card, and rent GPUs” — “it’s like the Maginot line, you just go around it.”

  • What everyone copies next: DeepSeek’s very sparse mixture-of-experts (~671B parameters, ~250 experts of ~2B each, only a handful active) plus synthetic-data retraining — Llama 3.3 70B already beat 3.1 405B via fine-tuning on better data. And DeepSeek restricting signups to Chinese phone numbers means they ran out of inference compute: “training scales with the number of ML researchers you have; inference scales with the number of end users you have.”

  • For the foundation-model complex: “pivot, get over it — just pivot.” The model is the engine, not the car; Perplexity is “perfectly positioned” for the moment hallucination rates drop (building on today’s models is like “trying to create Uber before we had smartphones”); and Europe’s prescription is 100 Station Fs by year-end, a thousand by next year.

  • 🔗 Original source & video: Jonathan Ross: DeepSeek Special - How Should OpenAI and the US Government Respond | E1253

Listen to full conversation →