Pioneers Insight Method Research Author
Cerebras CEO on the Future of Data Centres, Token Costs & Memory | Should US Companies Sell to China
Back to Episodes

Cerebras CEO on the Future of Data Centres, Token Costs & Memory | Should US Companies Sell to China

Summary

  • Feldman’s core anti-bubble argument inverts the historical analogy: rail in the 1880s and fiber in the late ’90s were “if we built it, they would come” — infrastructure ahead of demand — while AI is the exact opposite. “We can’t build data centers fast enough to keep up with demand”: Cerebras carries a $25B backlog, and Nvidia, AMD and others have their own. “We’re not building ahead of demand. We’re building behind demand.”
  • The demand inflection has a date: “somewhere in 2025 the models got smart enough to be really useful” — before that, AI “was like cool and then nobody used it.” Now it’s sweeping demographics from his 85-year-old father to his 11-year-old niece, and demand does not peak if AI continues to improve in usefulness — the one bear case he concedes against the “electricity into intelligence” thesis.
  • On the compute deal involving Elon: “They bought down rev gear” — H100s, not B200s, “a generation and a half, maybe two generations behind.” “This was not a great deal. It was a good deal for Elon” — forced action in an exponential, versus Sam Altman’s superpower of believing exponential demand data years out and contracting for power, data centers, and hardware ahead of everyone.
  • Memory is the No. 2 shortage after fab space, and it’s structural: HBM comes only from Samsung, Micron and Hynix, a new fab costs $40B and takes 5 years, so “we’re going to continue to see memory shortages for at least the next several years” if demand holds. Micron is printing 80–85% gross margins — “software gross margins on making memory.” Cerebras sidesteps it entirely: SRAM etched by TSMC, no HBM, no CoWoS, 5nm while 3nm is the oversubscribed node.
  • The speed thesis is absolute: “How big is the market for slow search? It’s zero”— you wouldn’t take $1,000/month to keep slow internet, so there will be zero market for slow inference. Cerebras posted Kimi K2.6 running 6.7x faster than the next-fastest GPU cloud “while one bozo at an analyst firm was on TV saying we couldn’t do it,” and is now digesting a $20B+ OpenAI deal — same concentration critique investors gave him at $1B with G42, one customer-size rung earlier.
  • Nvidia has “funded and backstopped and overallocated to the neo clouds” to create hyperscaler competitors — “a dependence, which is probably not healthy.” Neo clouds buy hardware carrying Nvidia’s 70–80% gross margin before adding their own; Google and Cerebras putting their own silicon in their own data centers don’t, though Google’s full-stack edge is capped by having only one TPU customer: itself.
  • On jobs: “most of the layoffs were AI-washed” — 90–95% of terminations are COVID over-hiring and old productivity gains finally harvested, “none of this is AI yet.” Meanwhile the real enterprise adoption blocker isn’t data cleanliness but lawyers and security (“no credit, no credit, failure, blame”), and Feldman sees no problem with software engineers using $50–100K/year in tokens — hardware engineers’ EDA tools are already closer to 15–20% of salary. At 47M software engineers, that is “$5 trillion just in software engineering token use.”
  • On China, arguing against his own book: leading-edge chips sold to China will be used by its military and its state-backed industry — “there is no debate on that point” — so don’t sell; “I’d like to keep my industrial adversaries more than down rev.” His one free policy wish: give TSMC and Samsung 20 years free of all local ordinances to build US fabs — “fabs are modern pyramids.” The IPO itself was unblocked when “we got a new government” and likely-CFIUS concerns over large customers “disappeared” — the largest semiconductor IPO ever, $185 to $311.

Deep dive

1. This is not a bubble — the build-out is behind demand

  • Feldman lived through the late-’90s fiber-optic build-out and rejects the comparison, along with the 1880s rail analogy economists reach for: those bubbles shared “a pension to believe that if we built it, they would come” — infrastructure way ahead of demand. AI is “in a strange way… the exact opposite”: Cerebras has a $25 billion backlog, Nvidia and AMD have backlogs, “because we can’t get data centers built fast enough.”
  • His test for bubble-ness: “when you are trying with your infrastructure to keep up with what people want today, not in the future” — and demands are still growing — that’s not a bubble characteristic. This is how he reconciles bubble talk with Jensen’s $3–4 trillion of AI infrastructure spend by 2030.
  • On Gavin Baker’s point that permitting delays helpfully throttle the market, Feldman agrees via two analogies: his first Vegas buffet (“you eat so much you feel sick for days”) and freeway meters — “metering makes the freeway traffic smoother” and avoids hiccups.

2. 2025 is when the models got useful — and believing exponentials is the superpower

  • The unremarked inflection: “somewhere in 2025 the models got smart enough to be really useful. Before that… these were sort of a novelty. AI was like cool and then nobody used it.” Now usage “is sweeping through demographic groups” — his 85-year-old father, his 11-year-old niece — on more and harder problems. Asked if demand peaks: “not if AI continues to improve in usefulness.”
  • Sam Altman’s brilliance, in Feldman’s telling: he saw the exponential, wasn’t afraid of it, and contracted for power, data centers and hardware while others’ minds hurt. “An ability to believe your data in an exponential growth environment out a year or two or three is a superpower.” Sam and maybe Elon can think at 100 or 500 gigawatts, “where everybody else’s brain shuts down.”
  • The counterexample — the compute deal involving Elon: “They bought down rev gear… They got H100s. They didn’t get the B200s… a generation and a half, maybe two generations behind. This was not a great deal. It was a good deal for Elon. He had them sitting around.” Forced action gets you the deal that’s available, not the one you want.

3. Memory shortages last years — and Cerebras doesn’t pay the toll

  • Memory is the No. 2 supply-chain pinch after TSMC fab space. HBM comes from exactly three makers — Samsung, Micron, Hynix — and they couldn’t keep up, so prices “shot through the roof”: Micron is producing 80–85% gross margins — “they’re getting software gross margins on making memory.”
  • Why it persists: capacity is lumpy. “You have to build a fab for $40 billion, and it takes 5 years… It’s a step function.” So if demand stays high, “we’re going to continue to see memory shortages for at least the next several years.”
  • Cerebras’s supply chain sidesteps every constraint he named: SRAM (no shortage, no separate maker’s margin — TSMC etches it into the chip), no CoWoS, and 5nm while “the 3 nanometer node is the most oversubscribed.” “We have been advantaged in this environment… others are paying the price.”
  • The long arc still deflates: everyone’s chips — “us, Nvidia, AMD, Qualcomm, ARM” — will produce more per watt and per dollar in three or four years. “The history of our industry is a massive reduction in the cost per unit compute.” Cerebras claims 15x speed for architectural reasons and “I believe the gap will widen.”

4. There is no market for slow inference

  • The signature riff: “How big is the market for slow search? It’s zero.” How big for dial-up? Even paid $1,000 a month you wouldn’t keep slow internet — “that’s how impossible it is to engage with an important technology slowly. Why do we believe that inference will be any different?”
  • For hard problems “there is no upper bound to how much faster you want to be”: solving in 3 minutes what takes a competitor 20, “imagine over a day or a week — you get smoked.” True in coding, agentic flows, everywhere.
  • The proof point, savored: Cerebras posted Kimi K2.6 running 6.7x faster than the next fastest GPU cloud “while one bozo at an analyst firm was on TV saying we couldn’t do it… If ever there was an example of being empirically proven dead wrong.” (“I’m a collector of examples of people being dead wrong. My wife has a list of when I’m dead wrong.”)
  • On the OpenAI contract — “one of the largest deals in the history of Silicon Valley,” $20-plus billion — he relishes the concentration critique: “I talked to you a year ago when I had a billion-dollar deal with G42. And you said you’re heavily concentrated… I come back with a 20-plus billion-dollar deal, and you tell me the same thing” — but with a different customer. His scaling law for customers: “The way to catch big customers is first catch one” — build the muscle, keep them happy, win the next.

5. Hyperscalers segment rather than commoditize — and Nvidia built an unhealthy dependence

  • His sharpest structural claim: “It has been Nvidia’s strategy to try and create competitors for the traditional hyperscalers. They have funded and backstopped and overallocated to the neo clouds. They have created a dependence, which is probably not healthy.”
  • Yet AWS and Azure don’t commoditize into utilities: security, credibility, Bedrock, SageMaker, your S3 data — “enormously valuable to most parts of the market, but not all.” The segment that says “give me cheap compute, I don’t care about anything else” makes the hyperscaler’s strength its weakness — his analogy: if you don’t want leather seats, you find the truck with Naugahyde seats. “Our business, just cuz it’s wrapped up in technology, is no different than any other business.”
  • On the Google-as-lowest-cost-token-producer thesis: real pros (land all the way up to tokens) but a historical con — “you can only sell your TPU to yourself,” constraining volume; Google stepping outside its own data centers shows it feels that constraint. Meanwhile anyone owning silicon in their own data center beats neo clouds, who buy hardware carrying Nvidia’s 70–80% gross margin before adding their own. CoreWeave he exempts from the overvaluation jab: “they’ve gotten paid for real innovation in financial thinking” plus genuinely rare rapid-deployment skill.

6. Multi-gigawatt is the new normal — delays are just what building is, but the industry was a bad neighbor

  • The habituation curve, via Sam: first GPT use is “this is amazing,” next day it’s “how come it’s not faster?” Same in power: 20MW was a lot, then 100MW, then a gigawatt — now “we’re running around looking for multi-gigawatt facilities” and 750MW gets a shrug. “Five years ago… that’d be delusional. Right now it’s like, ‘Oh, another one? Yeah, that makes sense.’”
  • On energy as the ultimate bottleneck, he keeps his distance from the Sam/Elon line that “we’re in the business of turning electricity into intelligence”: “I don’t know if I agree with that.” The alternative: “you bump into something else” — the thesis assumes models keep getting smarter enough to justify feeding them more energy. “That might be true. I don’t know.”
  • On 40 of 100 data centers stalling post-approval, his contractor analogy: “Have you built a kitchen? Was it built on time and on budget? No. Now imagine building something the size of 50 football fields” with municipalities, power companies, regulated industries, generators that literally fall off trucks. “Anybody who’s built anything big knows this is par for the course.”
  • But he doesn’t excuse the backlash: “our industry did a shitty job of engaging the community properly.” Brad Smith’s post should have been the template from the get-go: pay your own way, closed-loop water, upgrade substations and grids in full, don’t amortize power lines onto communities — “that’s BS.” When Harry likens it to Colombian cartels building churches, Feldman rejects it: these localities have unused power and depressed land; transparency, not buy-offs.

7. Layoffs are AI-washed, lawyers are the real blocker, and token spend goes to $5 trillion

  • The jobs call: “to date most of the layoffs were AI-washed. They were because we did boneheaded hiring during COVID” plus years of accumulated productivity gains now being harvested — the middle-management role of “information gatherers and presenters is being eliminated.” “None of this is AI yet… that is 90–95% of what the terminations have been about.” And the counter-position: “the list of things I want our engineers to do is 50 times as much as we have engineers… We’re going to hire more engineers, not less.”
  • On Benioff’s $300M/year Anthropic spend (3.8% of developer salaries, needing 20% to justify AI valuations): no concern. Hardware engineers’ EDA tools already run “much closer to 15 or 20%” of salary — “in software, we threw people at the problem rather than tools.” At $50–100K of tokens per engineer across 47 million software engineers, “that’s $5 trillion just in software engineering token use.”
  • Pushing back on Harry’s data-cleanliness thesis for slow enterprise adoption: “No, the biggest are lawyers” and the security apparatus — their payoff structure is “no credit, no credit, failure, blame,” so they’re structurally a drag on anything new, and lawyers can’t contract without precedent. He recalls (with a self-correction caveat) Jensen battling his own lawyers over Cursor and finally decreeing it. Only after legal/security relent does data matter — then Mayo Clinic’s 30-year data-organization quest and GSK become huge advantages.
  • The roles that don’t exist yet: CIO didn’t exist before Cisco’s mid-90s rise, CSO not before Palo Alto Networks-era security; the VP of telco infrastructure vanished with the desk phone. Coming: AI-governance roles (chief AI officers, maybe), while HR’s question-answering layer disappears — “AI can provide better answers, faster answers.”

8. Don’t sell chips to China — and give TSMC 20 lawless years to build American fabs

  • His China argument, stated against self-interest: strip out everyone in the chip industry — himself, Jensen, Lisa — and ask security people two questions. Will China’s military use leading-edge chips? “Everybody says yes. There is no debate on that point.” Will the state use them to advantage its industry against ours? Also yes. “That’s where I stop.” He grants the counter-arguments (keep them in our ecosystem, stop them building their own) “have real merit” — “I don’t agree with either.”
  • China is “at least today our industrial adversary” — see solar, lithium batteries, and Chinese cars displacing American ones worldwide — though he laments it: entrepreneurs he worked with at Baidu, Tencent and likely Didi are “every bit as good as anybody in Silicon Valley.” Even those who disagree say keep China down-rev; his line: “I’d like to keep my industrial adversaries more than down rev.” Choke points (TSMC, then ASML or Samsung) make the containment manageable.
  • On onshoring: America’s failure is “long-range policy… that endures more than a single administration” — China’s power infrastructure is extraordinary while the US grid is “a patchwork of 1950s technology if we’re lucky.” Losing fabs meant losing packaging and the whole surrounding strategic ecosystem. His one frictionless policy wish: TSMC and Samsung get 20 years free of all local ordinances to build fabs wherever they want in the US, using exactly their proven Taiwanese construction rules — “fabs are modern pyramids… the greatest things humans make in the manufacturing world by far.”
  • To Harry’s “should I be worried” about Europe: “You should be worried at the pattern” — a be-afraid, then regulate-and-tax mentality across chips, software and models, with pockets of excellence (likely Qwen? no) — including Cambridge, London and Stockholm, where 11 Labs and Lovable are doing interesting work. Harry’s nuance, which Feldman accepts: Europe’s application layer is world-class (11 Labs, Synthesia, DeepMind) but Europe lags on infrastructure, chips and models. The cultural gap: Silicon Valley’s “absence of a stigma if you try to do something extraordinary, crash and burn.”

9. The IPO: blocked by likely-CFIUS, cleared by a new government — and 18 months of daily failure behind it

  • The largest semiconductor IPO ever (Harry’s intro: $185 to $311, over $5.5 billion), with Feldman calling the timing deliberate but also luck and grit — and his account holds both truths. They didn’t know chips would run or that xAI/OpenAI couldn’t get out first; they did know “we had a chance to be the first and only AI pure play in the entire market. There’s only one, and that’s us.”
  • The earlier blockage: “we bumped into” what the captions render as “Sytheus”/“Sisyphus” — likely CFIUS — with “unnamed concerns that never got articulated… about some of our large customers. Then we got a new government and those concerns disappeared,” resolved on terms Cerebras had proposed a year earlier. Asked directly about the Trump administration: “unwaveringly better for business” — things he agrees and disagrees with, but unwavering on that axis. The builder’s lesson: “we kept building the business… You are always stronger if you keep building.”
  • The founder coda worth the whole episode: an 18-month stretch burning $8 million a month on a problem they couldn’t solve — failing at 2 seconds, then a year later at an hour, full failure analysis every time, never failing the same way twice. “There’s this myth that CEOs don’t doubt themselves… Of course you do.” “Nobody else to this day has solved it.” His kindest-thing answer, aimed at VCs: a board with empathy that knew “if the pressure doesn’t come from within, they bet on the wrong people.”
  • The human ledger: his last company made 100 millionaires, this one 800 — “if you don’t like delivering for your team, you’re not a real leader.” And the cost, verbatim: every CEO’s partner is “more lonely when you’re sitting next to them thinking about work than when you weren’t in the house.” “Emirates Airline sends me a Christmas basket… You know how frequently you have to fly for that to happen?”

Verification Notes

  • The IPO obstruction is not clearly identified in the captions; “likely CFIUS” remains an inference.
  • “Quan” may be Qwen, but the captions do not clearly establish that entity.
  • “DD” may be Didi, but the captions do not clearly establish that entity.