Pioneers Insight Method Research Author
Cerebras CEO, Andrew Feldman on Why Raise $1BN and Delay the IPO & Why NVIDIA’s Worried About Growth
Back to Episodes

Cerebras CEO, Andrew Feldman on Why Raise $1BN and Delay the IPO & Why NVIDIA’s Worried About Growth

Summary

  • Cerebras raised $1bn — the largest round ever in its category, at the highest valuation — led by Fidelity (“the Oxford or Cambridge of investing”) with Tiger Global, Valor and 1789, and Feldman says Cerebras still intends to go public: “We still have every intention of going public.” Feldman says the round was possible because Cerebras had margins while rivals shopping for money had negative ones — and he thinks the S-1 said the UAE accounted for about 75-80% of revenue, perhaps in H1 2024, with orders so large they “consumed all our manufacturing capacity.”
  • Nvidia is showing a worried giant’s tells: “use your balance sheet more and your technology less” — buying business rather than winning it (as Cisco did from 1992-2001), “predatory pre-announcing” B300s before anyone can get B200s, and silence on field failure rates “which are massive.” The $100bn OpenAI investment “was designed for nobody to understand it… it’s just not an analyzable thing” beyond Nvidia locking up a slice of OpenAI’s demand.
  • Two-year chip depreciation is “empirically wrong”: H100s are still earning past two years, A100s are at three-to-four and “could be as long as five or six.” The real depreciation variable is generational improvement — and apples-to-apples (8-bit to 8-bit) Feldman estimates 2-2.5x per meaningful generation, because memory bandwidth, not flops, is the binding constraint for inference on GPU architectures.
  • Demand is unknowable even to buyers — customers ask Cerebras for “between five and 40 million queries per second,” with Harry highlighting the uncertainty as an order of magnitude — so read megadeals as options on the future, and read “up to $100 billion over five years” as “the great CYA word in marketing history.” Chance he’s still underestimating demand: “100%. I’ve been wrong.”
  • On Mag7 concentration, the risk isn’t valuation (Nvidia at $4.5T: “maybe it’s too low”) — it’s mispriced diversification: the S&P “is not an index of the global economy — it’s 30% or 50% seven companies,” so holders carry sector risk they never signed up for. “Risk comes in financial markets where people fundamentally underestimate risk.”
  • On energy, the scarcity narrative is “strictly wrong”: “We have plenty of power. It’s in the wrong places” — West Texas gas and upstate New York hydro sit far from people and fiber, nuclear is reasonable but not unavoidable, and America’s patchwork permitting (a local fire ordinance set Samsung’s Texas fab back 8-10 months) is the real handicap versus China’s central planning.
  • Contrarian diffusion calls: Groq’s Jonathan Ross’s five-year AI labor-shortage prediction is “absolutely wrong — that might be true in 15 years”; AlphaFold won Nobels but “name a drug that’s resulting from it. Not one.” Productivity jumps only when society reorganizes around AI — the electricity-and-the-dynamo lesson — not when AI just replaces Google.

Deep dive

1. The $1bn raise: Fidelity’s stamp, and margins bought the ticket

  • Feldman calls it the largest raise ever done in Cerebras’s category, at the highest valuation, led by Fidelity — “the Oxford or Cambridge of investing… when they choose to lead a round it brings Wall Street a great deal of confidence” — co-led by “the treaties” (likely Atreides), with Tiger Global, Valor and 1789 participating. Harry’s corroboration from HubSpot’s CEO: getting Fidelity specifically, pre-IPO and at IPO, carries signal weight venture investors underrate.
  • Why not just go public? “We still have every intention of going public” — a pre-IPO round is standard “if you can get it done very quickly, if it doesn’t distract you.” And the raise was possible because of unit economics: “the reason we were able to raise at a higher valuation and from better investors and more money is because we had [margins] — and others who were out looking for money had negative margins.” Margins are “a really important part of moving from being an idea to being a real company.”
  • What the dry powder buys: manufacturing buildout, more data centers (five added in the US this year), and “more big ideas” — with a swipe at the competition: “make-believe gains achieved by dropping from 8-bit to 4-bit — those aren’t going to get us to the promised land in AI.”

2. Read the fine print: megadeals are options on the future

  • Feldman’s deconstruction of the headline numbers: “up to” is “the great CYA word in marketing history — up to $100 billion over five years… could be 30, could be 12, could be 40. It won’t be bigger.” And nobody audits the promises: “in eight months, has anybody got a little spreadsheet — nine jobs plus one factory? Who holds anybody to account? Nobody.”
  • What the announcements actually signal: demand so large that buyers can’t scope it. Customers ask Cerebras for “between five and 40 million queries per second” — with Harry highlighting the uncertainty as an order of magnitude — because “6, 8, 12 months out everybody’s unsure. It’s so fast. It’s so big.” The right frame: paid options on future capacity — “if the future moves against you, you lose the premium.” Planning in this regime is “brutal”: five-and-seven-year data center commitments demand “good planning with changing rules rather than good planning.” Odds he’s still underestimating demand: “100%” — pick the date OpenAI’s valuation became conceivable: “the day before.”
  • The demand proof point is the Gulf: Feldman thinks the S-1 showed the UAE at about 75-80% of revenue, perhaps in the first half of 2024, with orders so large they “consumed all our manufacturing capacity” — “you can be a professional salesperson in Silicon Valley for 20 or 30 years and not see a $500 million order.” They were “bold and they were early.”
  • Harry’s needle — don’t you have to say nice things about your biggest customer’s region? “Fair question. I went there to do business as a Jewish guy before we had any business done.” His quickfire contrarian belief follows: peace in the Middle East in our lifetimes, on the UAE’s evidence that moderation pays — “we’re too busy to hate… we’re too busy building.”

3. Nvidia is showing a worried giant’s tells

  • Asked whether Nvidia is unshakable (Groq’s Jonathan Ross told Harry it gets to $10T within five years), Feldman: “I hope he’s long on them then.” His read: “we are seeing some things that big companies do as they begin to worry about growth — use your balance sheet more and your technology less.” You “start buying business as opposed to winning business,” as Cisco did around 1992-2001.
  • The second tell is the “predatory pre-announce”: “you announce B300s before anybody can get B200s, you start talking about Rubin before B200s are technically finished” — while never mentioning “the field failure rate of your products, which are massive” — all to convince buyers to wait rather than choose technology “that’s better and present.”
  • The $100bn OpenAI deal “was designed for nobody to understand it… up to this amount, over an unspecified amount of time, at no valuation given — it’s just not an analyzable thing,” beyond the fact that “Nvidia has chosen to try and lock up a portion of OpenAI’s demand by investing in them.”
  • On margins: Nvidia’s are “some of the highest in history for a hardware company” — 78% gross, “might be on the high-end chips 85%” — and that’s precisely why AWS builds its own Trainium (likely; spoken as “cranium”) part. Resentment compounds: “when Intel stumbled, the number of people who came out of the woodwork to kick them when they were down was extraordinary.”

4. Two-year chip depreciation is “empirically wrong”

  • People are still getting value from H100s past two years, and from A100s — “closer to three or four years, and could be as long as five or six.” So: “if you say it’s a two-year depreciation, you’re empirically wrong.”
  • The actual question: “how much faster are future generations than the current generation.” Retiring a fully-paid-off chip only makes sense when the replacement is so much faster per watt that it earns more dollars from the same 50MW shell. If the industry stops delivering extraordinary generational gains, “they last longer. You depreciate them longer.”
  • And the honest generational gain, “a little bit of engineering digging” past the marketing: probably 2-2.5x per meaningful generation, apples-to-apples (8-bit to 8-bit). More flops are wasted if “your memory bandwidth didn’t improve more than 2x — you can’t get to them.” For inference, memory bandwidth “is the fundamental limiter for the GPU architecture.”

5. The wafer-scale bet: obvious-sounding, unsolved for 75 years

  • The memory trade: SRAM is blazing fast but tiny; HBM (a DRAM flavor) is big but slow — GPUs chose it because graphics rarely touches memory. Cerebras’s move: a dinner-plate-sized chip stuffed “to the gills with fast SRAM,” beating SRAM’s capacity limit with sheer silicon area. A normal-sized SRAM chip on a trillion-parameter model needs “four or 5,000 chips. What a mess.”
  • Harry’s pushback — isn’t that obvious? “It does, doesn’t it?” But nobody in the 75-year history of computing had built a chip beyond ~840mm²: Gene Amdahl failed, IBM failed, TI failed — “and after we did it, Elon tried at Dojo and they failed.”
  • The bet nearly died: roughly 15 months from 2017 to early 2019 when they couldn’t make one, burning $6-7M a month, running a formal failure analysis after every attempt. When the first wafer finally ran, the founders stared at the box for half an hour: “we have just solved a problem that for 75 years the smartest people in our industry have been unable to solve.”
  • On training versus inference: “No, we’re faster on both” — but training means porting GPU-native recipes, whereas in inference “nobody cares about CUDA. Nobody even cares about PyTorch… what they want is an API. It’s literally 10 keystrokes to move from a GPU-based solution on OpenAI OSS 120B to our solution.” That, plus inference users vastly outnumbering trainers, is why inference is the easier land-grab.

6. The dynamo lesson: no reorganization, no productivity jump

  • Feldman reaches for likely Robert Solow’s 1988 paradox (“computers everywhere except in the productivity statistics”) and Paul David’s “The Computer and the Dynamo”: electricity, adopted in manufacturing from 1880, produced almost nothing until shop floors were reorganized around it — then productivity leapt, just as it did in the mid-’90s once computers were networked. “If you use OpenAI the way you use Google, you’ll see a very modest jump… If we reorganize ourselves around AI, you’re going to see massive productivity gains.”
  • The demographic tell, echoing Sam Altman: older users treat ChatGPT as a Google replacement; younger users use it “as an operating system for life” — a consumption pattern that never existed before, and where the jump will come from. “What I know is that transition takes time.”
  • Inference growth itself is three multiplied variables — users × frequency × compute per use — “the problem is they’re all growing fast, and that produces some mind-numbing effects… we knew that going in, and it still takes your breath away.”

7. Diffusion is slow: no labor shortage, and where’s AlphaFold’s drug?

  • Ross predicted AI creates massive labor shortages within five years. Feldman: “Absolutely wrong. Economic dislocation isn’t resolved in very short periods of time. That might be true in 15 years.” AI “will nibble its way in” to the economy.
  • His evidence: AlphaFold solved one of chemistry’s hardest open problems and won its inventors Nobels — “name a drug that’s resulting from it. Not one.” And the X-ray crystallographers it theoretically displaced? “There’s more demand for them.”
  • What does change: education — “we’ve been educating children the same way since Alexander the Great was tutored by Aristotle”; AI can compare a student’s error patterns against thousands of others and prescribe the workbook that fixes that specific hole. And entry-level work at consultancies and banks — “being really good at spreadsheets and writing summaries of other people’s research: AI will be better at that” — which he always thought was a terrible use of 22-year-olds anyway.

8. “We have plenty of power. It’s in the wrong places”

  • The scarcity framing is “strictly wrong”: a ton of power in West Texas natural gas and upstate New York hydro — just not where the people, buildings, or telco fiber are. The problem is mismatch, not supply. Nuclear is “a very reasonable and cost-effective strategy” over decades, but not unavoidable — Canada’s falling water could be “the cheapest power on earth,” Finland and Iceland have geothermal.
  • China “thought long and hard about their power infrastructure” and planned strategically; America’s “decentralized form of government has left us with a patchwork of power infrastructures” — a local fire ordinance forced Samsung to redesign a Texas fab and set a multibillion-dollar project back eight-to-ten months.
  • Politics scorecard: “The Biden administration was misguided and afraid”; Trump on net “probably more to help,” having surrounded himself with smart AI people and relaxed painful regulation. On the US-China frame he resists the race narrative — “the arms race certainly didn’t help either the US or Russia” — though the realpolitik: “they’re better at making drones. They’re better at making robots,” and Beijing backstopped its AI venture funds’ losses. Feldman passed on a big China deal in 2019, before export controls, over how the technology would be used.
  • The obligation that comes with the wattage: “if we are going to consume this amount of power, the burden is on us to deliver value for it” — drugs, healthcare, aging. Likely Ghibli-image compute? “A market has a lot of bad ideas to get a few good ones” — the messiness is the mechanism; steer government dollars and permitting breaks toward projects that matter.

9. Mag7: the risk is mispriced diversification, not overvaluation

  • The concentration risk “is not that they are that much value — I think they’re that much value because the future economy will reward that.” The danger is the mental-model mismatch: “people continue to think the S&P is an index of the global economy — and it’s not, it’s 30% or 50% seven companies.” Holders “thought they were diversified and in fact they’re heavily dependent on a very narrow sector.”
  • The line that carries the episode’s risk framework: “Risk comes in financial markets where people fundamentally underestimate risk. When risk is priced properly, your outcomes are not surprising.”
  • On Nvidia at $4.5T: “the greatest company of the first quarter of the 21st century… I don’t know if 4 trillion is right, but a very big number — maybe it’s too low — is right.” He won’t pick public stocks — “you can lose money on good companies, you can make money on shitty companies, and that for me doesn’t sit well” — and betting on the biggest of big dogs has “no alpha” in that.
  • On whether the boom is sustainable: extrapolation has limits — “if Nvidia keeps growing at the rate they’re currently growing, 11 years from now everybody on earth works for them — do the math” — but a reorganized, AI-larger economic pie is “not only likely, it’s almost certain.”

10. Bottlenecks — and where investors will lose their shirts

  • Expertise first: universities “aren’t minting enough” AI practitioners, and US immigration challenges “don’t help”; Feldman argues for the J-1-to-H-1B path and says the US must take immigrant talent seriously, while universities are starved of compute. That scarcity is why the pay war doesn’t worry him: “No company ever went bankrupt by paying extraordinary people too much. If you want to go bankrupt, pay mediocre people too much.” (America paid Charlie Sheen $2.5M an episode, after all.)
  • Physical bottlenecks: TSMC and Samsung can’t build their $30-50bn fabs fast enough, capping everyone’s chip supply and keeping costs up. And the promised gigawatt data centers? “Everybody’s committing to them. Where are they? They’re not up yet” — Elon builds in six-to-eight months, “the rest of the world a year and a half, maybe longer.”
  • Wall Street loves data centers because “it looks like a bond… you get an investment-grade tenant” — CoreWeave’s financial engineering showed the way — but “building data centers is not for everyone”: the best build at $8M a megawatt; “if you’re spending 12 or 14, that’s how you lose money,” compounded by power access, permitting, cost control and tenancy.
  • On OpenAI or Anthropic building their own chips (Ross said definitely): “there is a long history of software companies failing to build chips” — Microsoft-scale companies failed, Google is the most successful “and they’re 10 years in”; wins are usually acquired (Apple/PA Semi, Amazon/Annapurna). Neither OpenAI (Azure-hosted for years) nor Anthropic (AWS plus Google) is vertical today. “Chip building is an MBA nightmare” — Intel had the world’s best architects and fabs in 2000-2010 and “proved completely unable to build a working cell phone part”; ARM won the century’s largest compute market. Silicon “is not a place for 25-year-old CEOs.” The underinvested corner: data cleaning and pipelines — “many AI projects fail… because the data was a disaster” — while data provision (Surge, Mercor, Invisible, Turing, Handshake) is “a very curious market”: clearly important now, “whether it’s durable, whether we get machines that do it every bit as well as people… could go either way.”
  • On silicon generally, Feldman rejects a 90% monopoly: Intel dominated x86 but had zero cell-phone share, while Broadcom dominated switching silicon; he does not expect the market to accrue to one or two companies.
  • On sovereignty, he says Mistral’s sovereignty strategy, combined with Cerebras delivering inference through what he calls the fastest hardware on Earth, makes its likely Le Chat product compelling; Europe otherwise has too few AI labs doing interesting work.