Benchmark's AI Bets: Cerebras, Sierra, Legora, Fireworks, Starcloud, Gumloop..
Benchmark's AI Bets: Cerebras, Sierra, Legora, Fireworks, Starcloud, Gumloop..
Summary
- The old inverse relationship between scale and risk no longer cleanly applies in AI, and the spreadsheet-investing playbook is breaking down. Randle notes exceptions for hard tech and some capital-intensive consumer internet businesses, but says AI now produces “businesses that are well over a billion dollars in revenue that haven’t proven out their unit economics” or durable differentiation. Impairment risk stays flat—maybe even rises—with scale. In the most popular AI categories, the old SaaS rules—70–90% gross margins, no services load, capital-light operations, and Rule-of-40 legibility—are almost inverted: “FDE is the new PLG,” and high gross margins can signal “no one’s using your AI features.”
- Inference is the business-model unlock behind many parabolic AI revenue curves. Randle’s advice to a monetization-stuck portfolio company: stop sitting on the riverbank—“over there, there’s a fucking waterfall… the waterfall’s inference.” Charging a margin on inference instead of dollars-times-heads is why companies can go “one to 30 to 300” instead of one-to-three-to-nine; agents represent a major product and business-model shift since SaaS. Developers were spending $3,000 per month each on Cloud Code—$36,000 per developer—turning a $50K SaaS ACV into a potential $20 million line item, or for some, $500 million a month.
- The frontier-lab outlook depends on two scenarios. If recursive self-improvement delivers “geniuses in a data center,” labs retain pricing power and may reaccelerate. If capabilities hit a ceiling and distillation makes open source 95% as good, “that’s a really scary situation for the frontier labs”—though not a death knell, because most of ChatGPT’s 900 million weekly active users “wouldn’t be able to tell you” if the model were swapped for 5.2 or something similar.
- Anthropic’s financing could create an unprecedented liquidity shock. Randle says that if Anthropic reaches a $1 trillion–$1.5 trillion public valuation, its $30 billion round at a $380 billion valuation would gross-return 35 times the Snowflake pre-IPO round in a single deal. He knows people with $3–4 billion invested in Anthropic. San Francisco housing already clears at “2x asking price… all cash” or lab equity, and he is unsure the ecosystem understands the impact of that much liquidity.
- Late-stage rounds can now have more upside than a Series C, while AI’s day-one capital needs remake venture. SpaceX, Randle’s first Kleiner investment at a valuation above $100 billion, was transformed by Starlink into a company whose S-1 business is mostly consumer and B2B broadband rather than launches. A neo-lab might need $2 billion of compute before knowing whether its thesis works, unlike Airbnb’s roughly $500K YC seed. Firms such as General Catalyst and Andreessen Horowitz increasingly operate as alternative asset managers, with venture as a product rather than the firm itself.
- Benchmark’s counter-strategy is founder-out, never theme-in. Its seemingly thematic portfolio—Lagora, Sierra, LangChain, Fireworks, HeyGen, Gumloop, and StarCloud—was not built from category selection. Chetan closed StarCloud weeks before Elon publicly professed enthusiasm for orbital data centers; the investment centered on Philip and his team, who already had a GPU working in space. “Great founders are always in style,” and it “would kind of suck to be the SaaS fund right now.”
- Open source and frontier models are not zero-sum—at least yet. Randle’s AI mom test says that 100% of his nontechnical mother’s AI needs can now be handled without a frontier model. He recalls Cognition publishing work on post-training an open-source model for low-complexity tasks, though he is unsure whether that is exactly what it did; the savings could reach 95%. Frontier demand is also growing rapidly after the Opus 4.5 coding breakthrough. Eric Vishria’s framing is yes to on-device inference, open-source inference, and proprietary models.
Deep dive
1. The scale–risk relationship has flipped
- Fresh off Benchmark’s AGM, Randle’s read on the market mood: the whole venture-growth ecosystem is in a phase of disorientation. Everyone feels like “down is up and up is down.” His explanatory framework: in the software era, scale and the risk of major impairment were inversely related, because scaling forced sequential de-risking—product-market fit, then unit economics, then TAM, then market leadership—and if you skipped a step, “you just stop growing.”
- What broke: “you can have businesses that are well over a billion dollars in revenue that haven’t proven out their unit economics… durable product differentiation.” Impairment risk—zero being the worst case, but even a markdown from the last round—now looks flat with scale, or “maybe there’s even a weird positive correlation between scale and risk.” That inversion, he argues, is why investors across the market feel the risk-reward paradigm no longer computes.
2. Every golden rule of spreadsheet investing is now almost inverted
- The old canon: 70–90% gross margins, pure software with no services or implementation load, capital-light operations with R&D leverage, and 90%+ gross retention—compounding into capital-light businesses producing durable free cash flow “at multiples of GDP.” That’s what justified paying high multiples of ARR, and it was so legible that some investors invested based only on a Rule-of-40 score.
- Now the hottest companies invert almost every rule. “FDE is the new PLG”—the “Palantirification of everything” makes sparkling implementation consultants a popular distribution strategy; high gross margins can be a bad sign because “if you have an AI product with high gross margins, that means no one’s using your AI features”; and the playbook for staying safe from the labs has app companies training their own models or at least post-training on user data—“unbelievably CapEx-intensive” versus old-era SaaS.
- Why it disorients: those golden rules “are the first-principles building blocks of a high-quality company.” When the most popular categories look less attractive from first principles than SaaS, “you have to kind of rethink a lot of how you invest.”
3. AI companies no longer “taste like chicken”
- Randle started at Vista, where CEO Robert Smith preached: “Software tastes like chicken, and that’s why it’s beautiful”—every SaaS P&L looked essentially identical at maturity. AI kills that: Fireworks, which Benchmark thinks is becoming the AI inference cloud, owns no data centers, leasing GPU capacity and getting paid for software that cuts cost and latency, while Crusoe acquires power, land, and permits to construct data centers for hyperscale counterparties. Ostensibly two inference companies, “they’re actually more different than they’re alike” in capital intensity, margins, and business model.
- His new taxonomy is P×Q×M. For an AI app company, Q is probably similar to SaaS because it sells to the same kinds of customers; M is “almost definitively lower—for I think 99% of AI app companies, it’s lower than 70%”; but P can be immense: inference platforms have nine-figure contracts with startups, while very few SaaS companies have nine-figure contracts with anyone. Who understands this best is case-by-case, but the edge goes to founders and investors who have thought through this new taxonomy.
4. Benchmark’s answer: founder-out, never theme-in
- Backing entrepreneurs at inception—“at 50 Post”—sidesteps the exit-multiple and capital-efficiency-at-scale questions: once a company has reached that level of scale and maturity, “you’ve sort of already done your job.” Business models cycle while “great founders are always in style”: “you weren’t supposed to touch hard tech, and now hard tech is apparently the only thing that’s safe from the foundation model labs.” It “would kind of suck to be the SaaS fund” rebranding as the SaaS AI fund—a tougher strategic position.
- Outside-in, the portfolio looks thematically engineered—vertical AI in Lagora, horizontal AI in Sierra, developer tools in LangChain, prosumer software in HeyGen, and orbital data centers in StarCloud. On joining, Randle discovered “there was literally no thought” of category selection; some companies were pivots, so their eventual categories were not even the categories Benchmark had originally invested in.
- The StarCloud specimen: Chetan closed the investment weeks before Elon began publicly professing enthusiasm for orbital data centers. Randle recalls an interview, on a platform he could not remember, in which half the discussion focused on orbital data centers; the partner group chat concluded, “this space is about to get really hot.” The team had the only company with a GPU working in space, but “the investment was about the people. It was about Philip and the team.” Great entrepreneurs “always pull rabbits out of hats” and end up pioneering big categories.
5. Get under the inference waterfall—agents are the biggest shift since SaaS
- His coaching image for a portfolio company with a huge developer base but no business model: you’re on a riverbank with a bucket, “and over there, there’s a fucking waterfall.” You do not know until you are underneath it whether the bucket has holes or how large it is; “the first thing you should do is get under the waterfall, and the waterfall is inference.” Fal, Baseten, Modal, and usage- and outcome-based app pricing are different derivations of monetizing inference, helping explain why companies go “1 to 30 to 300” instead of “1 to 3 to 9 to 20.” Some models thinly resell inference as brokers; others, like Sierra’s outcome pricing on completed support deflections, abstract it away.
- On the term itself, Randle cringes when it becomes over-marketed: “when private equity firms are telling all of their portfolio companies to say that they sell agents… the term is cooked.” Yet he calls it the greatest product and business-model innovation since the start of SaaS, possibly in the history of technology, because buyers shift from “I buy a license” to buying intelligence, white-collar activity, or economic output on tap.
- The moment Benchmark got “insanely excited”: developers spending “$3,000 per month each on Cloud Code,” back when it was still Cloud CLI. That’s $36,000 per developer against the SaaS era’s $50K total ACV. Instead of merely a $200K line item for the average company, this could become a $20 million line item—or, for some, $500 million a month.
6. Gumloop and the router thesis: independent vendors can navigate jagged models
- Gumloop is an independent third-party collaborative AI-agent and automation canvas for enterprises: every employee—not just developers, salespeople, or marketers—can build simple Zapier-style automations or full agents, triggered from Slack or Microsoft Teams or running in the background. Cloud Cowork and the news around OpenAI combining Codex and ChatGPT into a productivity product could produce big businesses here too, but the third-party position matters because “the models are actually pretty jagged”: Gemini is best at multimodal work, Claude typically has the best coding model—“although now a lot of people think GPT-5.5 is actually better than the latest Opus model”—and open-source models can offer substantial value for their inference cost. Meanwhile, some enterprises have employees checking the weather with Opus 4.8.
- His AI mom test: what does a nontechnical mother in rural-suburban Colorado need from AI that cannot be done by a very cost-effective open-source model? Two years ago, he did not know; now, he says there is nothing his mother asks of AI that needs a frontier model.
- Cognition has published related work. Randle does not know whether it trained or post-trained the model, though he thinks it post-trained an open-source model on low-complexity tasks observed within Cognition. For tasks that do not need frontier intelligence, that could save 95% on an action or query.
- But it is not zero-sum: frontier demand is also growing rapidly. “I don’t think we had unbelievable-quality coding models until Opus 4.5, until last winter”—the breakthrough behind Claude Code’s and Anthropic’s revenue growth. Partner Eric Vishria’s framing captured the breadth of demand: on-device inference, yes; open-source inference, yes; proprietary models, yes. “There’s a lot of demand for all of this stuff, and the demand is all going parabolic.”
7. The multi-trillion-dollar question: geniuses in a data center or a distillation squeeze
- Molly’s prompt was that OpenAI has enormous revenue, while Anthropic was last rumored around $45 billion and more recently speculated around $60 billion; what happens as the models become more efficient? Randle’s fork: if recursive self-improvement delivers “geniuses in a data center,” frontier labs get real pricing power and may reaccelerate. If capabilities hit “an absolute ceiling” and distillation keeps open source at 95% of that ceiling, “that’s a really scary situation for the frontier labs.”
- Not a death knell, though: “most of the users of ChatGPT would use ChatGPT whether or not there was a GPT model in there”—outside the San Francisco bubble, the vast majority of its 900 million weekly active users “wouldn’t be able to tell you” if the model were swapped for 5.2 or something similar. The products could retain their appeal; the pressure would instead fall on the premium margin frontier labs can charge for tokens.
- Randle’s caveat: “open-source inference also costs money… someone still has to run the GPUs, build the data centers, and operate the data centers.” The question is how much premium margin frontier labs can create.
8. Capital markets remade: day-one billions, late-stage rebirths, venture as a product line
- Beyond the familiar staying-private-longer trend, AI adds “massive day-one costs”: a new lab might need “$2 billion of compute” to test a research direction versus Airbnb’s “$500K or whatever they raised at YC” seed—“a completely different capital market” and funding mechanism.
- The strangest inversion: a late-stage company can now have much higher upside than a Series C. His first Kleiner Perkins investment was SpaceX at a valuation above $100 billion—peers asked if he was “gunning for a 2.5X”—but Starlink was “a rebirth of the company.” The S-1 business is mostly consumer and B2B broadband, not launch. Companies 15 years in can still be “only 1% done with their journey,” expanding what rounds and situations Benchmark considers.
- On industry nomenclature, Randle insists that General Catalyst and Andreessen Horowitz are alternative asset managers because they offer venture, growth, debt, health assurance, and wealth-management products. ICONIQ likewise has both a growth-stage venture practice and a high-net-worth wealth-management practice. “Venture, in many ways, is still the same, but it’s a product now for many of these firms. It’s not the firms themselves.”
9. The Anthropic liquidity shock: 35 Snowflake pre-IPO rounds in one round
- Randle’s quick analysis after the Anthropic discussion: four of the best pre-IPO investments over the last 10 years—Slack, DoorDash, Snowflake, and Nubank—involved $500 million–$2 billion rounds returning 2–5x over roughly four years. Snowflake’s roughly $500 million became “something like $2.5 billion,” and everyone went home happy.
- Randle discusses two Anthropic valuation references: a $30 billion round at a $380 billion valuation, and a statement that Anthropic had “just raised at about a $1 trillion valuation.” Using his $1 trillion–$1.5 trillion public-liquidity scenario, he says the $30 billion round at $380 billion would gross-return 35 Snowflake pre-IPO rounds in one deal. He knows several people with “$3 to $4 billion invested into Anthropic,” when before COVID “a normal growth fund was, like, a billion dollars.”
- The knock-ons are already visible in San Francisco housing—“everything’s going for 2x asking price… all cash” or lab equity—and the second-order questions compound: what do enriched lab employees fund, start, or stay for? If the liquidity arrives in under five years, it would be a shock affecting every aspect of life in Silicon Valley, unlike SpaceX, which took 20-plus years to distribute its wealth.
- Closing notes: his mentors are Founders Fund’s Napoleon Ta—“this beautiful, simple life” of work and family, with no podcasts or networking—and Eric Vishria, “a Mount Rushmore venture capitalist” at #3 on the Midas List who is “almost unassuming.” And the epistemic sign-off is worth keeping: “We can look back at this in a year, and I’ve probably been wrong about everything.”