Pioneers Insight Method Research Author
The $2B Company Cutting AI Costs By 60% | Tuhin Srivastava
Back to Episodes

The $2B Company Cutting AI Costs By 60% | Tuhin Srivastava

Summary

  • Baseten’s apparent overnight breakout is a 15-year founder journey, with the company founded in 2019 and still only 18 people at its late-2023 Series B. Srivastava says his daily inputs barely changed: he still checks support, investigates product feedback, and sells. “The only change was that a market arrived and we didn’t give up”; ample capital and a deliberately light organization preserved the option to catch it.

  • The breakout was a refocus, not a cold-start pivot: Baseten had already built model-serving infrastructure, then watched three market variables flip at once. Small models became large, internal use cases moved into production with real SLAs, and data scientists gave way to engineers empowered to change infrastructure. Stable Diffusion supplied the proof point: Riffusion unexpectedly climbed from Baseten’s estimate of five A10Gs to roughly 100–150, prompting Baseten to remake the product surface in six weeks.

  • Srivastava expects shared endpoints for vanilla open-source models to commoditize, but says dedicated capacity—roughly 99% of Baseten’s business—remains differentiable. The workload heterogeneity is material: inputs, outputs, traffic, SLAs, cloud regions, compliance, hardware, and runtime configuration vary by customer. Baseten competes through fault-tolerant infrastructure, “A+” speed, and software for customers typically managing three to 20 models.

  • Inference performance is a two-layer problem: scaling workloads across five to 100,000 GPUs, then optimizing how quickly a model runs on each GPU. Infrastructure must preserve KV-cache reuse, minimize geographic hops, and help aggregate scarce capacity—such as sourcing 2,000 B200s when one cloud offers only 500. Runtime work balances time to first token, time per output token, throughput, and cost per token, sometimes putting research “less than a week old into production.”

  • NVIDIA’s advantage is broader than raw chip speed: CUDA combines reliability, versatility, and a developer ecosystem while challengers must solve chips, manufacturing, and software simultaneously. Hardware takes six to nine months to turn around, yet Srivastava says, “I can’t project more than 60 days forward in this business.” At very high volumes—one customer processes hundreds of millions of tokens per minute—B200s can provide a stated 40–50% speedup on many video workloads if they can get them running.

  • Enterprise users move from closed models toward open or custom models for control, not ideology. Srivastava cites purpose-built behavior, lower cost, more control over reliability, dedicated infrastructure, privacy, and enterprise SLA requirements. Customers also care that data is not simply “piped” to a provider that will train models on it. He expects mixed estates and offers an intentionally uncertain endpoint: “40/60—don’t know which way.”

  • The inference thesis does not require imminent AGI: Srivastava thinks models are already good enough to unlock “a thousand-x” more economic value, while the next capability jump may be farther away. His constraint is blunt—“I think we’re out of data”—so another leap might require an architectural breakthrough. Meanwhile, tools, sandboxes, code execution, reinforcement learning, and the application layer all create additional inference demand.

  • Baseten’s capital-allocation doctrine pairs persistence with aggressive abandonment of sunk work. In 2022 it killed three of four products, including an application builder developed for two-and-a-half years and Blueprint after six people—roughly a third of the company—spent six months building it. The governing distinction is “don’t give up” versus “don’t be emotionally attached”: venture backing means avoiding local maxima and continuing to swing at the largest opportunity.

Deep dive

Not yet available upstream; scheduled sync will retry.