Baseten CEO Tuhin Srivastava on Custom Models, and Building the Inference Cloud
Summary
Baseten’s 30x growth reflects an AI-native application boom, while the larger enterprise inference market is still mostly absent. Sarah Guo said Baseten expects more than $1 billion in revenue this year, which Tuhin Srivastava acknowledged; he estimates application companies still represent 99% of inference count, so “the majority of the market hasn’t come online.”
The durable application moat is proprietary workflow feedback, not merely access to model weights. Abridge’s clinician edits and downstream EMR actions create a reward signal that frontier-model companies may not have access to; post-training on that signal could produce specialized, long-horizon agents. “The user signal that they can gather that only they can gather” is what protects the application layer.
Production inference is already overwhelmingly custom, linking deployment to an increasingly continuous post-training loop. More than 95% of Baseten’s tokens run through dedicated inference, and almost every customer modifies models for quality, performance or both: “No one is just running the vanilla open-source weights.” The sequencing matters—“no post-training pre-product-market fit”—because companies should first prove value with the best model, then specialize it to become “better, faster, and cheaper.”
Chinese open models are economically strategic, while Guo noted that closed-source U.S. labs still define the absolute frontier. Srivastava hedged that he “could be wrong,” but said that if these models are network-bounded, they will not “magically” cross network boundaries; he said DeepSeek could run at probably 20% of Anthropic’s production cost, with comparable or better latency and probably better reliability. His stance: U.S. open models are both necessary and inevitable, but ignoring today’s available intelligence would be “missing the forest from the trees.”
Capacity scarcity now reaches directly into contract duration, working capital and likely public-market timing. Baseten runs 90 clusters across 18 clouds at uncomfortably high utilization—close to, but not generally in, the mid-90s—and holds a daily 4 p.m. meeting to allocate supply. A 1,024-B200 block from a credible cloud can require a three-to-five-year commitment and probably 20% of TCV prepaid; asked whether that argues for going public earlier, Srivastava answered, “go sooner.”
Baseten’s moat thesis is software plus scarce compute, not bare GPU rental. Srivastava calls GPU-as-a-service a commodity, while inference software has delivered no churn among Baseten’s top 30 customers and roughly 400% annual NDR: “If we have all the compute, good luck running inference.” He expects specialized chips, but NVIDIA’s supply chain, CUDA and ecosystem mean infrastructure operators “can move fastest with NVIDIA today.”
Inference efficiency is behaving like Jevons paradox: lower unit costs produce longer agents and more total cognition, not demand saturation. Developers insert “a hell of a lot more intelligence” when it gets cheaper because better answers improve experiences and revenue. Srivastava therefore calls inference “the last market”—even with AGI, inference remains—and sees “concierge everything” for consumers but an “extinction moment” for workflow companies that fail to add intelligence.
Deep dive
Not yet available upstream; scheduled sync will retry.