Pioneers Insight Method Research Author
Open Models vs Frontier Models: Who Actually Wins? | The $100K Token Budget Every Engineer Will Need
Back to Episodes

Open Models vs Frontier Models: Who Actually Wins? | The $100K Token Budget Every Engineer Will Need

Summary

  • Clay Bavor’s central model call is that cheaper open weights will absorb yesterday’s frontier workloads, but they will not eliminate demand for the frontier itself. GPT-4-level intelligence from spring 2023 now costs roughly one-300th as much per equivalent token, creating an “assembly line” from frontier to fine-tuned open models. Yet coding, science, legal work, invention, and discovery could sustain “unbounded demand for, call it frontier levels of intelligence.”

  • Token prices will fall technologically without necessarily collapsing economically because reasoning consumes more inference while GPUs and power remain scarce. OpenAI’s o1 showed performance continuing to improve with test-time compute, while the founder of Nebius told Harry that even 10X more supply could sell out in a day. Bavor has heard and observed top engineers spending over $100,000 of tokens annually on a run-rate basis and would bet steady-state spend lands “much closer to 20%” of developer salary than 3.8%.

  • AI should produce smaller, higher-leverage teams, but Sierra’s enterprise motion argues against extrapolating a 149-person software company across every category. Its engineers estimate Claude Code, Codex, and Sierra’s internal tools make them three to 20 times more productive, yet customers representing 40% of the Fortune 50 still require integration, trust, regulatory understanding, and forward-deployed help. Bavor’s qualification is important: FDEs are not mandatory, but they are “an important catalyst.”

  • Sierra is becoming AI-native internally through a shared data gateway, a companywide agent, reusable skills, and a strategy model grounded in operating context. Pinecone can interrogate permitted Slack messages, documents, presentations, and reviews; Bavor uses a custom skill to scan every hiring packet. Sierra Brain adds a 20–30-page company primer, board letters, operating reviews, and strategic beliefs to create a “strategy thought partner” that knows the business deeply.

  • The product thesis is already moving beyond customer support toward an agent-mediated front office spanning discovery, sales, conversion, service, and marketing. Rocket deployments cover home discovery, refinance outreach, loan preparation, and servicing, while Next uses Sierra for personalized product recommendations and basket building. The destination is a world where “the conversation is the interface,” supported by reusable vertical expertise rather than an endless collection of bespoke projects.

  • Sierra runs governance and financing around milestones rather than calendar convention or maximal headline valuation. Its board alternates three-hour and 90-minute meetings every six weeks, using six-to-10-page memos because “writing is just thinking on paper”; the cadence let management react when Claude 4.5 and Codex 5.2 changed software development. Every financing round was inbound, and Sierra “guided to and took a lower price” than available to fund the next unequivocally higher watermark.

  • AI fluency is reshaping who creates leverage, how Sierra interviews, and what an entry-level advantage looks like. Some of its most effective employees are 22 or 23 and “completely AI-pilled”; engineering candidates now receive $150 for the coding agent of their choice and build on their own laptops before explaining the result. The broader hiring filter remains “smart, nice, intense,” coupling tool fluency with systems judgment, product thinking, and culture.

Deep dive

1. Sierra was founded at a platform reset without chasing pretraining

  • Bavor left Google after 18 happy years because late 2022 aligned three conditions: a longstanding desire to found a company, confidence in Bret’s competence and character, and language models reshuffling “the proverbial deck of cards” toward smaller companies. They had known each other for 20 years and had nearly worked together a couple of times before.

  • The key Google inheritance was a willingness to descend as far down the stack as the product required. Sierra anticipated language-model agents in April 2023, recruited the Princeton professor behind the ReAct paper as founding head of research, and built new agent frameworks and architectures when the capability “should be possible” but was not yet operationally possible.

  • Pretraining was considered and quickly rejected. Bavor called a frontier model “a highly perishable bag of floating point numbers” whose initial and continuing capital expense works for very few companies; Sierra instead slipstreams behind labs and hyperscalers, then builds proprietary fine-tunes atop open-weight models. His boundary: control your destiny, but do not invent a story that you must own more of the stack than necessary.

2. Open weights inherit yesterday’s frontier while intelligence demand climbs

  • Harry’s bear-case challenge was direct: if open models can handle an expanding majority of enterprise tasks, does the frontier inherit only increasingly difficult problems? Bavor’s answer starts with the current denominator—fully automated enterprise work remains “a rounding error,” and the gap could reflect model, application-layer, and organizational-diffusion problems alongside model capability.

  • Bavor separates routine capability overhang from high-value intelligence. Any software company would upgrade staff engineers to principal or distinguished engineers; returning shoes does not need the best available model, but coding, legal work, science, materials, invention, and discovery suggest that demand for useful intelligence has no obvious ceiling.

  • The economic pattern is an “assembly line”: GPT-4 in March, April, or May 2023 supported valuable workloads, and equivalent intelligence now costs about one-300th as much per token. Those workloads can migrate to fine-tuned open weights while frontier models remain aimed at tasks where greater intelligence earns its cost, leaving enterprises to mix both.

  • On China’s stronger open ecosystem, Bavor’s hedged explanation was scaled distillation of US frontier training runs. US labs would pressure their own hosted-model pricing by releasing comparable open weights; if another company cannot build the frontier itself, “maybe the next best approach is to distill them and offer them up.”

3. Reasoning and scarce compute keep token economics from collapsing

  • Bavor’s overlooked milestone was OpenAI’s o1 in late 2024. Its test-time-compute chart kept moving “up and to the right”—logarithmically, and therefore eventually flattening—as additional inference and thinking produced better performance. The implication is that cheaper tokens can enable models and agents to consume more tokens in pursuit of greater intelligence.

  • Hardware will produce more equivalent tokens per dollar, and some workloads will migrate to cheaper open weights. But if demand for intelligence is effectively unbounded while Blackwells, H100s, power, and GPU capacity remain limiting inputs, basic supply and demand creates a floor under token costs even when the underlying technology improves.

  • Harry relayed Nebius’s claim that 10X more capacity could still sell out in one day; Bavor believed it. Self-hosting removes part of the frontier provider’s margin stack, but not the constrained physical inputs of energy and compute.

  • Local models can improve consumer experiences but would not by themselves alleviate the server-side challenge. Phones encounter thermal limits; Bavor can imagine language-model-optimized devices or a mains-powered home compute appliance that might help alleviate some demand, but frontier work still requires “a giant rack of TPUs or GPUs in a data center somewhere.”

4. AI-native engineering moves the bottleneck above code

  • Sierra engineers who are deeply using Claude Code, Codex, and Pinecone estimate they ship three to 20 times more features. Bavor expects smaller, higher-leverage teams across functions, but treats the software-engineering and data-analysis gains as already unmistakable rather than speculative.

  • Sierra’s internal foundation is an MCP gateway aggregating its main systems and services into one permission-aware interface. Added to Claude, Codex, or Pinecone, it lets an employee reason across everything they are entitled to see—Slack, documents, presentations, and operating reviews—without granting access to another person’s private material.

  • Pinecone layers company-specific harnesses and a shared skills library onto that gateway. Pinecone knows how to build Pinecone; it has harnesses around its own engineering and around Sierra’s core agent architecture and Agent Studio, where agents are built and deployed. Bavor’s private “Clay scanner” encodes what he looks for in interview packets, accelerating his review and approval of every hire. The system is “approaching indispensable.”

  • Sierra Brain supplies any agent with a 20–30-page account of the company, organization, competition, strengths, and weaknesses, plus recent board letters, operating reviews, and other beliefs about the world. As code generation improves, Bavor expects the constraint to move from writing code, to reviewing it, to deciding what could exist and editing it into what should exist.

5. Enterprise AI still scales through people embedded with customers

  • Harry contrasted Sierra with Lovable, which had reportedly—and, in Harry’s phrasing, “I think”—reached $500 million in ARR with 149 people. Bavor accepted the direction toward leaner teams but rejected a universal template: 50% of Sierra’s customers generate over $1 billion of revenue, 30% exceed $10 billion, and the company works with 40% of the Fortune 50.

  • These regulated, technologically complex organizations are “snowflakes.” Success requires understanding business outcomes, integrating into heterogeneous stacks, building relationships, and earning enough trust to become a partner—not a vendor that “throws some software over the wall.”

  • Sierra borrowed Palantir’s forward-deployed model through early design partners including Olakai, SiriusXM, Sonos, and Weight Watchers. Engineers were embedded so deeply that founding engineer Mihai effectively became a Weight Watchers employee, even receiving performance-review emails. That proximity taught Sierra what deploying customer-facing AI actually demanded.

  • Bavor rejected the categorical claim that enterprise AI cannot sell without FDEs: Sierra’s platform is transparent, exportable, and usable independently. But customer-plus-Sierra deployment has taken Next from kickoff to phone and chat production in six weeks, and Cigna live in roughly 58 days—making FDEs an important catalyst for time to impact and result quality.

6. Sierra is expanding from support into the full customer lifecycle

  • Rocket illustrates the direction: Sierra worked with Redfin to rethink home search, contacted prospective refinance customers, built Rocket Assist to shape loans and gather information, and supported later servicing. Next uses personalized recommendations to assemble outfits and larger baskets. These are inbound and outbound sales motions, not merely support automation.

  • Bavor’s platform strategy is to build applications that inform the reusable platform, making the third, fourth, and fifth applications easier. Coding agents now make a truly unique Fortune 50 requirement economical to build, though his hunch is that apparent one-offs usually recur elsewhere and become shared vertical capability.

  • Competition is the price of a giant market. Bavor says customers are “voting with their feet,” with Sierra multiples larger than its nearest similar-vintage startup competitor and growing faster; because platform breadth and industry experience compound, his hunch is an Uber/Lyft-like market rather than an AWS–Google Cloud–Azure structure, with Sierra seeking the larger pole position.

7. Six-week governance trades ceremonial boards for fast repricing

  • Sierra alternates a three-hour board meeting with a 90-minute one every six weeks because “the AI time clock” outruns quarterly governance. The purpose is to absorb new evidence, update priors, and change course while the evidence still matters.

  • After winter break, Claude 4.5 and Codex 5.2 coincided with what Bavor saw as a fundamental step change in coding-agent capability. It altered Sierra’s software-development method and core product approach—exactly the kind of discontinuity the six-week rhythm is designed to surface.

  • Instead of decks, Bavor and Bret write six-to-10-page memos: “Writing is just thinking on paper,” and writing makes weak reasoning harder to hide. In their first eight quarters, even after beating forecasts, a typical letter names roughly seven areas of dissatisfaction, giving directors time to challenge real questions instead of being presented to and managed.

  • One early admission was that Sierra saw demand but failed to recruit quickly enough in early 2024, leaving additional customers it could have served unserved. Financing follows the same milestone discipline: raise enough to reach the next unequivocally higher watermark, remain sensitive but not maximally so to dilution, and accept a lower price than Sierra could have taken; every round cleared below the available price.

8. Craftsmanship and intensity are operating systems, not slogans

  • Craftsmanship means a great company is the sum of “thousands and thousands” of individually excellent things—people, processes, culture, and product. It also signals how Sierra will treat customers’ most precious asset: during its first Black Friday/Cyber Monday, a lead engineer, operations head, and either founder personally monitored every agent conversation in real time.

  • Intensity reflects competitive reality: customer interaction through sophisticated agents feels inevitable, but no company is entitled to win. Sierra hires against the difficult Venn diagram of “smart, nice, intense,” while Bavor and Bret try to remain the pacesetters asking why something cannot happen tomorrow rather than next week.

  • Founder mode is selective “direct applied force,” not indiscriminate involvement 17 layers deep. Japan exemplifies the method: asking what would make a major business possible this year implied needing 10 people locally, which led to acquiring Opera Technologies and building around Japanese service expectations such as omotenashi.

  • The counterweight is family, one of Sierra’s explicit values. Bavor rejects the “sometimes performative grind” of startup culture: people can turn on the afterburners while protecting children, parents, friends, or whatever matters beyond work. His warning on timelines applies equally—“work is like a gas” and expands into whatever space it receives.

9. AI fluency has become an entry-level unfair advantage

  • Bavor argues young graduates have an unusual opening, not merely a displacement risk. Four years of effectively unlimited time and disposable hours can produce mastery of AI tools that 1,000 companies want; some of Sierra’s most effective employees are 22 or 23, “completely AI-pilled,” and more comfortable with the tools than experienced colleagues.

  • Sierra rebuilt its engineering interview around that reality. Candidates choose an application, receive $150 for their preferred coding agent, bring their own laptop and tools, build, and explain their process. Architecture, systems design, product thinking, values, and “smart, nice, intense” still matter; Bavor wanted every interview to gain a strong AI-native component within two months.

  • In-person work supports the apprenticeship and mentorship that Bavor believes are important for learning. He invokes Richard Hamming’s advice to “find great people, work with them, and learn from them”: knowledge and hard work compound like interest, making early exposure to excellent practitioners potentially trajectory-changing.

  • Cybersecurity looks increasingly important because offensive capability has “ratcheted up five notches.” Bavor nevertheless leaves the product thesis open: the defensive winners might be specialist vendors, or the models and coding agents themselves.

10. Cofounder complementarity converts disagreement into judgment

  • Bavor and Bret recently split over whether a slow area required stronger process or different leadership. They interrogated both views and concluded it needed some of each; their shared test is not ownership of an argument but the truth-seeking phrase, “This is correct.”

  • Instead of dividing the company, they assign majors and minors. Bret majors in sales and software engineering, with instincts on system design and architecture that Bavor trusts deeply; Bavor majors in operations, finance, legal, and running the company. Consequential contracts require both “nuclear keys,” preserving overlap where mistakes would matter most.

  • From Sundar, Bavor learned “dynamic range”—moving from five-year strategy to pixels, drop shadows, sounds, and textures without losing humanity. His broader Google lesson is that an enduring mission, smart people, a truth-seeking culture, and well-directed people caring for many experiments can make an organization feel capable of solving almost anything.