Pioneers Insight Method Research Author
The Open-Source AI Reality | How Token Costs Will Fall 10X & Usage Will Explode 100X | Lin Qiao
Back to Episodes

The Open-Source AI Reality | How Token Costs Will Fall 10X & Usage Will Explode 100X | Lin Qiao

Summary

  • Lin Qiao’s founding thesis is a direct challenge to AGI maximalism: “if you think intelligence is a derivative of data,” the majority of the world’s data is private, “locked inside enterprise,” and will never be shared — so the frontier is specialized, private intelligence, and the endgame is “millions of specialized models — one per application, per use case,” not one model that rules everything. Fireworks already processes 40+ trillion tokens a day, the majority from customized models, not off-the-shelf ones.
  • On Harry’s are-OpenAI-and-Anthropic-overvalued question, Lin reframes frontier labs as “power lines” — essential infrastructure that won’t replace what’s built on top — and leaves the investor question standing: are power lines good businesses when open and closed models have both “crossed the quality threshold” and open weights can be tuned on small proprietary data to beat general models on your eval? “What I don’t want to see is there’s only one company owns intelligence.”
  • The headline economics call: token costs haven’t fallen yet “because of supply chain constraint,” but competition will deliver 10x cost reduction in the next three years, driving 100x usage — and Fireworks’ own token count could go “anywhere ranging from 20 to 100x” by end of next year. “We’re at a very early stage of S-curve explosion.” Harry’s inference — a capex bubble thesis is “ridiculous” — she endorses, with the caveat that the real bottleneck is the bottom of Jensen’s five-layer cake: energy, chips, “we’re bottlenecked by small parts — transistors.”
  • “Scaling to bankruptcy” is the mechanism pushing enterprises to open weights: PMF and durable business have decoupled because inference isn’t a commodity like CPU was in SaaS — incumbents with huge traffic can’t get AI features past the CFO, so they must own and tune their own models. Her three-year non-consensus: “every single company will own their own intelligence as a must-have. It’s not optional.”
  • On app companies training their own models (Harvey vs Legora, where Harry disclosed his Legora position): Cursor pioneered tuning and “now almost all coding companies tune their own models” — the tool-calling harness “needs to be co-trained with the model powering it.” Harry said Fireworks’ CTO Dima was embedded at Cursor for months; Lin described decoupled RL infrastructure running across five-six data center regions on scattered GPUs instead of an InfiniBand supercluster.
  • The business: $800M in AR, expecting to “at least double” by year-end with just 200 people. The 30-40% gross margin (vs SaaS’s 80%) is framed as a hyper-growth choice, not the new normal — “constraints slow down innovation.” Hard line: “we absolutely are not going to move into application layer”; data centers are “always on the table — the question is timing”; chips are out because workloads are too dynamic and hardware depreciation math has broken (“within a year from one vendor alone we have three SKUs”).
  • Chinese open models (the top six on OpenRouter) get a pragmatic answer: guardrail every model, open or closed, because every provider “will infuse their own judgment, their own taste” into training. If China restricts access, it’s “a big impact in the short term” — but the US “will be able to build that open system by ourselves, and we should” (Nvidia is training a model likely called Nemotron amid that supply-chain gap).

Deep dive

1. The founding thesis: most of the world’s data will never touch a frontier model

  • Lin’s core argument, in response to Harry’s question: “If you think intelligence is a derivative of data, then the majority of the data is actually not used for training a general intelligence model.” Training corpora are public internet plus labels — “a very small corpus” against the world’s data, most of which is “private, locked inside applications, locked inside enterprise — it will never get shared with anyone else because this is the company’s proprietary IP.” Fireworks exists to activate it: “the frontier of the intelligence is actually private intelligence, specialized intelligence.”
  • The backstory explains the conviction: PhD in distributed systems, LinkedIn, and by 2015 a business proposal and a co-founder list — but she paused because “I don’t think I have the skill set on people.” She joined Facebook “secretly planning to learn for one year or two and leave”; she stayed seven, then founded Fireworks at 48. Eric Vishria broke his no-big-tech-directors rule after his advisor warned “how many big tech executives have you seen being successful? Very few.”
  • Harry’s own marker on the company: a $10 million check after a 15-minute meeting — “one of the easiest investment decisions that I’ve made in a 10-year investing career.”

2. AGI maximalism vs “an army of robots”

  • Harry’s sharpest early push: isn’t activating private enterprise data exactly Anthropic’s thesis with Claude co-work? Lin’s reframe — Anthropic “fully believes in AGI,” defined as one model that solves all problems, which by definition means never specializing. Her rebuttal is civilizational, not technical: regions differ in values, policy, taste, and “if our future world is going to be ruled by one standard, a taste dictated by one company, we turn ourselves into an army of robots.”
  • The line she keeps returning to came from Jensen Huang after his GTC keynote: “There’s no specialized general company” — every company exists on a unique belief, baked into product design, data, and understanding of user intent, “not learnable or assured by another company sitting outside.”
  • So why do Dario, Sam, Larry and Sergey talk AGI as inevitable? Lin’s answer: they’re “building power lines to distribute a really great source of intelligence” — vital, like electricity enabling her beloved coffee machine — “but is this power line going to replace everything we do? I don’t think so.” Harry’s investor translation, left deliberately open: “are power lines good businesses?” On government stakes (Sam’s 5% offer to the administration), she cites the PG&E precedent but lands on: “what I don’t want to see is there’s only one company owns intelligence.”

3. Open source crossed the threshold — and “scaling to bankruptcy” pushes everyone there

  • The founding bet — build on open models when they were “almost at infancy” — came from PyTorch roots: “openness gave control.” It paid off twice: both open and closed models “crossed a quality threshold” solving real problems, and open models became easy to steer — with a small amount of unique company data “you hill-climb towards your eval” and “often the end result is: to solve your unique problem with your data, you are better than a general purpose model.” Fireworks itself runs open models for recruiting, finance, and internal coding agents.
  • Her signature framing of the economic forcing function: “Have you heard about scaling to bankruptcy?” In SaaS, product-market fit and durable business were equivalent because CPU was a commodity; now they’re separate concepts. Startups with real PMF “could scale into bankruptcy” — and it’s worse for digital-native incumbents whose CFOs can’t justify rolling AI features out to a decade of accumulated traffic. The alternative: “have control over your open weights model.”
  • Harry’s counter: doesn’t Sam Altman’s newly released, dramatically cheaper models solve this? “It could be,” she concedes, but open weights have no acquisition cost while frontier labs must recoup R&D — and “you just cannot customize those general purpose models… with open model you have full control.” At billions of users, “even 5% of cost reduction means a lot — let alone what we have seen in the past, five times to ten times.”

4. Chinese models, sovereignty, and the millions-of-models future

  • On the national-security question — “the top six models on OpenRouter today are Chinese” — Lin refuses the framing: guardrail every model, open or closed, because any provider “will infuse their own judgment, their own taste into the model training process; you cannot guarantee it matches yours.” Then the episode’s biggest call: “It may be scary, but I think that’s true. It will be millions of specialized models — one per application, per use case.”
  • If China restricts open-model access (reports surfaced last week): “a big impact in the short term,” but open ecosystems attract many parties — “in terms of talent density and resources, I do believe US will be able to build that open system by ourselves, and we should.” Nvidia is training a model likely called Nemotron amid this supply-chain gap: a missing US-native open model “is a supply chain problem.”
  • On sovereign AI, Harry cites Fable being briefly banned by the administration for 19 days as Europe’s wake-up call. Lin extends the power-line metaphor: “every country should own their own power line… you don’t want any single person to cut you off. That’s an extremely scary moment” — and the same holds for every company.

5. Should app companies train their own models? Harvey, Legora, and the Cursor precedent

  • Harry disclosed his Legora position before asking: Harvey committed to its own model, Legora didn’t — a year ago the non-builders looked right, “now it looks like they’re wrong.” Lin’s context: the application lifecycle has collapsed from “tens of very strong product engineers and PMs, multiple quarters” to “one person, a few weeks” — so implementation is no longer the moat, and competition has moved elsewhere.
  • Harry’s pushback, worth keeping: enterprise legal is multi-year relationship sales with bespoke deployment — “it’s not like 11 Labs where you pick it up and go.” Lin concedes legal is uniquely unforgiving (“lawyers are usually more conservative… legal is not tolerant at all on errors — that’s why lawyers get paid”) but insists both firms own proprietary workflow knowledge, and the orchestration harness deciding which tools to call “needs to be co-trained with the model powering it.”
  • The proof point: “in the coding space, Cursor probably is one of the pioneers” tuning their own model — “now almost all coding companies tune their own models.” Maybe Harvey-vs-Legora is just a timing question.

6. Inside the Cursor build: distributed RL on scattered GPUs

  • Harry said Fireworks’ CTO Dima was embedded at Cursor for months building RL infrastructure. Both companies being “capital conscious,” they broke reinforcement learning into trainer and rollout — instead of hyperscaler-style 10,000-100,000 chips on InfiniBand (“extremely expensive and really hard to find”), they run “fully distributed across five, six data center regions globally, tapping into scattered GPUs.” The hard part is syncing fresh weights across regions so rewards don’t go stale: “if it’s too stale, then you are too off.”
  • Her adoption-curve frame for why this partnership matters: early adopters are hackers — “they have researchers from frontier labs and they want to control every single thing” — while the late-stage mass market needs little control. Fireworks targets the latter, but partnering with the pioneer teaches “what is required to get there.”
  • On Cursor concentration risk after the SpaceX acquisition, an honest answer: “everyone’s concerned — the whole entire industry” is shaped by a handful of escape-velocity apps, and “all model companies were concentrated on Cursor.” Since then: “last year is the year of coding, and this year is the year of co-work” — general-purpose deep research plus legal, finance, customer support, recruiting, healthcare — and an uptick of consumer companies rethinking recommendation systems with genAI.

7. 40 trillion tokens a day — and why the capex-bubble thesis fails

  • The numbers: “we today process more than 40 trillion tokens a day,” the majority from customized models. End of next year? “Anywhere ranging from 20 to 100x could be possible… we’re at a very early stage of S-curve explosion.” Harry draws the conclusion — then the capex bubble idea is ridiculous — and she agrees, redirecting to Jensen’s five-layer cake: “we are bottlenecked by the lower part” — energy, chips, manufacturing lines never designed for 100x scaling. “We’re bottlenecked by small parts — transistors.”
  • Why haven’t token costs fallen yet, per Harry’s challenge? “Supply chain constraint — but we are living in a free economy”: shortage invites competition, competition compresses cost. Her prediction: “10x cost reduction in the next three years, and this 10x cost reduction will drive a 100x usage.” (The demand anchor Harry offers: an investor says Salesforce spends ~3.8% of developer salaries on Anthropic and Claude Code.)
  • The nuance she insists on: “not all tokens are equal” — evaluate token economy per task, since a model 2x cheaper but 2x more verbose solves the same task at the same cost. As quality improves, “being precise is going to be part of the optimization.”
  • The under-discussed bottleneck: “we don’t have a great system designed for 10-trillion-parameter models today” — solving it needs co-design from model through serving platform down to systems of chips, and that’s where she sees the remaining infrastructure innovation.

8. The business: premium on quality, 30-40% margins, and where the stack ends

  • On Together being cheaper: “we’re probably not comparing apples to apples” — most Fireworks traffic is customized models, optimized quality-first to the point of “zero KLD”: bit-equivalence between training and inference systems, “we do not lose a bit of accuracy.” Otherwise “you pay your training investment by discounted quality — why do you do that?” One-size-fits-one deployment, backed by an applied-ML (FDE-style) team now building agents to automate deployments.
  • On 30-40% gross margins vs SaaS’s 80%: “I don’t think that’s the new normal” — it reflects hyper-growth. “Margin optimization is a constraint problem… and constraints slow down innovation”; you optimize the heck out of a system only once you know you’ll scale it a thousand-x. The disaster case: hill-climb margin to a high number and stop growing. Hard boundaries: “we absolutely are not going to move into application layer”; data centers “could always be on the table, but the question is timing.”
  • Chips are categorically different: “I know building a chip is extremely hard.” Meta’s MTIA dates to ~2018 and serves ranking/recommendation workloads; you tape out only “once your workload stabilized,” and today’s AI workloads are “very, very dynamic.” Pre-genAI accelerator startups were serendipity bets — the SRAM-heavy designs happened to fit memory-hungry models (she notes Nvidia’s recent acquisition of Groq, whose SRAM-intense chips pair naturally with flops-intense GPUs: prefill on one, generation on the other — a heterogeneous data-center design she finds genuinely interesting).
  • Depreciation math is breaking build-vs-buy: hardware used to depreciate over six years against three-year release cycles; “now within a year from one vendor alone we have three SKUs,” models peak weekly, and “the newer model usually runs the best on the newest hardware.” After three years, “do you still want nine-generations-older hardware running a three-year-old model? That’s questionable.”

9. $800M in AR doubling, George Hu, and the Jensen operating system

  • The trajectory: $800M in AR, “we think we can at least double” by year-end, at 200 people (50 a year ago). The formative customer story: Cursor signed when they were “single-digit million dollars… only two years ago — they grew by 100 to a thousand-x over two years,” choosing a partner for platform R&D and focusing on product. Harry’s context for the era: Slack’s 1-to-10M in 18 months was once venture’s golden child.
  • On hiring ex-Salesforce president George Hu: a year ago she told him “we’re probably too small for you” — he spent that year helping interview executives before joining as growth made it serious. Her people filter isn’t competence: “whether they are really built for extreme ownership… we are not putting people in boxes and stacking the boxes into a tower.”
  • The Jensen lesson, from his one-minute email replies: “Leadership is just judgment. It’s not privilege.” In a high-velocity space you can’t wait for information to cascade through layers — “not knowing what exactly is happening and having the position of making judgment makes bad leadership.” Her own change of mind this year: shedding the fear that fast headcount growth kills agility. Her admitted mistake: waiting on marketing — “marketing is not about flows, it’s about education… clarity.”
  • The closing three-year call: “every single company will own their own intelligence as a must-have. It’s not optional” — the same logic by which every company owns its software stack. And the industry’s next phase: “token maxing is just a thing in time — we’ll quickly move into ROI maxing, which is all about running a business.”