Pioneers Insight Method Research Author
Cohere Founder, Nick Frosst: How To Compete with OpenAI & Anthropic, and Sam Altman’s AI Disservice
Back to Episodes

Cohere Founder, Nick Frosst: How To Compete with OpenAI & Anthropic, and Sam Altman’s AI Disservice

Summary

  • Nick Frosst’s central call is that AGI timeline and existential-threat rhetoric are misleading, not evidence of imminent AGI: “I don’t think Sam Altman has done a service to the world by talking about how close AGI is… he has made several predictions now that are wrong and that were obviously wrong at the time he made them.” Altman’s world tour warning leaders of existential threat was “academically disingenuous and I think did a disservice to the technology he loves.” Harry’s sharper version — doom rhetoric correlates with funding need — gets a qualified reply: “the correlation you’re pointing out exists,” but Frosst says he does not know if that was the strategy, adding: “we’re a venture capital funded company and we need funding. And we don’t say that.”
  • Frosst is skeptical that more compute alone guarantees continued progress. When Harry asks how much better GPT-5 was than GPT-4, Frosst answers, “I actually think it was worse”; Harry then says that tells you something about throwing more compute at the problem. Frosst attributes the product problem to slower, cumbersome model auto-selection. His AGI definition — “a computer that you treat like a person” — is not reached by current language-model technology: “I don’t think this technology gets us there.”
  • Cohere’s counter-positioning is capital efficiency plus enterprise focus: fewer than 20 LLM companies exist worldwide; the fundraising discussion references roughly $600M and a $6.8B price; Cohere has spent “truly orders of magnitude less” than rivals, and trains its Command A models to fit on two GPUs because enterprise customers are bottlenecked on GPU access. The mantra: “ROI not AGI.”
  • On the labor question the two openly clash — the best exchange of the episode. Harry: 25-26-year-old marketing managers and SDRs “are not brilliant… They will be replaced” within 12 months. Frosst’s rebuttal is structural, not hopeful: “There has been no independent breakthrough that an LLM has made… The breakthroughs are still people” — and no, it’s not a matter of time, because “that’s fundamentally the way that sequence models work.”
  • Benchmarks are not a reliable measure of enterprise utility: “They’re a reflection of how much the model have been trained on those benchmarks.” A math-reasoning benchmark and ARC-AGI (“a pixel manipulation challenge”) map to nothing enterprise customers ask for; the leaderboard cycle matters for consumer hype, not for “did I buy LLMs, deploy them and then get ROI.”
  • Model sovereignty is an investable theme: language models are infrastructure “like power plants,” countries should have their own infrastructure and models, and America “has shown that they’re willing to turn off access to tech based on political reasons” (he cites a 10% Intel stake). Being Canadian is now a commercial asset for Cohere — and his pick for a trillion-dollar AI company outside the US: “maybe north of the border.”
  • Boldest 2026 prediction is deliberately unglamorous: you’ll open North and say “file my expenses” and the model will execute the full multi-step workflow — “I know that’s not very bold… and yet getting it to actually work and be a thing you can rely on is still not in every company.” On China: seven new model providers released seven models in one week about a month earlier; that made him say “wow,” but “I don’t think they’ve made models that are beating the other models out there.”

Deep dive

1. The AGI indictment: hype was “obviously wrong at the time”

  • The episode’s spine is Frosst’s charge about likely Sam Altman: “I don’t think Sam Altman has done a service to the world by talking about how close AGI is. I think he has made several predictions now that are wrong and that were obviously wrong at the time he made them” — including allusions that “AI is going to kill the whole world in two years.”
  • The world tour is his best specimen: Altman “spoke to every major leader the world over to tell them, hey, this technology poses an existential threat. And I think that was academically disingenuous and I think did a disservice to the technology he loves.”
  • Harry’s pushback is the sharpest moment: doom rhetoric correlates with funding need — he says Demis and Zuck did not need dollars and stayed measured, while “other people who do need funding have to be much more provocative.” Frosst concedes “the correlation you’re pointing out exists,” but says he does not know whether that was the strategy. He lands his counter: “we’re a venture capital funded company and we need funding. And we don’t say that.”
  • Harry then says even the leaders he names have changed their rhetoric about the coming labor changes, “which does make me worry.” Frosst grants the technology is “fundamentally transformative the same way the personal computer” was, but insists “there’s also a whole lot of not legitimate things people are spending their time on.”

2. More compute alone isn’t enough — and GPT-5 is the exhibit

  • Harry asks how much better GPT-5 was than GPT-4. Frosst answers, “I actually think it was worse.” Harry then says that tells you something about the nature of just throwing more compute at the problem. Frosst explains that model auto-selection is “slower and more cumbersome… I just want a quick answer and it suddenly goes into deep research. All right, PhD, calm down.”
  • His AGI definition is behavioral, not benchmark-based: “a computer that you treat like a person… people do not treat language models like they treat people.” And the categorical call: “I don’t think this technology gets us there.” He claims the consensus is quietly with him — ask university CS students whether more compute gets to AGI and “most of them say no.”
  • Crucially, he distinguishes plateau from progress: the work of making a model file your expenses — read emails, find receipts, cross-reference policy, hit the internal API, get approval — “is not plateauing. That’s more modeling work, more product work… building better connectors.” The architecture, though, is stagnant: “we’re approaching 10 years of the same model architecture,” still transformers from 2017, still next-word prediction with new training steps added.

3. Cohere’s wedge: one of fewer than 20 companies, with enterprise focus

  • The market structure as Frosst counts it: “some number less than 20 in the whole world” building LLMs — most in America, a handful in China, Cohere in Canada, one in France. Cohere is distinctive in its “singular focus on bringing this technology to enterprise” — no consumer app, no engagement metrics, “we’re not trying to get anybody to spend $200 a month on something for their personal lives.”
  • Enterprise focus changes the training data, not the architecture: Cohere now generates synthetic enterprise environments — “fake companies and fake emails between people at these fake companies and fake APIs within those fake companies” — and trains models to be useful inside them. But data remains a bottleneck: “you need real world data in order to start a process of synthetic data,” and Cohere still makes data in-house with annotators. Between compute, algorithms, and data, algorithms are least constrained — quality data and quality synthetic data derived from it is the gate.
  • His personal-life framing explains the strategy: “I actually don’t want to respond to text messages from my mom faster… whereas in my work life there’s a ton of stuff I don’t want to do.”

4. Benchmarks measure training on benchmarks

  • Frosst’s history lesson does the work: LM1B (predict the rest of a newspaper article), then Hello Swag circa 2022 — “no one’s talking about that anymore” — now a math-reasoning benchmark and ARC-AGI. “None of our customers ask the model to do math reasoning… ARC-AGI is like a pixel manipulation challenge. That’s not a thing any of our customers have ever asked the model to do, nor do I think they will.”
  • The verdict, delivered flat: benchmarks are “not an accurate reflection of the utility value of models. They’re a reflection of how much the model have been trained on those benchmarks.” Can they be gamified? “Oh, you can definitely gamify them.” Do the big players do it? “I don’t know” — a hedge he leaves hanging.
  • The enterprise substitute metric: “did I get to production? Did I buy LLMs, deploy them and then get ROI on that?” Leaderboards are “cool” for consumer apps; his customers don’t care.

5. Models and apps converge — but on a spectrum, not at the extremes

  • On whether foundation labs stay AWS-style commodity layers or eat the application layer, Frosst rejects the dichotomy: “if you want to make the best model for a given interface, it’s best to be training the model on that interface” — which is why Anthropic challenges Cursor and why Cohere builds North.
  • The load-bearing insight is the spectrum: 2015-era ML meant one model per task (train a cat-counter on cat pictures); transformers are “a little over here” — “train a model that is generally good at all language and refine it on the type of stuff you want to do with it.” Anthropic’s models are generically good at code, not split into refactoring and debugging models; Cohere’s are refined for enterprise tool use and massive documentation.
  • On the “your evals are worse” criticism, he’s blunt: “what we care about is if a customer gets a copy of our model and they try to do something with it, we care that it works as easy as possible… None of those are reflected in the various benchmarks that cycle through every year.”

6. The labor clash: “they will be replaced” vs. “the breakthroughs are still people”

  • Harry, invoking a Salesforce guest’s “same human plus agent” line from two days earlier, goes for the jugular: “Most 25, 26-year-old marketing managers or SDRs — I’m sorry to say it. They’re not brilliant. They do not love the craft. They are not better than a phenomenal agent will be in the next 12 months. They will be replaced.”
  • Frosst’s rebuttal is his most quotable technical claim: “There has been no independent breakthrough that an LLM has made… nobody asked an LLM, hey, solve this problem no one’s solved before and got the answer. The breakthroughs are still people.” Is that just a matter of time? “No, it’s not. That’s fundamentally the way that sequence models work” — they’re statistical models of text, and the marketer’s real work — “understanding the culture, understanding the zeitgeist… using their intuition” — “is not in the data set of text from the internet.”
  • Harry’s Evian-campaign counterexample (prompt for three storylines, pick the best) gets absorbed rather than refuted: “that’s a good use case… but that’s not where the work ends. That’s the beginning of the work. You’re now starting with things to go for instead of a blank page.”
  • Frosst’s 5-10-year company: you sit at a computer, use language to offload anything where “the information is out there, doesn’t require creativity or insight… and doing it is kind of boring,” and spend your time talking to people and judging whether what the model did was good.

7. The industrial revolution frame: policy decides whether AI helps or hurts inequality

  • His answer to “does AI help or hurt income inequality” is conditional, exactly as hedged: “it depends on policy. If there’s good labor policy, I think it could help. If there’s bad labor policy, it could hurt.” The precedent: everyone now endorses the industrial revolution (farm labor from ~90% to under 5%), but it produced kids in coal mines before it produced unions and workers’ rights — “a lot of those were from public policy… created in unison between businesses and governments.”
  • This connects back to his Altman critique: existential-threat discourse “made it harder to talk about the real things, you know, like income inequality” — which was already rising before language models were popular, and which he worries the technology “has the potential to exacerbate without being deployed correctly.”
  • The Adam Smith detour is a genuine disagreement: Harry argues the invisible hand prices Frosst’s old minimum-wage grill job as “definitively less valuable” than Cohere’s work touching millions of Fujitsu users. Frosst holds the middle: the economy is “decent at figuring that out… but I don’t think it’s perfect” — and running for extra potatoes in a broken-AC kitchen “was challenging and rewarding work.”

8. The two-GPU business model: “not AGI, ROI”

  • The fundraising discussion references roughly $600M. Harry floats a last-round price of “like 6.7 or something”; Frosst corrects it to $6.8B. Cohere has spent “truly orders of magnitude less on creating foundational models than some of the other foundational model companies” — a lineage that goes back to founding-era papers on training with “the scraps of GPUs in data centers.”
  • The tactical core: Command A, the just-released Command A reasoning model, and Command A vision are all “trained to fit on two GPUs” — because enterprises “were bottlenecked on deploying because they don’t have enough GPUs,” and two turns out to be the sweet spot of performance, cost, and actual GPU access. Consumer companies can “be losing a ton of money on every inference call” because engagement pays; enterprise can’t.
  • Distribution follows: forward deployed engineers are “a crucial component” of getting customers to production (a good idea, though “I don’t know if that’s true for every business”), reference customers are RBC, Fujitsu, and LG, and the weights are released for non-commercial use — a middle ground that builds credibility (“you can validate, hey, do they work on my problem?”) while forcing commercial users into a paid relationship. “I’m surprised there aren’t more foundational models taking that approach.”
  • On talent-war headlines ($100M Meta packages), studied skepticism: “I read as many stories of those as I read of people leaving the next day.” He confirms hiring Joelle Pineau (likely) from Meta, and would pay $5M for a researcher “if they were bringing in the right value” — but people stay for “stability… purpose and value alignment.”

9. Sovereignty: models are power plants, and Canada is now a selling point

  • The infrastructure analogy carries the whole argument: “having a language model that speaks the language of your country is like building infrastructure for the people of your country… I like that Canada has several nuclear power plants.” Frosst says countries having their own model infrastructure is a good idea — using a model built in China or America “might not set your country and your economy up as well” as one with the language, dialect, and cultural fluency.
  • The geopolitical kicker: “America has shown that they’re willing to turn off access to tech based on political reasons… the connections between American tech and the American government is less clear as time goes on” — citing a 10% stake in Intel. Result: “there’s a lot of companies in Canada and around the world interested in working with non-American tech companies, and I would say that’s been an asset for us.”
  • His regulation nightmare mirrors his benchmark critique: the worst move would be regulators who think “what we’re building is digital gods” picking “this random benchmark we think represents AGI” — gameable in either direction — and shutting down development on it.

10. Quick-fire: the lapsed optimist, China, and two junior chickens

  • Frosst no longer calls himself a technological optimist — hasn’t for 10 years. The conversion story, as told: he loved Google Glass, then “I got on a bus one time and somebody was wearing one… everybody clocked it immediately.” On VR: “I actually don’t want to strap a computer to my face. I want to be engaged in the world more.” But he holds Harry’s loneliness worry alongside the Greek-philosophers-bemoaning-writing pattern: “you have to hold both of those two conflicting views in your mind.”
  • China: not worried. “They’ve made good models… but I don’t think they’ve made models that are beating the other models out there” — though seven new model providers released seven models in one week about a month earlier, drawing a genuine “Wow.” Boldest 2026 prediction: “file my expenses” actually works end-to-end — “I know that’s not very bold… and yet that becoming a ubiquitous way of using a computer is crazy.” Trillion-dollar AI company outside the US? “Maybe north of the border.” If not Cohere: Google DeepMind.
  • On being wrong: he once believed human progress was monotonic — “life expectancy will go up, income inequality will go down… that’s not true.” And a technical mea culpa from 2020: “I remember being like, you can’t make a small data set of feedback from people and make a model better” — RLHF’s data efficiency proved him wrong. M&A offers have come (“we’ve been a company for 5 years — yeah, we have”) but the pull is “building something that outlasts us,” even knowing the Ozymandias ending: “one day all that will be left are two legs in the desert.”
  • The ritual, confirmed: after signing every round — five now, maybe four — the three co-founders go to McDonald’s. “I normally get two junior chickens.”