Cognition CEO Scott Wu on Acquiring Windsurf: The Process, The Deal, The Rationale
Cognition CEO Scott Wu on Acquiring Windsurf: The Process, The Deal, The Rationale
Summary
- Cognition’s Windsurf acquisition was a weekend deal built on a cold call: Scott Wu learned of the Google deal Friday “the same time everyone else did,” reached out cold that evening, got verbal agreement Saturday, hashed terms and legal Sunday, and signed by Monday morning. His model for the urgency: “the bank goes into receivership Friday night — you got to have an answer by Sunday and by Monday morning.” Consideration was a mix of stock and cash; he wouldn’t disclose the split.
- Against the “deteriorating husk” read (the scale comparison), Wu argues Google left a treasure chest behind: “the entirety of the customer book,” all the code, a lot of the data and proprietary IP, plus GTM/marketing/finance teams that complement Cognition’s engineering/product focus. On founders, he’s pointed: “there’s an unspoken covenant that as a founder you go down with the ship… it’s changed a bit over the last year and I think it’s a bit disappointing to be honest.”
- Wu defends Meta-scale talent spending as “quite reasonable” because at least 100 and certainly fewer than 10,000 people — “probably a good bit less” — are helping determine AI’s trajectory, and even frozen at today’s capabilities AI would already be “bigger than the internet.” Competition for talent is “actually greater even in the application layer” than at the foundation labs.
- On the Anthropic dependency that burned Windsurf once already, Wu deflects — value accrues “wherever you are able to establish real differentiation” — and flatly refuses to say what percent of Cognition’s revenue goes to Anthropic: “I have to pass on that one.” He describes competition in the foundation layer, including Kimi and Grok, as the natural equilibrium.
- The capability call: RL is “the biggest breakthrough of the last year and a half” — “you roughly can solve any benchmark” given clean environments and success criteria — and the data story has shifted from mass quantity to small curated sets (“a lot of compute, not a lot of data”), illustrated by Cognition’s Kevin 32B outperforming other models on CUDA kernels.
- Productivity math for the value chasm: AI makes engineers 1.5–2x faster today, “no reason that shouldn’t be a 10x” in 3 years, with 10x more code written (Jevons paradox) — and Devon prices usage “about 10x cheaper than the value of your time.” Devon usage grew 5–10x since January, and Wu claims growth roughly matching the 0-to-80-million ramp Stebbings cited for Lovable over 6–8 months, Windsurf excluded.
- The portfolio bet: private foundation labs (OpenAI, Anthropic, X, SSI, Thinking Machines, ~$500B combined) and the app layer ($50–100B: Perplexity, Sierra, Decagon, Harvey, Cognition, Cursor) both go “way up” in 5 years. Forced to pick one: “Anthropic has quadrupled their revenue since they were valued at 60 billion — seems like a fair one to pick” over OpenAI at $300B. He sees 3–5 surviving foundation players; Stebbings insists on two.
Deep dive
1. Cold call Friday night, signed Monday morning
- The deal’s origin was pure opportunism executed at bank-run speed: Cognition found out Friday “the same time everyone else did” that Google was involved and what would happen next with the team, and reached out cold to Jeff Friday evening — first call that night. Wu’s frame: “a month long is not the way to do this… even a week long” — customers scrambling, the whole team in limbo, “everyone in Silicon Valley reaching out saying hey, you want to come interview here.”
- The sequencing: “get to a verbal agreement on Saturday, hash out all the details and the terms and the legal on Sunday, get everything signed and done on Monday morning.” His analogy — “the bank goes into receivership Friday night, you got to have an answer by Sunday” — with the concession that diligence was necessarily shallow: “we are not going to have time to go as deep into the diligence and details as we would like.”
- The strategic fit as Wu tells it: Cognition is “especially focused on the core engineering and product team,” while Windsurf built “an amazing go-to-market team, marketing team, finance, operations” — plus “a very naturally complementary lean” in products. Consideration was a mixture of stock and cash; the split is undisclosed.
- Windsurf’s team had only a few days’ notice of the previous development, and Wu credits their handling: the options were operate independently, raise a fresh venture round (“now that there are actually no investors”), or find the right partner — fast in every scenario.
2. The “husk” was a treasure chest — and Google may have blundered
- Stebbings invoked a scale comparison — a husk “deteriorating in real time.” Wu’s rebuttal to the “all the best researchers left, there’s nothing left behind” discourse: “I don’t think so is basically how we felt.” What remained: “an amazing product that a lot of people use… the entirety of the customer book… all of the code, a lot of the data and the proprietary IP, and a really incredible team” — “a somewhat incomplete piece” whose missing pieces Cognition happened to hold.
- On the question (posed by, likely, Ruchi at South Park Commons) of whether Google blundered by not seeing the asset’s true value: “I think there’s some real truth to your point that often there are actually a lot of really valuable pieces that get left behind.”
- Asked whether this structure becomes the new norm, Wu turns moralist: “There’s an unspoken covenant that as a founder you go down with the ship… for better or for worse, it’s changed a bit over the last year and I think it’s a bit disappointing to be honest.”
3. The talent war is “quite reasonable” — a small number steer AI’s trajectory
- Wu’s self-described “maybe crazy, controversial opinion” on Meta’s hiring spree: it’s rational. AI is “the greatest technology shift in our lives” — and “even if you froze all the capabilities today… I think AI already would be bigger than” the internet. “And by the way, I don’t think it’s going to freeze. I think it’s going to keep moving even faster.”
- Pressed on the quantum of people who actually matter: “at least 100 folks… certainly less than 10,000, and probably a good bit less than that if I had to guess.”
- Stebbings’ follow-up — do application-layer companies even need those people? Wu’s counter: “the level of talent and the fierceness of competition is actually greater even in the application layer, which is crazy to say because the competition at the foundation labs is extremely strong.”
4. The Anthropic dependency — and the revenue question he wouldn’t touch
- Stebbings pushed on the obvious tail risk: Windsurf was reliant on Anthropic, got cut off when the OpenAI deal materialized — is the reliance now greater than ever? Wu’s answer stays structural: “the boring but true answer is [value] occurs wherever you’re able to establish real differentiation in your space.”
- The hardest question landed flat. What percent of revenue goes to Anthropic? “I have to pass on that one.”
- On whether he wants foundation-model commoditization for leverage: he describes Kimi and Grok launching “even in the last week or two” as the natural way of things — competition in both layers, and “folks in these layers want to collaborate… as long as that holds.”
- Why Anthropic doesn’t just own the category: focus. Cognition teaches a specific model “how to go to Datadog and pull up the logs, how to debug a front end live, a representation of the codebase which we’re learning and iterating over time” — versus solving for generally smarter base models. “I think the truth is you really need both.”
5. RL: “you roughly can solve any benchmark”
- Wu’s capability thesis against the self-driving-plateau analogy: RL is “the biggest breakthrough of the last year and a half,” and its converging property is stark — “if you have a clean enough set of here are exactly the behaviors that I want, here are the environments, here’s what it means to succeed or fail, you can just train a model that does that.” The open question per application is simply “what is the benchmark” — his example: for an accountant, an IRS audit finding is the RL fail signal.
- His stated change of mind: 24 months ago “the story was all about data… quantity of data”; now it’s “a small set of highly curated data for exactly the use case that you care about.” Exhibit: Kevin 32B, Cognition’s released model that used RL on CUDA-kernel agent trajectories and, Wu said, was much better than other models on the CUDA-kernel task. “It’s a lot of compute, not a lot of data… quality of data over quantity.”
- And the floor even if progress stops: “In AI code, to be truly honest, if there was zero progress, the world would still be entirely different… you are just slower as a software engineer if you’re not using AI. That is the truth.”
6. 1.5–2x today, 10x in three years — and 10x more software to write
- On claims from a Salesforce executive and Vlad at Robinhood that “50% of net new code is AI,” Wu calls the metric mushy — “if it’s just a ton of protobufs, that’s one thing, versus the core business logic” — and prefers output-per-hour: “1.5 to 2x feels right to me today in aggregate. In 3 years there’s no reason that shouldn’t be a 10x.”
- The Jevons argument, as told: he tracks every time software fails him daily. The best-made products — YouTube, TikTok, Instagram — represent “hundreds of millions of hours of engineering time,” and each order of magnitude down (your bank, your insurer) the difference shows fast. “All the software can be 10x better, and I think there actually is 10x more of it to write.”
- On Stebbings’ “value chasm” (a $300k engineer made 1.5x faster is ~$150k of value): Devon is usage-based, priced by the hour, “about 10x cheaper than the value of your time.” With 30M software engineers going 10x, whether companies capture 5% or 30% “is actually less [the point] than getting the technology to where everyone is going a lot faster.”
- The under-discussed frontier: deep context — reusing what someone asked Devon a month ago, knowing what a front end is supposed to look like, how a bug was found. “That’s the difference between code and software engineering… not the kind of thing you can solve in a sandbox.”
7. The zeitgeist gap: Devon grew 5–10x while Cursor owned the brand
- Stebbings’ bluntest push: “Devon fell out of the zeitgeist… was that a marketing failure or a product misstep?” Wu’s counter with numbers: “between January and now, the usage of Devon has actually grown something like 5 to 10x,” across both self-serve and enterprise — driven by real engineering teams tagging Devon in Slack and Linear, reviewing its pull requests in GitHub.
- On the Lovable comparison (Stebbings: “basically zero to 80 million”): “even with the Windsurf deal aside, we’ve done roughly that as well in the last 6 to 8 months.” His distinction: Replit and Lovable have “a much more consumery lean,” which is why Twitter hears about them; Devon serves engineers on engineering teams. Stebbings stands corrected — then doubles down: “you guys should be better at selling yourself.”
- Wu takes it: “Perhaps we should be. The great news is we’ve just now inherited a great marketing team.” His structural excuse: IDEs became no-brainer obvious a year before agents did; “fast forward 6 to 12 months and people will be familiar with agents to the same level.”
8. Endgame: “Tony Stark does not pull up his laptop”
- Asked whether the market splits Cursor-bottoms-up / Cognition-top-down, Wu refuses the frame: “it’s far too early to call… none of us are that close to the future of software engineering.” The real product being built is “the next generation of human-computer interface” — code is just “the language your computer happens to speak,” and eventually intent replaces it: “Tony Stark does not pull up his laptop. Tony Stark goes and talks to Jarvis.”
- His timeline arithmetic: “there are 10 or 20 levels of what the product experience looks like until we get there, and every level is like 2 or 3 months… we’re going to be there in a few years, if this pace keeps going.” The engineer becomes “a technical architect, a technical product manager”; the durable skill is “deciding what solution to build and how exactly to architect it” — which is also why he still tells students to study CS: “the degree in CS is more a degree in how to think.”
- The first thing he wants to build post-deal: the combined IDE-plus-agent experience — plan a task in the IDE using Devon’s retrieval and wiki, hand off to the agent for the bulk, review locally: “synchronous to asynchronous to synchronous.” Near-term, both products keep their philosophies. His competitive north star is a Jensen line: “when you have figured out a way for your company to win that means no one else has to lose, you have found your path” — and he “truly” believes code has more than one winner.
9. The bets: labs at ~$500B go “way up” — and he’d pick Anthropic at $60B
- Channeling Sam Altman’s decade-old “bubble theory” post (bet against the bubble-callers, with money), Wu offers his own: take the private foundation labs — OpenAI, Anthropic, X, SSI, Thinking Machines — “what is that, like 500B? I think that’s going to go way up in the next 5 years.” Same for the application layer (Perplexity, Sierra, Decagon, Harvey, Cognition, Cursor — “probably worth 50 to 100 billion”): way up in aggregate, “obviously it doesn’t mean every single company.”
- Forced to choose OpenAI at $300B or Anthropic at $60B, he first ducks — “I think both are great investments at the price” — then commits: “Anthropic has quadrupled their revenue since they were valued at 60 billion. So it seems like a fair one to pick.”
- On consolidation, Wu hedges to “two to six players… three to five seems pretty reasonable.” Stebbings’ pushback — worth keeping: just two, “ChatGPT and OpenAI on consumer, Anthropic on enterprise,” everyone else “the duck duck go of AI.” Wu half-concedes on consumer share (“it is always a power law — number one is 75%, number two 20%, number three 4%”) but holds that enterprise choice-seeking keeps a few players capability-competitive.