Mike Krieger, Instagram CoFounder & Anthropic CPO: Where Will Value Be Created in an AI World?|E1265
Mike Krieger, Instagram CoFounder & Anthropic CPO: Where Will Value Be Created in an AI World?|E1265
Summary
- Krieger’s core answer to “where will value be created”: companies with differentiated go-to-market, differentiated domain knowledge, or proprietary data — “ideally two or even three of those” — in sectors like finance, legal, and healthcare, where the unsexy upfront legwork is exactly what makes the position durable. The winning loop is to sell into places you uniquely understand “and then get better for being deployed there over time.”
- His contrarian call on commoditization: “I think models over time get more different rather than more similar” — “there is something claudy about Claude and there is something GPT about GPT.” Model-layer moats are three: talent density, model character deepened by traction (coding traction feeds the next generation of RL), and being “an AI partner, not just AI models.” Fail on any one and “I think you’re in trouble.”
- DeepSeek had “almost no impact” on Anthropic’s go-to-market — enterprise relationships aren’t swapping input tokens for output tokens — but it was a marketing and shipping wake-up: it went from unknown to “in many circles better known than Claude” (“likely my great-aunt was calling me about DeepSeek”), and being late to first-party product hurt “significantly.”
- For startups riding model improvements: don’t wait. “My startup was not a startup until Claude 3.5 Sonnet” is a pattern he hears from multiple people; the winners of each model-generation shift are the ones already beating against the wall — Cursor iterated repeatedly before breaking through — “when the model arrives you’re not starting from square zero.”
- The biggest blocker to progress isn’t compute, data, or algorithms — it’s training environments and evals that match real multi-step work. SWE-bench undersells what a software engineer does; nobody evaluates office professionals well; the missing eval is “I show up to a new job, quickly understand my role, who is who” — the gap between models “extremely good at extreme slices” and generally helpful collaborators.
- Software engineers become delegators and code reviewers “a year from now,” not three — agents that try three approaches in a browser and run vulnerability tests before asking one question. But “figuring out what to build is still the hardest part,” at least three years from being solved — which is why he’s “really bullish” on startups, where alignment is “a coffee conversation.”
- The most sobering line for anyone underwriting AI app adoption: Anthropic’s own traction “is ahead of their actual true product-market fit because they are still the best ways of getting the models — I don’t think that’s durable over time.” And on usage broadly: “we are still in day one around is AI an indispensable part of most people’s work — and I think the answer is no.”
- Underappreciated risk at the coming agent-to-agent intersection: discernment plus privacy — his 5-year-old metaphor, a child who can’t yet distinguish family secrets from checkout-aisle chat. “Models fundamentally want to be helpful and that is not always what you want them to be.”
Deep dive
1. Value accrues to differentiated GTM × domain knowledge × proprietary data
- Krieger fields “what can I build that won’t be in the lane of an Anthropic” from entrepreneurs often. His answer: places with differentiated go-to-market, differentiated industry knowledge, or data only you have — “ideally two or even three of those” — in finance, legal, healthcare. Healthcare is “a tremendously complex ball of yarn” whose upfront legwork can’t be done in an accelerator, “and it is the legwork that you’ve put in” that makes it durable.
- The durable loop: pull in what’s great from foundation models, fine-tune if needed, but win by selling into places you uniquely understand “and then get better for being deployed there over time.”
- Incumbents vs. net-new startups: both can win, with inverted risks. Startups get license to overpromise — early adopters are “kicking your tires” — while an incumbent that announces “we’ve added AI” and delivers “you said it could do these 30 things, it does like two of them well” breaks trust.
- Startups’ problem is the mirror image: no data, no relationships yet. Their differentiation “is not the established relationships, it’s painting the future” and landing lighthouse customers willing to take that bet.
2. Don’t wait for the perfect model — be the one beating against the wall
- Krieger hears a recurring line from founders: “my startup was not a startup until Claude 3.5 Sonnet” (or the second 3.5 Sonnet) — a generational leap taking accuracy “from 95 to 99, or from 70 to 90,” suddenly clearing an industry’s bar.
- His best specimen is Cursor: someone showed him the founders’ Hacker News front-page submissions over time — it finally broke through, “but that was not their first product or their first iteration.” The companies that benefit from model shifts aren’t the ones that start that day; they’re the ones with accumulated context about what goes wrong in the space, so “when the model arrives you’re not starting from square zero.”
- The succinct version: “be frustrated by the current generation of the models and then be very aggressively trying the next one so that you can finally deliver on the thing that you saw in your head.”
3. Three moats at the model layer — miss one and “you’re in trouble”
- Harry’s challenge: with releases coming this thick, is there value in the model layer at all? Krieger’s three: first, talent — “talent begets talent” around a cohesive mission; Anthropic’s research team lands “some new significant hire” monthly, but people are free agents, so the attractor must be maintained.
- Second, divergence: “I think models over time get more different rather than more similar — there is something claudy about Claude and there is something GPT about GPT.” Coding wasn’t an accident, and it compounds: seeing companies rely on Claude for code “inspires the next generation of what you want to do from a reinforcement learning perspective.”
- Third, partnership. DeepSeek’s go-to-market impact was “almost no impact” because enterprise relationships aren’t “they send it for the API, they want to just exchange their input tokens for output tokens at some rate” — it’s “I want to be your long-term AI partner, I want to help co-design products with your applied AI team.”
- The failure mode, inverted: resting on laurels, believing incremental benchmark gains are enough, “treating the API as just a way of exchanging money for intelligence.”
4. The biggest blocker isn’t compute — it’s environments and evals that match real work
- Where Alex Wang and Groq’s Jonathan Ross give Harry different answers, Krieger’s is neither compute nor data: it’s getting training environments to match real-world, non-single-shot challenges. SWE-bench undersells the job — a software engineer understands requirements, negotiates timelines with PMs, ships and iterates: “there’s no eval for that.”
- Office professionals — a use case Anthropic thinks about heavily — “nobody’s really evaluating that well.” The missing environment: “I show up to a new job, I quickly understand what my role is, who is who in the organization… and then be in the run loop of the business.” That’s the blocker between models “extremely good at extreme slices” and generally helpful collaborators.
- On synthetic vs. human data: it “absolutely has to be a mix” — seed with original human data, then generate synthetic environments to explore. Claude playing Pokémon is his example of many runs through one game; it gets much harder “when the problem space is less well-defined than did you make it out of Viridian Forest.”
- The underappreciated data problem is vibes: character has no regression testing. Going from Claude 3.5 to 3.7, “people will say oh, Claude seems friendlier but more terse… I wish it was better at creative writing — these things are not easily evaluable.”
5. Today’s AI products are “extraordinarily leaky abstractions”
- Harry’s bet: in three to five years you won’t select models any more than you select “which Google you use.” Krieger agrees — current AI product design is “an extraordinarily leaky abstraction”: “why should you choose Opus, Haiku or Sonnet? Most people don’t understand the difference… we suffer from this problem as well.”
- Memory is the second leak — his coworker analogy: you might have different email threads, “but it’s still one coworker behind all of that” who doesn’t forget your favorite sports team between conversations. Third is prompting, which should become “absolutely transparent”; the good-vs-bad prompter gap “closes generation to generation, but we need to collapse it even further.”
- Model quality and UX “can’t be separated anymore”: you’re designing “a scaffold and a product around a fundamentally non-deterministic system,” where whether Claude asks follow-up questions or reasons longer is a product decision. Without regression-tested evals, a product can degrade and you can’t tell “is it the model, is it the product design… the system prompt got longer” — “in many ways the most complex product development work I’ll ever do.”
- Ship cadence now varies by surface: the API demands predictability (prompt caching launched behind an opt-in beta header), consumer tolerates experimentation, and enterprise AI “is still an early adopter product,” so Anthropic ships far faster than Salesforce’s two-to-three releases a year — but “it’s an active topic of conversation.”
6. Product marketing as Crossy Road — and why nobody switches on evals
- At Instagram the big rocks were known (don’t launch WWDC week); now “it reminds me a little bit of Crossy Road — the car’s going by, all right, there’s a gap.” Claude 3.7 Sonnet launched Monday; the blog post was locked Sunday 9 p.m. and press briefed that Sunday. The comparison table included Grok 3, released just a week prior.
- The psychology he coaches: “it’s so over / we’re so back — that is, you have to live that in AI.” Sometimes state-of-the-art lasts two or three months, sometimes a week; he shows sales the trajectory chart from Anthropic’s founding — “trust that you are going to continue to make improvements.”
- Why churn runs cooler than leaderboards suggest: customers do fine-tunes and bespoke work on a model, and you’re “one of three or four options within a model selector” — switching daily on evals “would be an insane thing to do to your user base.”
- Brand is real: he endorses Harry’s “I’m a Claude person / I’m a ChatGPT person” framing, citing Ben Thompson’s discussions with Nat Friedman and Daniel Gross. His Instagram-era “fake formula” — format + audience + vibes — has an AI analogue: model personality + scaffolding prescriptiveness + vibes.
7. DeepSeek changed the playbook
- On whether the West underestimates China: “the DeepSeek piece — people seemed surprised that there were cutting-edge research teams there, and if you were paying attention, that part should not have been the surprising piece.” He watched a parallel startup world emerge after Instagram was blocked; WeChat solved scale problems “of the same scale of challenges that Facebook was doing.” Dismissing China as replication is “a pretty Western-centric view” — they can train at the frontier, “especially if they get access to compute.”
- What DeepSeek did that Claude hadn’t: broke through on narrative. The cheaper-training story — “whether that was exactly true or not” — landed into January, a new presidency, and China relations: “likely my great-aunt was calling me about DeepSeek. I’m not even joking.” His self-criticism: “I don’t think we tell the Claude story well enough” — Claude 3 was state-of-the-art trained by “a team that was much, much, much smaller than any other lab.”
- On product it was “stronger than a nudge — a shove”: ship ideas to market quicker, because “sometimes the novelty of experience is itself valuable — it was the first time most people experienced the live chain of thought… I wish we had done that sooner.” (Anthropic had already planned to show CoT; the distillation risk may push labs to obscure it later.)
- On DeepSeek’s staying power, Harry notes that emerging-market usage retains while Western usage doesn’t; Krieger is skeptical but humble, and for everyone, “I still think we are in day one around is AI an indispensable part of most people’s work — and I think the answer is no.”
8. When a model provider becomes an application provider: generalizability, not verticals
- Anthropic’s product team is roughly a tenth of the company yet supports Claude Code, the API, Claude, and Claude for Work — so the filter is generalizability: “I don’t anticipate us building a lot of verticalized experiences that are fairly bespoke to a given workflow.”
- Harry probes horizontal categories — translation, transcription, customer service. Krieger’s counter: workflow knowledge defends the power user — ElevenLabs’ console is “very clearly for people translating hours of content,” and Descript is “some of the best product design in AI… clearly built by people who are day in, day out sitting in this workflow.” Their synthesis: professional workflows hold value; on consumer, basic AI “gets good enough” — a $10 monthly translation subscription “feels iffy.”
- Claude Code embodies the strategy: built internally first “because we just wanted to accelerate our own team,” shipped after months of dogfooding, and deliberately not an IDE — other companies “wake up and go to bed every night thinking about how do we make a great IDE,” low-latency autocomplete and the VS Code plugin ecosystem. Anthropic’s lane is the agentic loop between the IDE and likely Cognition’s Devin-style full delegation, because models today “still need hands on keyboard.”
- Proof point from his own hands: two pull requests last week, his first code since joining Anthropic, in a codebase he’d never opened — “Claude Code is very good at finding the file that has the right piece.”
9. Engineers become delegators within a year — but deciding what to build stays human
- The engineer’s role is already shifting to knowing what to build: “many, maybe even most of our good product ideas come from our engineers… prototyping.” Code review changes too — his own PR drew comments of “yeah, Claude Code does this sometimes, we don’t actually use default arguments in this case,” so models must learn idiomatic patterns from codebases and reviews.
- The end state: “from mostly code writers to mostly delegators to the models and code reviewers” — with a static-analysis comeback, AI-driven vulnerability checks, and computer-use agents testing UIs. His scenario: you return to an agent that tried three approaches in a browser, vulnerability-tested the winner, and asks you to review one critical section — “empowered to be more of a manager and delegator.” Harry: “three years sounds ridiculous, a year would be much more realistic.” Krieger: “I agree.”
- The bottleneck that survives: he started the year auditing where Anthropic’s own process is “cloudified” — Claude drafts PRDs, codes, synthesizes disagreements — but “driving alignment and actually figuring out what to build is still the hardest part,” best resolved in a room or in Figma, and “probably more than a year away” for models — he later says at least three. That’s why he’s “really bullish” on startups: “alignment is a coffee conversation in an afternoon rather than steering the ship of a large company.”
10. First-party strategy faces hard constraints
- First-party products teach fastest: within a week of internal Claude Code deployment they found a tool the model under-used — the fix went “directly into 3.7 Sonnet.” He’s changed his mind on this in the past 12 months (“how much first-party stuff is important”), admits being late hurt “significantly,” and names two underinvestments: first-party iteration speed (“my current obsession”) and API abstractions “beyond tokens in, tokens out” — agentic planning, knowledge repositories, tool use, memory transcending conversations. Instagram was “95% product, 5% API”; Anthropic is roughly an even split.
- Harry’s sharpest jab — do you and OpenAI have too much money? — draws the episode’s most honest concession: “the adoption that we’ve gotten of our products is ahead of their actual true product-market fit because they are still the best ways of getting the models, and I don’t think that’s durable over time… we’re under-serving people.” The fix: stop running “a larger company playbook,” ignore calcified org boundaries, spend calendar on product review not administration.
- Quickfire: OpenAI has been better at “shipping v1s faster, even ahead of where the model is sometimes,” worse at “personality and having the features they build be cohesive.” Rebuilding from scratch, he’d tear down the projects-vs-artifacts-vs-chats information architecture — Claude.ai—and probably ChatGPT.com—were “initially just built to be showcases of the models.”
- The critical unspoken challenge: discernment × privacy at the agent-to-agent intersection. His metaphor is his 5-year-old, who can’t yet distinguish family secrets from checkout-aisle chat: “do you trust your Mike agent or your Harry agent to be out in the world and not be jailbreakable?… Models fundamentally want to be helpful and that is not always what you want them to be.” On AI friends, he won’t dismiss Alex Wang (“I don’t think he’s wrong”) but insists AI practice is “absolutely insufficient” versus real interaction — the person who only read about red in a black-and-white room, versus seeing it. And on Dario’s live-to-150 optimism: likely Huma cut clinical trial reports from ~15 weeks to 20 minutes with Claude, and Arc Institute’s cell foundation models could cut the discovery loop itself — “the smartest minds of my generation were working on serving more targeted ads; a lot of them today are working on models.”