Mercor Head of Product on Revenue Concentration from Frontier Labs
Mercor Head of Product on Revenue Concentration from Frontier Labs
Summary
- Open source needn’t cannibalize the frontier data business — it raises the floor. Osvald’s core defense of Mercor: data is “most valuable on the frontier of model performance,” each lab buys eval and training sets to fill gaps in current capabilities, and “open source models just mean that nobody’s buying anything that K3 can already do.”
- He rejects the “90% of enterprise workflows can be automated” framing on two grounds: Mercor’s Apex benchmarks show top models scoring only ~50% on long-horizon workflows, and there’s a whole class of latent demand nobody is even attempting yet — months-long procurement agents you check once a week. For legal arguments or medical advice “the percentage framing is just totally off and we need to be thinking more about continuous uncapped rewards.”
- Mercor’s token spend already exceeds salaries. Benioff’s $300M/yr on Anthropic works out to ~3.8% of developer salaries; Osvald thinks that percentage rises above 3% macro — and confirms Mercor already spends more on tokens than on salaries: “100% sounds reasonable.” No enterprise ROI problem yet, just “a period of exploration and experimentation” with some “tightening of the screws.”
- On revenue concentration and the “it’s not revenue” peanut gallery: “we end every week with so much more money in the bank… the cash flow is insane” — the business can’t spend money fast enough to service demand. The concentration fix is moving downmarket via self-serve projects and AI project managers, since lab projects are white-glove and operationally brutal.
- RL environments are the fastest-growing data type — high-fidelity simulations of apps (a mock that “acts exactly like Salesforce”) plus rich start states of hundreds or thousands of files. It’s also what he changed his mind on most: he thought environments wouldn’t scale, “but then it did.” His 3-year call: robotics/physical data, with progress perhaps taking the form of “a Waymo robo-taxi/Cruise-type moment rather than a ChatGPT moment.”
- The AI-era product org inverts: PM-to-engineer ratios rise as engineering stops being the bottleneck, teams fight to reduce surface area, Mercor is moving off Figma toward cloud design, and “skill issues have almost gone away. So now it’s all about judgment.” His warning: “don’t delegate your decision-making… to a model because you’re going to lose that ability and then you’re going to get psychosis.”
- Competitive read: rivals are a “cottage industry of founders doing annotation themselves” — VC-subsidized, “totally mispriced,” beloved by labs, and unable to scale 10x. Other competitors often copy Mercor soon after; Surge is described as “a bit out there… very secretive.” And frontier labs eating the app layer? Google Plus and Threads say big companies “lose to companies that have intense focus on their market.”
Deep dive
1. Open source raises the floor — it needn’t eat the frontier data business
- Harry opens with the migration-to-open thesis (Kimi’s new model in hand): does open cannibalize Mercor? Osvald’s answer is the episode’s anchor: data is most valuable on the frontier of model performance — customers buy eval and training sets to fill gaps in current capabilities, so “open source models just mean that nobody’s buying anything that K3 can already do.”
- The mechanism: open models “just raise the floor of what people are interested in.” As long as customers still want new capabilities, the business grows — commoditization below the frontier is irrelevant to where the spend sits.
2. The “90% of enterprise workflows” number is measuring the wrong thing
- Osvald isn’t convinced 90% of enterprise workflows can be handled by open or frontier models. Those calculations may be based on existing demand — there’s latent demand nobody is attempting yet, most commonly long-horizon tasks: “setting up a procurement agent to fully automate your procurement team for months on end. You only check on it maybe once a week.”
- The number he’ll stand behind: on Mercor’s Apex benchmarks, top models score around 50% of long-horizon workflows.
- His sharper distinction: sufficiency-based workflows (“updating a CRM — you couldn’t really get much better at it”) versus uncapped ones like legal arguments or, to an extent, medical advice, where you could always improve. There, “the percentage framing is just totally off and we need to be thinking more about continuous uncapped rewards.”
3. Enterprises fear labs on core work — and every company may get its own model
- On Alex Karp’s claim of enterprise skepticism toward sharing data with frontier labs: it depends how core the workflow is. HR and procurement are less sensitive, and enterprises are more open to putting those workflows on proprietary models; the sensitivity is around differentiating work — “the actual memos” a law firm writes, the advice it gives clients. Harry’s irony — sensitive data on open-source, most likely Chinese models, HR on closed ones — draws the rebuttal that open weights mean inference “can happen in multiple places… you have more control.”
- On the specialized-models-for-every-company thesis Harry attributes to Lynn Qual, the founder of Filecoin: “I buy it” — and he concedes it’s self-serving for Mercor too, since every specialized model needs enterprise-specific eval and training data. The market’s size “depends on the value that customers can get from the specialized models,” and those ROI-justified cases will increase over time.
4. No ROI problem — Mercor’s token spend already exceeds salaries
- “I don’t think there’s an ROI problem right now. I think we’re in a period of exploration and experimentation” — more tolerance and patience while token prices and performance are moving too fast for the calculation to settle, even as spend screws tighten “here and there.”
- His budgeting framework: coding-agent spend for engineers “could still be giving you compounding gains” — it isn’t COGS. A customer-service agent whose token spend exceeds the revenue from the customers served puts you “definitely in a bad position.”
- On Benioff’s $300M/yr to Anthropic (~3.8% of developer salaries): Osvald wants better accounting of outcomes per token, different spend profiles per team — but macro, “the percentage will increase over time to more than 3%.” When Harry reveals Brendan told the show it would hit 100% and that Mercor already spends more on tokens than salaries: “Yeah, we do… 100% sounds reasonable.” Headcount is up more than 10x since he joined, revenue commensurately — “we can’t spend money fast enough to service all of the demand that we have.”
5. Services bridge the gap — focused companies can beat big labs
- His self-declared hot take on the services boom (Microsoft’s services arm, Palantir): it’s the future in the short term, while the knowledge of how to deploy and eval agents — currently “a concentration of a bunch of people in San Francisco” — disseminates through industry. Eventually, it may become a job function like software engineering: “maybe a decade-long change.” Against Matt from Factory’s line that services are “just an excuse for crap product,” he holds the frame: it’s a talent-availability problem, not a product one.
- Do good engineers want to be forward-deployed? The ones who communicate well and cut through to the source of a problem do — and “those engineers are also great fits to eventually become founders,” which is why Mercor bleeds alumni to startups. “It’s a lot better to lose someone to starting a company than to taking another job.”
- On Harry’s Legora worry — that Anthropic is the real competition: look at Google and Microsoft. “You remember Google Plus, right?… That didn’t go anywhere.” Large companies diversify but “lose to companies that have intense focus on their market.”
6. The AI-era product org: shrink the surface, hire PMs, judgment is the job
- Faster engineering makes product management harder, not easier: people push multi-thousand-line PRs because they can, so the team is “constantly in this battle to try to simplify our product surface area.” His biggest regret runs the same direction — the annotation platform tried “to serve every ask,” ended up with hundreds of heterogeneous projects (“that’s just chaos to manage”), and guardrails should have come much sooner. Determining enduring demand is leadership judgment from constant lab contact — “ultimately, it’s kind of a guess.”
- The org consequence: a higher ratio of PMs to engineers, because engineering is less bottlenecked and understanding user workflows and what drives revenue “becomes the bottleneck now.” Two product areas — the expert marketplace and the Studio annotation/eval platform — run pods of four or five, with the PM:eng ratio set to keep rising.
- What makes a great PM has changed twice over: fewer tools (“even Figma — we’re moving away from it in favor of cloud design more and more”), and everyone up-leveling to business impact. “Skill issues have almost gone away. So now it’s all about judgment and am I doing what is going to drive the most business value?”
7. Hiring senior, testing judgment — and refusing to delegate your brain
- Hiring now biases toward senior candidates in their prime — “25 to 35” — who grok what drives revenue faster; junior can-you-use-the-tool skills are becoming less relevant. Take-homes are down to one AI-fluency check, then whiteboarding on experiments, statistics, and systems design — because with AI tools “it’s very easy to offload a lot of thinking” and he wants people who don’t “just regurgitate what comes out of Claude.”
- His brain-rot rule draws the line between judgment and execution: models “make you think that it’s doing the right thing, but you have to be paranoid with them still… don’t delegate your decision-making, like your actual job, to a model because you’re going to lose that ability and then you’re going to get psychosis.” His analogy: phones “kind of fry your brain, they turn it into goop” — learn where the boundary is.
- What he misses in bad hires: “it’s really hard to assess agency and ownership in the interview process.” Talented jerks are tolerable — “personalities can change… it’s really hard to make someone give a shit.”
8. The business mechanics: insane cash flow, concentrated revenue, deliberate downmarket push
- To the “it’s not revenue” crowd: “we end every week with so much more money in the bank… the cash flow is insane.” He’s seen “interesting financial engineering” elsewhere; this isn’t that.
- On frontier-lab concentration: the direction is moving downmarket — self-serve human data projects and AI project managers — because lab projects are white-glove and operationally intense: continually surfacing edge cases, “insanely fast alignment” between customers and experts, and “a huge amount of paranoia” that every data point is perfect. There are many more enterprises than labs, and the shift has already reduced concentration.
- Supply-side secret: expert experience — paid on time, well, transparently — powers referrals, plus a sourcing team for spiky skill demand. It’s not topping competitors’ “crazy bonus payouts”; it’s visibility into future work and skill growth. Margins are decided after the fact, and project costs now split roughly equally between paying experts and LLM spend on synthetic data and automatic quality control.
- Labs negotiate, but the setup is enviable: eval sets represent what their customers want, training sets hill-climb them, so “as long as the amount of money they’re spending on data is less than the revenue that they’re going to get… people want to crank the lever harder and harder.”
9. Environments are the frontier data type; the cottage industry can’t scale; robotics is the 3-year call
- Environments are the fastest-growing data type and Mercor claims category leadership: simulations of apps plus a rich start state (“hundreds of files, thousands of files”) so training data “looks a lot closer to what they see in deployment” — a mock that “acts exactly like Salesforce.” It’s hard now the way SFT and preference ranking once were; “labs are figuring it out. Eventually… enterprises can do it, too.” It’s also his biggest change of mind: he thought environments wouldn’t scale — “we just kept going at it and it ended up working.”
- The competitive field is “a cottage industry of founders doing annotation themselves” — smart ex-technical founders running “VC-subsidized work that labs love” because it’s “totally mispriced.” It breaks when customers want 10x throughput. Other competitors often copy — “we write a blog, someone else writes a blog like a week later that’s the exact same thing” — while Surge is described as “a bit out there… very secretive.” Cyber is the other growth vector: adversarial, “AlphaGo-type situations” with uncapped rewards where “the goal posts are always going to move.”
- The missing revenue line in three years: real-world, physical data for robotics, a market nascent relative to GenAI and AVs. Against Harry’s rant about 15-minute water-fetching demos, his precedent is Cruise-to-Waymo — “now I take Waymo more than I take Uber” — so any inflection might be “more of a Waymo robo-taxi/Cruise-type moment than a ChatGPT moment.”
- The $200B bull case: a tech-enabled services company where “evals serve as the PRD… but also the optimization objective for better performance.” As long as better models are valuable to the economy, eval and training demand persists — plus a growing agent-deployment enterprise arm.
Verification Notes
- The raw captions mention “Kimi” and separately render “K3”; “Kimi K3” is not clearly spoken as one name.
- The raw captions render the $200B bull-case company as “Brex”; the surrounding context suggests Mercor, but that name cannot be confirmed from raw alone.