Mercor CEO: Why Application Layer Companies Have No Moat & The Cost of Hiring AI Researchers
Mercor CEO: Why Application Layer Companies Have No Moat & The Cost of Hiring AI Researchers
Summary
- Mercor CEO confronts the hack rumors head-on: there was an incident — the attacker used “a swarm of coding agents” to gain access — but the claim that revenue flatlined is false. Mercor added $300 million in net new ARR in the last 60 days, engaged Mandiant immediately, and added security as a seventh company value. The Twitter narrative, he says, included “one person that’s very prominent who’s invested in multiple competitors” tweeting an untrue claim that Chinese actors accessed the data.
- The revenue is real revenue, not GMV: customers buy tasks (e.g. $1,000 per task delivering model improvement) at a 30-40% gross margin, with Mercor running the full stack — expert sourcing, platform, AI project management, quality checks. The business is “very profitable,” has over $500M in cash, more cash than it has ever raised, and has “almost 4x’d” since the $10B round at ~$400M run rate in fall 2025.
- The core investment call: infrastructure upstream of OpenAI/Anthropic beats application layer downstream over the next 12 months, because “the model is the product” and app-layer defensibility is increasingly difficult. 2025 was the year a model makes a PR; “2026 is the year of how do you get the model to clone Slack end-to-end” — those capabilities land in models within 12 months. The litmus test for surviving SaaS: network effects (Salesforce integrations, Slack Connect, Carta) — companies without them face severe difficulty.
- Token spend will exceed headcount spend at the average enterprise within 5 years — Mercor is already there: “right now, we’re spending more on tokens for our internal agents than we are on employee headcount.” Benioff’s $300M Anthropic spend is only ~3.8% of Salesforce developer salaries, showing how early this shift is. Every Fortune 500 will need a per-workflow eval “system of record” — which commoditizes the API layer (zero switching costs, new frontier model every 2 months) while stickiness lives in workflows.
- The guest would put at least one of OpenAI/Anthropic above $10 trillion in 5 years — a change of mind: he used to doubt labs could hold pricing power, but “the sheer revenue ramp of these businesses” convinced him they’ll be the most valuable companies in the world. Yet he simultaneously expects the majority of inference in 5 years to run on open-source, fine-tuned or distilled models, not frontier ones; per-workflow evals are often “a 10x lever on price performance.” Nvidia may lose its monopoly to a multi-chip future but even at 30-40% share of “the largest market in the world by far” remains the most valuable company.
- Training agents is “the fastest job category ever created in history”: Mercor pays out $3 million a day to its 5M+ talent network, estimated to roughly triple, perhaps quadruple, in 12 months. All knowledge work converges on training agents because it’s “structurally more efficient to do something once” — and on Mercor’s Apex benchmark the frontier model scores ~40% versus o1 at 1% just 12 months ago.
- Pricing elasticity: Harry cited Nebius raising prices 30% with zero demand impact, while Mercor “has the demand to double overnight” but lacks capacity — yet the guest says pricing must balance near-term optimization against competition because “high margins invite competition.”
- The talent market is dislocated: one candidate held a $20M/year liquid offer from Meta’s superintelligence group, top AI researchers cost “tens of millions of stock per year,” and demand outstrips supply 10:1. Europe, he says, has lost the model race to talent network effects and should accept it — labs will simply “hire 10,000 people in France to teach the models French law,” gutting the sovereignty argument.
Deep dive
1. The hack, mythbusted: $300M net new ARR in the last 60 days
- Harry opens with the rumor — a hack, revenue flat since. The CEO’s answer: “There was an incident. All of the other parts are false.” Mercor engaged Mandiant and other security firms immediately, communicated proactively with customers, and has since “expanded our relationships with all of the Frontier Labs and added $300 million in net new ARR in the last 60 days.” The company added security as a seventh value “to make sure it’s very ingrained in the culture.”
- The Twitter storm was worse than the reality: he cites “one person that’s very prominent who’s invested in multiple competitors and just made this tweet about how all of our data was getting accessed by China when it was totally untrue” — and lawyers advised against firing back explicitly. His crisis takeaway: it “definitely wasn’t close to the most stressful” moment in the company’s life, and deep customer relationships plus a full internal picture beat the “echo chamber on X.”
2. The attacker used an agent swarm — a golden age of AI cyber is coming
- The mechanics of the breach are the tell for the next market: “it was the attacker that used a swarm of coding agents to help get access to the system.” A human attacker reviews code at human speed; a swarm is “very exhaustive in reviewing the entire code base,” which can let attackers move much more quickly.
- The guest’s call: “an enormous boom in AI security engineering tools” — customers are focused on improving models’ cyber-defense capabilities toward “the best AI security engineer that is able to defend every enterprise,” and the waves of incidents “are just getting started.”
3. Customers and rumors: OpenAI “stronger than ever,” Meta paused, no Micro 1 offers
- Did Mercor lose OpenAI and Meta in the hack? “False. Our relationship with OpenAI is stronger than ever.” Meta is the exception — “currently the relationship is still paused.” The CEO says there are other things happening and mentions the Scale acquisition as one reason Meta may naturally work more with Scale. Every other frontier lab has grown its relationship since. He flatly denies Harry’s theory that Handshake’s parabolic revenue is Meta spend shifting from Mercor: “That’s not true” — but won’t elaborate.
- The Micro 1 poaching story with “signing packages in the millions”: an employee sent outbound messages floating a $500,000 signing bonus and some recipients took first meetings — “we have not extended a single offer to someone from Micro 1.” The press framed messages as legal offer letters.
- On the Amazon acquisition rumor at $13B: “That one is false.” Would he sell at $30B? “No… I could walk away with billions of dollars in cash and that’s just not what motivates me” — the mission is “how humans fit into the economy,” and “our probability of executing on that vision wouldn’t be as high if we weren’t an independent company.”
4. Revenue is real, margins are 30-40%, and the business has never really burnt cash
- Against the “it’s just GMV” critique: customers buy tasks end-to-end — “they’ll pay $1,000 for this task that delivers model improvement” — and Mercor does everything from expert sourcing to the AI project manager to automated quality checks, at a 30-40% gross margin. “We’re powered by a talent network in the same way that Uber is powered by a driver network, but that’s not the end product.” Revenue is “dramatically higher” than the ~$1B posted publicly.
- The vertical integration compounds: downstream quality signal informs upstream expert onboarding, and data value is power-law distributed — “out of a data set of 10,000 tasks, the top 2,000 tasks will create the majority of the value,” which is why quality confers pricing power.
- Financial posture: “We burnt half a million dollars after our seed round and from there we’ve pretty much been profitable ever since… we have more cash than we’ve ever raised” — over $500M in cash, positioned deliberately for a market correction: frothy funding lets anyone run negative margins now, and “when markets come back to earth… that’s when there’s periods of consolidation.”
5. Training agents is the fastest job category in history
- The guest’s answer to the layoff wave (Intuit 16,000, Meta 8,000, ClickUp 22%) is the lump of labor fallacy: 250 years of 25x productivity growth — “equivalent to automating about 96% of someone’s job” — produced more jobs, not fewer. Harry’s counter is the speed: past revolutions took decades; “with Nano Banana Pro, I can get rid of all designers in my media company pretty much overnight.” The guest concedes displacement will be “very significant” but argues the economy now creates job categories more effectively too.
- Exhibit A: Mercor pays out over $3 million a day to its talent network — “the fastest job category ever created in history” — which he expects to roughly triple, “maybe quadruple,” in 12 months. On the Apex benchmark (Mercor’s AI productivity index across consultants, bankers, lawyers, engineers), the frontier model now scores ~40%; twelve months ago o1 scored 1%.
- The five-year new job: agent trainer. “All knowledge work is converging on training agents because it is structurally more efficient to do something once” — the support rep trains an agent instead of redundantly answering hundreds of tickets. The human contribution that survives is tacit knowledge: “there’s just an enormous amount of context that lives in people’s heads” that models can’t get elsewhere — whereas data cleaning itself gets done by the models as reasoning improves.
6. Horizontal aggregation beats niche data vendors — labs want one flexible partner
- Against the unbundling thesis (surgeons with head-cams selling medical data): “the kind of data shapes that we would build for a lawyer are often very similar to the kinds of data shapes that we would build for a doctor.” With a 5M-person talent network that refers friends, finding the marginal doctor is easy — so labs prefer one horizontally capable vendor over “100 different vendors that they have to train for the same data shape in 100 different domains.”
- What labs actually need is all-encompassing: “the full distribution of everything that you could pass into Google Workspace and everything that you could want out on the other side in every job category throughout the economy” — and experts who are also “power users of ChatGPT or Claude that are able to find where the model makes mistakes.”
7. Chopper, Ferraris, warship: $23M post to $10B in two years
- The round-by-round tape: seed September 2023 at ~$1M run rate — General Catalyst term sheet within 36 hours, $2.3M at $23M post. Series A: Benchmark’s Victor got a refused second meeting until “have you ever been in a helicopter?” — $250M post at ~$2.5M revenue. Series B: Felicis lured the founders onto a private jet to race Ferraris at the Vegas F1 track — $2B at $20M revenue, 100x. Then $10B at ~$400M run rate (September/October 2025), ~25x — “and the business has almost 4x since then.” Next round: “probably a much higher valuation,” unhurried because the company is profitable. Harry’s tally of transport modes: “We need a warship now for the Series D.”
- The defense of the crazy prices: 50% month-over-month growth for six straight months, sustained another 12+ — “I was projecting 50 million in revenue run rate by the end of the year and 500 million by the end of next year… and we beat the projections.” Most uncomfortable round in hindsight: the Series B — “it’s very different to be 100 times the revenue at 2.5 million versus at 20 million.”
8. The model is the product — app-layer defensibility is increasingly difficult
- The tweet that anchors the episode: the next 12 months will be “dramatically better for infrastructure companies upstream of Anthropic and OpenAI than for application layer companies downstream.” Reason one: “over the last 2 years everyone has increasingly realized that the model is the product” — end-to-end trained models beat every stitched-together abstraction, drag-and-drop agent builder included. Reason two: software recreation speed — “2025 was the year of how do you get a model to make a PR in a code base. 2026 is the year of how do you get the model to clone Slack end-to-end,” capabilities he expects in models within 12 months. It’s “not a far leap for Claude CoWork to add capabilities across medical and legal.”
- Harry pushes back with his Legalese position: deep lawyer-specific workflows, GTM, CS teams — “the defensibility is there. Argue back.” The guest’s rebuttal: the moat isn’t pre-sales GTM but the forward-deployed motion — a savvy customer paying $1M/year for SaaS “could just tell Claude to copy it,” but an agent trained on a company’s tacit knowledge is “incredibly differentiated and hard to recreate.” Hence the Sequoia line that “services are the new software.”
- The litmus test for incumbent SaaS: network effects. Salesforce’s integration marketplace, Slack Connect, Carta’s cross-company graph — those companies can 10x product velocity on top of a real moat. “The companies that don’t have network effects are going to struggle very significantly… that is the litmus test that determines whether this company is going to become worthless.”
- The services thesis, live inside Mercor: an AI project manager just completed its first project end-to-end — hiring experts, answering questions, building the annotation tool with its own coding tools — work previously done by a ~100-150 person delivery org. “The experts all had a really good experience reporting to the AI project manager.” “We’re seeing in real time that services are getting automated.”
9. Token spend passes payroll — and evals commoditize the API layer
- The headline stat, delivered casually: “Right now, we’re spending more on tokens for our internal agents than we are on employee headcount.” His five-year call: “the average enterprise spends more on compute than headcount” — versus Benioff’s $300M Anthropic spend today, which pencils to just ~3.8% of Salesforce developer salaries. Token costs rising despite efficiency gains is “a fascinating case study in Jevons paradox.”
- The mechanism enterprises will use: a per-workflow eval as “system of record” — Mercor runs one for each internal agent (interview agent with 5M+ interviews done, candidate ranking, accounting, fraud detection) that dictates model choice on the Pareto frontier of price-performance. Enterprises will use these to “commoditize the model layer… they want perfect competition with zero switching costs.” The API layer commoditizes — a new frontier model every 2 months, hot-swappable by eval score — but workflow stickiness survives: “I have all of these routines running in Claude Code and I probably wouldn’t put in the time to move those.”
- A workflow eval is “often a 10x lever on price performance” via distillation to open-source models — so his split forecast: OpenAI and Anthropic are “incredible investments” with four-to-five orders of magnitude more demand coming, yet “the majority of inference in 5 years is going to be using an open-source or custom fine-tuned or distilled model, not a frontier model.” Valuation call: “I could definitely see one of them being a $10 trillion company… at least one of them worth more than $10 trillion.” His quickfire change of mind: the labs’ revenue ramp converted him from doubting their pricing power to “immense conviction that they will be the most valuable companies in the world.”
- On just buying Nvidia instead: “not a crazy idea,” but a multi-chip future is forming — Cerebras executing, Etched, in-house lab silicon — so the monopoly may fade. “Even if they only have 30 or 40% market share in the largest market in the world by far, that is the world’s most valuable company.” Harry cited Nebius raising prices 30% with no demand impact; Mercor “has the demand to double overnight” but not the capacity — and the guest says pricing must balance winning the decade against competition, because “high margins invite competition.”
10. $20M cash offers, Europe’s lost race, and eliminating income tax for the bottom half
- The talent market in one anecdote: a candidate the guest was hiring held an offer of “twenty million dollars in cash per year from TBD” — Meta’s superintelligence group, stock but liquid. Top researchers cost “tens of millions of stock per year”; demand outstrips supply ten to one. He expects escalation to continue at the very top but supply of lab-trained people to normalize “the ninety-ninth percentile.”
- On Europe (Harry: Mistral places “like the Eurovision Song Contest — kind of at the bottom”): the model race is lost to talent network effects — brilliant French researchers aggregate at OpenAI, Anthropic and DeepMind, compounding into “one of the largest not only economic but geopolitical advantages that the US has.” His advice: accept it, keep some post-training and application capability, don’t “lean aggressively into competing head-to-head with Anthropic.” The sovereignty argument is limited too: “the labs are just going to hire 10,000 people in France to teach the models how to be better at French law” — transfer learning does the rest.
- His freshman-year essay, now Bezos-retweeted: eliminate income tax for the bottom half of Americans — it’s only ~3% of government revenue, and “the largest positive externality in the economy is jobs,” yet we tax exactly that. He points to capital gains, especially short-term gains, and carbon as alternative taxes: “it’s crazy to me that instead of taxing carbon, we tax the bottom half of Americans.” Harry’s pushback is ferocious — “I say this with the nicest respect. It’s just wrong… you f* off to somewhere that doesn’t have capital gains, and then you lose all the tax revenue completely” — and the guest agrees any scheme needs “sensitivity analysis” on capital flight.
- Quickfire residue worth keeping: about half of data-provider competitors are “just transactional talent marketplaces”; the rival he most respects is Surge’s Edwin for staying close to research; IPO “in the next few years” but not this one or next; the 996 rumor is false — “we’ve never mandated hours,” though he and Adarsh “work from when we wake up until we sleep.” And the kindest thing: the Prod nonprofit community — weekly meetings, working capital, a first big customer — “they took no equity… Mercor wouldn’t exist if it weren’t for any of those individuals.”