Mercor’s Secret to Hypergrowth and the Smart People Behind It|A Conversation with Mercor’s First Chinese Employee, 虞快
Summary
Mercor’s core business is no longer low-end “data labeling,” but organizing experts for AI labs, defining evaluations, and turning those evaluations into the models’ “PRDs.” Experts may label data or define labeling standards; today, most of the work is defining standards. As models improve, non-expert labeling becomes less valuable. Once standards are established, they can be used directly for Reinforcement Learning, though customers may still need other platforms for large-scale volume.
Mercor targets the “most painful and best-funded” customers: AI labs may urgently need 50 lawyers or 200 software engineers at once, while the economic return from improving a model far exceeds the cost of experts. 虞快 says platform experts earn more than $90 an hour on average, software engineers typically earn $100–200, and dermatologists can command $400; customers care more about finding qualified people and improving model quality.
Mercor’s moat is not just finding people, but coordinating screening, matching, verification, quality control, time tracking, and global payments. Candidates take 20–30-minute AI video interviews and are ranked against clients’ predefined hard and soft criteria; multiple people in the same cohort can be given the same question for comparison, and underperformers may be removed mid-project. “10,000 people can come apply”—the key is scientifically identifying who can actually do the work.
虞快 estimates the market at roughly $5B–$10B today, and says data spending will continue expanding as long as core products at companies such as OpenAI and Anthropic remain models. His logic is that algorithm talent can be poached, OpenAI may not beat Google on compute, while data can still generate differentiation through quantity, quality, and cleaning methods; competitive pressure will push giants such as Google and Meta to invest as well.
The market may consolidate around each AI lab working with two or three vendors, rather than running ten platforms in parallel. After Meta acquired Scale AI, customers became more aware of the risks of relying on a single vendor that could be acquired; 虞快 no longer sees Scale as a major growth threat, but considers Surge AI a threat, calling it larger than Mercor and “very low-key and pragmatic.” 曲凯 summarizes the endgame as a competition over expert quality, speed, and output quality.
Mercor reportedly went from a $1M to a $100M run rate in just 11 months—one month faster than Cursor, according to 虞快—and growth itself has become a flywheel for talent and execution. The company had about 50 people when he joined and now has more than 100 full-time members; he says monthly revenue is growing 50%, and the team believes that what is promised on Saturday can generally be delivered by Friday. “As long as you get moving, it’s better than staying still.”
This is not a retreat from the AI recruiting-platform story, but an effort to build a project-based labor market in advance by serving the most urgent customers. Mercor believes full-time work may decline in the future as more people move to project-based work; its evaluation, matching, management, and payment capabilities can also be reused in other industries. Whether China can replicate the model depends mainly on model companies’ budgets and willingness to buy, as well as the value of independent specialists for confidential hiring needs.
Coding agents will reduce the importance of raw coding speed while increasing the value of engineers who can choose the right direction, understand customers, and push work forward independently. 虞快 calls this capability agency: how clearly a manager must specify the task, and how long someone can operate independently afterward. Organizations may become flatter; 虞快 emphasizes proper incentives and managers who can recognize talent, while 曲凯 adds that founders must be able to sell their vision to talent before the company has a valuation or growth halo.
Deep dive
1. As Models Improve, Low-Skill Labeling Fails First
虞快 defines Mercor’s business as helping the largest AI companies recruit part-time experts, spanning doctors, lawyers, investment bankers, and consultants down to Swift engineers and Russian biologists. The traditional model of sourcing low-skill labor from less-developed regions can no longer cover models’ emerging needs.
These experts do more than follow pre-existing rules; they may also help define the rules. Most of the work today is defining standards: AI labs do not have enough medical or legal expertise internally to determine where a model is good or bad, so they need experts to systematically map the boundaries of its capabilities. The episode’s shorthand was: “If you’re not an expert, you can’t label anymore—the model already knows what you know.”
曲凯 asked whether Mercor simply sets the standards first, after which customers still need a Scale AI-type platform to complete the labeling. 虞快 said clearly that once the standards are established, they can be used for Reinforcement Learning; the same pool of experts typically defines the standards and performs the work against them. But if a customer needs 10,000 lawyers at once, Mercor may not be able to fill the entire order alone, and other platforms may still be needed for additional volume.
2. Mercor Compresses Sourcing, Interviews, and Verification into an Automated Funnel
Expert acquisition comes through advertising, outbound sourcing, and referrals, with more than half of sign-ups coming from mutual referrals. Some people earn enough through referrals to leave their full-time jobs, effectively becoming external recruiters who understand Mercor’s future needs and continuously source for the platform; extremely obscure roles require targeted searches in professional forums or offline events.
After submitting a résumé, candidates can apply to multiple projects and then take a roughly 20–30-minute AI video interview. Questions are generated from both the résumé and the role requirements, while recruiters can manually add questions and scoring criteria; the recording comes with subtitles, and clicking the transcript jumps to the relevant point in the video, making it possible to inspect pauses and clarity of expression.
Clients do not need to watch every video themselves. When creating a role, they define hard requirements and soft objectives, and the system ranks candidates based on their résumés, interview performance, and those criteria. 虞快 says AI’s core role comes down to two things: assessing a person’s true quality and matching the right person to the right job.
Verification does not necessarily require generative AI: the system can cross-check IDs, LinkedIn, and GitHub. 虞快 gave an extreme example of scammers selecting real names from the IMO winners’ list, impersonating the winners, and mass-generating résumés; those checks are ordinary automation, not AI.
3. Customers Buy a Manageable Global Expert Workforce, Not a List of Names
A project may simultaneously require 100 software engineers distributed around the world, each working 20 hours that week. Mercor also has to handle time verification, quality assessments, renewals and removals, cross-border payments, question support, and disputes such as “the person reported 15 hours, but the system recorded only 10.” 虞快 believes AI labs do not want to take on this operational minutiae themselves.
曲凯 asked whether Mercor was simply making one-time introductions or paying experts centrally, and raised the risk of disintermediation. 虞快 said the U.S. market had not presented that problem so far, then returned to the more fundamental management challenge: customers need to manage large numbers of contractors continuously, not merely find them once.
In 虞快’s view, Mercor differs from traditional outsourcing firms not in mission or vision, but in having “a way to assess someone’s quality.” The platform does not currently support universal self-serve; most customers are AI labs, with some AI startups as well, and roles are primarily contractor positions with some full-time work.
4. Expert Hourly Rates Range from $21 to $400, with Pricing Set by Scarcity
虞快 says the platform’s average hourly rate is above $90; software engineers typically earn $100–200, while dermatologists earn around $400. A successful referral of one such doctor can generate a $5,000 referral fee, showing that the cost of acquiring the rarest experts is also high.
At the other end are audio trainers: most Americans can do the work as long as their English is good enough and they can read from a script, earning roughly $21 an hour. The gap between the two ends shows that pricing is not a generic “data-labeling rate,” but reflects professional barriers, the breadth of supply, and project urgency.
Mercor looks at prices for similar roles in the past and how many days it took to recruit how many people, then adjusts for the urgency of the current request and negotiates with the AI lab. Customers naturally want lower prices, but they also do not want to save money by hiring unqualified people and end up with no model improvement; delivery quality on prior projects therefore affects trust and repeat business.
5. The Model Race Turns Data Budgets into a Necessity for AI Labs
Mercor focuses on AI labs because they may suddenly need 50 lawyers or 200 software engineers, while the return from improving a model far exceeds the recruiting expense. Whether the gain shows up in benchmark scores or in bringing the model closer to replacing a category of full-time worker, 虞快 believes customers will pay if the result is real.
虞快 breaks model progress into three paths: algorithms, compute, and data. Top researchers command extremely high salaries and may be poached by a competitor tomorrow; on compute, he believes OpenAI may not beat Google; data, however, can still accumulate advantages through quantity, quality, and cleaning methods. He cites Anthropic’s coding performance as an example, attributing it to its data and how that data was cleaned.
His market estimate is $5B–$10B, and he believes it will continue growing as long as OpenAI and Anthropic still depend on the models themselves to win. Both companies are weaker than Google, Amazon, and Microsoft on distribution and are even less able to tolerate falling behind on foundation models; once they increase spending, competitive pressure will push Google and Meta to invest as well.
曲凯 compared this with reports that Meta was recently willing to pay $100M for a single person and inferred that the impact of high-quality data on a model would be far greater than that. This was 曲凯’s judgment, not a standalone budget estimate from 虞快.
6. The Scale Acquisition Highlights Single-Vendor Risk, and Competition May Consolidate Around Two or Three Players
曲凯 noted that after Meta acquired Scale AI, the market reportedly believed Scale’s business had declined sharply because other model companies did not want to hand sensitive data to the Meta ecosystem. 虞快 stressed that he had no inside information and was only speculating: Zuckerberg may primarily have wanted Alex Wang to lead the Superintelligence team; he also guessed that the roughly $15B arrangement could satisfy the original investors while ensuring Alex Wang would receive support for a future startup.
虞快 therefore does not view Scale as a major growth competitor, but considers Surge AI a threat. Surge is larger than Mercor, and its external style is the opposite of Scale’s publicity: “very low-key and pragmatic.” 曲凯 suggested the financing could also serve the talent competition; 虞快 added that strong candidates look at signals from investors such as Sequoia, Benchmark, and Andreessen Horowitz, rather than relying only on a bootstrap company’s own account of how well its business is doing.
AI labs likewise do not want to manage ten vendors at once: processes and settlement systems differ, and when the same model has a problem, accountability becomes difficult to trace. But Scale’s acquisition may teach customers not to place all their bets on one provider; 虞快 believes the more rational structure is at most two or three vendors, with one primary supplier and one secondary supplier.
7. Evaluation Is the Model’s PRD—and Mercor’s Higher-Value Position
虞快 rejects the label “data-labeling company,” saying Mercor is closer to an eval provider: “Evaluation is essentially the PRD—the product requirement document—for these models.” Experts first define tasks the model cannot currently perform but ultimately should, and researchers then work to push the model toward that target.
This position sits closer to the start of R&D than executing an established standard, because labs often do not even know how to evaluate performance in vertical fields such as law and medicine. Mercor supplies not just answers, but the ability to answer “what counts as good, and where does the model fall short today?”
Evaluation remains with Mercor and can be reused within limits. 虞快 cited Humanity’s Last Exam, which Scale helped create, as an example of a public benchmark; data and personnel from specific projects are “absolutely kept separate,” and AI companies sometimes explicitly require that experts recruited for them not be made available to competitors.
8. The Current Vertical Wedge Has Not Rewritten the Original Vision of Matching People to Every Kind of Work
Mercor’s original goal was simple: any company could write down its hiring criteria, and the platform would find the corresponding people. 虞快 says that goal has not changed; the most urgent, best-funded, and most willing large-scale buyers today simply happen to be AI labs, so the team chose to go deep in that market first.
The company believes full-time work may decline in the future as more people shift to project-based work: someone might work 40 hours a week for one company this month, finish the project next month, and then work for another company for two months. The evaluation, matching, management, and payment capabilities being built for AI customers are also investments in that labor structure.
He extends the use case to every kind of “selection” scenario. A VC facing 1,000 people who want a conversation could define preferred questions and answers, have candidates speak with an AI first, and rank them based on the results. AI interviews could therefore become a reusable people-evaluation tool beyond recruiting.
9. China’s Constraints Are Budget, Independence, and Willingness to Buy—not Technology
For China to fully replicate Mercor, 虞快 says the first requirement is how much the largest AI companies are willing to spend on expert data. He personally believes “any amount of spending is worthwhile,” but the outcome ultimately depends on those companies’ awareness and willingness to buy.
曲凯 pointed out the asymmetry in valuation and budgets: Chinese model companies may tell a $3B–$4B story, while leading U.S. companies tell stories in the $300B–$400B range. Even when the nature of demand is identical, the data budgets available to third parties may be completely different.
The counterargument to “big tech can build it internally” is specialization and confidentiality. Alibaba and similar companies can obviously build the technology, but it is unclear whether they would invest 100–200 people in optimizing interviews and matching; moreover, an internal system can serve only itself, while competitors would never hand it confidential hiring needs. If Chinese customers are genuinely willing to pay, 虞快 still believes someone will build the business.
10. A Hundredfold Run-Rate Increase in 11 Months Turns a Young-Founder Story into a Growth Flywheel
虞快’s standard for choosing a startup is that it must be “a little special” and have something to talk about, rather than being decent at everything without a defining strength. Mercor’s 3 founders were high-school classmates and debate teammates, all Thiel Fellows who started the company and dropped out at 21; after speaking with them, he believed that if he built a product and gave these people responsibility for selling it, they could answer any customer question convincingly enough to win them over.
The more decisive signal was growth. Mercor reportedly went from a $1M to a $100M run rate in 11 months, which the founders believed was the fastest in history at the time; 虞快 compared that with Cursor’s 12 months. Momentum attracts better people, and better people increase efficiency, creating the positive loop he values most.
When 虞快 joined around June, the company had roughly 50 people; it now has more than 100 full-time members, mostly in the U.S., with some in India. He clarified that he was not the only Chinese engineer, but possibly the first engineer born in China. The team’s average age was around 22 when he joined; half had founded companies, and one colleague had sold a product to U.S. News before turning 22.
曲凯 preserved a common domestic counterargument: experienced founders have spent money and seen the pitfalls, and might execute faster with the same $10M. 虞快’s response was not that youth always wins, but that Mercor found the right problem at the right time and recruited early core members from Scale; the lower career cost of startup failure in the U.S., along with a larger pool of success stories, also makes more talented young people willing to enter the funnel.
11. Speed Comes from Correctable Intuition and a Weekly Rhythm of Delivering What Was Promised
虞快 attributes rapid growth to two kinds of speed: making decisions quickly and executing quickly once a decision is made. Mercor is not extremely data-driven, relying more on the intuition of the founders or the core members they recruit; even after passing the technical interview, a candidate may still be rejected by one or two founders, who are evaluating judgment.
曲凯 asked whether those intuitions were actually right most of the time. 虞快’s answer was that “it actually doesn’t matter”: as long as a decision is not disastrously wrong, the team should move first, then use metrics to see “what’s working, what’s not working” and correct course quickly. The team specifies what result it expects to see the following week and what to do if it does not appear.
Execution means predictable delivery, not simply working longer hours. After planning the following week’s work on Saturday, 虞快 believes that most of it can be completed by Friday, and that other teams can make the same commitment; only when this mutual trust exists can the company promise customers an outcome.
The company tells candidates directly during interviews: “Our company is 996.” 虞快 says he generally starts around 7:30 a.m. and finishes at 1 a.m., with roughly 95% of colleagues on a similar schedule and some voluntarily working until 4 a.m. He believes this is not unusual at a 10-person company; maintaining it past 100 people is rare. The strongest incentive is monthly revenue growth of 50%: “All these considerations about whether something is worth it feel meaningless right now.”
12. Mercor Hires for Agency, Not People Who Only Know How to Answer Standard Questions
虞快 defines agency operationally: how clearly he needs to explain a task, and how long the person can operate independently afterward. If a manager has to prescribe every step and every day’s work, agency is low; hiring is not dogmatic, however, and when talent is extremely scarce, a diligent “workhorse” with other standout qualities can still join. The team simply cannot consist entirely of such people.
To judge intelligence, he prefers introducing a new concept live and watching whether the candidate can quickly understand, internalize, and transfer it. The strongest candidates build analogies: starting with MCP versus API, for example, they identify similarities before differences. He also looks at Google autocomplete and product-comparison discussions on Reddit, using an existing knowledge tree to locate a new concept quickly.
The parking-function interview question tests this ability. Candidates who simulate cars parking one by one can easily fall into quadratic time; the better approach is to ask, “Under what conditions is parking guaranteed to be impossible?” Once they see that order does not matter, they sort the preferences and require the first preference to be ≤1, the second ≤2, and so on. The point is to step back from forward simulation and find the structure, not necessarily to identify its relationship to Catalan numbers.
One candidate spent half an hour implementing multiplication for two strings, only to say at the end that he had forgotten the rules of long multiplication. 虞快 did not care about the lapse itself; the real question was, “Why did you waste half an hour over there? If you didn’t remember, you should have asked me at the beginning.” Similarly, a software engineer whose computer has no IDE would trigger a major red flag within minutes.
13. Coding Agents Are Pushing Engineers Toward Customers, Judgment, and Flatter Organizations
虞快 says his full management framework did not appear immediately after graduation; it took shape after roughly 4–5 years of managing at an earlier startup and working under a senior manager from Two Sigma. He had studied computer science and math, worked as a software engineer at Google for about 1.5 years on Chrome, and then moved through Two Sigma, Citadel, and startups.
A formative turning point came in his second month at Google. Working on a gRPC project that even his team did not fully understand, he complained to Takely about the lack of support. Takely replied, “If someone already clearly knew how to do this, why hasn’t it been done yet?” 虞快 realized that senior people do not automatically know the answer; a junior’s job can be to figure it out first and become the first person on the team able to teach others.
In choosing a career, he believes the sector matters more than “big company versus startup.” The strongest people around him once prioritized finance; now they prioritize AI. Asked which sector is more likely to grow 10x or 100x over the next decade, finance or AI, his answer is unequivocally AI. The real question is not company size, but “Do you want to work on AI or not?”
Coding agents have reduced the time engineers spend generating code and increased the time they spend brainstorming ideas and speaking with customers. 虞快 can send 3 requests to Cursor Bot while riding in a car and arrive home to find that, even if the output is not necessarily correct, someone has “gotten me started.” 曲凯 noted that company hierarchies are already flatter than before; 虞快 emphasized the need for proper incentives and managers who can “recognize what they’re looking at,” while 曲凯 added that founders must sell the vision to key talent before the company has a valuation or growth halo.