Pioneers Insight Method Research Author
Scale AI CEO on Meta's $14B deal, scaling Uber Eats to $80B, & what frontier labs are building next
Back to Episodes

Scale AI CEO on Meta's $14B deal, scaling Uber Eats to $80B, & what frontier labs are building next

Summary

  • Scale AI remains independent after Meta invested a little over $14 billion for 49% of its non-voting stock, according to new CEO Jason Droege. Meta received no new board seat or preferential data access; roughly 15 of Scale’s 1,100 employees moved, while Alex Wang joined Meta and retained his Scale board seat. Droege says Scale’s two major businesses each generate hundreds of millions in revenue, the business has grown every month since the deal, and it recently signed two $100 million government contracts.
  • Enterprise AI’s delivery gap is measured in reliability and implementation time, not an absence of economic value. Proofs of concept often reach 60%–70%, but closing the remainder resembles adding successive “nines” of data-center uptime: legal, policy, regulatory, accuracy, and change-management work makes important automation a six-to-12-month project. Droege’s summary is “easy to learn, hard to master.”
  • Frontier-model training has moved from quick preference rankings to hours-long demonstrations by elite professionals. A task that was choosing between two short stories 18 months ago can now require a top developer to build and explain an entire website or a PhD to teach nuanced cancer knowledge; 80% of Scale’s expert network has at least a bachelor’s degree and roughly 15% has a PhD. Behind seemingly magical models is persistent “operational chiseling”: compute, model improvement, and increasingly specialized data all improve together.
  • The next enterprise bottleneck is digitizing institution-specific judgment rather than ingesting more raw data. Scale’s healthcare example turns 200–300 pages of mixed-format records into five to ten considerations and once surfaced an allergy that conflicted with a planned medication. Off-the-shelf models, RAG, and fine-tuning can only go so far because identical words may carry different importance across companies; the local experts themselves must increasingly label “what good looks like.”
  • Models are moving from knowing things to doing things, making reinforcement-learning environments a critical infrastructure layer. Agents must learn inside realistic sandboxes—navigating a configured Salesforce instance, handling business data, completing goals, and escalating uncertain decisions—while labs seek training tasks generalizable enough to avoid collecting “forty-five trillion combinations.” Droege expects the technology could become close enough within two to three years to force difficult change-management and policy decisions.
  • Droege rejects a near-term white-collar apocalypse while openly acknowledging Scale’s incentive to keep humans involved. He does not think the transformation will happen in the next year and calls it within two years “very far-fetched,” though “nothing’s impossible here”; longer term, he argues that if these systems are to work for people, humans will need to remain in the loop for consequential decisions. His deeper thesis is that labeling has a “history of new beginnings”: as old needs fade, new human knowledge and skills become valuable.
  • Droege’s company-building framework combines independent insight, buyer urgency, structural economics, and survival. Uber Eats reconstructed restaurant economics, initially charged 30% before the market settled around 25%, and pursued incremental demand that, if demand tripled while labor stayed fixed and only ingredients scaled, could carry 70%–80% incremental gross margin. The business then grew from zero to roughly $20 billion in four and a half years. Founders still need a reason they uniquely see the opportunity, the willingness to spend five to ten years on it, and the discipline to remember that “not losing” is a prerequisite to winning.

Deep dive

1. Meta took a large minority stake, not control of Scale

  • Droege’s central clarification is categorical: Scale remains “a fully independent company.” Meta invested a little over $14 billion for 49% non-voting stock, received no new board seat, and gained no preferential access to Scale’s products, customer data, or confidential information.

  • Alex Wang now works at Meta, where he leads its super intelligence team, rather than at Scale, although he continues to fill the same Scale board seat. Droege says the board and governance remain largely unchanged, as do the privacy and data-security boundaries governing Meta’s longstanding customer relationship with Scale.

  • Only about 15 people crossed to Meta in the transaction, leaving Scale with roughly 1,100 employees. Its data business and its applications-and-services business each generate hundreds of millions in revenue—what Droege calls “two unicorns inside the company today.”

  • Droege directly challenges reports implying post-deal deterioration: the business has grown every month since the transaction. It has 250 open roles and recently signed two $100 million contracts in one month, alongside growing federal, enterprise, and international-government operations.

2. Training data moved from simple preferences to expert work

  • Droege calls competitors’ suggestion that Scale remains tied to low-skill labeling “just bogus.” His historical through-line starts with autonomous-vehicle labeling in 2016, moves through computer vision and a Department of Defense relationship in 2020, and reaches generative AI: as models improve, their data requirements change.

  • Eighteen months ago, a representative task asked a contributor which of two short stories was better and how one might be edited. Today, a single task could require one of the world’s best developers to build a complete website or an expert to explain a nuanced cancer topic.

  • Those newer tasks take hours, demand professional judgment, and often require advanced credentials. About 80% of Scale’s expert network has a bachelor’s degree or higher and roughly 15% has a PhD; some PhDs earn significant amounts contributing their expertise.

  • Scale does not merely wait for laboratories to specify work. Droege says its teams identify model weaknesses, assemble a suitable cadre of experts, and approach model builders with the offer: “We noticed that this is a problem,” and this data could help fix it.

3. The expert network compounds through reputation and referrals

  • Experts are genuinely difficult to recruit, and Scale uses referrals, campus programs, professors, students, LinkedIn, and other channels. The highest-quality contributors usually arrive through “grassroots and referral networks,” which only compound if existing experts receive a strong experience.

  • Money matters—some contributors can earn hundreds or thousands of dollars—but it is not the entire incentive. Specialists also enjoy correcting models that frustrate them and contributing to how AI handles a subject they consider important.

  • Training artifacts vary with the research objective. A model might receive a finished website, an annotation explaining “I made this decision for this reason,” an explanation of why an alternative was not chosen, or a broken site with an expert diagnosis; the deliverable is not merely code, but the judgment behind it.

4. Agents need environments before they can reliably do things

  • When Rachitsky raises the claim that reinforcement learning could eventually occupy much of the economy, Droege endorses RL’s importance without endorsing that sweeping labor conclusion. He redirects the discussion toward RL environments: realistic sandboxes in which agents learn to accomplish goals.

  • A Salesforce environment, for example, contains customer data, company-specific configurations, and business processes. An agent must understand all three, operate at high reliability, and recognize when uncertainty is high enough to “pop it up to a human being” for guidance.

  • The permutation problem is enormous: software products vary by configuration, data type, scale, user count, and complexity. Labs therefore need training tasks generalizable across many situations, rather than separately collecting “forty-five trillion combinations” of actions and conditions.

  • Droege’s simple specimen is finding the interview with Lenny on his calendar. Valuable training should generalize beyond that one retrieval—to other searches, potentially other calendar actions, and ultimately broader families of digital tasks.

5. Enterprise AI must digitize local judgment, not merely ingest data

  • Scale’s second business sells applications and services into healthcare, insurance, government, and other institutions. Its healthcare example involves specialists who handle rare cases, face a large backlog, and want both more patient capacity and fewer revisits caused by incomplete first diagnoses.

  • A physician may receive 200–300 pages of records combined into one document but stored in different formats. Human teams scan, delegate, and prioritize imperfectly; Scale’s system reads the material and highlights the five to ten factors most relevant to diagnosis and treatment.

  • In one case, the system surfaced a non-obvious allergy that conflicted with a medication the patient might otherwise have received. Droege uses it to show the potential endpoint: the AI can identify a correlation that would be difficult even for a highly trained, time-constrained human.

  • Yet the off-the-shelf model eventually reaches a ceiling. RAG, historical records, and fine-tuning do not automatically capture how this particular institution exercises judgment, so its own experts must label decisions—the emerging bottleneck Droege calls “digitizing judgment.”

6. Evals define “good” for probabilistic systems

  • Enormous data volume is not equivalent to useful training data. A bank may ingest hundreds of petabytes, but much of it does not explain how its bankers synthesize evidence, apply incentives, or make decisions differently from peers at another institution.

  • For enterprise and government customers, Scale’s work is mostly evals: comprehensive benchmarks establishing “what good looks like.” When Rachitsky asks why Droege says “good” rather than “correct,” Droege notes that these are probabilistic systems making recommendations under incomplete information.

  • His workflow-selection heuristic is asymmetric. If humans currently achieve only 10%–20% and AI can reach 50%–80%, “you’re in the money,” provided uncertain cases reach people; expecting AI to improve an already 98%-accurate process through the final 2% is “not totally there yet.”

7. Human involvement keeps shifting rather than simply disappearing

  • Droege describes data labeling as “a history of new beginnings.” Autonomous vehicles require less labeling than before, but better models expose new frontiers where different skills, contributors, environments, and judgments become valuable.

  • A world requiring no external human data would imply that no new human knowledge or skill is important enough to add to a model. Droege calls that level of advancement “almost unfathomable,” while acknowledging that Scale is financially incentivized to believe humans will remain involved.

  • His personal argument is conditional as well as commercial: if these systems are to work for people, people will need to remain in the loop on the decisions the systems make. Scale’s operational task is continually discovering which human capabilities have newly become useful to models.

  • On employment, Droege is deliberately practical: he does not think the transformation will happen in the next year, and considers it within two years “very far-fetched,” though he preserves the hedge that nothing is impossible. Technology changes work, but history suggests people adapt.

8. Enterprise automation is a six-to-12-month reliability project

  • Rachitsky presses with failed-pilot reports and findings that AI tools can sometimes slow engineers down. Droege concedes “there’s a lot of hype,” but argues that effortless prototyping inflates the denominator: companies can initiate far more experiments than they could with earlier technologies.

  • Many proofs of concept reach 60%–70%, after which the human mind assumes the rest is easy. Droege compares the remaining work to data-center uptime: moving from one “nine” to five looks numerically small but demands orders of magnitude more reliability engineering.

  • He calls the widely circulated 95% pilot-failure figure somewhat clickbait—“it tells the right story,” but hyperbolically. Serious programs need experienced builders plus legal, policy, regulatory, accuracy, and change-management work before an important process can be automated at an acceptable level.

  • The realistic timeline is six to 12 months, “months, not minutes.” Once deployed, the outcome can still be startling—even an elite doctor may see a finding they would have missed—but the work beneath the result is closer to laying broadband than performing magic: “Someone’s gotta dig up the road.”

9. Scour taught Droege that every business rule is negotiable

  • Droege and Travis Kalanick were 19 or 20 when they ran Scour from dorm-room computers at scour.cs.ucla.edu. They expected trouble for parking the domain on UCLA infrastructure; instead, the computer-science department was excited, an early lesson that assumed constraints may not be real.

  • Financing made the lesson harsher. Terms swung from a few million dollars at a $5 million valuation to demands for 50%, 75%, and finally 80% of the company under same-day pressure. Droege concluded that “there is no way to do things”—only outcomes people can negotiate by aligning incentives.

  • Scour was later sued for $250 billion by entertainment-industry associations, which settled for $1 million. The disparity revealed the tactical objective: drive the company out of the market. Even established institutions, Droege learned, can invent numbers rather than follow a stable playbook.

10. Customer discovery starts with incentives and reconstructed economics

  • Droege does not take customer statements literally; he investigates incentives, including ego, career advancement, and an executive sponsor’s need to trust a vendor with a risky project. Product adoption depends on what the buyer must hear and achieve personally, not only the nominal financial ROI.

  • When restaurateurs would not disclose reliable economics, the Uber Eats team ordered meals, weighed the ham, cheese, bread, and lettuce, and matched them against a supplier catalog. It triangulated this independent ground truth with restaurant claims and what “site guys” were saying about restaurant economics.

  • The reconstruction suggested ingredients consumed roughly 20%–30% of a meal, labor another 20%–30%, and real estate around 10%. Uber initially charged restaurants 30%; despite comparisons to Groupon, the eventual clearing price was around 25%, close enough to validate the model.

  • Incremental demand was the key incentive: if demand tripled while labor stayed the same and only ingredients scaled, the added meals could carry roughly 70%–80% incremental gross margin. Restaurateurs disliked that simplified framing because reality was more complicated. A valuable but non-urgent product still fails—if it is not high on the buyer’s daily agenda, “you’re just gonna have a long road to a small TAM.”

11. A new business needs proprietary insight and structural quality

  • Droege frames entrepreneurship as a search for market alpha: “Why am I so lucky to have this insight?” In a world containing a million smart entrepreneurs trying ideas, a founder needs a credible reason they see something others do not—and why they are uniquely positioned to act.

  • The second test is endurance: “Why do I wanna work on this problem for 5 to 10 years?” Discovering one customer problem is insufficient. Founders need a burning desire to keep questioning themselves and must avoid falling in love with an idea at the expense of the customer mission.

  • His most important success factor is a founder who remains “a force of nature over a long duration of time,” preserving the energy to pivot through years of difficulty. But force applied to a structurally poor market still carries avoidable odds.

  • Droege therefore filters for recurring revenue, stickiness, network effects, lock-in, and businesses that become more valuable at large scale. Studying what could plausibly become a $100 billion company eliminates weak ideas before passion selects among the survivors.

12. Wide exploration made Uber Eats’ signal impossible to miss

  • Droege keeps the aperture open until evidence coalesces. One experiment put 250 convenience-store SKUs into ten vans in Washington, DC; demand was below “crickets” because the team misunderstood that cigarettes, beer, and Slurpees drew the traffic supporting everything else.

  • Grocery’s picking-and-packing economics frightened him, while generalized point-to-point delivery lacked meaningful consumer demand in 2014. After roughly 15 variations, food delivery stood apart: demand was rising, unit economics worked, and non-prime real estate could compete through “prime food.”

  • Uber Eats launched in Toronto in December 2015 and recorded about $20,000 in sales within two hours. Droege describes its trajectory as zero to roughly $20 billion in four and a half years; during COVID it rose from about $20 billion to $50 billion in a year and later approached $80 billion.

  • Droege refuses to turn that result into pure foresight: “Luck is part of the game.” His point is not to begrudge fortunate timing, but to combine experimentation with enough operating readiness to capitalize when demand suddenly accelerates.

13. McDonald’s converted strategic stubbornness into distribution leverage

  • Uber Eats initially positioned itself as an ally to independent restaurants, so Droege rejected McDonald’s approach as inconsistent with the product’s “vibe.” He delayed for four or five months even after McDonald’s emphasized its roughly 80 million daily consumers.

  • His team eventually told him he was being irrational. The delay nevertheless appears to have helped Uber secure an exclusive relationship and major customer acquisition; with a $17 basket, Droege’s instruction was essentially “figure it out” through delivery radius, pricing, and other economic levers.

  • Uber activated McDonald’s globally in about six months while Uber Eats itself was less than two years old, using a makeshift organization to satisfy an 80-year-old company’s process expectations. Three months later, Droege says, growth hockey-sticked again at a higher level.

14. Gross margin exposes differentiation, while survival preserves the upside

  • Droege treats high gross margin plus healthy churn curves as a coarse test of value creation. When presented with a proposed 40% margin, he asks, “Start at a 60% gross margin. Why does that not work?” The answer rapidly reveals alternatives, pricing power, and differentiation.

  • If the alternative is a mature offshore provider earning 20%, an aspiring entrant’s 40% will likely compress faster than forecast. A defensible answer might be that competitors can copy the product now but will not be able to after two years of fast execution; absent that, margin erosion is the base case.

  • Rachitsky’s Costco counterexample is worth preserving: low margin can itself build a moat. Droege agrees that Costco and Walmart use price to absorb demand, establish habits, deepen supplier relationships, and develop operating expertise until competing at 8% against an incumbent at 10% becomes punishing.

  • “Not losing” precedes winning because founders must survive long enough for timing, insight, and product to converge. Droege favors asymmetrically positive decisions over reflexive “go for it” culture: risks remain necessary, but one enterprise-compromising bet can remove every future opportunity.

15. Adaptability is the enduring operating advantage

  • Droege learned risk discipline from a profitable used-golf-club business that reached a couple million dollars in revenue and paid dividends. Early eBay margins inspired the hubristic ambition to buy every used club in America, but easy entry crushed the logic; insufficient thinking upfront created years of pain.

  • His hiring view has become more nuanced: perhaps 5% of roles require exact, current expertise or customer relationships because speed precludes training. For most roles he tests three things—curious problem-solving articulated clearly, humility and cross-functional collaboration, and leadership.

  • Uber Eats’ management team was composed as an “organism of strengths,” with members compensating for one another’s weaknesses. Much of it remained intact from zero to $20 billion, because mutual knowledge and the capacity to learn mattered more than whether each person had previously operated at that scale.

  • Droege applies the same adaptability personally: he uses voice-mode AI as a commute-time tutor and asks it to identify the most important point in internal documents before verifying the answer. His sustaining motto is “the end is never the end”—an imperfect next step usually exists, and tomorrow preserves another chance to act.