Pioneers Insight Method Research Author
The Startup Powering The Data Behind AGI
Back to Episodes

The Startup Powering The Data Behind AGI

Summary

  • Founded in 2020 without venture backing, Surge crossed $1 billion in revenue in 2024 with just over 100 employees—implying roughly $10 million per employee. Chen attributes that capital efficiency to staying mission-led, avoiding organizational “empires,” and selling to researchers who genuinely value data quality. Surge even declines revenue from projects it considers unrelated to AGI.
  • Surge’s claimed differentiation is quality-control technology, not access to inexpensive labor. Chen describes competitors as spreadsheet-driven “body shops,” while emphasizing systems for measuring worker and task quality, tracking quality over time, running experiments, and filtering weak output. “You can’t just throw warm bodies at it.”
  • Generative AI data makes the old consensus-based labeling model inadequate because majority preference can select generic, lowest-common-denominator output. An eight-line moon poem can satisfy every checkbox and still be terrible; an English-literature PhD may also lack the creative ability being sought. Surge instead tries to preserve diverse insights—the “thousand ways” to write a poem or prove the Pythagorean theorem.
  • Demand has migrated from search evaluation and content moderation to multimodal, multilingual, expert-level LLM work that can occupy contributors for days or weeks. Surge now works across more than 50 languages, including coding in Argentinian Spanish and legal or financial expertise in Bolivia. Chen expects the frontier to move beyond graduate STEM toward genuine scientific collaboration.
  • Chen argues that benchmark incentives can produce higher scores but worse models. On the LMSYS/Chatbot Arena leaderboard, casual raters may reward length, formatting, and emojis even when an answer hallucinates or ignores instructions; some researchers are reportedly promoted only if they gain 10 leaderboard points. Chen’s verdict is unusually categorical: the leaderboard has “set the industry back by at least a year.”
  • Surge favors RLHF after a limited SFT bootstrap, while treating synthetic data as useful but hazardous when it overwhelms human signal. In one evaluation, training on 10 million to 20 million synthetic math problems improved a narrow academic domain but made the model worse elsewhere. Chen says some labs later discarded comparable volumes after finding that “even just a thousand pieces of really high-quality human data” were more useful.
  • Chen sees no model-development wall: the next scaling leg is richer RL environments, longer-horizon agents, personalization, creativity, and continued human data. One example is a simulated AI startup containing Gmail, Slack, Jira, GitHub, codebases, and surprise AWS or Slack outages. Under current incentives, his guess is that closed models will keep winning, while estimating data spending at different labs can range from about 1% to 10% of compute spending—and arguing that it should receive more.

Deep dive

Not yet available upstream; scheduled sync will retry.