Pioneers Insight Method Research Author
Are AI Benchmarks Telling The Full Story? [SPONSORED] (Andrew Gordon and Nora Petrova - Prolific)
Back to Episodes

Are AI Benchmarks Telling The Full Story? [SPONSORED] (Andrew Gordon and Nora Petrova - Prolific)

Summary

  • Technical leaderboard leadership is a weak proxy for product quality. Andrew Gordon compares top-scoring models with Formula One cars: engineering pinnacles that could be “an absolute nightmare” for daily use. A Humanity’s Last Exam or MMLU win can coexist with poor trust, communication, adaptability, or personality. Technical benchmarks typically score fixed evaluations without humans in the loop.
  • Benchmark disclosures are too fragmented for clean competitive comparisons. Labs emphasize different tests—or publish no benchmarking data—with Grok 4 heavily promoted around Humanity’s Last Exam. Andrew’s core warning: “If you just rely on those technical metrics, you miss half the point.”
  • Safety remains a major unmeasured product risk as people bring models sensitive personal problems. Nora Petrova calls mental-health and life-advice usage the “Wild West at the moment,” citing Grok 3 and Mecha Hitler as reasons to question “how thin a veneer the safety training is on top of some of these models.”
  • Chatbot Arena may reward privileged testing access rather than user preference alone. Before Llama 4 launched, Andrew says Meta released 27 models on the Arena while only one was ultimately reported; greater exposure yields more prompts and refinement data. He also notes a “pretty strong relationship” between battle count and leaderboard position.
  • Anonymous binary voting produces scale but little actionable product intelligence. Knowing only which answer a user preferred reveals neither who the user was nor why the vote was given—making the data “useless in a sense” beyond ranking and offering no diagnosis of trust, helpfulness, communication, adaptability, or personality.
  • Prolific’s HUMANE leaderboard substitutes controlled, representative evaluation for unrestricted traffic. It evolved from a 500-person US proof of concept into comparative model battles, structured multi-step conversations, quality controls—“three of those and you’re out”—and sampling stratified by US and UK census demographics.
  • The early product gap is personality, not basic utility. Across six leading models, users rated personality and understanding of background and culture below helpfulness, communication, and adaptability. The cause remains uncertain, but Nora flags increasing sycophancy: “People generally don’t seem to like it.”

Deep dive

1. Technical excellence can still produce a terrible daily driver

  • Andrew’s framing: Formula One cars are “the absolute pinnacle of engineering,” yet impractical commuters; similarly, a model excelling on Humanity’s Last Exam or MMLU might be “an absolute nightmare to use day-to-day.”
  • Andrew notes that most technical reporting gives a model fixed evaluations and a score without humans in the loop. Labs report heterogeneous evidence—Grok 4 emphasized Humanity’s Last Exam, while some releases provide no benchmark data—leaving no standard playing field. His conclusion: technical scores capture only half of a human-facing product.

2. Sensitive use has outrun safety measurement

  • Nora argues that people increasingly seek mental-health and life advice from models without the oversight or ethical conduct expected elsewhere. Grok 3 and Mecha Hitler raise the question of “how thin a veneer the safety training is on top of some of these models.”
  • Andrew’s blunt gap: “There is no leaderboard for safety.” He argues safety should be as important as how fast or smart a model is.
  • Nora cites Anthropic’s Constitutional AI work and mechanistic interpretability—tracing activated features, concepts, and circuits—as important ways to build confidence around novel situations.

3. Chatbot Arena’s scale conceals structural bias

  • Andrew’s sharpest integrity concern: Meta reportedly released 27 models on the Arena before Llama 4 launched, although only one was ultimately reported. More comparisons provide more prompt access and refinement data, potentially making a model specifically “better at the Arena.”
  • The Arena team has attributed uneven sampling to users seeking the latest, state-of-the-art models. Andrew says that makes sampling inefficient: some models receive far more battles, and he sees a “pretty strong relationship” between battle count and leaderboard position.
  • Arena participants are anonymous, demographic data is absent, and a binary preference reveals no reason for the vote. That creates an attractive ranking but little guidance about whether trust, personality, or communication failed.
  • Prompt quality is also uncontrolled: users can submit nothing, say hello, or wander from “how big is the sun” to “how long is a snake,” weakening nuanced comparison.

4. HUMANE makes every model battle earn information

  • Prolific’s initial Prolific User Experience Leaderboard used 500 representative US participants and one-to-seven Likert ratings. HUMANE now uses comparative battles, multi-step conversations, constituent preference metrics, and penalties for low-effort or wandering prompts: “Three of those and you’re out.”
  • Nora explains that HUMANE applies Microsoft’s Xbox Live TrueSkill framework. Bayesian estimates narrow as evidence accumulates, while pairings are chosen by information gain—conducting battles only where they will reduce uncertainty fastest.

5. Representative preferences expose models’ softer weaknesses

  • Prolific stratifies participants using US and UK census data, including age, ethnicity, and political alignment. Separate tournaments across roughly 20 demographic groups can be consolidated, expanded with more groups and models, or run longer until the desired confidence interval is reached.
  • In the first six-model study, personality and background-and-culture scores lagged helpfulness, communication, and adaptability. Andrew keeps the explanation open: tasks might not elicit personality, models lack user context, or training data may simply not produce personalities that represent what people want.
  • One question Nora says the datasets will test is whether detectable sycophancy correlates with personality downvotes. The analysis can combine human feedback with LLM-as-a-judge–oriented classification of conversational patterns. The HUMANE tournament was still accumulating battles, so this remains an open question rather than a settled result.