Artificial Analysis
Key Views & Dialogues
Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
- 🗓️ Date:
2026-01-08| 🎙️ Show:Latent Space
Artificial Analysis makes independence its product, selling standardized reports and private benchmarking while keeping public placement free of pay-to-play incentives. Reliable rankings require repeated runs, mystery-shopping lab endpoints, and evolving tests because measured capabilities quickly become optimization targets.
View Dialogue Notes & Key Takeaways
Artificial Analysis has turned independence itself into the product: no one pays for public placement, while enterprises buy decision reports and AI builders buy private benchmarking. The company now has just over 20 people and two customer groups, pairing free public data with subscriptions and bespoke work. swyx’s shorthand is the “presumptive new Gartner of AI,” but the commercial firewall is the thesis: “There’s no use doing what we do unless it’s independent AI benchmarking.”
Reliable evals cost far more than firing off a question set because prompts, output parsing, answer ordering, sampling variance, endpoint manipulation, and repeated runs can all move rankings. Artificial Analysis targets roughly ±1 point at 95% confidence for its Intelligence Index, making its real costs higher than the one-repeat figure shown publicly. Private lab endpoints are checked through a “mystery shopper policy” using unidentified accounts, because the served model might differ from the endpoint supplied for testing.
The Intelligence Index has to keep changing because yesterday’s frontier tests are now saturated and successful benchmarks rapidly become optimization targets. V1 tasks such as HumanEval-style Python functions would be close to trivial for many current models; V3 blends 10 datasets spanning Q&A, agents, long-context reasoning, and use cases. The core warning is that “things that get measured become things that get targeted,” so improving competition-math scores need not equal broader intelligence.
The frontier has shifted from an apparently unassailable OpenAI lead to a market that Micah-Hill Smith described as “strictly more competitive every quarter.” At recording, Gemini 3 Pro High led, followed by Claude Opus 4.5, GPT-5.1 High, and Kimi K2 Thinking. DeepSeek’s decisive signal arrived with the open-weight DeepSeek V3, described as a 61.1B MoE, on Boxing Day the previous year—before R1 made the broader world pay attention.
Artificial Analysis is expanding “intelligence” beyond percentage-correct scores by explicitly penalizing hallucination and rewarding “I don’t know.” Its Omniscience metric runs from -100 to +100, subtracting a point for a wrong factual answer; Claude models showed the lowest hallucination rates, while general intelligence had little correlation with knowing when to abstain. The trade-off is contextual: on Critical Point’s research-level physics problems, where the top score was only 9%, researchers deliberately raise temperature because exploratory hallucination can be useful.
Agentic performance depends as much on the harness as the underlying model, creating a new competitive layer above model APIs. GDPval-AA turns 44 white-collar tasks with spreadsheets, PDFs, presentations, audio, and video into a model-agnostic benchmark; every tested model performed better in Artificial Analysis’s harness than in its corresponding consumer chatbot. That minimalist harness—web tools, filesystem, code execution, context management, and image viewing—was released as Stirrup, with the advice to “let the models work as long as they want.”
The infrastructure paradox is that GPT-4-level intelligence is at least 100× cheaper, yet total inference spending can still rise by orders of magnitude. Larger sparse frontier models, reasoning tokens, long agentic workflows, and repeated turns overwhelm declining unit costs; one startup was reportedly spending $5,000 per employee on coding agents alone. George Cameron expects both trends to continue: another order-of-magnitude reduction in comparable intelligence cost and another order-of-magnitude expansion in feasible consumption.
The next efficiency battleground is not “reasoning versus non-reasoning” but whether models spend the right tokens and turns on each problem. Earlier in the year, reasoning models averaged 10× the tokens per Intelligence Index query; now model-to-model token efficiency itself varies by more than an order of magnitude. In τ²-bench Telecom, GPT-5 could be cheaper overall than smaller open models despite pricier tokens because it resolved the customer’s problem in fewer turns.
🔗 Original source & video: Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith