Terence Tao – How the world’s top mathematician uses AI
Terence Tao – How the world’s top mathematician uses AI
Summary
- Tao’s core call for anyone underwriting AI-for-science: idea generation now costs ~zero, verification is the bottleneck. “AI has driven the cost of idea generation down to almost zero, in a very similar way to how the internet drove the cost of communication down to almost zero. It’s an amazing thing, but it doesn’t create abundance by itself” — journals are already flooded, human peer review doesn’t scale to a thousand AI theories a day, and “that’s not something we know how to do at scale.”
- The Erdős-problem scoreboard is the cleanest live benchmark: ~50 of ~1,100 solved with AI assistance, then a plateau. Three separate attempts to have frontier-model AIs attack every problem simultaneously produced no new pure-AI solutions; systematic studies show a 1–2% per-problem success rate — “they can buy scale, and you just pick the winners.” Expect the same publicity dynamic on prestigious open problems, and standardized benchmarks are needed rather than relying on AI companies to disclose negative results.
- Tao’s own productivity number is a caution against headline multipliers: ~5x on auxiliary tasks (plots, literature search, formatting), little change on the core — “the core of what I do, actually solving the most difficult part of a math problem, hasn’t changed too much. I still use pen and paper for that.” Papers are “richer and broader, but not necessarily deeper.”
- His diagnosis of the missing ingredient: current systems are artificial cleverness, not yet the kind of intelligence he describes — jumping machines that leap two meters but “can’t jump a little bit, reach some handhold, stay there, pull other people up.” No cumulative learning across sessions; hence his timeline call that “hybrid human plus AIs will dominate mathematics for a lot longer,” requiring “additional breakthroughs beyond what we already have.”
- The bullish flip Dwarkesh lands and Tao accepts: AI breadth is qualitatively new — once models reach a waterline they clear every problem at that height simultaneously, which humans can’t do. Tao: redesign science around breadth, run experimental math on thousands of problems at once — “the idea of doing mathematics at scale is at its infancy… and then science will be unrecognizable after that, I think.”
- A concrete tail-risk worth knowing: mathematicians believe the Riemann hypothesis via the “random model of the primes,” and if RH were false, “I think we would very rapidly abandon any cryptography based on the primes” — one unknown pattern would probably imply more, and patterns mean exploits.
- Track record check: Tao’s 2023 prediction that by 2026 AI would be “a trustworthy co-author if used correctly” is “looking pretty good in retrospect.” Forward call: within a decade, a lot of what math students currently do—and a lot of what goes into papers today—can be done by AI — but as with human computers and $1,000 genome sequencing, “we moved on. You move to a different scale.”
Deep dive
1. Kepler was a “high-temperature LLM” — but Brahe’s data made him
- Dwarkesh’s opening gambit: Kepler spent twenty years trying random relationships — Platonic solids inscribed between planetary spheres, astrology, musical harmonics (Earth’s misery explained because its note is “mi-fa-mi”) — and the third law appears as an aside in a book on world harmonics. “Kepler was a high-temperature LLM.” Tao’s extension: as long as there’s a verifiable data bank like Brahe’s, random generation plus verification drives deep progress, and LLMs can do this kind of thing.
- Dwarkesh’s rebalancing: idea generation is the prestige step we celebrate, but “it has to be matched by an equal amount of verification, otherwise it’s slop.” Brahe’s decades of naked-eye observations were ten times more precise than anything prior — “that extra decimal point of accuracy was essential for Kepler to get his results.”
- Tao’s statistical caution, as told: the third law was regression on six data points, and Kepler got lucky. Johann Bode fit the same data to a shifted geometric progression, predicted a missing planet — Uranus and Ceres fit exactly, everyone got excited — “but then Neptune was discovered, and it was way off. Basically it was just a numerical fluke.”
2. Idea generation is free; verification doesn’t scale
- The episode’s most quotable economics: “AI has driven the cost of idea generation down to almost zero, in a very similar way to how the internet drove the cost of communication down to almost zero. It’s an amazing thing, but it doesn’t create abundance by itself. Now the bottleneck is different.”
- The institutional problem: peer review was the wall built against amateur theories of the universe, and it’s breaking — “many journals are reporting that AI-generated submissions are just flooding their submissions.” Per paper, scientists can debate to consensus “in a few years. But when we’re generating a thousand of these every day, this doesn’t work.”
- Dwarkesh’s meta-observation on method: science has partly reversed — classically you hypothesized then collected data; now you collect big data first and extract hypotheses. Tao’s counter: Kepler already fits that pattern — his 1595–96 Platonic theory was wrong, and only Brahe’s dataset plus twenty years of trying random things produced the empirical regularity.
3. You can’t reinforcement-learn the value judgment
- On identifying the next “bit” among a million AI papers, Tao’s honest answer: “a lot of it’s the test of time.” Deep learning was a niche, controversial area for years; the transformer didn’t have to win; base ten beats Roman numerals but “there’s nothing special about ten” — value depends on future adoption and culture, “so it may never be something that you can just reinforcement learn.”
- Dwarkesh’s pushback — worth keeping: correct theories often look worse at first. Copernicus was less accurate than Ptolemy’s millennium-tuned geocentrism; Newton was stunned by inertial-gravitational mass equivalence and Leibniz attacked action at a distance — all resolved only by Einstein. How would AI peer review recognize progress in a falsifiable-but-flawed theory? Tao’s addition: progress often means deleting assumptions, not adding theories — geocentrism survived because Aristotelian physics said objects want to rest.
- On Darwin vs Newton (Origin 1859, Principia 1687, and Huxley’s “how stupid not to have thought of that”): Tao credits communication — Darwin wrote persuasively in plain English; Newton wrote in Latin, invented new math to explain himself, and hid insights from rivals. “How can you score how persuasive you are?… Maybe that will be forever the human side of science.”
- The frame he offers for the moment: “we’re going through a cognitive version of the Copernican revolution” — human intelligence isn’t the center of the universe, and “our assessment of which tasks require intelligence… has to be reordered quite a bit.”
4. Erdős scoreboard: fifty down, then a plateau
- Status check on his own post: ~50 of ~1,100 Erdős problems solved with AI assistance, ~600 open, and the one-shot era has stopped — “not for lack of trying. I know of three separate attempts to get frontier model AIs to just attack every single one of the problems simultaneously,” yielding minor observations or literature finds, no new pure-AI solutions.
- What actually fell: “almost all of the 50 problems that were solved by AIs were ones for which there was basically no literature” — the solution was combining one obscure technique with an existing result. “That’s the median level of what AI can accomplish, and that’s really great.”
- His signature analogy: cliffs in the dark, and AI tools as “jumping machines that can jump two meters in the air, higher than any human” — they cleared all the low walls in one exciting period, but “they’ve been really bad at creating partial progress or identifying intermediate stages.”
- The base-rate warning: “whenever we do a systematic study, on any given problem an AI tool has a success rate of maybe 1% or 2%. It’s just that they can buy scale, and you just pick the winners.” Hence the push for standardized challenge sets rather than relying on “AI companies to only publish their wins and not disclose their negative results.”
5. Breadth × depth: redesign science for math at scale
- Dwarkesh’s bull case, which Tao accepts: the bearish read (walls lower than humans reach) and the bullish read are the same fact — once AIs hit a waterline they clear every problem at that height simultaneously. “We can’t make a million copies of you and give each of them a million dollars of inference compute.”
- Tao’s synthesis: “they excel at breadth, and humans excel at depth… very complementary.” New fields could be explored by broad, moderately competent AIs mapping the terrain and making easy observations, then “identify certain islands of difficulty, which human experts can then come and work on.”
- Math has been almost purely theoretical; AI enables an experimental side — take a thousand problems and test which methods actually work, like a software company hunting workflows that scale. “The idea of doing mathematics at scale is at its infancy… We don’t even have the paradigms to really take full advantage of it. But we will, and then science will be unrecognizable after that, I think.”
6. Tao’s own uplift: 5x auxiliary, little core
- Pressed for a “2x more productive by what year” prediction, Tao refuses the one-dimensional frame: his papers now have far more code and plots, and today’s papers “would definitely take five times longer” without AI — “but I would not write my papers that way.” Reproducing a 2020-style paper “actually hasn’t saved that much time, to be honest.” Richer and broader, “not necessarily deeper.”
- On error-checking against frontier models on tasks he can do himself: “sometimes they pick up errors that I make. Sometimes I pick up errors that they make. It’s about a tie right now.” But when standard techniques leave holes, chasing AI suggestions “wastes more time than it saves.”
- His temperature check: “the progress is simultaneously amazing and disappointing” — and people acclimatize fast, as with Google search. “2026-level AI would be stunning in 2021.” His 2023 prediction — a “trustworthy co-author if used correctly” by 2026 — is “looking pretty good in retrospect.”
7. Cleverness isn’t cumulative intelligence: no handholds, no session carryover
- His definition by example: real collaboration is adaptive — prototype a strategy, test, fail, modify, cumulatively map what works. The jumping robots “can jump and fail, and jump and fail. But what they can’t do is jump a little bit, reach some handhold, stay there, pull other people up, and then try to jump from there.”
- Dwarkesh’s sharpening, confirmed flatly: if Gemini 3 or Claude 4.5 solves a problem, its own understanding of math has not progressed — “you run a new session and it’s forgotten what it just did.” Dwarkesh’s caveat: “maybe what you just did is 0.001% of the training data for the next generation.”
8. Incomprehensible proofs don’t scare him; strategies need a language
- Could a Lean proof of the Riemann hypothesis be assembly-code gobbledygook? “We don’t know.” The four color theorem was brute-forced and “we have still not found a conceptually elegant proof… maybe we never will.” But RH is prized precisely because “we’re pretty sure a new type of mathematics has to be created” — though “it could be false actually,” in which case a computed zero off the line “would be very disappointing.”
- Why he’s “not as worried” about incomprehensible proofs: Lean lets you study each lemma atomically, and post-processing already happens on the Erdős site — an AI generates 3,000 lines, other AIs summarize, humans rewrite. He foresees “entire professions of mathematicians who might take a giant Lean-generated proof and do some ablation on it.”
- His wish (not plan): a semi-formal language for strategies and plausibility, doing for conjecture what Lean did for proof — critically with no backdoors, “because reinforcement learning is just so good at finding these backdoors.” His worked example: Gauss computing ~100,000 primes and conjecturing the prime number theorem, maybe the first important statistical conjecture, which grew into the “random model of the primes” — heuristic, non-rigorous, “extremely accurate,” and the reason everyone believes the twin prime conjecture and RH. If RH failed, “I think we would very rapidly abandon any cryptography based on the primes.”
- The data problem underneath: we have one timeline and ~100 turning-point stories. “If we had access to a million alien civilizations,” we might formalize what progress is — his practical proxy: evolve small AIs on simple problems as “little laboratories” (e.g., smallest network that does 10-digit multiplication).
9. Timelines, serendipity, and advice for the unpredictable era
- Dwarkesh’s autonomous-AI bet: if a Millennium Prize Problem is solved this year, he would put 95% odds that an AI did it autonomously. Tao says “hybrid human plus AIs will dominate mathematics for a lot longer”… It will require some additional breakthroughs beyond what we already have. Within a decade, a lot of what math students currently do—and a lot of what goes into papers today—can be done by AI — and, as with 19th-century equation-solvers, human computers, and genome sequencing that went from a full PhD to $1,000, “we moved on. You move to a different scale.”
- A hedge with teeth: “it’s also possible that by destroying serendipity we actually inhibit certain types of progress.” His evidence is personal — library browsing that surfaced adjacent papers, hallway coffee encounters lost to COVID scheduling, and a year at the Institute for Advanced Study where he ran out of inspiration after several months: “you actually do need a certain level of distraction in your life.”
- Career advice: adaptable mindset, keep credentials — but note the access shift: “now it’s quite possible at the high school level… you could get involved in a math project and actually make a real contribution because of all these AI tools, Lean, and everything else.” Closing register, hedged exactly as he hedged it: “It’s a scary time, but also very exciting.”