What Today’s Best Models Still Can’t Do in Math
What Today’s Best Models Still Can’t Do in Math
Summary
- Daniel Litt’s favorite fully autonomous result is the mid-May solution to the Irish unit-distance problem, which he found “in some ways a little bit creative” rather than a last-mile cleanup of human work. Lisha notes that the community thought the statement was true, but a counterexample emerged using classical 1960s techniques that were new to point configurations in the plane. Mathematicians then used those ideas to find counterexamples to other open questions, including the sum-product conjecture over the real numbers. Litt’s test is post hoc: did the new ideas prove useful elsewhere or improve understanding?
- The capability frontier is sharply lopsided: models are “very, very good at applying known techniques”—grinding computation and pulling together technical ideas from many papers—but weak at intuition, big-picture philosophy, and theory-building. Neither Fable nor ChatGPT 5.6 is good at building theory autonomously, though with hints they can do something interesting. Litt’s hedge is that if they can do it with “a hundred bits of hints,” perhaps they can do it without hints in six months.
- AI math proofs read as human, not alien—“it’s not like there’s some ‘move 37.’” Lisha points out that labs appear to scale natural-language reasoning rather than Lean-verified proofs. Litt says this suggests that, although people often explain math progress through verifiability, the techniques may generalize to other domains; he labels that conclusion “just my guess.”
- Short clever proofs are a verification ceiling, not simply a style choice: “the reason they’re not producing long complicated proofs is that they cannot”—models generating long outputs “might not know they’re wrong.” OpenAI’s recent list of 10 problems was formalized in Lean, while Litt says an 800-page AI-generated claimed proof of resolution of singularities in positive characteristic is almost certainly wrong. Harnesses that elicit 250-page papers trade away reliability.
- The incentive structure is already at risk: postdocs can produce papers by “playing the slot machine” until a model emits a hopefully correct proof. Litt told Codex to find five recent algebraic-geometry conjectures and prove them; with some back-and-forth, it produced three “quite bad, but correct papers” in an hour. He and Lisha discuss cases where 3–5 papers used “the exact same proof of the exact same theorem” within days. The systemic risk is “one mathematician duplicated a thousand times” rather than a million people pursuing different curiosities.
- Anthropic and OpenAI are “pretty neck and neck,” with Claude catching up around Opus 4.5 or Opus 4.6 after being unhelpful for research math for a long time. A new rank-30 elliptic-curve result is interesting but unevaluable without knowing its methods; some recent results have involved Levent Poge, with unclear degrees of autonomy. Historically, records of this kind are cool but “not an Annals-level result.”
- Even if models become “really robustly superhuman,” Lisha says “we still want human mathematicians.” Litt argues that optimality does not guarantee broad, diverse fundamental research: autonomous systems might pursue a direct path instead. A community with broad interests can keep human control, push models toward varied research, and preserve the pipeline that develops people able to think mathematically.
- Litt’s central principle is that “the goal of mathematics is not to produce mathematics papers; it’s to produce some kind of understanding.” Model-weight understanding would be unsatisfying to him. Human mathematicians should keep developing their faculties and use AI to deepen understanding rather than relinquish the thinking to it.
Deep dive
1. The one autonomous result that clears Litt’s bar: the Irish unit-distance counterexample
- Litt’s taxonomy of AI results: some are autonomous, some semi-autonomous, and in some the AI contribution is “not at all clear.” Many are “last mile”—deep human work with the AI taking the final step. His favorite fully autonomous result is the solution to the Irish unit-distance problem, announced in mid-May, because it seemed “in some ways a little bit creative.”
- What made it creative: Lisha says people working in the area thought the statement was true, but a counterexample was found. The result brought in techniques “from another area”—“classical ideas from the 60s,” not deep or new, but new to the study of point configurations in the plane. It was fruitful: mathematicians then used the ideas to find counterexamples to other open questions, “for example the sum-product conjecture over the real numbers.”
- Litt’s scoring function is post hoc: look at whatever new ideas were introduced and ask whether they were “useful to do other things” or “improve our understanding of something.”
2. The proofs are human, and natural-language reasoning may generalize
- Against Lisha’s “inhuman feats” framing, Litt’s correction is that the released chain-of-thought summary “was very recognizable”—something like a human mathematician’s reasoning. Across the results he has studied, “it’s not like there’s some ‘move 37’”; it is “like a human mathematician doing certain types of math.” The more inhuman aspects are that models do not get tired and know a lot.
- Lisha observes that labs appear to scale reasoning in natural language rather than a large corpus of Lean-verified proofs. Litt says people often argue that math is a verifiable domain, but, “because they’re primarily scaling in formal reasoning,” he guesses the techniques may generalize to other domains. He emphasizes that this is “just my guess.”
3. Lab scoreboard: 5.6 vs. Fable, neck and neck at a jagged frontier
- Litt sees the two labs solving “a very similar collection of problems”—OpenAI drops a solution and Anthropic says it knows how to solve it too—and that collection is “a relatively small portion of what human mathematicians do.” He mostly uses ChatGPT partly out of inertia: “for a long time the Claude models were just not useful for research math,” then around Opus 4.5 or Opus 4.6 they more or less caught up.
- Lisha’s anecdotal split is that 5.6 gives a clearer theory of mind about what she does and does not know, while Fable explains something trivial and then jumps ahead. Litt’s verdict is that both are “pretty bad at theory.” He has tried to get both Fable and ChatGPT 5.6 to build theories; “they definitely are not good at it autonomously,” though hints can produce something interesting. He adds that it is hard to tell which part comes from the model and which part comes from the person providing the hints.
4. What mathematicians actually do: open problems as benchmarks, analogies as engines
- Litt self-identifies as a problem solver rather than primarily a theory builder, but says an open problem is meant to measure a failure of understanding: “It’s kind of like a benchmark.” His example is the Grothendieck p-curvature conjecture, which measures a failure to understand differential equations.
- Another mode is philosophy-driven: his work is motivated by an analogy between the homology of algebraic varieties and representations of fundamental groups. “Any phenomenon that appears on one side, you can find an analog on the other.” Pursuing that analogy led to decades of mathematics by Carlos Simpson, Takuro Mochizuki, and others. “That philosophy is super nonrigorous, actually. It’s not symbol pushing at all.”
- Finding the right question is often the hard part. The Birch and Swinnerton-Dyer conjecture was “the first big-data conjecture”: Birch and Swinnerton-Dyer collected statistics on elliptic curves in the 1960s, graphed them, and noticed a relationship between a slope and an algebraic invariant.
- On AI’s fit to these activities: “the more vague a phenomenon is, or the less you have a precise question in mind, the less useful it happens to be.” For projects he has pursued for 3–5 years, AI is “primarily a substitute for Google.” But because Litt is bad at coding, AI has also unlocked projects involving massively parallel example-hunting—asking it to work through 1,000 examples in parallel, or 10 at a time with 10 subagents—that he otherwise might have postponed for months.
5. Against beauty: “win by any means necessary”
- Litt tries not to be motivated by aesthetics and names a failure mode among young mathematicians: abandoning a proof because it “feels really ugly.” His response is, “what if you’re wrong and it’s not ugly? Why limit yourself?” He says, “You should win by any means necessary.”
- His preferred compass is “doing kind of physics, except with concepts”: asking what is fundamental and what will open up further understanding, rather than trying to produce mathematical art. He acknowledges that many mathematicians see themselves as closer to poets.
- His broader sociological point, echoed later in the AI debate, is that progress comes from people pursuing personal curiosity. Many people with different ideas of what is interesting let “a thousand different flowers bloom”; the frontier expands, creating opportunities for new ideas to cascade into answers to old questions.
6. Why theory-building resists the models—and how failing to grind produced a better theorem
- Litt thinks one reason models struggle with some of his hard conjectures is that the conjectures are generally believed to be true and fit into broad theoretical frameworks. Unlike a false conjecture that can be refuted by a specific construction, these may require resolving other conjectures and developing “very serious new ideas.” He stresses: “I’m not saying the models won’t be able to do this; it’s just that so far they seem not to.”
- He is not skeptical of continued capability growth. His guess is that it is “probably totally doable” and perhaps requires a different RL environment. The difficulty is that theory-building is fuzzier to reward: it is easier to reward a proved conjecture than an intermediate improvement in understanding.
- His strongest case study came from a paper where Gemini Deep Think, then still on the frontier, was useful for proving lemmas. None of the frontier models could prove one lemma. Litt worked through many examples, realized why a better statement might be true, and then the models quickly proved that improved statement.
- Today, ChatGPT 5.6 Pro can prove the original lemma with “the worst proof you’ve ever seen”—10 pages of brutal calculation with no insight. Litt realized that such a grind proof would work but could not bring himself to do it, so he looked for a conceptual argument. “Our inability to just grind is kind of important to our ability to make discoveries.”
7. Slop economics: the slot machine and mode collapse on arXiv
- The incentive problem is that a postdoc on the market may, “for the next couple years before the community adapts,” produce papers by “playing the slot machine until the model produces a hopefully correct proof.” Litt’s experiment was to tell Codex to “find 5 recent conjectures in algebraic geometry, and prove them.” With some back-and-forth, it produced three “quite bad but correct papers” in an hour; they were sitting on his hard drive, waiting for him to contact the relevant people.
- The tell of low quality is not necessarily wrongness but absent human engagement: “no evidence that a human being is actually engaged with it,” and no development of human capital or understanding. Lisha cites cases where “3, 4, or 5 papers with the exact same proof of the exact same theorem” appeared within days. She calls this a kind of mode collapse; Litt says ChatGPT is consistently finding the same thing.
- Litt does not think the problem obviously disappears as models improve. If mathematical exploration is subordinated to what the models pursue, the field might get “one mathematician duplicated a thousand times” rather than “a million different mathematicians doing a million different things.”
8. Short proofs are a checking ceiling, and models cannot yet unit-test arguments
- Lisha relays Mark Sellke and Mehtab Sawhney’s observation that AI proofs are often charmingly short. Litt’s explanation is that “the reason they’re not producing long complicated proofs is that they cannot”: the ability to check correctness is not yet there, and a model producing a long argument “might not know they’re wrong.”
- He guesses OpenAI and Anthropic have solved more problems than they have released, but some are not formalizable because prerequisite results are not yet in Mathlib, while longer proofs are harder to check. OpenAI’s recent list of 10 problems was formalized in Lean, which is “very good evidence” that they are true.
- Exhibit A is an 800-page AI-generated claimed proof of resolution of singularities in positive characteristic. Litt has not read it, but says “there’s no way it’s correct,” that no human has read it, and that current models cannot check such an argument.
- Harnesses built to elicit long proofs can decrease reliability. Someone trying to produce a 250-page paper may be pushing the model out of its careful mode; “any harness that can elicit a 250-page paper is probably not being very careful about what it’s producing.”
- Litt’s favorite unpassed test is an unnamed wrong paper whose precise error was subtle, but whose overall structure immediately suggested to experts that it proved something too strong. Humans stress-test long arguments globally—asking whether they would imply something known to be false or work in a special case. Models still cannot perform this “fuzzy unit-testing of a proof” reliably.
9. Even robustly superhuman AI leaves humans with a job—if the pipeline survives
- Litt’s central claim is that “the goal of mathematics is not to produce mathematics papers—it’s to produce some kind of understanding.” Some understanding might reside in model weights, but “to me, that’s pretty unsatisfying.”
- Frontier researchers rely on “thousands, millions, or billions” of people learning to think mathematically. Existing incentives may instead reward postdocs for producing many papers, including papers generated by models, without developing human capital or understanding.
- Lisha says that even if models become robustly superhuman, “we still want human mathematicians.” Litt responds that even if autonomous research were optimal, there is no guarantee models would pursue broad fundamental research rather than a direct path toward instrumental goals. The easiest way to help ensure diverse research is to maintain a community with broad interests that pushes the models and helps shape society.
- Lisha and Litt also discuss the risk that cheap, slightly worse outputs displace high-quality work. Litt thinks AI can improve quality along every dimension, but doing so requires thoughtfulness and institutional redesign. Lisha describes a classroom “bimodal distribution”: some students learn to use the tools, while others let AI do their homework and then fail elsewhere.
- On parenting, Litt says his three-year-old daughter is starting to add; he jokes that “icosahedron” was one of her first words, while clarifying that she learned the Platonic solids early. Lisha says her own daughter can count to about 30 reliably and 50 semireliably, and describes finding her under the blankets saying, “Oh, I’m doing some math.” Lisha says she teaches addition and subtraction in the context of a general group; Litt says group theory is next. Their shared educational thesis is that math helps people think clearly and understand the world, even when extremely capable AIs exist.