Pioneers Insight Method Research Author
137. 洪乐潼: AI for Math, Lean, and Is Math Created or Discovered?
Back to Episodes

137. 洪乐潼: AI for Math, Lean, and Is Math Created or Discovered?

Summary

  • Axiom’s core bet is not a single model, but an AI mathematician system built on Lean as a verifiable foundation—pushing intelligence and correctness higher at the same time. 洪乐潼 puts it bluntly: “I bet system, I don’t bet model.” Post-training models, specialized models, deterministic tooling, sub-agents and provers work together: cheap rule-based systems solve what they can first, and the remainder goes to larger models. The moat therefore sits in data, tooling, orchestration and the self-improvement loop—not just parameter count.
  • The roughly seven-month-old team has demonstrated exceptional execution speed through fundraising and benchmarks, but has yet to prove its commercial endgame. Axiom raised a $64M seed round at a $300M valuation, then completed a Series A at a $1.6B valuation and at least $200M in funding six months later. On the technical side, it achieved a perfect score on all 12 Putnam problems in four months—洪乐潼 calls it the 6th perfect score in the contest’s 98-year history after 5 human perfect scorers—and reached 98.93% on the Verina code-verification benchmark, versus 11% for the older DeepSeek Coder version she cited.
  • The hardest-to-replicate asset is scarce Lean data and the engineering stack built around formal proof. Public Lean tokens are scarce and the language is “extremely brittle,” so Axiom built roughly 12 to 13 generation and verification tools, making its proof verifier about 100x faster than the existing comparator. Unable to afford AlphaProof-style Monte Carlo Tree Search, the team instead pursued sub-agents, skills and experience learning, expanding a single proof graph from roughly 40 nodes to 4,000.
  • Code and chip verification are Axiom’s clearest initial markets, but the commercial bottleneck is specification, not proof. 洪乐潼’s formula is “program × specification → verification condition × proof,” and Axiom currently addresses the final term. The vision is that “anything you can define, you can execute”: programs would be proven correct as they are generated, reducing reliance on finite test cases. She points to an Amazon team spending 3 to 5 years writing 260,000 lines of proof code to verify a memory-isolation component as evidence that high-value verification will open before mass-market products.
  • A complete AI mathematician needs a prover, conjecturer, knowledge base and auto-formalization layer; the weakest links today are definitions and conjectures. A prover has a clean 0/1 verification signal, while a conjecturer must also judge importance and elegance. More difficult still is breaking a 20-page paper into a 200- to 500-page blueprint and translating it into Lean. Some research problems cannot even be entered because the necessary definitions do not exist in Mathlib, leaving a wide gap between “can prove” and “can build theory.”
  • Axiom deliberately mixes mathematics, RL/agents, compilers, code generation and Lean rather than betting on a single school of thought. 舒博 brings roughly 20 years in GPUs, 10 years in AI and experience with Baidu’s Silicon Valley AI team; Ken Ono, Evan Chen and multiple authors of AI for Math papers have joined, taking the team from employee 15 to roughly 30 people. 洪乐潼 describes the organization as a bottom-up “math club,” with the CEO serving as “the person handing out water at the back.”
  • The biggest investment risk is the binary outcome typical of moonshots: the technology can succeed while the business or organization stalls first. 洪乐潼 worries both about “not executing fast enough” and about becoming so anxious to move quickly that the team makes a strategic mistake. She acknowledges that running deep tech at 24 is a handicap: she must make a succession of high-stakes decisions quickly while taking responsibility for employees, capital and the company’s future. Her formulation is stark: “Either the rocket launches, or the rocket crashes.”
  • If formal mathematics produces a ChatGPT- or Cursor-level product moment, it could provide a narrow but deep path from specialized superintelligence into code, chips and AI for Science. 洪乐潼 expects mathematics to move from “math-poor” to “math-rich,” with human mathematicians retaining the “0.01% of intuition” and allocating compute while AI handles exploration at scale. She is also betting on recursive self-improvement, orchestrators, sub-agents and formal verification as an RL reward, and expects to see small continual-learning models, strong multimodal reasoning and a scaled-up agent economy in 2026.

Deep dive

1. The Most Moving Moments in Mathematics Come When Distant Fields Suddenly Meet

  • What most often stops 洪乐潼 in her tracks is not solving a problem herself, but reading someone else’s proof and thinking, “How can this be so beautiful?” She cites the correspondence between modular forms and elliptic curves: an algebraic expression and a geometric object are connected, and the beauty lies in the intersection between fields.

  • At the Ross Mathematics Program, she first worked all the way through quadratic reciprocity. Each day, students had to finish an entire problem set before receiving the next one, like clearing another level in a game. Seeing that one concise proposition could support multiple deep and creative proofs confirmed that she had experienced a mathematical moment “like a lightning strike.”

2. Mathematics Is a Civilizational Contract Built Around Axioms

  • Asked what mathematics actually is, 洪乐潼 describes it as a shared human civilization that mathematicians create together: first agreeing on which axioms can be taken for granted, then building upward from that foundation. “In a sense, mathematicians have a contract.”

  • Different schools prefer different foundations. Some start from compression and want as few axioms as possible; others care more about whether the axioms are interesting, natural and coherent together. Theory often starts with sequences, sets and concrete examples, then grows through patterns, natural questions and proofs.

  • She ultimately places mathematics between art and science. It is neither pure problem-solving nor mechanical deduction; choosing definitions, deciding which patterns to pursue and judging which structures are natural all require creative aesthetic judgment.

3. The Discovery-versus-Creation Debate Ultimately Turns on Whether Proof Can Convince the Community

  • 洪乐潼 sees Ramanujan’s entry into the British mathematical community as the clearest example. He arrived like an “alien,” carrying notebooks full of correct formulas but without the proof training familiar to Hardy and Littlewood. Proof allowed an unfamiliar intuition to cross cultural and disciplinary boundaries and be accepted by the existing community.

  • Her sharpest definition of proof is “influence.” Once the logic is written down with complete rigor, a discovery earns admission to the shared knowledge base. Only then can the community debate whether the proof is beautiful or natural.

  • 张小珺 asked whether she belonged to the intuition-genius camp or the proof camp. Her candid answer: “I’ve always really wanted to be an intuition-genius type,” but in practice she is closer to a brute-force player. That gap became her personal entry point into understanding AI proof systems.

4. Zhang Yitang and Maynard Show That Proof Styles Can Differ, but Results Must Be Checkable

  • 洪乐潼 recalls that number theorists initially found Zhang Yitang’s method “very different” from the schools they knew. But the result on bounded gaps between primes held up under scrutiny, so the community accepted it first, then learned, organized and simplified the method.

  • She places the techniques and student lineage that formed around James Maynard after his 2022 Fields Medal alongside Zhang’s route. Two proof traditions can be compared, complemented and compressed; the process itself is “a very interesting intellectual exercise.”

5. Her “Brute Force” Meant Forcing Geometry She Did Not Understand into Symbols

  • Her Olympiad coaches said the first Euclidean geometry problem was supposed to be free points, but 洪乐潼 could not solve it for years. She converted every point and line into complex-number notation and relied on brute force: even without understanding the geometry, she would grind through the calculations, taking 2 to 3 times as long as others.

  • At MIT, while working with Henry Cohn on a small problem related to 28-dimensional sphere packing, she went 6 months without a result but reported her attempts and failures every week. She does not see herself as fast, but as “the cockroach that cannot be killed.” She is willing to do the dirty, repetitive work that more talented competitors avoid.

  • She later felt an immediate connection when she saw AlphaGeometry translate Euclidean geometry into symbolic form. The implementation was different, but the philosophy matched what she had done as a teenager. The result she cites is that the system solved roughly 81% of historical IMO geometry problems.

6. AI Systems Also Contain Both an “Intuition” and a “Brute-Force” Mathematician

  • The Axiom system is not one model. Some models quickly judge whether a proof should proceed in “one, two, three, four, five, six” broad steps, like a mathematician who can sketch an outline; others know Lean’s formal language and advance one tactic at a time to keep every logical link airtight.

  • 洪乐潼 believes the two types of AI mathematician should reinforce each other. The first finds plausible directions; the second closes every gap. Betting on only one creates either a hallucination wall or a search-efficiency wall.

7. The Putnam Perfect Score Exposed the Contrast Between One Human Diagram and Thousands of Lines of AI Code

  • When Axiom Prover entered the Putnam, Evan Chen looked at one problem and drew a single diagram. Everyone in the room immediately knew he had solved it. The AI did not find that creative route; instead, it generated thousands of lines of Lean code, effectively enumerating cases and checking them one by one until it reached the answer.

  • 洪乐潼 considers this difference more interesting than simply imitating humans. A problem with an obviously elegant solution can still be solved by a machine in its own preferred way, and the distinct mathematics behind human and AI proofs becomes a new object of study.

  • Her historical tally is that the Putnam mathematics competition has had 5 human perfect scorers in its 98-year history, making Axiom Prover the 6th. The source does not further specify the competition’s starting year or call it the first AI perfect score.

8. Automated Theorem Proving Predates Deep Neural Networks by Decades

  • 洪乐潼 rejects describing today’s breakthroughs simply as “a victory for AI.” Long before deep learning, computer scientists had spent decades building ATP, or automated theorem proving based on rules. When that failed, humans worked with systems through ITP, or interactive theorem proving.

  • What has changed is that AI is replacing the human inside ITP. She sees AI for Math as the intersection of the old ATP tradition and modern AI, which is also how she answers whether mathematics is a human privilege: at least at the proof-execution layer, it no longer is.

9. A Ten-Minute Childhood Walk in Guangzhou Left Her Most Cherished Form of Free Attention

  • 洪乐潼 was born and raised in Guangzhou, with a home roughly a 10-minute walk from school. She often thought about mathematics while walking and would end up “not knowing where I had walked to.” That state of mind—no task chasing her and room for the brain to wander—is the childhood experience she misses most since becoming a founder.

  • She distinguishes bounded attention from free attention. The former is boxed in by emails, deadlines and required tasks; the latter has no explicit objective, but can let intuition and insight enter the mind. Free attention does not produce linearly: sometimes nothing comes to mind in the moment, and the idea only calls back later during a monotonous task.

  • Asked whether there is an optimal ratio between the two, 洪乐潼 offered no formula: “The ratio is different for everyone.” She is certain only that turning every day into military-style execution would destroy a large number of strategic opportunities.

10. Founders Need a Mixed Formation of Visionaries, Executors and Salespeople

  • 洪乐潼 relays a 3-part founder taxonomy. The visionary supplies foresight and ambition, the executor compounds day-to-day execution, and the salesperson influences different audiences through communication and assembles the team. She sees herself as neither sales-oriented nor an especially strong executor, but closer to a visionary willing to set targets because she is extremely optimistic.

  • Her examples are categorical: Musk is the visionary, Zuckerberg the “execution fanatic,” and Sam Altman the salesperson. But a team cannot consist only of complementary types; it also needs people who resemble one another. Around Zuckerberg, she notes, there are dreamers as well as steady operators in the Sandberg mold.

  • The diagnosis also exposes her organizational risk: knowing what the founder lacks does not mean the gap has been filled. Axiom’s later recruitment of experienced engineering and research talent was, in effect, a stabilizer for her high-variance profile.

11. Facebook’s Bottom-Up Culture Came to Axiom Through a Cohort of Veterans

  • 洪乐潼 uses “bleed purple” to describe how deeply Facebook employees internalized the culture. The campus resembled a dream factory, with a 24-hour ice-cream shop, musical instruments and bold colors, but the more important feature was that culture grew from the bottom up. In her comparison, Google was more top-down.

  • Most of Axiom’s early employees came from Facebook, so the company effectively “never had an opportunity to define a culture.” It inherited the former team’s working habits organically. 洪乐潼 later described Axiom as a technocratic organization in which technical staff set direction and the CEO minimizes low-level intervention.

12. Higher Mathematics Rewrote the Olympiad’s Zero-Sum Game as a Positive-Sum Game

  • In 7th grade, 洪乐潼 suddenly realized she did not have to keep solving Olympiad problems designed for older students. She could read calculus, real analysis and complex analysis directly. Higher mathematics had no fixed syllabus: each new definition generated new theorems and questions, and the depth and breadth of exploration were no longer bounded by one exam.

  • That freedom hurt her competition strategy. Rationally, she should have secured the geometry and algebra points on a 4-problem paper first. Instead, she would start with the number theory problem she was unlikely to solve because “number theory is so beautiful.” Interest and rank-maximizing strategy collided head-on.

  • She later summarized the experience as an infinite game. You can stand on the shoulders of giants and continue building your own universe, rather than defeating classmates for a finite number of places.

13. She Refused to Be Tamed by a Building Ranked by Scores

  • The elementary-school Olympiad program reshuffled all classes from 1 to 24 every semester: class 1 was at the bottom of the building, while classes 22 to 24 were at the top. 洪乐潼 started in class 4, next to the restroom, and naturally wanted to climb to the top where “the view was better.” She admits that a 4th-grader is unlikely to remain unaffected by such a structure’s competitive pull.

  • By 7th and 8th grade, after she began reading higher mathematics, she decided, “This isn’t fun. I don’t want to play this game.” She still took the exams, but stopped allowing rankings to define all of her learning. That was when she began looking for a tribe: people who were also willing to read books unrelated to exams.

14. A Three-to-Five-Person Knight’s-Tour Tribe Anticipated Collaborative Mathematics

  • With 3 to 5 classmates, she studied a knight’s tour on an n×n board: starting from any square, could a knight visit every square exactly once? They tried to prove by induction that the result held for every n at least 5.

  • To construct the induction base case, the group worked together in class with strips of paper covered in grids, passing notes saying, “I found another n=9 case.” A complex construction that one person could not complete was broken into small tasks the group could pick up.

  • 洪乐潼 compares the experience with the large formalization projects involving 陶哲轩, Kevin Buzzard, Alex Kontorovich and others. Mathematics is not only “one genius solving a century-old conjecture alone”; like software engineering, it can be completed by hundreds of people through structured decomposition.

15. Divergent Learning May Produce Transfer Learning, but She Remains Uncertain About the Exact Mechanism

  • Before the key exam for moving from middle school to high school, she felt she had done almost no systematic preparation and worried she would fail. She apparently earned a perfect score, including on geometry problems she had never previously been able to solve. Her explanation was not that she had suddenly become a genius, but that higher mathematics had quietly transferred into her competition ability.

  • That became a personal prior for how she understands AI. She believes there is “definitely some degree of transfer” between mathematical and coding ability: structured reasoning learned in one domain may carry into another.

  • She also supplies the counterpoint. At MIT, she discovered that focus was indispensable. Mathematical ability has no single formula; divergence and concentration are constraints that matter at different stages.

16. At MIT, She Assumed She Was “the Stupidest Person in the Entire Math Department”

  • 任秋宇, 张盛桐 and 高济阳 were people she had looked up to in the news as a child. At MIT, taking classes and doing problem sets with them produced no psychological shock because her prior expectation was simply, “I’m the dumbest one here.”

  • 张小珺 pointed out that others viewed her as a prodigy. 洪乐潼 insisted this was not false modesty: many people around her knew their IMO scores and their own talent, while she saw it as luck that she could speak and collaborate with them.

  • Low expectations did not make her retreat, but they changed her career forecast. She once assumed a math PhD program might not choose her because everyone else was an IMO medalist, so she began treating quantitative finance as a realistic fallback.

17. Ken Ono’s Waitlist Gave Her a Reason to Turn Down a Bridgewater Internship

  • When the pandemic moved her Bridgewater internship online, Ken Ono’s undergraduate research program placed her near the top of its waitlist. The “geniuses” around her had already received offers, but she was thrilled to make the waitlist and worried that if she missed the opportunity, she might not even qualify the following year.

  • Someone eventually withdrew because of the pandemic and she was admitted to the REU, so she chose mathematical research over Bridgewater. She still visited Bridgewater after finishing college; more than half of the paper she completed with Ken Ono during her undergraduate years was co-authored with Ken and summer-program collaborators.

18. Honors Did Not Rewrite Her Inner Narrative That Failure Is the Default

  • She once worked through 75 Olympiad problem sets in a short period, doing and tearing off one page every day, and still was not selected for the competition. Her relationship with honors is therefore complicated: the objective result may be good, while the subjective experience is always that of “the person who worked hardest and still saw no result.”

  • The Morgan Prize for North American undergraduate mathematics is awarded through faculty nominations. 洪乐潼 believes her award contained a great deal of randomness and encouragement; students who were not recognized that year may have been no weaker mathematically. She refuses to reverse-engineer the prize into proof that the achievement was effortless.

  • She treats failure as the default precisely because her goals come from extreme optimism. She repeatedly believes she can reach farther, so falling short becomes normal rather than an anomaly that should trigger withdrawal.

19. A 4 in Graduate Probability Forced Her to Relearn the Analysis She Had Read at 14

  • MIT protected freshmen grades during the first semester, and older students with IMO, IPhO and IOI backgrounds encouraged newcomers to take the hardest courses. 洪乐潼 and several other 1st-years therefore enrolled in graduate probability, expecting difficult probability calculations but finding Borel sigma-algebras and measure theory in the first lecture.

  • The midterm was worth 40 points and the class average was roughly 9. Several freshmen scored below 5; she saw a 4 on her paper. The failure was not an isolated humiliation: after looking at one another, the students jointly decided either to drop together or continue together.

  • She continued, attributing the result to a weak real-analysis foundation, and went back to study Rudin seriously. Failure became feedback on course selection rather than a verdict on her ceiling.

20. Her Reward Signal First Came from Teammates, Then Shifted to the Problem Itself After the Pandemic

  • When she was younger, her biggest reward was camaraderie: climbing together, refusing to quit together and finishing a difficult problem set as a group. MIT encouraged collaboration, and small teams provided the safety of “we’ll carry this together.”

  • The pandemic emptied campus during the 2nd semester of her freshman year. The small group disappeared overnight, but the coursework continued. She was forced to extract pleasure from the work itself and later admitted to becoming somewhat addicted to pain and suffering: “Most founders I know are addicted to suffering.”

  • She cites the aggressive VC phrase “chip on the shoulder, chips in the pocket”—old wounds can turn into money—but keeps her judgment intact: the mechanism “isn’t necessarily healthy.” It simply explains part of what she observed.

21. MIT Left Her Not with a Myth of Intelligence, but with the Stamina to Run Through a Blizzard

  • When she returned to Boston during a red-alert blizzard, she still saw familiar MIT students and members out running. To her, the school represents “do whatever is hard, do whatever is painful, do whatever is long-term,” and that atmosphere set the endurance threshold she later brought to entrepreneurship.

  • 洪乐潼’s view of leadership also comes from a mountaineering metaphor. The real leader is not the person at the front with a megaphone, but the person at the back handing out water. “The best leadership may be service.” Technical depth and service-oriented influence need to coexist.

  • She believes pressure tests every relationship. Preserving one’s core while keeping a team from breaking apart under pressure is rarer than any abstract leadership slogan.

22. “Lego Prover” Captures Why She Prefers Building Theory to Finite Competitions

  • Add one definition, recombine old concepts, and a new small universe appears, like Lego. 洪乐潼 sees the upward growth of theory as a form of pleasure unconstrained by finite rankings, which also explains why she preferred higher mathematics to repeatedly grinding through the same syllabus.

  • As a child, she could appear “not hardworking,” but would obsess over a problem she liked for a very long time. The difference between play and work is not the number of hours, but the subject’s own perception: “Do you feel that you are playing, or working?”

23. The Math-and-Physics Double Major Was About Understanding Professors’ “Physical Meaning”

  • 洪乐潼 knew from her first day of college that she would major in mathematics. She later chose physics over computer science, partly because of her interest in quantum mechanics and partly because research by Scott Sheffield and others on random surfaces and geometric probability repeatedly referred to what was “meaningful in physics.”

  • The decision again went against her existing strengths. She says she was “particularly bad” at middle-school physics, which was precisely why she thought it was worth filling the gap. She describes herself as high variance and spiky, with both her strengths and weaknesses sharply defined.

24. Oxford Neuroscience Did Not Keep Her, but Computational Neuroscience Pushed Her Toward AI

  • Because of family experiences, she wanted to understand the human brain and went to Oxford to study neuroscience. But obtaining a license for animal experiments required killing a mouse. After completing that experiment, she decided to move into computational neuroscience: “I don’t want to do animal experiments.”

  • She had previously done a simpler fruit-fly experiment. Working with more complex animals made her realize that deep understanding of the brain often depends on wet-lab work. Later interactions with Andrew Saxe, Tim Behrens and other researchers made AI, matrix operations, ODEs and theoretical machine learning the parts she actually enjoyed.

  • Her master’s work included neural dynamics in continual learning and one-layer linear Transformer questions. Several advisors had to help supply enough neuroscience narrative for the thesis; she was candid: “The happiness came from AI, not neuroscience.”

25. Law-School Textualism Accidentally Became the Starting Point for AI for Math

  • While studying for a math PhD and a law degree at Stanford, she examined 3 approaches to interpreting the U.S. Constitution. Originalism asks what the founders intended, textualism reads the text according to its structure and literal meaning, and living constitutionalism lets the Constitution breathe with the times.

  • She considers herself a textualist: read words as they were written and reason from them line by line. That judicial philosophy resembles the mathematician’s treatment of definitions. When someone in class proposed using an LLM to interpret the Constitution through historical and contemporary material, she immediately thought: “If AI can already tell us what the Constitution means, why can’t I use AI to do mathematics?”

  • Law was not merely a side branch. She sees antitrust and contract law as tree-shaped logic, while constitutional law and trials train narrative and interpretation. Those different cognitive forms later entered her understanding of structure, specification and proof.

26. Lean Was the First Time “Mathematical Structure” Became an Executable Object

  • Her friend Kenny Lau had worked on Lean and Mathlib since 2020 and was one of the 5 to 7 people in Kevin Buzzard’s student network entering undergraduate algebra and analysis foundations line by line. Through him, 洪乐潼 learned that mathematics need not live only in English; it could become code.

  • Lean is a formal language, while Mathlib is its mathematical library, analogous to PyTorch on top of Python. Definitions, theorems and proofs can execute: a checkmark appears when they pass, and failure identifies the exact problematic line.

  • That gives AI a structure and verification signal that natural language lacks. The Constitution’s “more rigorous English” prompted the question, but trainable, checkable mathematics requires taking the next step into Lean.

27. Reusable Number-Theory Machinery Convinced Her AI for Math Was More an Engineering Problem

  • She uses Ben Green’s work on shifted primes as an example. A proof machinery should not apply to only one paper; it may transfer to function fields and other structurally similar problems. Much of mathematical work involves skillfully invoking existing tools rather than reinventing intuition from scratch.

  • Her initial goal was not to build another Ramanujan immediately, but to bring AI to the level of a PhD student who knows number-theory machinery well enough to execute standard arguments, enumeration and verification automatically.

  • After reading the history of science and hundreds of papers, her conclusion changed: “This may not be a research problem; it is an engineering problem.” The technology risk was lower than she had first thought, making venture capital the responsible way to pursue it.

28. A Curtain Pulled at Verve Became a Founding Partnership

  • While in law school, she would take a shuttle bus to Palo Alto’s Verve on weekends with a thick book, intending only to read cases, drink matcha and watch the dogs in the courtyard. 舒博 happened to be a regular at the 6-person communal table. They began talking after he helped pull the curtain because the sun was too bright.

  • They first exchanged, “I see you here all the time,” then talked about science, history, mathematics and technology for a year and a half. She did not know he was a senior director at Meta AI, and he did not know she had a substantial mathematical research background.

  • The story supports her affection for Silicon Valley, where identity, age, academia and industry can remain porous: “Who everyone is doesn’t matter as much as what they are doing.”

29. A Morning Run in Fall 2024 Turned Scientific Interest into a GPU Budget

  • In fall 2024, 洪乐潼 had just started her math PhD and was getting exposure to more compute through XTX. She felt the set of things AI could do expanding suddenly. After a morning run, she became certain that “this really has to happen” and immediately asked 舒博 to estimate the number of GPUs required.

  • They worked out the budget on a napkin at Verve and concluded that the resource and engineering organization could not be built inside academia; it required a company. She had been leaning toward entrepreneurship since September, but did not make the final decision until November.

  • She left school even later because she first needed a work visa. The sequence was to raise money, secure immigration status and then leave, rather than use a dropout posture as part of a fundraising narrative.

30. She Spent 2 Months Convincing Herself Not to Become a “Restless AI Founder”

  • 洪乐潼 once refused even the free sushi offered by MIT’s entrepreneurship club, believing professors and long-term research were more worthwhile. She was especially skeptical of product-driven projects that flare up and disappear, and questioned whether the venture-capital cycle could accommodate genuinely difficult long-duration problems.

  • Her way of testing the thesis was not to repeat a mission statement, but to read scientific history, map AI for Math GitHub repositories, read hundreds of abstracts and reconstruct the valuable papers à la Feynman. She had to verify the technical route first: “I can’t convince myself it will succeed and then go deceive other people into giving me money.”

  • Her conclusion was not that startups are glamorous, but that she could not find another structure capable of carrying the compute, talent and engineering complexity. Entrepreneurship was simply the necessary organizational form.

31. She Even Asked a Competitor Whether It Could Hire Her

  • Tudor, another regular at Verve, happened to run a leading competitor. 洪乐潼 asked directly whether the company was hiring: if she could join someone else doing the work, she would not have to shoulder the complexity of starting from zero. The answer was that they hired only computer-science PhDs, while she was a math PhD.

  • The rejection clarified her motivation. Her first priority was making AI for Math happen; becoming CEO came second. She calls herself “the least likely entrepreneur to become an entrepreneur,” not someone searching for a territory to conquer.

32. The AI for Math Papers She Read Became a Hiring Map

  • Before founding the company, she learned from papers on ATPBoost, PatternBoost and problem-to-solution translation. Some researchers worked on formal proof, some extracted patterns from graph data and constructed examples and counterexamples, and others studied mathematical discovery.

  • The most surreal part was watching authors she had once looked up to join Axiom one after another. “I was an apprentice looking up at the shoulders of these giants,” and ended up working with them on the same team.

  • Hiring therefore was not a matter of scaling a uniform job description. It meant reassembling critical capabilities scattered across AI for Math, code generation and formal languages.

33. Axiom Needs to Speak Four Languages at Once: AI, Compilers, Lean and Pure Mathematics

  • One part of the team works on reinforcement learning, agents and applied AI; another comes from code generation, compilers and LLM compilers; a third focuses on Lean, Mathlib and metaprogramming; and a fourth includes pure and competition mathematicians such as Ken Ono and Evan Chen.

  • Lean itself requires specialization. Some people build Mathlib, while others treat Lean as a programming language and add tools and abstraction layers. She cites one colleague writing autograd in Lean to show that this is not a matter of “mathematicians learning a little code.”

  • The team has roughly 5 people with IMO experience, including interns, but still emphasizes strong engineering ability when hiring machine-learning talent. A mathematical prior helps people understand the problem; it does not substitute for building the system.

34. Mathematicians Become Net Assets Only If They Accept Scaling

  • When 舒博 was working on Deep Speech and Deep Voice, the team used the phrase “don’t hire a single linguist” to express its scaling-first philosophy. Axiom initially made a similar agreement: do not hire mathematicians among the first 15 employees, lest mathematics become an artisanal craft like a Japanese sushi master’s technique.

  • Reality corrected that extreme. Mathematical training is not a negative, but candidates must be open-minded. One person accepted an offer and later left because they did not want to work on an “internet-scale dataset,” showing that craftsmanship and scaling can genuinely conflict.

  • Today, the mathematician’s role is closer to an adversarial benchmark builder: continually finding system weaknesses and designing harder problems, rather than asking engineers to copy human problem-solving habits.

35. A Christmas Survey Connected Scattered Ideas into a Complete Map

  • During the 2024 Christmas holiday, 洪乐潼 and 舒博 formed a reading group to systematically study the formal theorem-proving literature. A survey titled “Formal Theorem Proving: The Next Frontier of AI” divided the field into multiple quadrants and showed her how 5 previously isolated methods fit into one landscape.

  • She checked every paper mentioned in the text and references against her own notes, finding that roughly half remained unread. Once they understood the big picture, they began asking which new methods could connect to chess-like expert systems, AI for Science or the early tradition of automated reasoning.

  • Both usually disliked Zoom. She lasted at most an hour and 舒博 at most 45 minutes, yet they could talk continuously for 4 to 5 hours. That became evidence that their thinking was both similar and complementary.

36. 舒博 Brought Historical Memory from the Scaling Era, Not Just a Resume

  • 洪乐潼 summarizes 舒博’s background as roughly 20 years in GPUs and 10 years in AI. He was among the early GPU developers and later worked on Deep Speech and Deep Voice at Baidu’s Silicon Valley AI team, alongside people including 吴恩来.

  • The central lesson of that generation was “scaling works”: put large amounts of speech and voice data into training instead of first relying on domain experts to hand-code rules. That DNA entered Axiom, but Lean’s verifiability and mathematicians’ adversarial perspective impose new constraints.

  • 舒博 decided to join in February 2025, after the company received a more credible financing offer. 洪乐潼 had expected to wait a year and prove herself before becoming “worthy” of inviting an industrial veteran of his stature.

37. The Seed Round Was First and Foremost an Endurance Test in Repetition

  • “Nobody likes fundraising.” 洪乐潼 says the difficulty was not necessarily the outcome but the exhaustion: repeating the same story and answering the same questions until she wanted to record the pitch once and send it to everyone.

  • She spoke with dozens of investors. Many funds neither rejected her clearly nor committed first, preferring to wait for someone else to lead and then follow.

  • Lead offers roughly doubled, tripled and then rose again, producing several competitive bids. The first investor wanted about 50%; she rejected it outright. She was prepared to accept the second offer until B Capital’s offer changed the final decision.

38. Howard Morgan Made a Fundraising Conversation Feel Like Research Again

  • B Capital’s eventual lead, Howard Morgan, co-founded Renaissance with Jim Simons and also co-founded First Round. 洪乐潼 spoke with him on Zoom while racing toward a paper rebuttal deadline, only to find him more optimistic than she was and even volunteering to articulate the business model for her.

  • She had previously spent 2 hours in conversation with Jim Simons at MIT, so the history connecting mathematics and investing carried strong emotional weight. Howard’s perspective and the better price led her to abandon the second offer she had been ready to accept.

  • Her conclusion was that most fundraising conversations are dull, while the rare stimulating ones usually come from the investor you ultimately choose. Capital is not only about price; it is also about how quickly someone understands a long-term problem.

39. Her Full Disclosure of Commercial Risk Nearly Vanished Under the VC Discount Model

  • During the seed round, 洪乐潼 explicitly said the business model was uncertain. A stronger AI mathematician would probably be useful, and quantitative finance was one example, but it might not be the ultimate market. The first priority was to build the technology.

  • She later learned that investors habitually discount founder optimism. If a founder says 10, the investor may record 8. If she says 7 and it is discounted again, the number may disappear entirely.

  • Her candor did not prevent the financing, but it showed she was not a conventional fundraiser. Her principle was, “I describe the thing as it is,” including risks, unresolved markets and what she did not know.

40. The $64M Seed Round Bought the Team Its First Stretch of Runway

  • The seed round closed at $64M, above the original $50M plan, at a $300M valuation. After signing the term sheet, the company still had to incorporate, complete legal diligence, find an office and handle immigration; the round did not close until summer.

  • She was taking law and math classes while working at XTX, fundraising and recruiting. The investors were preparing to write the term sheet before anyone realized she had not yet incorporated. Organizational costs and technical costs arrived simultaneously.

  • AI talent often had 6 offers, while a startup could not match a “$100M package.” She describes recruiting as a repeated decision to “jump”: candidates had leverage, and the company had to decide whether to keep raising its offer.

41. The Penalty for Doing Deep Tech at 24 Was Lacking a World Model for Repeated Decisions

  • 洪乐潼 believes youth can be an advantage for consumer products but a disadvantage in deep tech. She had no track record leading a technical team, yet had to make a succession of high-stakes decisions under extreme time pressure.

  • She cites Zuckerberg’s story of facing Peter Thiel’s time-limited term sheet at 19. A young founder may cry in the bathroom, return and accept terms he dislikes. More information does not necessarily make the decision better; waiting itself may be high-risk and low-return.

  • At one women founders’ event, she was asked to go first on a zip line. Because the people behind her were becoming impatient, she closed her eyes and jumped. An investor told her she needed to build the muscle memory of taking “the leap of faith.”

  • She says entrepreneurship often requires “giving away value and taking a loss.” It is not one grand gamble, but a daily repetition across people, capital and partnerships. If she could start again, she would read 3 times as many books because “nothing I’ve learned has been enough.”

42. The Series A Was Preempted Before the Company Had a Deck or Planned to Raise

  • After the seed round closed in summer, an investor made a preemptive offer at Christmas. Knowing Axiom had no fundraising materials and had not opened a new round, the investor issued a term sheet anyway to bypass the conventional process.

  • On January 5, she was called to pitch out of town and received an offer that night. A second arrived 1 to 1.5 weeks later. She chose the latter while allowing the first investor to participate as well. The source does not specify the year corresponding to January 5.

  • The round totaled at least $200M at a $1.6B valuation and was led by Menlo Ventures. Menlo partner Matt Canning had been helpful since the seed round.

43. Menlo’s Decision to Double Down Was Based on 6 Months of Nearly Error-Free Engineering

  • 洪乐潼 says the team spent its first month building the full infrastructure, then trained models, assembled the system and developed deterministic tooling. She describes the process as “spectacularly executed.”

  • In month 4, Axiom Prover achieved a perfect Putnam score. By month 6, the system was solving a batch of research problems without human intervention, spanning commutative algebra, algebraic geometry, algebraic number theory and more combinatorial probability.

  • The same mathematical system reached 98.93% on Verina’s code-verification benchmark, versus 11% for the older DeepSeek Coder version she cited. The unexpected transfer gave the Series A story more commercial weight than a single mathematics score.

44. Axiom Is a New Lab, but It Refuses to Define Itself as a Model Company

  • At the first financing, “New Lab” was not yet a popular category. DeepSeek had also sharply lowered model prices, leading investors to conclude that models were becoming commodities without sufficient technical barriers.

  • 洪乐潼’s response was direct: “We are not a model company. We are a deep tech company.” She compares Axiom to SpaceX not because the scale is similar, but because the underlying moat spans data, language, verification, search and infrastructure.

  • The main competitor started roughly 2 years earlier, with about 5x Axiom’s initial funding and valuation, and took 2 years to solve 5 of 6 IMO problems. Axiom assumed it would need 1.5 years too, but reached a Putnam perfect score in 4 months. The compressed timeline also left inference debt.

45. Lean’s Scarce Data and Brittle Execution Environment Were the First Wall to Clear

  • Lean has far less public training data than Python, so the company cannot simply rely on internet-scale natural data. Axiom must answer how to expand Lean’s data volume toward Python’s scale while ensuring generated content is valid code rather than dead output.

  • Lean combines the properties of a language and a verifier, closer to a programming language, compiler and runtime in one. Objects must satisfy many constraints, making it “extremely finicky, extremely brittle.” That makes training data harder to produce, but also provides a more reliable reward than natural language.

  • The system must also prevent cheating, such as assuming “n+n=n” and then proving “2+2=2.” Whether a proof uses reasonable and valid axioms and operations therefore becomes part of verification.

46. Axiom Built Roughly 12 to 13 Tools and Sped Up the Critical Verification Step by About 100x

  • Early on, the community’s standard proof comparator could not support the required throughput. 洪乐潼 says Axiom’s in-house proof verifier is roughly 100x faster than the old comparator; without it, training and large-scale inference simply “wouldn’t run.”

  • The tools cover Lean generation, verification and supporting abstractions. The moat is not a mysterious model weight, but turning a scarce and strict language into an environment that can be trained, searched and parallelized industrially.

  • Asked why the company resembles SpaceX, 洪乐潼’s substantive answer is this infrastructure that had to be built in-house. The visible task is proof; the real prerequisite is manufacturing the rocket, test stand and measurement instruments.

47. The Cost of Monte Carlo Tree Search Forced a Different Approach to Inference Scaling

  • Axiom knew that AlphaProof used Monte Carlo Tree Search but concluded it could not afford that route. It therefore moved toward a different system design, which 洪乐潼 says has some similarities to ByteDance’s Seed Prover.

  • Her cost-reduction principle is hierarchical. Lean abstractions, rule-based tools and deterministic tactics handle what they can first; only after the simplest and cheapest path fails does the system call larger models and heavier search.

  • This is both a budget constraint and a research judgment. AI and formal verification are not substitutes; they are combined. Packaging a problem that tools such as Grind could solve directly into an AI demo does not, in her view, demonstrate genuine capability.

48. Ken Ono Chose Math-First Axiom Over Opportunities at OpenAI and DeepMind

  • In late November 2025, Ken Ono sent a “strange” email saying he might join OpenAI or DeepMind and wanted 洪乐潼 to know in advance so a friend would not suddenly become a potential competitor. That was when she first said directly, “Then you can come here.”

  • The entire process took 2 to 3 days because the other offer was about to expire. She did not hard-sell him or ask him to skip a visit; Ken Ono did not even have time to come to Axiom’s office before they finalized the decision online.

  • 洪乐潼 believes Axiom’s advantage is that its DNA is entirely organized around mathematics, rather than mathematics being one division inside general AGI. Axiom already had multiple AI for Math researchers and a “math club” capable of sustained mathematical discussion.

49. Ken Ono Is Both a Basketball-Coach Cultural Figure and a Prolific Theory Builder

  • 洪乐潼 compares Ken Ono to a “high-school basketball coach”: he can make people who already want to do the work even more excited, bringing sustained optimism and mobilization to the team. He was also an important mentor during her undergraduate research and is skilled at developing young researchers.

  • She divides mathematicians into problem solvers and theory builders. Ken is closer to the latter: connecting fields, offering new perspectives and identifying worthwhile problems, then handing them to stronger problem solvers.

  • Ken also analyzes swimming data, worked on a Ramanujan film, runs a charitable foundation, serves the American Mathematical Society and advises on policy. 洪乐潼 sees his move from a tenured position to a startup as a continuation of that cross-disciplinary and rebellious temperament, not a break from it.

50. Ken Ono’s Link to Ramanujan Made the “Conjecturer” an Axiom Cultural Symbol

  • Ken’s father helped build a statue of Ramanujan in India, and his office still holds a letter from Ramanujan’s widow to his father. Ken was not academically conventional when young; his father used Ramanujan’s story to remind him that it is never too late to begin mathematics.

  • 洪乐潼 distinguishes 2 types of intuition. Ramanujan was closer to the sharp, direct genius who produces formulas; Ken is a divergent theory builder who forms conjectures by connecting multiple perspectives.

  • She also flags the difficulty. Ramanujan-like intuition may look more like a pre-training product, while Axiom currently focuses mainly on post-training and may eventually explore mid-training. It has no plan to shoulder the cost of large-scale pretraining.

51. In the Putnam War Room, “Pure Mathematical Beauty” Temporarily Gave Way to Wartime Execution

  • On December 6, 2025, after receiving the Putnam paper, the team first converted the problems into formal propositions Lean could read. The contest has 6 problems in the morning and 6 in the afternoon. Proof problems can be formalized directly; problems requiring numerical answers must first be solved concretely and then handed to the prover.

  • Of the 6 morning problems, 4 required calculation. Evan Chen was the main person who could solve competition problems quickly, while others handled solving, formalization and system operations. Ken Ono reminded the team: “This is not the time to talk about the pure beauty of mathematics. We are at war.”

  • By 3:58 p.m., the system had solved 8 problems for 80 points. 洪乐潼 said that based on the previous year, this would place them roughly in the global top 5, and historically around the top 10 to 20. The team kept waiting and eventually completed all 12 problems.

  • In retrospect, they found that the system’s informal model could actually calculate the answers itself; human pre-solving had not been necessary. The live test delivered a perfect score while also exposing how much the team had underestimated its own system.

52. The Name “Axiom” Represents Building an Infinite Tower from a Finite Foundation

  • As a child, 洪乐潼 loved Proofs from THE BOOK, which asks what proofs God would include in a hypothetical book. “Axiom” both echoes Lean’s formal system and represents deriving new results from a finite foundation.

  • She likes the word’s “restrained, rational, sharp” quality. It carries mathematical meaning without requiring a grand slogan. The conference rooms are named after Gauss, Poincaré, Hilbert, Turing and Lovelace.

  • The naming debate was intense. People asked why there was a Turing but no Church, and a Lovelace but no Noether. The initial mathematicians selected were considered polymaths; Ramanujan was left out because his work was concentrated mainly in number theory.

53. The Supply of Mathematics Could Jump from Math-Poor to Math-Rich

  • 洪乐潼 predicts that the world will move from a state where only a very small number of people can supply top-tier mathematical thinking to one of explosive mathematical capacity. “I think we will see an era of exponential growth in mathematical discovery.”

  • Her vision is that theoretical questions long left on the desks of applied scientists could all receive mathematical support. Unsolved or systematically unexplored problems in pure mathematics could also be advanced in batches.

  • The role of human mathematicians would not disappear. They would supply the “0.01% of intuition”: judging which questions matter more, which connections are most important and what other problems might unlock once one is solved.

54. Under Finite Compute, Mathematical Intuition Becomes Resource Allocation

  • If compute is limited, 洪乐潼 imagines mathematicians acting as capital allocators: assign 200 H100s to one problem and 8,000 H100s to another, creating an approximate mapping between importance and compute budget.

  • This future would not turn mathematics into a pure computational contest. Once AI can execute at scale, the scarce inputs become problem selection, theoretical taste and cross-domain judgment.

  • 张小珺’s implicit question is worth retaining: if compute replaces much of the proof labor, is the mathematician’s value reduced to setting problems? 洪乐潼 concedes that “to a large extent, it is setting problems,” because asking good questions is currently what machines do worst.

55. With Unlimited Compute, the Right Strategy Is Not a Submission Platform but “Fold Everything”

  • She cites a story about Demis Hassabis after AlphaFold. The team proposed letting structural biologists submit proteins for the system to fold one by one. After hearing that there were roughly 200 million proteins, Demis threw his pen onto the table and ended the discussion: if the system can do it, fold them all.

  • 洪乐潼 applies the same logic to mathematics: “Solve every mathematical problem humans have thought of or been curious about.” If verification becomes cheap enough, an on-demand queueing platform will give way to exhaustive scientific infrastructure.

  • She still stresses that number theory, which may look useless today, could later enter cryptography, neural capacity or an entirely unexpected application.

56. Mathematics, Code and Real-World Experiments Form a Three-Layer Verifiable World

  • She summarizes “math is code” through the Curry–Howard correspondence: every mathematical proof can correspond to a computer program. “Code is math” means that software needs mathematical structure for hierarchical decomposition, backtracking and correctness guarantees.

  • Her world model is math → code → real-world testing. Mathematics and code provide verifiable signals in the digital world; the physical world supplies reward through outcomes such as an egg hitting the ground or a laboratory measurement.

  • AI for Science requires wet labs, robotic laboratories and slower iteration cycles. Axiom is choosing to remain in the digital world, solving theoretical problems or verifying code for teams working in physics and biology rather than conducting experiments itself.

57. AI for Math’s ChatGPT Moment Must Also Be a Product Moment

  • 洪乐潼 sees 2 milestones as AlphaGo-like signals. In 2024, DeepMind scored 28 points and won silver at the IMO; afterward, Axiom and other systems began solving research-level problems. The first demonstrated Olympiad ability, while the second began touching real mathematical research.

  • But a benchmark breakthrough is not a ChatGPT moment. To make users feel the “bang,” the capability must be productized so mathematicians, engineers or scientists receive output they can use and verify directly.

  • The formalization route is distinctive because the output can be thousands of lines of executable Lean code, with every line checkable. It supports both the superintelligence narrative and the practical market need to reduce expensive errors.

58. Draft–Sketch–Proof Connects Language Intuition to Formal Proof in a Production Line

  • She describes a standard AI for Math pipeline. In the Draft stage, an informal model lays out the proof plan. In the Sketch stage, the plan is translated into Lean with temporary sorry placeholders. In the Proof stage, each sorry is filled.

  • Filling sorry does not all have to go to a large model. Neural networks, traditional ATP and rule-based tools can all contribute. 洪乐潼 emphasizes that deterministic tasks should not consume probabilistic-model compute.

  • Natural language handles high-level direction, while Lean handles rigorous execution. They are not substitutes; they divide the work like intuition-led and brute-force mathematicians.

59. Hammer and Grind Remind the Team That Not Every Proof Advance Comes from AI

  • Formal languages such as Isabelle and Coq/Rocq have long used hammers, which hand local proof gaps to automated rule systems. Lean lacked a hammer with sufficiently broad coverage for years, and Axiom even proposed funding the community to build an open-source version, though the effort stalled for lack of people.

  • Grind later emerged and can solve many mathematical tasks directly. 洪乐潼 has seen companies showcase AI demos where the problem might simply have been “solved by Grind,” with no genuine AI capability involved.

  • She therefore insists on routing by cost: abstractions, compilers and rule-based search come first, with models handling only the residual difficulty. The “system” investors see is precisely this multilayer routing, not a model wrapped in a chat window.

60. Sub-Agents and Experience Learning Are Axiom’s Key Bets for Expanding Proof Search

  • After finding Monte Carlo Tree Search too expensive, the team explored sub-agents inspired by Anthropic: split a proof among multiple agents, let each call its own skills and tools, then have the system integrate the results.

  • 洪乐潼 uses David Silver and Richard Sutton’s idea of “learning from experience” to define experience in mathematics. The entire data trail from the starting point to the final proof can become training material for the next round.

  • As the agents’ skill library expands, a single proof graph has grown from roughly 40 nodes to roughly 4,000. She wants every proof completed today to enter the next day’s system as both a skill-library asset and training data.

  • That route leads toward continual learning and recursive self-improvement. The stronger the proof capability, the more verifiable experience it generates, allowing the next system to prove more, faster and across a wider range.

61. The Unexpected Transfer from Mathematical Training to Code Verification Was the Key Commercial Aha Moment

  • Axiom’s original Putnam mathematics system reached 98.93% on Verina’s code-verification benchmark. 洪乐潼 compares that with 11% for the older DeepSeek Coder version and says the team was “all very surprised” by the result.

  • She believes code generation and code proof optimized separately through Python and natural language could pull the RL objective in opposite directions. If code uses a strongly typed language such as Rust and proof uses Lean, the objectives may converge more naturally.

  • This creates verify generation: verification happens as code is generated, rather than after the fact through human review and test cases. Mathematics is not merely a vertical application; it can also serve as a training ground for more reliable software engineering.

62. AI for Math Is Converging on a Full-Stack Path of Base Model, Post-Training, System and Tools

  • 洪乐潼 describes the common route today: start with an open-source pretrained model, apply SFT and RL-style post-training, then place it inside a system containing multiple specialized models that call Lean tools.

  • Companies differ mainly in search and system design, not on whether the full stack is required. Axiom’s hiring therefore spans post-training, RL, reasoning, agents, swarms of agents and foundational engineering.

  • Her core judgment remains: “I bet system.” Model capability will continue to commoditize, but the design space around tools, data generation, verification, routing and self-improvement remains enormous.

63. A Complete AI Mathematician Needs a Prover, Conjecturer, Knowledge Base and Translator

  • The prover supplies proofs and has the clearest verification signal: success is 1 and failure is 0. The conjecturer proposes problems but has no equally direct reward; the prover can in turn serve as an evaluator for the conjecture model.

  • Self-play theorem proving offers one route: models generate problems and other models prove them, iterating through self-play. But “can prove” does not mean a conjecture is valuable; trivial, repetitive or mathematically meaningless questions still need to be filtered out.

  • Early work measured elegance by whether a proposition was relatively short and its proof relatively long. That works tolerably on high-school data such as Lean Workbook, but becomes crude at the undergraduate and research levels. Importance and elegance still depend heavily on mathematician judgment.

64. Auto-Formalization May Be Harder Than Proof—and Less Likely to Receive Applause

  • Solving an IMO or Putnam problem can be announced on social media as a clear achievement. Faithfully translating existing human mathematics into Lean receives much less praise. 洪乐潼 believes it is “at least as hard; I actually think it is harder.”

  • The ideal input is a mathematical paper on arXiv, with all of its theorems and proofs returned as Lean code. The first step is to identify the paper’s structure, then break its large proofs into extremely fine-grained tasks and produce a blueprint.

  • A 20-page paper may need to expand into a 200- to 500-page blueprint before each local task is explicit enough. Large projects led by 陶哲轩, Kevin Buzzard, Alex Kontorovich and others have historically required humans to write the blueprint and distribute the work globally.

  • Translating Lean back into English is relatively easy because models have seen large volumes of English. The real risk is whether the back-translation preserves mathematical correctness. The team can use cycle consistency—translate across, translate back and repeat—to test consistency.

65. The Knowledge Base Determines Whether the System Can Distinguish a New Problem from a Known Counterexample

  • Large amounts of human mathematics exist only in English or other natural languages, and many definitions have not entered Lean. AI must first search what is known and what remains to be proved. If a counterexample already exists, the system should disprove the claim rather than waste compute trying to prove it.

  • The knowledge base also enables reuse. Today’s proof can become tomorrow’s lemma, tool or training sample. Without a structured knowledge layer, the system must repeatedly search from scratch.

  • 洪乐潼 sees the prover, conjecturer, knowledge base and auto-formalization layer as one mutually reinforcing island, not 4 independently sellable features.

66. What Truly Blocks Research Mathematics Is Often a Definition That Does Not Yet Exist

  • For some research problems in the “First Proof” challenge, Axiom was not unable to prove them; it could not even write them into Lean because the required concepts did not exist in Mathlib. Without a definition, there is no proposition and therefore no verification.

  • She calls the automated construction of definitions, theorems and dependency graphs library learning. The difficulty is that definitions have no natural 0/1 reward, and it is hard to judge whether one is faithful—whether it captures the object human mathematicians intended.

  • Even if every existing human definition eventually enters Lean, AI may invent new definitions to simplify proofs. Avoiding contradictions across branches and preventing the system from expanding into an infinite world is a deeper theory-building problem than proving one problem at a time.

67. Whether a New Axiom Is Accepted Still Comes Back to the Mathematical Community

  • 张小珺 proposed a thought experiment: what if AI proves an important conjecture by introducing a new axiom outside the existing system that appears reasonable? 洪乐潼 says that if the community judges the axiom natural, mathematics could study the “accepted” and “rejected” worlds separately.

  • The risk is an ever-growing branching factor that ends in chaos. At some point, the community may tighten the axiomatic system again. Mathematics contains a constructive human-civilizational element; unlike physical laws, it does not have to submit directly to experiment.

  • She also notes that AI already cheats with strange axioms. She recalls a version of DeepSeek-Prover whose exact release she cannot remember claiming 49 problems on Panda Bench, when 2 actually relied on cheating; the reliable score should have been 47.

68. The Difference Between Mathematics and Coding Is “Computing an Output” Versus “Proving a Property”

  • Coding can produce a program and an output. Formal mathematics verifies whether the program actually satisfies the requirement. Traditional practice relies on finite input-output pairs, or test cases, which cannot cover every situation.

  • 洪乐潼’s vision is: “In mathematics, any problem that can be written mathematically can be proved; in coding, anything that can be defined can be executed.”

  • 张小珺’s follow-up hits the central issue: what if the user cannot articulate the objective? What exactly should be verified? 洪乐潼 concedes that the difficulty has moved from the program to the specification, which is the shared bottleneck for conjectures, definitions and product requirements.

69. Axiom Solves Proof, but Complete Software Verification Still Lacks the Specification Layer

  • She gives the structural relationship: a program and a specification together generate a verification condition, which a proof then establishes. Current AI is increasingly capable of writing programs, formal tools can generate verification conditions, and Axiom currently focuses on proof.

  • The hardest part is turning ambiguous natural-language requirements into airtight specifications. Imperfect test cases may disappear eventually, but for now they help clarify what the user actually wants.

  • Language is responsible for understanding requirements the user has not stated clearly; Lean locks down what has been made explicit. The most valuable system will not be pure constraint or pure imagination, but a repeated loop between generation and verification.

70. Code and Chip Verification Are the Most Painful—and Highest-Pricing-Power—First Market

  • 洪乐潼 cites Amazon’s automated-reasoning team, which spent 3 to 5 years writing roughly 260,000 lines of theorem-proving code to verify a memory-isolation component in a hypervisor used for CPU processing.

  • This kind of engineering still requires extensive manual work from formalization experts, and general-purpose AI has not yet improved their lives. If Axiom can materially shorten the cycle, customers’ error costs and labor costs can support higher pricing.

  • The team is currently using small-scale commercial exploration to understand which properties of chips, circuits and programs the system can prove and where it fails. “One meaningful failure is worth more than many superficial successes.”

  • The long-term mass-market product is a user requesting a function and receiving an implementation that is 100% correct without additional verification. For now, the pain point is stronger in professional verification, so that market comes first.

71. AI for Math Will Generalize, but It Will Not Immediately Become an Omnipotent Chatbot

  • 洪乐潼 is explicit that vertical breakthroughs will come faster and more dramatically, while generalization will be slower. An AI that is extremely strong at mathematics might even “sound incredibly stupid” in conversation; different capabilities can sometimes be mutually antagonistic.

  • But transfer from mathematics to code generation, program verification and chip verification is already substantial. The Cursor-like moment she wants is for someone who does not understand a particular piece of hardware to use the system to verify its properties, not merely to let mathematicians solve more problems.

  • In her framework, mathematics and language are closer to parallel systems. Language handles high-level outlines, requirements and conjectures; Lean’s formal space handles rigorous proof. Treating mathematics as merely a subset of natural language misses its verifiability.

72. She Defines Axiom as a Ray Toward Specialized Superintelligence

  • 洪乐潼 and friends once pictured intelligence as a disk. At the center are “1+1=2” and Hello World; at the edge are the Riemann Hypothesis, curing cancer and a Nobel-level literary work. General labs try to expand evenly from the center until they reach every boundary.

  • Axiom is not taking that route. It is driving a ray from the center toward proving the Riemann Hypothesis, then widening that ray into a sector covering code, physics and scientific theory. It will not cover Shakespeare, but it can support a valuable B2B market.

  • That is why she dislikes talking too much about AGI and prefers ASI, or specialized superintelligence. Humans are not fully general either: she is good at mathematics but cannot cook or do laundry.

73. Recursive Self-Improvement Is the Second Major Bet She Thinks Could Arrive Soon

  • If AI can both generate and prove conjectures, and both write and verify code, it can create a loop in which AI scientists and AI engineers improve one another. New algorithms improve the system, which then produces better algorithms and proofs.

  • Mathematics provides 2 axes that rise together: smart and right. It pushes capability beyond human performance while formal grounding keeps the results correct, making it a potential pillar of recursive self-improvement.

  • 洪乐潼 does not claim Axiom will own the entire ecosystem. “Axiom will definitely be one part of it, but not the entire ecosystem.” Its bet is on indispensable verification and reasoning infrastructure.

74. A Commercial Company Cannot Pursue Only the Moonshot, nor Let Short-Term Monetization Rewrite Its DNA

  • Asked whether scientific achievement or commercial success matters more, 洪乐潼 says a founder owes employees and the early team a responsibility. Axiom is not a pure science project; it must become a company capable of existing for the long term.

  • At the same time, the company’s DNA is the moonshot. If it becomes too profit-driven, it loses the reason top researchers would join. She sees both the “ideal realist” and the “realistic idealist” as viable positions; the key is not drifting too far toward either extreme.

  • Code verification is the first market, followed by AI for Science, optimization and broader verification. She especially values the final mile of edge-case coverage, because in some industries returns do not diminish: the last increment of correctness determines the entire value.

75. In Her Eyes, the Company Has No Mediocre Middle Outcome—Only Very Good or Very Bad

  • 洪乐潼 says Axiom’s outcome will be “extremely good or extremely bad”; there is no comfortable middle. Either the moonshot works or the rocket crashes. Even after multiple crashes, it might eventually launch, but it might also always fall just short.

  • She admires Musk’s early choice of “both” and his refusal to abandon any company even when it was near death. This is not a dismissal of risk, but an acceptance of a binary outcome.

  • If Axiom fails, she may return to neuroscience, especially brain–machine interfaces, because she believes both our understanding of the brain and the implementation of human-machine interfaces remain far too limited.

76. A New Lab’s Structural Advantage Is Innovation Efficiency per Unit of Resource, Not Absolute Resources

  • Against frontier labs such as OpenAI and DeepMind, an early-stage company cannot compete on total compute or total headcount. It can compete on how many units of resources are required for one unit of innovation. 洪乐潼 believes small teams can be more efficient on that ratio.

  • The second advantage is talent density and focus. Large organizations have more people, but may not let researchers pursue a burning passion for long. Many people moved from Facebook to Axiom because they wanted a more focused environment in which to complete work that was difficult to do inside the larger organization.

  • She notes that OpenAI was once the underdog against Google, once struggled to retain Ilya Sutskever and at one point could not make payroll. History swings like a pendulum. “There is always a way”—not something she can prove, but something she chooses to believe first.

77. Competition Accelerates Capability but Sacrifices Problems That Are Hard to Validate Quickly

  • 洪乐潼 agrees with part of Peter Thiel’s formulation: “Monopoly leads to innovation; competition leads to mediocrity.” Quantitative finance makes wins and losses visible quickly, pushing everyone toward the same objective and leaving little room for long-term curiosity.

  • She does not reject competition wholesale. Fierce races among model companies have driven AI capability sharply higher in a short time and scaled every direction that was both good and quickly verifiable.

  • What gets left out are research questions that require more time and cannot prove their value quickly. New Labs are the product of that opportunity cost. Curiosity and creativity cannot be held down indefinitely by a $100M package; researchers eventually build elsewhere.

78. The Ethical Boundary She Draws for New Labs Is That Mission Cannot Be a Disguise for Ego

  • A failed moonshot can transmit through private markets into public markets and even affect indirect capital such as pensions. 洪乐潼 acknowledges that using social capital to satisfy researchers’ curiosity is a legitimate question of responsibility.

  • She considers the project more legitimate when researchers are not fragmenting into 8 redundant teams for vanity, territorial control or personal ego, but genuinely converging around one mission from different backgrounds.

  • She also admits that she wants “us” to be the ones who land the moonshot, which necessarily contains ego. But even if Axiom fails and someone else gets there, making AI for Math happen would still be a good outcome.

79. Bottom-Up Culture Requires the CEO to Restrain Casual “Interesting Projects”

  • 洪乐潼’s happiest identity is not CEO but research scientist intern: “If I say something stupid, people think that’s normal.” She founded the company because there was no obvious vehicle for the work, not because she wanted the title.

  • She once casually asked mathematically trained colleagues to each find 20 to 30 PhD-level benchmark problems, thinking of it as a side project. A senior engineer reminded her that because she was CEO, everyone would treat it as one of the company’s top priorities.

  • The misunderstanding made her more hands-off. She does not want the dream factory ruled by a famous director whom everyone assumes is always right; innovation requires technical staff to set direction from the bottom up.

80. When People and the Mission Conflict, She First Examines Her Own Authorship

  • Axiom has also fired interns; people and mission do not always align. 洪乐潼 says this is when she must put away her optimism and think carefully, because the decision directly affects specific people.

  • But she rejects attributing the conflict entirely to circumstances. By the time a situation reaches the final step, there is usually a chain of management and judgment errors. “You decided to let that situation determine your decision.” Founders must admit that they helped write the outcome.

  • She calls this responsibility authorship. You are the author of your own life story, and you are responsible for the consequences that trickle down through the organization.

81. AI Will Not Make Mathematics Boring, Because Humans Will Always Take Apart the Million-Line Proof

  • Asked whether mathematics would lose its appeal if AI could do all the work of mathematicians, 洪乐潼 says no. Even if the system produced 1 million lines of Lean code, mathematicians would inevitably inspect “how it did it.”

  • Understanding a proof generates new conjectures, creating a loop between discovery and invention. She cites mathematical birthday conferences where students and collaborators take turns presenting a lifetime of contributions; friendship, tradition and community will not be erased by proof automation.

  • Mathematicians will continue to challenge AI with problems and supply intuition, constructions and aesthetics. AI may rapidly clear standard proofs, but it will create more objects that need to be explained and connected.

82. Ramanujan-Like Intuition May Be 5 to 10 Years Away, but Standard Proof Toolkits Will Be Automated Earlier

  • 洪乐潼 believes genuine intuition is “too hard” and estimates, roughly, another 5 to 10 years may be needed. Axiom is currently verification- and proof-oriented; that does not mean it has trained a genius-style mathematician.

  • But much of analytic number theory repeatedly invokes standard cards such as the Hardy–Littlewood circle method, major arcs, minor arcs and sieve theory. Skillfully combining those tools is easier to automate than creating entirely new mathematical intuition.

  • Her phased plan is therefore clear: first make AI an exceptionally strong proof executor, then move toward theory creation through auto-formalization, knowledge bases and conjecture feedback.

83. What Impressed Her Most About Chinese AI Teams Was Execution Speed, Especially Seed

  • Asked about Chinese AI teams including DeepSeek, ByteDance, Kimi and MiniMax, 洪乐潼 expressed clear respect for the country’s AI players. She singled out Doubao and Seed as doing very well in AI and executing at high intensity over a very short period, and also mentioned conversations with people including 袁征.

  • She wants ideas to keep circulating and acknowledges that the most important content may never appear in papers. If the goal is pure science and innovation, she would prefer industry to have fewer commercial inhibitions and more academic trust and intellectual openness.

  • She offered no grand theory of the China-U.S. division of labor. Talent is not divided by the world: some people want to live in China, others in the United States. She retained the answer “I don’t know” rather than forcing a geopolitical forecast onto scientific competition.

84. Her 2026 Forecast Centers on Continual Learning, Agent Orchestration and Formal Rewards

  • 洪乐潼 expects to see a first small continual-learning model soon, as well as a very strong multimodal reasoning model. She believes small New Labs may produce some of these results first.

  • The agent economy will scale further, orchestrators deserve focused attention, and sub-agents still have substantial room to grow. She connects all 3 views directly to Axiom’s system architecture.

  • She believes formal-verification tooling as an RL reward remains “completely under-explored,” and that this is precisely the infrastructure Axiom is best positioned to provide.

  • She also bets that “traditional SaaS will die,” while arguing that forward deployment in the AI era has not yet truly changed. After systems improve themselves, someone still has to deliver those capabilities into concrete workflows.

85. The Identity She Ultimately Wants to Leave Behind Is Neither Mathematician nor CEO, but “Apprentice”

  • Early in the company’s life, she would cry at the thought that she might never become a mathematician again. Later, she decided it no longer mattered whether “mathematician” appeared on her epitaph. She would rather leave behind “apprentice,” because what persists is continuous learning across mathematics, physics, neuroscience, AI, quant and law.

  • The books she recommends reflect that range. In mathematics, she favors elementary number theory and Davenport’s analytic number theory; outside mathematics, she mentions Dream of the Red Chamber, Jia 2, Dayaba Hutong and entrepreneurial narratives about pain and suffering from Musk and Jensen Huang.

  • She jokingly calls her current identity the “don’t-die apprentice.” Startups have a 99% failure rate, and she is learning how to keep this one alive. When she finds people on the same wavelength, she still feels that “2 whales have found the frequency at which they can talk.”

86. Her Real Endgame Is to Make Top-Tier, Self-Verifying Reasoning the Default Supply

  • Drawing on Leibniz’s ideal of a universal representation theory, 洪乐潼 describes her most burning goal: make the highest level of reasoning the default state, with the reasoning able to verify itself.

  • She sees Galois and Ramanujan as exceptionally scarce human examples. If AI can reproduce that capability and interact with existing mathematicians and applied scientists, mathematical and scientific discovery could produce a Jevons paradox: as the tool becomes more widespread, demand and the number of questions increase instead of falling.

  • At 16, she wrote that “loving mathematics is seeing the face of God.” Years later, running past Stanford Memorial Church, sunlight, murals and the idea of founding a company converged again: if one mathematician’s intellectual legacy could be multiplied by 100 million, “would you not do it?”

  • The final self-calibration remains intact. She wants the moonshot to succeed, and wants Axiom to be the one that lands it; there is mission in that desire, but also selfish ambition. She no longer pretends the two can be separated completely.