Pioneers Insight Method Research Author
🔬How GPT‑5 derived new results in theoretical physics and quantum gravity — Alex Lupsasca, OpenAI
Back to Episodes

🔬How GPT‑5 derived new results in theoretical physics and quantum gravity — Alex Lupsasca, OpenAI

· Source link · AI Summary archive

Summary

  • Alex Lupsasca argues that frontier physics has crossed from AI assistance into superhuman performance on bounded research tasks. ChatGPT o3 cut a calculation from days to 11 minutes; GPT-5 reproduced one of his best papers in under 30 minutes; and recent models solved a problem that experts had pursued for a year. His call is categorical about the milestone but conditional about its breadth: “at least in some directions AI has become superhuman.”

  • The flagship result replaced a factorially exploding gluon calculation with a formula whose size grows only linearly with particle count. Humans found the loophole showing single-minus gluon tree amplitudes need not vanish when particles become collinear in a special spacetime setting; GPT-5.2 Pro simplified the five- and six-point cases, conjectured the general formula, and a stronger internal model independently rediscovered and proved it in 12 hours. “We solved the problem before he even got off the plane.”

  • A second paper showed that the first result was not a one-off: public ChatGPT-5.2 Pro generalized the approach from gluons to mathematically different graviton amplitudes. The model selected the directed matrix-tree theorem, performed checks, and produced much of a paper through a 110-page exchange. The team says it could have released the paper roughly three days after the first; publication took three weeks because humans checked it carefully. “Most of the time was spent verifying the answer, not writing.”

  • Research economics are shifting from producing calculations to choosing questions and validating abundant output. Lupsasca says that, with proper steering, models could probably produce a paper a day for problems that are not especially hard or resemble known calculations, while ten parallel chats can scout competing approaches and rapidly eliminate dead ends. The scarce human contributions become taste, expert steering, and verification—not raw symbolic labor.

  • AI currently resembles an extraordinarily capable graduate student more than an autonomous scientific agenda-setter. Give it a sharp, well-posed problem and it can perform frontier calculations, but “the difference between a good physicist and a great physicist is knowing what is the right question to ask.” That moat is narrowing: when asked for three follow-ups to the gluon paper, GPT proposed essentially Lupsasca’s own top three.

  • The same capability that grants researchers “AI superpowers” threatens training and information quality. Traditional six-to-twelve-month research problems can now be “crushed,” weakening the rites of passage through which students build technique and confidence; meanwhile, poorly framed prompts can yield polished “AI slop” for arXiv. Lupsasca’s response is to raise the publication bar, not spend the year producing 30 incremental papers.

  • Verification, formalization, and scientific communication are emerging infrastructure opportunities. Natural-language reasoning became good enough to make formal proof look less urgent, but mass parallel generation puts the burden back on humans and makes tools such as Lean more valuable again. Lupsasca also doubts static papers survive another 20 years, imagining interactive research objects that can explain the big picture, expose derivations, and expand detail on demand.

Deep dive

1. Frontier physics moved from email utility to research-grade reasoning

  • A little over a year earlier, Lupsasca regarded AI as useful for email but irrelevant to the “important theoretical physics calculations” that distinguished his work. ChatGPT o3 overturned that view by completing in 11 minutes a calculation that would have taken him days and that, as far as he knew, ordinary scientific software could not perform.

  • GPT-5 caused the sharper conversion. With one warm-up hint, it reproduced in under 30 minutes one of his best and most difficult papers—a calculation he believes only a handful of people could execute. “Oh my god, this changes everything,” he remembers thinking; joining OpenAI during his sabbatical became an obvious response.

  • Lupsasca says GPT-5.4 represents another large jump, even if consumers judging email quality barely notice it. His complaint about GPT-5’s initially lukewarm reception was blunt: “GPT-3 could write email. How much better can it get at writing email? That’s not the point.” The frontier-science capability was moving much faster than the visible consumer experience.

  • Evidence now arrives unsolicited from research groups. One correspondent reported that Codex produced a simulation of the technically demanding SYK model in ten minutes after physicists had struggled to set it up. After RJ noted that many physicists also have coding skills, Brandon stressed that capable researchers would have attempted this and concluded that “Codex is just really good now.”

2. Scattering amplitudes compress the information in a quantum field theory

  • Lupsasca begins from the tension quantum field theory must reconcile: relativity imposes the absolute prohibition against faster-than-light information, while quantum mechanics makes physical quantities intrinsically fuzzy. Quantum field theory is the twentieth century’s framework for accommodating both principles while describing forces through probabilistic outcomes.

  • Those probabilities come from squaring complex quantities called quantum amplitudes. Scattering amplitudes specify what can happen when particles with given energies, momenta, and polarizations collide and emerge in another configuration; knowing the n-point amplitudes for arbitrary n amounts, “more or less,” to knowing the entire theory.

  • Polarization introduces helicity: a photon’s transverse arrow can wind right-handed or left-handed as it travels, producing plus or minus helicity. Gluons, which mediate the strong force binding nuclei, carry analogous labels. Tree amplitudes provide the leading contribution, while loop diagrams add interactions—and therefore powers of the coupling constant—as progressively smaller corrections.

3. A collinear loophole reopened an amplitude textbooks had declared zero

  • If every gluon has plus helicity, the tree amplitude vanishes: the interaction is forbidden and cannot happen. Textbooks extended essentially the same argument to a single-minus configuration, where one gluon has the opposite helicity, and treated that amplitude as zero too.

  • The next case, with two minus-helicity gluons, famously does not vanish. Parke and Taylor’s 1980s calculation summed a formidable collection of terms that almost entirely canceled, leaving a half-line formula. These became “maximally helicity violating,” or MHV, amplitudes—though Lupsasca now prefers the less theory-laden description “double minus.”

  • Alfredo Guevara, David Skinner, and Andrew Strominger had recognized a loophole: the zero argument assumes particles arrive from generic directions. When some are exactly collinear, and the problem is considered in the required unusual spacetime signature—described conversationally as “two dimensions in space and two dimensions in time”—single-minus amplitudes can be non-zero.

  • The humans could represent the answer but could not expose its simple form. Three particles gave one term, four gave two, five gave eight, and six produced 32 complicated terms; for general n, the Feynman-diagram count grew factorially. For a year they searched for the single-minus counterpart of Parke–Taylor’s miraculous simplification.

4. GPT turned a factorial expansion into a linear formula

  • Before Strominger arrived at OpenAI for a planned collaboration, the group fed the five-point, eight-term expression to public ChatGPT Pro. The model identified a phase-space region—one particle carrying a different frequency sign from the others—in which the expression collapsed to a product of three terms. It reportedly wrote Python, checked roughly 5,000 possibilities, and selected the simplification itself.

  • The six-point test was more dramatic: 32 terms, each containing nested products, collapsed to a product of four. “Whoa, okay. That is really nice,” was the researchers’ reaction. They had no seven-point expression because expanding one by hand would have been prohibitive.

  • Asked to infer the all-n pattern, GPT-5.2 Pro proposed a formula whose complexity grows linearly rather than factorially: doubling the number of particles merely doubles the number of terms. Lupsasca regards this as the single-minus analog of the Parke–Taylor formula, although the public model could conjecture it but not prove it.

  • A stronger internal model received the sharply formulated problem without the candidate formula or low-point answers. After 12 hours, it independently rediscovered the same expression and supplied a three-step derivation. The remainder of the paper after the result is stated is essentially the proof produced by that model.

5. The scientific claim stands apart from the unusual means of discovery

  • Lupsasca carefully divides credit: humans discovered the collinear loophole and established that the amplitudes need not vanish; AI found the concise general expression and its proof. The paper’s title, “Single-minus gluon tree amplitudes are non-zero,” foregrounds the physics rather than presenting the work as an AI demonstration.

  • The authors omitted AI from the abstract and confined provenance to a paragraph noting the GPT-5.2 Pro conjecture and internal-model proof. Lupsasca’s analogy: a reader consulting a decades-old result does not care which MS-DOS version or stack of floppy disks powered an indispensable computation. Method matters historically, but it should not eclipse a durable physics result.

  • A host’s pushback—worth keeping—is that the progression from the displayed five- and six-point formulas to the general expression may look natural to a graduate student. Lupsasca answers that the proof came from a fresh session with no limiting cases, giving an independent check; he nevertheless refuses to rank the paper prematurely because real impact is measured by decades of follow-on work.

6. Public GPT generalized the discovery from gluons to gravitons

  • Gravitons are the hypothesized indivisible quanta of gravitational waves, although no experiment has directly measured them and the correct theory of quantum gravity remains unknown. In field-theory language, gluons are spin-one particles while gravitons are spin two, doubling aspects of the polarization data and making the mathematical problem materially different.

  • Three weeks after the gluon paper, the team released “Single-minus graviton tree amplitudes are non-zero.” They say they could have released it about three days after the first paper because that is how quickly ChatGPT produced the answer; the delay came from checking it carefully, writing a polished paper, and adding context. “Most of the time was spent verifying the answer, not writing,” an inversion Lupsasca would have considered absurd a year earlier.

  • This time no internal model was required. The team gave public ChatGPT-5.2 Pro the gluon paper, emphasized its appendices, described two changes needed for gravity, and said, “Good luck. You’re a brilliant theoretical physicist.” The model recognized that a directed matrix-tree theorem could organize the new calculation—known mathematics that the human experts had not thought to apply there.

  • Across a 110-page chat, the model proposed next steps, carried out sanity checks, handled sums over trees and reduction formulas, and repeatedly asked permission to continue. The humans’ replies were often just “go ahead” or requests for one explicit verification. Lupsasca calls the resulting mode “vibe physics.”

7. Steering still determined which result became meaningful science

  • Human judgment remained visible in what the model did not originate. Strominger wrote the broader introduction; the generic AI draft lacked his account of why the result mattered. The researchers also initiated a separate investigation of how the amplitudes transform under symmetries relevant to celestial holography and quantum gravity.

  • From section three onward, however, Lupsasca says the published graviton paper remained close to the model’s draft, with the mathematics derived by public ChatGPT Pro. His calibrated claim is therefore substantial: it is “a real solid result in quantum gravity” produced largely by AI, but with experts selecting the problem, steering the work, checking every step, and supplying scientific perspective.

  • The boundary has moved, not disappeared. AI has now solved a problem that several field experts pursued for a year, but Lupsasca explicitly says it has not yet solved one that defeated an entire community for decades. That is the next threshold he wants to see.

8. AI compresses confusion and turns parallel chats into research scouts

  • Physics education traditionally uses arduous calculations as rites of passage: students learn technique, but also prove to themselves that they can survive difficult work. Professors keep tractable six-to-twelve-month problems “in their pocket” to bridge coursework and research. Lupsasca suspects current models can “crush” many of those assignments, making old training structures unstable.

  • Yet AI can also help students cross what he calls “the desert” between graduate classes and the research frontier. It can unpack any fact at the needed level, connect unfamiliar concepts, and prevent months of unproductive confusion. The unresolved question is how students develop independent confidence when the strongest tutor is also capable of doing the assignment.

  • In Lupsasca’s own work, time spent confused has collapsed. After a calculation, he can ask how it reconciles with another known fact and immediately hear what he forgot or framed incorrectly, replacing days of walks, other projects, and mental blockage with rapid course correction.

  • AI also changes exploration. Instead of mentally committing scarce energy to one route from A through B to C, he can launch ten chats as fast-moving “scouts” along different paths. Even imperfect scouts identify promising terrain and place signposts; following a partially mapped route is much easier than being first into the unknown.

9. Technical competence is scaling faster than scientific taste

  • Graduate students often enter physics for questions such as why space has three dimensions, what happened at the Big Bang, or what lies inside a black hole. Professional maturity means recognizing that some questions sit too far beyond the current edge of knowledge to attack fruitfully; progress comes from choosing something just beyond that edge.

  • Lupsasca distinguishes competence from greatness accordingly. A competent physicist can acquire whatever mathematics, code, or calculational tool a problem demands. “The difference between a good physicist and a great physicist is knowing what is the right question to ask”—the hardest scientific skill, and usually the last one learned.

  • Current models resemble exceptionally skilled graduate students: once handed a precise question, they may be superhuman at the computation. They do not yet consistently decide which question deserves pursuit, so human taste remains central—and the academic skill of matching a question, detail level, and framing to a student transfers directly to prompting an AI collaborator.

  • That advantage may be temporary. When Lupsasca gives ChatGPT Pro a page of the gluon paper and requests the three best follow-ups, it returns approximately his own top three. He declines to disclose internal work on autonomous frontier discovery, but says question selection is already becoming surprisingly strong.

10. The black-hole “Love” test made the trajectory personal

  • The hosts challenge whether the apparent creativity is merely recombination. Lupsasca replies that human invention may also be “recombination of a stack machine” and that GPT feels like a creative collaborator. He preserves a dissenting benchmark: in his understanding, Terry Tao has traced apparently creative AI proofs back to obscure references and has not yet seen a truly impressive novel mathematical move.

  • Lupsasca’s decisive GPT-5 experiment used his then-unseen paper on why black holes have no “Love”—Love numbers measure tidal response, and black holes’ zero response suggests a protecting symmetry. Because the model’s training cutoff preceded publication, he supplied only the governing equation and asked for its symmetries. It initially answered, incorrectly, that none existed.

  • He then offered the obvious flat-space warm-up. GPT-5 Pro solved that 200-year-old problem correctly in nine minutes, identifying three generators; primed in the same chat, it returned to the black-hole equation and found the new symmetries in 18 minutes. “That was my Move 37,” he says—the moment a model reproduced his prized calculation in under half an hour.

11. Cheap paper generation raises the bar and makes verification scarce

  • With expert steering, Lupsasca believes models can already produce human-quality papers, perhaps one per day for calculations adjacent to known work. Bad questions can equally generate polished nonsense, contributing to “AI slop” and inundating arXiv. The community can no longer treat professional presentation as evidence that the underlying question or derivation is sound.

  • His response is not to publish 30 variants of the single-minus papers. The standard for significance must rise as production gets cheaper. He wants to use the new amplitudes as a line of attack on progressively harder quantum-gravity questions, aiming eventually at problems that have resisted whole communities for decades.

  • Static papers themselves look inefficient: AI performs a calculation, humans compress it into terse notation, and the result is then put back into an AI to expand it again. Lupsasca imagines an interactive paper that can explain the big picture, zoom into a derivation, or expose the intuitive pictures mathematicians currently leave off the page—while retaining writing’s useful discipline of clarifying thought.

  • Verification may become this year’s dominant bottleneck. Lupsasca once thought formal proof systems essential, then natural-language reasoning improved enough to resemble human mathematical discussion; now models can attack thousands of questions and return too many proofs for people to check. Formalization in systems such as Lean, automated checking, and clearer model confidence signals therefore become valuable again.