đŹHow GPTâ5 derived new results in theoretical physics and quantum gravity â Alex Lupsasca, OpenAI
Summary
Alex Lupsasca argues that frontier physics has crossed from AI assistance into superhuman performance on bounded research tasks. ChatGPT o3 cut a calculation from days to 11 minutes; GPT-5 reproduced one of his best papers in under 30 minutes; and recent models solved a problem that experts had pursued for a year. His call is categorical about the milestone but conditional about its breadth: âat least in some directions AI has become superhuman.â
The flagship result replaced a factorially exploding gluon calculation with a formula whose size grows only linearly with particle count. Humans found the loophole showing single-minus gluon tree amplitudes need not vanish when particles become collinear in a special spacetime setting; GPT-5.2 Pro simplified the five- and six-point cases, conjectured the general formula, and a stronger internal model independently rediscovered and proved it in 12 hours. âWe solved the problem before he even got off the plane.â
A second paper showed that the first result was not a one-off: public ChatGPT-5.2 Pro generalized the approach from gluons to mathematically different graviton amplitudes. The model selected the directed matrix-tree theorem, performed checks, and produced much of a paper through a 110-page exchange. The team says it could have released the paper roughly three days after the first; publication took three weeks because humans checked it carefully. âMost of the time was spent verifying the answer, not writing.â
Research economics are shifting from producing calculations to choosing questions and validating abundant output. Lupsasca says that, with proper steering, models could probably produce a paper a day for problems that are not especially hard or resemble known calculations, while ten parallel chats can scout competing approaches and rapidly eliminate dead ends. The scarce human contributions become taste, expert steering, and verificationânot raw symbolic labor.
AI currently resembles an extraordinarily capable graduate student more than an autonomous scientific agenda-setter. Give it a sharp, well-posed problem and it can perform frontier calculations, but âthe difference between a good physicist and a great physicist is knowing what is the right question to ask.â That moat is narrowing: when asked for three follow-ups to the gluon paper, GPT proposed essentially Lupsascaâs own top three.
The same capability that grants researchers âAI superpowersâ threatens training and information quality. Traditional six-to-twelve-month research problems can now be âcrushed,â weakening the rites of passage through which students build technique and confidence; meanwhile, poorly framed prompts can yield polished âAI slopâ for arXiv. Lupsascaâs response is to raise the publication bar, not spend the year producing 30 incremental papers.
Verification, formalization, and scientific communication are emerging infrastructure opportunities. Natural-language reasoning became good enough to make formal proof look less urgent, but mass parallel generation puts the burden back on humans and makes tools such as Lean more valuable again. Lupsasca also doubts static papers survive another 20 years, imagining interactive research objects that can explain the big picture, expose derivations, and expand detail on demand.
Deep dive
1. Frontier physics moved from email utility to research-grade reasoning
A little over a year earlier, Lupsasca regarded AI as useful for email but irrelevant to the âimportant theoretical physics calculationsâ that distinguished his work. ChatGPT o3 overturned that view by completing in 11 minutes a calculation that would have taken him days and that, as far as he knew, ordinary scientific software could not perform.
GPT-5 caused the sharper conversion. With one warm-up hint, it reproduced in under 30 minutes one of his best and most difficult papersâa calculation he believes only a handful of people could execute. âOh my god, this changes everything,â he remembers thinking; joining OpenAI during his sabbatical became an obvious response.
Lupsasca says GPT-5.4 represents another large jump, even if consumers judging email quality barely notice it. His complaint about GPT-5âs initially lukewarm reception was blunt: âGPT-3 could write email. How much better can it get at writing email? Thatâs not the point.â The frontier-science capability was moving much faster than the visible consumer experience.
Evidence now arrives unsolicited from research groups. One correspondent reported that Codex produced a simulation of the technically demanding SYK model in ten minutes after physicists had struggled to set it up. After RJ noted that many physicists also have coding skills, Brandon stressed that capable researchers would have attempted this and concluded that âCodex is just really good now.â
2. Scattering amplitudes compress the information in a quantum field theory
Lupsasca begins from the tension quantum field theory must reconcile: relativity imposes the absolute prohibition against faster-than-light information, while quantum mechanics makes physical quantities intrinsically fuzzy. Quantum field theory is the twentieth centuryâs framework for accommodating both principles while describing forces through probabilistic outcomes.
Those probabilities come from squaring complex quantities called quantum amplitudes. Scattering amplitudes specify what can happen when particles with given energies, momenta, and polarizations collide and emerge in another configuration; knowing the n-point amplitudes for arbitrary n amounts, âmore or less,â to knowing the entire theory.
Polarization introduces helicity: a photonâs transverse arrow can wind right-handed or left-handed as it travels, producing plus or minus helicity. Gluons, which mediate the strong force binding nuclei, carry analogous labels. Tree amplitudes provide the leading contribution, while loop diagrams add interactionsâand therefore powers of the coupling constantâas progressively smaller corrections.
3. A collinear loophole reopened an amplitude textbooks had declared zero
If every gluon has plus helicity, the tree amplitude vanishes: the interaction is forbidden and cannot happen. Textbooks extended essentially the same argument to a single-minus configuration, where one gluon has the opposite helicity, and treated that amplitude as zero too.
The next case, with two minus-helicity gluons, famously does not vanish. Parke and Taylorâs 1980s calculation summed a formidable collection of terms that almost entirely canceled, leaving a half-line formula. These became âmaximally helicity violating,â or MHV, amplitudesâthough Lupsasca now prefers the less theory-laden description âdouble minus.â
Alfredo Guevara, David Skinner, and Andrew Strominger had recognized a loophole: the zero argument assumes particles arrive from generic directions. When some are exactly collinear, and the problem is considered in the required unusual spacetime signatureâdescribed conversationally as âtwo dimensions in space and two dimensions in timeââsingle-minus amplitudes can be non-zero.
The humans could represent the answer but could not expose its simple form. Three particles gave one term, four gave two, five gave eight, and six produced 32 complicated terms; for general n, the Feynman-diagram count grew factorially. For a year they searched for the single-minus counterpart of ParkeâTaylorâs miraculous simplification.
4. GPT turned a factorial expansion into a linear formula
Before Strominger arrived at OpenAI for a planned collaboration, the group fed the five-point, eight-term expression to public ChatGPT Pro. The model identified a phase-space regionâone particle carrying a different frequency sign from the othersâin which the expression collapsed to a product of three terms. It reportedly wrote Python, checked roughly 5,000 possibilities, and selected the simplification itself.
The six-point test was more dramatic: 32 terms, each containing nested products, collapsed to a product of four. âWhoa, okay. That is really nice,â was the researchersâ reaction. They had no seven-point expression because expanding one by hand would have been prohibitive.
Asked to infer the all-n pattern, GPT-5.2 Pro proposed a formula whose complexity grows linearly rather than factorially: doubling the number of particles merely doubles the number of terms. Lupsasca regards this as the single-minus analog of the ParkeâTaylor formula, although the public model could conjecture it but not prove it.
A stronger internal model received the sharply formulated problem without the candidate formula or low-point answers. After 12 hours, it independently rediscovered the same expression and supplied a three-step derivation. The remainder of the paper after the result is stated is essentially the proof produced by that model.
5. The scientific claim stands apart from the unusual means of discovery
Lupsasca carefully divides credit: humans discovered the collinear loophole and established that the amplitudes need not vanish; AI found the concise general expression and its proof. The paperâs title, âSingle-minus gluon tree amplitudes are non-zero,â foregrounds the physics rather than presenting the work as an AI demonstration.
The authors omitted AI from the abstract and confined provenance to a paragraph noting the GPT-5.2 Pro conjecture and internal-model proof. Lupsascaâs analogy: a reader consulting a decades-old result does not care which MS-DOS version or stack of floppy disks powered an indispensable computation. Method matters historically, but it should not eclipse a durable physics result.
A hostâs pushbackâworth keepingâis that the progression from the displayed five- and six-point formulas to the general expression may look natural to a graduate student. Lupsasca answers that the proof came from a fresh session with no limiting cases, giving an independent check; he nevertheless refuses to rank the paper prematurely because real impact is measured by decades of follow-on work.
6. Public GPT generalized the discovery from gluons to gravitons
Gravitons are the hypothesized indivisible quanta of gravitational waves, although no experiment has directly measured them and the correct theory of quantum gravity remains unknown. In field-theory language, gluons are spin-one particles while gravitons are spin two, doubling aspects of the polarization data and making the mathematical problem materially different.
Three weeks after the gluon paper, the team released âSingle-minus graviton tree amplitudes are non-zero.â They say they could have released it about three days after the first paper because that is how quickly ChatGPT produced the answer; the delay came from checking it carefully, writing a polished paper, and adding context. âMost of the time was spent verifying the answer, not writing,â an inversion Lupsasca would have considered absurd a year earlier.
This time no internal model was required. The team gave public ChatGPT-5.2 Pro the gluon paper, emphasized its appendices, described two changes needed for gravity, and said, âGood luck. Youâre a brilliant theoretical physicist.â The model recognized that a directed matrix-tree theorem could organize the new calculationâknown mathematics that the human experts had not thought to apply there.
Across a 110-page chat, the model proposed next steps, carried out sanity checks, handled sums over trees and reduction formulas, and repeatedly asked permission to continue. The humansâ replies were often just âgo aheadâ or requests for one explicit verification. Lupsasca calls the resulting mode âvibe physics.â
7. Steering still determined which result became meaningful science
Human judgment remained visible in what the model did not originate. Strominger wrote the broader introduction; the generic AI draft lacked his account of why the result mattered. The researchers also initiated a separate investigation of how the amplitudes transform under symmetries relevant to celestial holography and quantum gravity.
From section three onward, however, Lupsasca says the published graviton paper remained close to the modelâs draft, with the mathematics derived by public ChatGPT Pro. His calibrated claim is therefore substantial: it is âa real solid result in quantum gravityâ produced largely by AI, but with experts selecting the problem, steering the work, checking every step, and supplying scientific perspective.
The boundary has moved, not disappeared. AI has now solved a problem that several field experts pursued for a year, but Lupsasca explicitly says it has not yet solved one that defeated an entire community for decades. That is the next threshold he wants to see.
8. AI compresses confusion and turns parallel chats into research scouts
Physics education traditionally uses arduous calculations as rites of passage: students learn technique, but also prove to themselves that they can survive difficult work. Professors keep tractable six-to-twelve-month problems âin their pocketâ to bridge coursework and research. Lupsasca suspects current models can âcrushâ many of those assignments, making old training structures unstable.
Yet AI can also help students cross what he calls âthe desertâ between graduate classes and the research frontier. It can unpack any fact at the needed level, connect unfamiliar concepts, and prevent months of unproductive confusion. The unresolved question is how students develop independent confidence when the strongest tutor is also capable of doing the assignment.
In Lupsascaâs own work, time spent confused has collapsed. After a calculation, he can ask how it reconciles with another known fact and immediately hear what he forgot or framed incorrectly, replacing days of walks, other projects, and mental blockage with rapid course correction.
AI also changes exploration. Instead of mentally committing scarce energy to one route from A through B to C, he can launch ten chats as fast-moving âscoutsâ along different paths. Even imperfect scouts identify promising terrain and place signposts; following a partially mapped route is much easier than being first into the unknown.
9. Technical competence is scaling faster than scientific taste
Graduate students often enter physics for questions such as why space has three dimensions, what happened at the Big Bang, or what lies inside a black hole. Professional maturity means recognizing that some questions sit too far beyond the current edge of knowledge to attack fruitfully; progress comes from choosing something just beyond that edge.
Lupsasca distinguishes competence from greatness accordingly. A competent physicist can acquire whatever mathematics, code, or calculational tool a problem demands. âThe difference between a good physicist and a great physicist is knowing what is the right question to askââthe hardest scientific skill, and usually the last one learned.
Current models resemble exceptionally skilled graduate students: once handed a precise question, they may be superhuman at the computation. They do not yet consistently decide which question deserves pursuit, so human taste remains centralâand the academic skill of matching a question, detail level, and framing to a student transfers directly to prompting an AI collaborator.
That advantage may be temporary. When Lupsasca gives ChatGPT Pro a page of the gluon paper and requests the three best follow-ups, it returns approximately his own top three. He declines to disclose internal work on autonomous frontier discovery, but says question selection is already becoming surprisingly strong.
10. The black-hole âLoveâ test made the trajectory personal
The hosts challenge whether the apparent creativity is merely recombination. Lupsasca replies that human invention may also be ârecombination of a stack machineâ and that GPT feels like a creative collaborator. He preserves a dissenting benchmark: in his understanding, Terry Tao has traced apparently creative AI proofs back to obscure references and has not yet seen a truly impressive novel mathematical move.
Lupsascaâs decisive GPT-5 experiment used his then-unseen paper on why black holes have no âLoveââLove numbers measure tidal response, and black holesâ zero response suggests a protecting symmetry. Because the modelâs training cutoff preceded publication, he supplied only the governing equation and asked for its symmetries. It initially answered, incorrectly, that none existed.
He then offered the obvious flat-space warm-up. GPT-5 Pro solved that 200-year-old problem correctly in nine minutes, identifying three generators; primed in the same chat, it returned to the black-hole equation and found the new symmetries in 18 minutes. âThat was my Move 37,â he saysâthe moment a model reproduced his prized calculation in under half an hour.
11. Cheap paper generation raises the bar and makes verification scarce
With expert steering, Lupsasca believes models can already produce human-quality papers, perhaps one per day for calculations adjacent to known work. Bad questions can equally generate polished nonsense, contributing to âAI slopâ and inundating arXiv. The community can no longer treat professional presentation as evidence that the underlying question or derivation is sound.
His response is not to publish 30 variants of the single-minus papers. The standard for significance must rise as production gets cheaper. He wants to use the new amplitudes as a line of attack on progressively harder quantum-gravity questions, aiming eventually at problems that have resisted whole communities for decades.
Static papers themselves look inefficient: AI performs a calculation, humans compress it into terse notation, and the result is then put back into an AI to expand it again. Lupsasca imagines an interactive paper that can explain the big picture, zoom into a derivation, or expose the intuitive pictures mathematicians currently leave off the pageâwhile retaining writingâs useful discipline of clarifying thought.
Verification may become this yearâs dominant bottleneck. Lupsasca once thought formal proof systems essential, then natural-language reasoning improved enough to resemble human mathematical discussion; now models can attack thousands of questions and return too many proofs for people to check. Formalization in systems such as Lean, automated checking, and clearer model confidence signals therefore become valuable again.