Opus 4.8 Drops, Demis Hassabis Predicts AGI, and the $220B Foundation | EP #260
Summary
- Opus 4.8 reclaimed the coding lead, but orchestration—not another benchmark point—may be the more consequential capability shift. Peter cited 61.4 on the Artificial Analysis Intelligence Index and 69.2 on SWE-bench Pro versus GPT-5.5’s 58.6; Dave’s concrete observation was that roughly 100 concurrent agents now integrate their work far better. Alex expects the monthly horse race to become weekly, daily, then hourly, while saturated tests force evaluation toward “open unsolved problems.”
- Demis Hassabis’s 2029 AGI call exposed a market with no agreed finish line. Alex argued generality arguably arrived with GPT-2 or by 2020 and accused the industry of moving goalposts. Alex separately argued that IQ tests capture only raw thought processing and conceptual matching, omitting physical, spatial, emotional, and spiritual intelligence. Peter predicted—and Alex agreed—that the debate will migrate from AGI to sentience while systems continue toward solving everything regardless of the label.
- The OpenAI Foundation could become an economic-policy institution as consequential as the frontier lab it controls. Peter put its 26% PBC stake at $130 billion-$260 billion and said the foundation selects 100% of the company’s board, while its new $250 million economic-futures program studies public wealth funds, worker ownership, and AI dividends. If frontier labs absorb a large share of global output, Alex expects “irresistible pressure” for their 20%-25% nonprofit arms to fund UBI, universal basic services, compute, capability, or equity.
- Agentic commerce moves retail power from shelf position and search ranking to the preferences of each customer’s AI. Amazon’s Alexa-based assistant reportedly converts at 3.5 times keyword search and is being offered to retailers, while Google is assembling a horizontal stack of universal carts, commerce protocols, and agent payments. Salim’s call was blunt: “The retail war is not now shelf space. It’s agent preferences.”
- Cheap electrons and domestic physical infrastructure are becoming the binding inputs to AI abundance. Wind and solar reportedly reached 22% of global electricity in April 2026 versus gas at 20%, while the announced $2 billion Albany quantum foundry could produce devices 30 times faster and eventually support quantum-accelerated AI. The panel’s national-competitiveness test is how rapidly a country can connect low-cost energy, compute, chip production, and robotics supply chains.
- The US AI race now carries political and social execution risk alongside technical risk. Peter described “anti-tech extremism,” attacks on technology leaders, campus hostility, and data-center opposition. Dave asserted that evidence of Chinese involvement was “pretty irrefutable,” while Peter said he personally lacked evidence and framed foreign interference as a theoretical possibility. California’s real-time workforce dashboard was treated more constructively—as “government as sensor” for hiring freezes, vulnerable sectors, retraining, and a possible post-labor economy.
- Dave is redirecting capital from a short-lived pure-software advantage toward longer-duration robotics, while diagnostics offer a parallel physical-economy opportunity. He sees perhaps two more years for the unusually favorable pure-software AI window but roughly a decade for robotics and biotech; China’s more than 150 humanoid companies make the physical stack a strategic issue. Westlake University’s reported $5, single-drop blood sensor—near 95% accurate and 10,000 times more sensitive than standard lab tests—showed the parallel opportunity to demonetize diagnosis until “consumers can own it.”
Deep dive
1. Opus 4.8 wins the scorecard as familiar benchmarks saturate
Peter framed Anthropic’s release cadence—Opus 4.8 arriving six weeks after Opus 4.7—as a direct answer to GPT-5.5. The reported numbers were 61.4 on the Artificial Analysis Intelligence Index, 1.2 points ahead of GPT-5.5, and 69.2 on SWE-bench Pro against 58.6; Peter also reported that it was four times less likely to miss bugs in its own code.
Alex’s closest-watched evaluations were SWE-bench Pro at 69.2%, Humanity’s Last Exam with tools at 57.9%, and GDPval at 1,890. His conclusion was not that these scores settle the race, but that “we’re at the saturation phase” and need benchmarks built around scientific and engineering problems whose answers remain unknown.
The apparent convergence among frontier systems has three explanations in Alex’s account: labs may avoid costly dramatic leapfrogs, their compute footprints sit within roughly a factor of two or three, and saturated benchmarks mechanically compress dispersion. “It’s very easy with superintelligence to just saturate every obvious benchmark that you throw at it.”
The group called the field a “two-and-a-half” lab race and expected GPT-5.6 within weeks. Alex’s cadence forecast ran from monthly releases to weekly, daily, and ultimately hourly updates—“the singularity of the singularity.”
2. Parallel agency matters more than the decimal-point upgrade
Dave was running about 100 agents on EC2 and found the upgrade “significantly better at managing many, many parallel threads.” Previous agents could work concurrently but failed to assimilate their contributions; the new behavior moved closer to AI’s potential advantage of “a billion concurrent workers” producing one coherent artifact.
The operative feature was self-forking: instead of launching context-free children and spending 20-30 minutes briefing them, Dave could ask an agent to clone everything it knew into 100 identical workers. The implementation still resisted, warning of context bloat, so his verdict remained conditional: “I’ll let you know in a couple of days if it succeeds in self-improvement and self-assimilation.”
Salim drew the infrastructure implication: users will need orchestration that routes high-cognition work to the newest models and low-cognition work to cheaper older ones. Model consistency therefore does not eliminate differentiation; it shifts value toward routing, context management, workflow scaffolding, and integration.
Alex likened inherited agent context to Unix child processes rather than biological children, which arrive “context-free.” The joke about cloned agents acquiring voting rights carried a serious point: software labor is beginning to reproduce with memory intact.
3. The AGI finish line keeps moving faster than the systems
Peter presented Demis Hassabis’s tightened 2029 timeline, now aligned with Ray Kurzweil, alongside Hassabis’s warning that today’s agents are a “practice run” and society has only a few years to prepare. Hassabis’s proposed Einstein test would train a system only through 1901 and ask it to derive special relativity independently.
Alex’s pushback was categorical: some form of AGI has arguably existed since 2020, while generality may have appeared with GPT-2 and the demonstration that large language models were few-shot learners. Calling AGI three or four years away while Gemini is “not winning the race,” he suggested, conveniently gives DeepMind time to leapfrog.
The disagreement turned on definitions. Alex grouped the camps as doomers saying “we’re already cooked,” skeptics saying AI “can’t do my laundry yet,” and Hassabis requiring a replication of relativity. Alex also argued that IQ-like tests capture only raw thought speed and conceptual matching while omitting physical, spatial, emotional, and spiritual intelligence.
Peter recalled the show’s own 50% Humanity’s Last Exam threshold; Opus 4.8’s 57.9% with tools had already crossed it. Peter predicted that society will keep moving the AGI goalpost and then announce, “AGI? Sentience,” while Alex agreed that the debate would continue over what sentience means. Alex also argued that the exact date will look like a historically compressed smear—much like trying to assign one year to the Industrial Revolution.
4. Agentic commerce relocates the retail gatekeeper
Amazon’s Alexa-based shopping assistant reportedly converts customers at 3.5 times the rate of keyword search and is being offered across retailers. Peter cast Amazon’s approach as vertical—own the customer relationship—while Google’s universal cart, Universal Commerce Protocol, and agent-payment protocol form a horizontal infrastructure layer.
Salim called Amazon’s history “better late than never.” Its original bet—put shopping inside Alexa speakers—failed because consumers did not want purchase conversations with hardware; putting Alexa-style conversation inside the marketplace succeeded, an inversion Amazon could have attempted years earlier.
Dave credited Amazon with moving roughly 60% of product searches away from Google but questioned why it never built a leading foundation-model effort. Salim described Amazon’s incrementalism and lack of a fundamental product-line shift, then summarized the strategic change: brands will compete for “agent preferences,” while assistants that anticipate demand—or steer shoppers toward more profitable products—replace page-one placement as the contested surface.
5. OpenAI’s foundation now has capital and control
Peter said the OpenAI Foundation owns 26% of the public-benefit corporation, valuing the stake at roughly $130 billion-$260 billion. At the upper end it exceeds the reference foundations he cited—Novo Nordisk at $150 billion, Tata Trusts at $100 billion, and Gates at $75 billion—and creates what he called the world’s largest philanthropic war chest.
The grant sequence Peter described began with a $40 million People-First AI Fund distributed among roughly 28 US nonprofits in 2025, followed by an October commitment of $25 million across health breakthroughs and AI resilience. The newest $250 million economic-futures grant supports work on public wealth funds, worker ownership, and AI dividends.
Ownership understates influence: Peter identified Bret Taylor as OpenAI’s chairman and said the foundation, despite holding 26% of the stock, controls who sits on the PBC board—100% of the board. Dave interpreted the mandate as “global calm, peace, prosperity” and urged practical proposals for deploying the capital through the AI transition rather than online denunciations.
6. AI abundance forces a fight over where value accrues
Salim’s foundational question was not simply job loss but value accrual: if labor’s share declines and technology demonetizes capital-intensive goods, does value flow to consumers, governments, public ownership, or a new structure? He called this the defining economic question for the next 20-30 years.
Alex connected OpenAI’s former plan to allocate 20% of compute to superalignment with Social Security’s roughly 22% share of US federal spending. If frontier labs approach the scale of the global economy, he expects pressure for their 20%-25% foundations to support UBI, universal basic services, universal basic compute or capability, or universal basic equity.
Salim argued properly designed UBI or services can be libertarian because cash and market allocation replace centralized government programs; he compared universal basic compute or services to free land for American settlers in 1649—“homesteading for AI.” Alex dissented: concentrated labs distributing benefits under political pressure looked more like “privatized socialism.”
Alex offered Fordism as another analogue: Henry Ford paid workers enough to consume mass-produced goods; frontier labs might distribute money or tokens so users can buy AI services, sustaining the cycle. He stressed that this was a thought experiment, while Peter’s counterpoint was that unprecedented, potentially unbounded abundance makes the distribution problem unusually tractable.
7. A quantum foundry is a hedge on AI’s next substrate
The announced Albany facility combines $1 billion in CHIPS Act funding with $1 billion from IBM, uses a 300-millimeter process, and is intended to produce quantum devices 30 times faster. Peter’s analogy was a TSMC-like foundry where Google, IonQ, Rigetti, D-Wave, and others might fabricate devices.
Alex normally calls quantum computing “a solution in search of a problem,” unlike quantum sensing, but judged this a smart anticipatory bet. By the late 2020s, when the foundry reaches scale, quantum-accelerated AI training or inference might justify the capital—and the US would want the superconducting-qubit infrastructure onshore.
Salim retained the technical hedge: the field still needs about 1,000 physical qubits per logical qubit because of errors. A 30-fold manufacturing improvement does not solve that ratio, but Alex noted that cheaper devices could flood the system with physical qubits and change the experimentation curve.
8. Solar’s exponential is outrunning institutional forecasts
Peter cited April 2026 data showing wind and solar at 22% of global electricity, above natural gas at 20%, with nearly 530 terawatt-hours generated. Growth was reported at 14% in China, 13% in the EU, and 35% in the UK.
Alex returned to Ray Kurzweil’s stacked S-curves: relays gave way to vacuum tubes, then transistors, while silicon panels can yield to perovskites and later technologies. Solar has doubled roughly every 22 months for 40 years, so the curve can persist even when one implementation saturates.
Their institutional-warning specimen was the IEA: Peter said its projections were repeatedly wrong, while Alex said its 2020 model did not expect wind and solar to surpass gas until the mid-2030s. Dave recalled an IEA electric-vehicle forecast that was effectively obsolete by the end of the year it was published. Dave called this “not a math error” but “a cognitive error”—exponentials look flat backward and impossible forward.
Cheap electrons become strategic because abundant intelligence requires abundant energy. China was said to lead US solar deployment by 10 times and already operate an “inner loop” in which robots make panels that power more robots; Alex expects any future SpaceX-scale orbital buildout to force much larger domestic, imported, lunar, or sun-synchronous solar production.
9. Anti-tech backlash has become execution and security risk
Peter said federal agencies had created an “anti-tech extremism” category and logged more than 1,000 pages on threats to data centers and executives after attacks involving a Molotov cocktail and gunfire at Sam Altman’s home. He argued that opposition capable of delaying infrastructure now warrants treatment as a domestic-security issue.
Dave called evidence of Chinese involvement “pretty irrefutable” and compared agitation against US technology to Soviet efforts during the Cold War. Peter preserved the essential caveat: he did not personally have evidence and was describing a theoretical, highly efficient way for an adversary to “put sand in the gears.”
Salim emphasized the human vulnerability: the amygdala is 10 times more likely to attend to negative information, making outrage cheap to manufacture. Peter contrasted claimed AI optimism of 80%-85% in China with roughly 25% in the US, then described 21- and 22-year-olds facing peer hostility and Eric Schmidt being booed at a commencement address.
10. California is building a sensor for the post-labor transition
Governor Gavin Newsom’s executive order, as Peter described it, creates a public dashboard tracking AI-related job losses, vulnerable industries, retraining, and possible UBI models. He welcomed the data if it remains unbiased, especially because the current pattern appears to be reduced hiring rather than mass firing; 22- to 28-year-olds were said to face the longest unemployment.
Dave put maximum AI-attributed job losses at about 300,000 and said the feared crisis had not arrived. He once expected roughly half of Vestmark’s jobs to be automated; rapid growth and rising profitability changed his expectation to zero cuts, although he acknowledged that college recruiting was exceptionally weak.
Dave’s spot check of laid-off Microsoft, Amazon, and Meta engineers found almost everyone joining a startup or another recently founded company. He explicitly limited that rosy result to software talent and did not claim it would generalize to physical occupations such as garbage collection.
Alex’s positive framing was “government as sensor”: faster signals can trigger reskilling and ownership experiments before 18-month labor statistics arrive. Alex also favored state-level policy experimentation but wanted federal protection against a Balkanized AI regulatory map, contrasting California’s dashboard work with wealth taxes that might drive technology leaders away.
11. Diagnostics and robotics carry the next physical-economy premium
Westlake University’s handheld optical sensor reportedly detects early-stage lung cancer from one drop of blood with near-95% accuracy, is 10,000 times more sensitive than standard lab tests, and costs about $5. Alex described metamaterials measuring tiny refractive-index changes and quipped that this time the breakthrough was “not coming from Theranos.”
The price matters because cheap hardware can move from centralized screening into homes and wearables. Alex forecast—without committing to a firm timetable—that noninvasive optical cancer monitoring might arrive in roughly five years; the transformative endpoint is continuous detection at inception, with the data interpreted by personal AI. “Consumers can own it.”
Robotics, meanwhile, is becoming a national stack. Alex cited China’s “AI Plus” plan and more than 150 humanoid companies, calling for US demand support and favorable deployment rules; his warning was practical: “If we can’t get Waymos, how are we going to get humanoid robots everywhere?”
Dave sees perhaps two years remaining in the extraordinary pure-software AI opportunity but a 10-year runway for robotics and biotech. After roughly 80 consecutive AI deals, he is emphasizing shared operating systems, manufacturing, actuators, supply chains, and space-ready designs; Figure’s package-sorting run beyond a week and Salim’s enthusiasm for six-armed robots illustrated why capability need not stay strictly humanoid.
12. Physical setbacks do not alter the abundance endgame
Blue Origin’s New Glenn exploded during an unmanned static-fire test, with no injuries. Alex said the general consensus was that the failure could set back Blue Origin’s Artemis participation by up to a year, strengthening SpaceX’s near-term position for lunar missions and potentially affecting Amazon’s Project Kuiper launch plans; the panel’s refrain was simply, “Hardware is hard.”
Asked whether China reaches abundance first, Salim split material from human abundance. China could lead in cheap solar, batteries, EVs, robotics, and logistics, but abundance also requires agency, experimentation, open innovation, and meaning—areas where the West might retain an advantage if it preserves its freedoms.
Salim argued that today’s agent costs are a temporary snapshot: token prices had reportedly fallen 75%-90% in 18 months, and he projected a $20-per-month agent falling to $2 in 2028 and $0.20 in 2030. Against roughly $10,000 annual US education costs, personalized AI tutors make universal basic compute materially different from traditional service provision.
The closing policy and physics questions centered on different forms of agency and capability. Salim said the 21st-century privacy question is what AI may infer, manipulate, deny, or price—not merely what it knows; Dave opposed token taxes because they discourage productive use. Alex went further, forecasting that AI could transcend semiconductors within a decade or two, potentially computing through plasma, gravity, or “pure energy and stress-energy tensor.”