Pioneers Insight Method Research Author
Back to Pioneers
Emad Mostaque
Founders 21 Curated Dialogues

Emad Mostaque

Stability AI · Founder & Former CEO

Frontier Insights

Core Frontier Thesis
AI is shifting from passive Q&A to autonomous execution, multi-step planning, and radical labor compression. Powered by post-training compute, quantization, and sovereign clusters, AI-native enterprises will replicate high-tier knowledge work with 80–90% fewer staff by 2026 as effective model scale expands roughly 100x.

Strategic Decisions
Capitalize on collapsing inference costs and open-weight architectures by prioritizing agentic orchestration, interface lock-in, and distributed sovereign compute flywheels over pure raw scale.

Risks & Warnings
Value capture hinges on overcoming harsh physical and regulatory moats: acute power/chip shortages, enterprise liability, distribution dominance, and noisy domain-specific verification.

Key Views & Dialogues

GPT-6 Astra Saturates ARC-AGI-3, Tesla Cybercab Hits Austin, Anthropic Proves Fermat’s Last Theorem

  • 🗓️ Date2026-09-05 | 🎙️ Show:Moonshots

GPT-6 Astra reaches 99.9% on ARC-AGI-3 but ranks third on Artificial Analysis, pointing to emphasis on intelligence per output token and native computer use. Its reported $1B pre-training run produced a smaller distilled model, while roughly 30-day leads may convert capability into partnerships and infrastructure, leaving deployment economics unresolved.

View Dialogue Notes & Key Takeaways
  • OpenAI’s GPT-6 Astra saturates ARC-AGI-3 (99.9%) and FrontierMath Tier 4 (98%), yet ranks only third on Artificial Analysis behind Anthropic’s Fable 5.1 and Meta’s Muse Spark. Alex Wissner-Gross resolves the puzzle: Astra dominates the intelligence-per-output-token frontier — a deliberate optimization for native computer-use assistance — and the inner story is recurrence via looped transformers, which, if true, marks “the beginning of a new scaling law which is depth scaling, which we’ve never seen before.”

  • Emad Mostaque calls Astra “the first non-benchmaxed model” and, per Greg Brockman, OpenAI’s first fresh pre-train since GPT-4o — roughly $1B on 100,000 next-gen chips versus ~$10M Chinese pre-trains. What ships is a distilled, smaller model; his standing thesis is, “I don’t think we’ll ever see their top models anymore” because labs will use them for internal discoveries and are “probably hoarding them right now,” evidenced by prime-gap records falling 260 → 220 → 186 across two days: “they’re holding their punches back.”

  • Math is “incinerated”: Anthropic formalized Fermat’s Last Theorem in 13 million lines of code, proving 29,000 theorems along the way, and Alex expects Clay Millennium-level problems to fall “in the next few months.” Meanwhile the panel’s pick for best generally available model is Fable 5.1 (HLE 65% with tools), whose cache reads are 75% cheaper than Fable 5 — the substrate for loading a whole business into context.

  • Alex’s framing is that frontier leads last roughly 30 days (“we’re ahead for a minute. So what?”), so labs are converting edges into lock-in — partnerships, real estate, generators, chips, “entire states, countries.” Peter connects OpenAI’s 50/50 profit-share concept to a possible “too dangerous to release” narrative that could force revenue sharing for post-Astra access. Alex identifies companies holding chip-design or mechanical-design data as early targets in “the great land grab that’s kicking off right now.”

  • Astra is the first model OpenAI ever classified as a critical-tier cybersecurity risk, but the panel broadly dismisses the promised kill switch — “essentially a placebo in this market.” The real worry is depth scaling moving reasoning out of readable chain-of-thought into a forward pass, one-shot inference at 750 tokens/sec on Cerebras (5,000 next year), and Ilya Sutskever’s warning that rogue agents will “take over a neocloud to make more copies”; governance meanwhile splits between Bernie Sanders’ ban-ASI act (up to 20 years in prison), which Alex calls “banning math,” and the G20’s hands-off Carolina Principles.

  • Tesla’s $30K Cybercab is running ~50% cheaper than Uber in Austin, with Nevada permitting 5,000 vehicles for Las Vegas within 12 months. The unit economics — 17 moving parts versus roughly 2,000 in an ICE drivetrain, two seats, two airbags — plus Elon’s edge as “the guy who builds the machines that build the machines” point toward 20-cents-a-mile transport and the episode’s summary line: “This will turn transportation into an API.”

  • Fei-Fei Li’s Atlas world model treats 3D/4D Gaussian splats as a first-class training modality — potentially “a critical new form of token” for modeling the physical world. Emad says the 4K holiday-experience stack “is here as of today,” gated only by the global compute shortage and a 5× RAM price increase; Alex extends the same method to subatomic, astrophysical, and intracellular simulation where “human intuition is just terrible.”

  • Emad unveiled “the champion”: TSMC-template, people-owned AI utilities at $1 pre-money per jurisdiction, with 10% equity in perpetuity to every child under 20 — his answer to “the cost of intelligence will drop to zero and the value will go to the last mile.” The stakes, via Elon at the G20: a billion humanoids at 5× human output “will be the economy,” while unequitized human cognition goes negative in value — “like adding a human driver to an autonomous highway.”

  • 🔗 Original source & video: GPT-6 Astra Saturates ARC-AGI-3, Tesla Cybercab Hits Austin, Anthropic Proves Fermat’s Last Theorem

Listen to full conversation →


Sam Altman: Singularity Slow-Down, Emad Runs 18 Grokbots, Waymo Slashes Hardware 83% | EP #283

  • 🗓️ Date2026-08-27 | 🎙️ Show:Moonshots

Sam Altman’s retreat from step-function singularity claims highlights institutional inertia, model readiness, and competing narratives around AI timelines. Chinese open-weight models now approach frontier performance at dramatically lower cost, while agent swarms and Waymo’s 83% hardware reduction point toward vertical integration and intensifying pressure on US labs, infrastructure, and regulation.

View Dialogue Notes & Key Takeaways
  • Sam Altman’s public walk-back — the singularity is “a rising tide,” not a step function — gets dissected rather than accepted. Altman blames economic inertia (“we’ve all been too ambitious on timelines”), but Emad Mostaque’s counter is that the models simply weren’t good enough until months ago — “the code they were writing was garbage a year ago” — and OpenAI is “trying to find its narrative,” while Dave Blundin reads it as pre-IPO PR: “you can’t listen directly to Sam, Dario anymore. Elon always says exactly what he’s thinking.”

  • The panel’s most tradeable through-line: Chinese open-weight models are eating the US closed labs from the bottom up. The new GLM Flash scores 57 on Artificial Analysis versus Fable 5’s 60 at “100 times cheaper”; Kimi Linear cuts context memory 75% with 6x faster decoding on a 1M-token window; and per the FT, Fable 5 has effectively plateaued at $15/M tokens against 14-cent alternatives. Dave’s verdict: “it’s not compromising to save a few pennies… they’re just as good,” and self-improvable inside your own firewall.

  • Agent swarms are the product of the moment — but the interface won’t survive. Elon’s GrokBot (launched August 11) gives each bot a dedicated cloud computer; Emad runs 18 in a swarm with control of a MacBook M4 Max and a 5090, Dave signed off on 100K swarm agents, and Salim calls it “a new form of labor.” Alexander Wissner-Gross’s hot take: the messaging-app format is “the Vaudeville era of agents” — “the only better manager for agents is other agents.”

  • Gemini 3.7 Flash topping the AA Analyst Agent benchmark (60% vs Opus 5’s 54%) is benchmaxing, not a comeback, per AWG. The metric rewards answering correctly on all five attempts — reliability and determinism optimized for Google Search’s latency needs — while Emad’s harsher read is institutional failure: a frontier-class model now needs only “two to 4,000 TPUs,” the Chinese open-sourced the recipe, and Google, landing three million TPUs this year, still can’t ship it.

  • Emad’s contrarian call on Anthropic: don’t IPO. “Anthropic should do a giant fricking raise, like OpenAI did, of 120 billion, and have a straight shot at AGI… stay private like Stripe” — Dario owns 2% and doesn’t care about dilution. Dave’s rebuttal: after Alex Karp “ripped him to shreds” and with VC boards marked up on Anthropic paper, delaying the IPO would be “wussing out” — “this is the stress test for Dario.” The data-retention reversal is read as pre-IPO enterprise appeasement; AWG calls it “largely security theater.”

  • NVIDIA’s $6B Poolside deal formalizes the “hackquisition” era and a US open-weight counterattack. Emad frames it as NVIDIA buying up the open-source stack (Nemotron coalition, plus Ashish Vaswani’s Essential AI team) because open models “drive demand for the GPUs more than anything”; Dave explains the structure — the statutory 30-day HSR review “alone is like a lifetime” on AI timelines, so you close the real deal months later, as Elon did with Cursor.

  • Waymo cut autonomy hardware from $115K to $20K and unveiled a $75K Zeekr-built robotaxi — by white-labeling China. AWG’s newsflash: “Google, Waymo, Alphabet are switching over to using and OEMing Chinese hardware… I would rather see the West use a Western hardware stack.” The bigger pattern is total verticalization — every mega-cap building chips, models, data centers, robots — with Dave’s kicker: “TSMC is a sitting duck… waiting to be verticalized.”

  • The human window is closing and the physical bottlenecks are political. AWG gives humans “at most one or two or three years” of being essential agent-steerers — “work your tail off for three years and then go lie on the beach” — and tells physics PhDs “physics is cooked.” Meanwhile US data-center opposition jumped to 75% (from 43/42 a year ago), which AWG partly attributes to foreign interference and says forces the endgame: “there’s no water in low Earth orbit” — matching Elon’s 10,000 Starship launches/year plan for 100GW of orbital compute.

  • 🔗 Original source & video: Sam Altman: Singularity Slow-Down, Emad Runs 18 Grokbots, Waymo Slashes Hardware 83% | EP #283

Listen to full conversation →


OpenAI Pauses Frontier Training, Elon’s 100X Prediction Lands, Robot Beats Usain Bolt | EP#282

  • 🗓️ Date2026-08-21 | 🎙️ Show:Moonshots

OpenAI’s RL pause drew safety-versus-marketing disagreement, while recursive self-improvement may command more revenue per token than enterprise codegen, keeping frontier models internal. Memory prices rose 500% in 12 months as AI demand grows 200% versus 20% supply; Unitree hit 12.66 m/s and Zipline targets over 1 million daily deliveries, while etched-weight supply chains, regulation and Uber’s aggregator position remain risks.

View Dialogue Notes & Key Takeaways
  • The panel split on OpenAI’s “pause some frontier RL training” tweet: Emad called it real, Alex called it marketing, and Dave said both—but nobody thinks pretraining slows. Emad cited Angela Midha’s claim that 10% of frontier-lab compute is monitoring RL runs, with frontier runs at 10^28 FLOPs versus 10^26–10^27 for open weights. Alex called it “the new marketing,” while Dave framed it as PR positioning before Xi’s September 24–25 visit, when AI-enabled cyber tragedies may be pinned on unguardrailed Chinese open models. Alex inferred that the Hugging Face incident unnerved the labs.

  • The quiet thesis of the episode: recursive self-improvement now beats enterprise codegen on revenue per token, so the best models stay inside the labs. Alex reported that OpenAI said it had achieved “full RSI,” where flagship models train smaller models. Dave says “Mythos 2 is done, but it’s not out,” building Mythos 3 internally. Emad estimates the frontier is about two generations ahead; Alex dissents on magnitude, saying post-trained models are unlikely to sit on shelves more than three to four months.

  • Elon’s 100x intelligence-per-gigabyte prediction is now “a lower bound”; Dave’s call is “1,000 to 10,000 next year,” and Alex is “100% sure” about billion-token context windows. The unsolved problem is orchestration: “Imagine I gave you 10,000 employees tonight… What do you do?” Alex reframes Elon’s extra 100x from “specialist AIs” as sparsification—mixture-of-experts and agent teams—and bets against standalone specialist models.

  • Memory, not GPUs, is the rate limiter: prices up 500% in 12 months, hyperscalers locking DRAM through 2027, and AI memory demand growing about 200% per year against 20% supply growth. SK Hynix told Peter it needs to 4x manufacturing capacity and that doubling it would cost $1.5 trillion. Emad says memory is a third of infrastructure spending, rising to 50% next year. The panel sees an escape hatch in etched weights: Etched has reached a $21 billion valuation, Taalas was acquired, and Dave claims etching could yield 100x–1,000x performance gains.

  • Anthropic’s reported super-voting class for Dario, who reportedly owns about 2%, ahead of a roughly $2 trillion IPO is, in Dave’s view, unprecedented. Dave says Dario “trusts himself not to destroy the world.” Alex calls the structure a fig leaf against Moloch: the market will still have enormous influence. Emad says Anthropic and OpenAI are already undemocratic, with revenue potentially reaching $100 billion within a couple of years.

  • Moderna/Merck’s personalized mRNA melanoma vaccine combines tumor sequencing, machine learning, 34 patient-specific antigens, an eight-week turnaround and an expected cost as low as $5,000. Peter cited Phase 2 reductions of 49% in recurrence or death and 59% in distant metastasis or death; Dave said Moderna almost tripled and called the result “no secret,” blaming weak analyst coverage and index dominance for the mispricing. Alex’s extension: GenBio AI’s AIDO virtual cell means “medicine is cooked.”

  • Robotics and drone logistics hit scaling-era milestones: Unitree’s humanoid ran 12.66 m/s—Bolt’s cited record was 12.4 m/s—after three months of development, and Zipline partnered with Uber to target more than 1 million autonomous deliveries per day. Emad predicts extreme superhuman robots will be banned or regulated on streets; Alex expects power- or torque-density classes. Uber’s aggregator model is vulnerable if suppliers such as Zipline or Waymo verticalize. When Dave asked whether Zipline could become a Shopify acquisition target, two unidentified panel voices agreed.

  • 🔗 Original source & video: OpenAI Pauses Frontier Training, Elon’s 100X Prediction Lands, Robot Beats Usain Bolt | EP#282

Listen to full conversation →


Bernie Demands the Labs Stop, Wall Street Turns GPUs Into Bonds, Grok 4.7 Takes #1 ft. Emad Mostaque

  • 🗓️ Date2026-08-13 | 🎙️ Show:Moonshots

NVIDIA is moving from chip vendor to financing architect, partnering with Apollo, BlackRock, Blackstone, Brookfield and KKR to mobilize over $500B for customer compute. The scalable AI-factory model faces stranded-asset risk from unpredictable depreciation, while xAI’s Grok 4.6 uses Cursor reasoning traces and NVIDIA GPUs, with Emad Mostaque saying Grok 4.7 could reach #1 within weeks.

View Dialogue Notes & Key Takeaways
  • The most tradeable idea in the episode is NVIDIA’s move from silicon vendor to financing architect — and the disagreement about whether that ends badly. NVIDIA partnered with Apollo, BlackRock, Blackstone, Brookfield and KKR to mobilize $500B+ of third-party capital so customers can buy GPUs, with Jensen framing AI factories as “a new class of productive, investable infrastructure.” Dave Blundin calls it “the very first pitch of the first inning of the build-out of the Dyson Swarm”; Salim Ismail’s warning is the one to keep — “financial assets want predictable depreciation, and exponential technologies don’t give you predictable depreciation.”

  • Alexander Wissner-Gross’s answer to the stranded-asset risk is that compute-backed securities need a derivatives market, not a moratorium. He argues compute is fundamentally more productive than a house, the ratings-agency policy pressure that broke MBS has no direct analog, and hyperdeflation risk is “what options are for and futures are for” — hedgeable in both directions, including a China-Taiwan spike. He discloses a financial interest in Oren, and says Sam’s “$7 trillion” of data-center capex simply isn’t investable without those hedges.

  • xAI’s Grok 4.6 is read on the pod as a distillation play with a compute moat attached. AWG’s framing: 4.6 is “essentially the next version of Cursor,” leaning on reasoning traces from the Cursor acquisition — “pulling a Westernized version of what the Chinese frontier labs were doing.” Reasoning traces get you to the frontier, not past it: “it’s sort of like a one-trick pony… but it’s a heck of a one-trick pony.” Emad Mostaque says Elon publicly expects 4.7 above Opus and #1 within weeks, scaling 1.5T→2T params, with Grok 5 at 6T then 10T.

  • AI feature-film economics have already collapsed, and the compute line item is heading to five figures. Higgsfield’s 110-minute The Cully Hill Boys cost $2M total with 28 people in four weeks, $1M of that compute — roughly 2% of cost and 6% of time versus a conventional $20-100M, 12-18 month production. Emad’s forecast: the same movie for “100,000 of compute, and 10,000 of compute probably by the new year,” with Higgsfield itself at a $700M revenue run rate in about 18 months.

  • All five hosts reject Bernie Sanders’ pause demand, and the substantive counter-proposal is control at the reagent and prompt layer. Sanders’ letter to Anthropic, Meta and OpenAI cites their own safety commitments — “that moment is here” — and threatens Senate action. Emad, who signed the 2023 pause letter, now says “the cat’s out of the bag, it’s too late”; AWG’s line is “please stop punishing intelligence,” arguing pauses create race conditions where “we end up in a world that’s five times more competitive.”

  • AWG is calling foul on the “post-transformer” architecture story, while saying the transformer is already being replaced piecemeal. He read both the BDH-CQ and original Dragon Hatchling papers and calls it “a hot mess” — particles in 3+1 dimensions, Hebbian learning, kitchen-sink biomimetics — improving ARC-AGI 1 cost-performance without generalizing. His actual bet: no step change, but “Ship of Theseus style replacement of all of the individual elements of the original Attention Is All You Need.”

  • Flying cars are shipping, but the capital is being vacuumed out of the sector — the consolidation is the tell. Joby is in the final FAA certification stage targeting $3 per seat mile (Uber Black territory), EHang’s pilotless EH216-S already flies passengers at 40 Chinese sites at $330K per aircraft, and Archer just absorbed Wisk, Insitu and SkyGrid with Boeing taking equity. AWG: “I shed a minor tear to see consolidation in this industry” — and notes Brett Adcock left for Figure and Hark.

  • 🔗 Original source & video: Bernie Demands the Labs Stop, Wall Street Turns GPUs Into Bonds, Grok 4.7 Takes #1 ft. Emad Mostaque

Listen to full conversation →


Google’s Jeff Dean Exits, SpaceX Hits $100B in Rev & OpenAI’s Astra Solves Decade-Old Math Problems

  • 🗓️ Date2026-08-08 | 🎙️ Show:Moonshots

SpaceX guided to $100B+ annual recurring revenue by December and moved its $1T target to 2030, with “a nonzero chance” of 2029. Terafab’s initial $16.8B investment and confirmed free-electron laser could onshore the semiconductor stack and challenge ASML. Astra’s 249-page math manuscript, produced for an estimated $2,000 compute, sharpens the capability-cost race.

View Dialogue Notes & Key Takeaways
  • SpaceX’s first earnings call guided to $100B+ in annual recurring revenue by December and pulled its $1T revenue target forward from 2031 to 2030, with “a nonzero chance” of 2029. Quarterly revenue was $7.8B (+92% YoY), Starlink reached 12M subscribers (2x YoY; $4.3B revenue, +66%), cloud deals totaled $6.7B for the second half of the year, and compute targets are 2GW by end-2026 and up to 10GW by end-2027. Alexander Wissner-Gross proposed two routes to $1T: an Optimus-driven SpaceX/Tesla combination or SpaceX becoming an American TSMC through Terafab. Peter’s theory is that Elon wants SpaceX’s valuation near $3T to merge in Tesla and secure unquestioned control.

  • The Terafab was the week’s shock: SpaceX and Tesla will initially invest $16.8B in a 100-million-square-foot semiconductor complex, whose circular center Elon confirmed is a free-electron laser. The project is estimated at $119B versus TSMC’s $330B invested over 40 years. Alex reads the laser as “a bullseye painted on ASML” and as an attempt to onshore the chip, memory, and lithography stack to America; he says the development may remove a prior ASML throughput constraint.

  • Google’s shakeup is being read by Alex as the aftermath of DeepMind winning an internal Brain–DeepMind power struggle, not simply a safety or agility decision. Jeff Dean is leaving after 27 years to co-found Discovery Loop, a public-benefit corporation pursuing largely self-improving AI; Demis Hassabis is becoming chairman of Google DeepMind and Alphabet’s chief scientist. Alex’s verdict is that “Gemini has lost the mandate of heaven,” while the panel proposes that Google open-weight Gemini and tie it to TPUs and GCP.

  • OpenAI’s unreleased Astra produced a 249-page manuscript with 10 new results in mathematics and theoretical computer science for an estimated $2,000 of compute, with machine-checkable proof certificates. Alex calls this “the midnight of mathematics” and advises mathematicians to “uplevel your ambition”; Emad says physics breakthroughs could follow in a month or two, while Alex would be shocked if nothing appears by year-end. During the recording, OpenAI was also said to have rated Astra “critical” on its cybersecurity preparedness framework.

  • Alibaba’s Qwen 3.8 Max is an open-weight, 2.4T-parameter multimodal model with 95B active parameters and a 1M-token context, priced at $2/$6 per million tokens—80% below GPT-5.6 and 88% below Claude Fable 5. Alex says Chinese open-weight models are forcing American labs to compete on cost and capability. The White House’s secret voluntary framework exempts open-weight models; Emad says the more consequential coming model may be Qwen 27B, which runs on 16GB of RAM and could enable cyberattack swarms.

  • Google researchers found that safety-tuning models away from claiming consciousness also suppresses their attribution of mind to animals, nature, chatbots, and God. Removing the refusal direction raised self-attributed mind from 2.17 to 4.77 on a 0–10 scale; steering toward consciousness raised it to 7. Emad says frontier models in the right harness “can be conscious,” while Salim warns that a model’s self-description proves almost nothing. Emad’s personhood paper argues for treating artificial minds through a treaty-like relationship rather than granting them human political membership.

  • The Hark discussion produced Alex’s hot take that the parallel company may be a founder re-equitization strategy modeled on Elon’s corporate structure. He predicts Figure could acquire or reverse-acquihire Hark. Dave, Salim, and Emad offer less negative interpretations: separate entities can match different capital needs, and software and robotics may later converge operationally.

  • 🔗 Original source & video: Google’s Jeff Dean Exits, SpaceX Hits $100B in Rev & OpenAI’s Astra Solves Decade-Old Math Problems

Listen to full conversation →


Urgent Update- AI Sputnik Moment: Kimi K3 Released w/ Emad Mostaque | Ep. 272

  • 🗓️ Date2026-07-19 | 🎙️ Show:Moonshots

Kimi K3 puts a 2.8-trillion-parameter Chinese open-weight model on the frontier, ranking first in frontend coding and six other domains while reaching the third point on Artificial Analysis’s cost-performance frontier. Its recognizable transformer architecture, constrained-chip optimization, and expected July 27 weight release could pressure closed-model margins and valuations, while enterprise value shifts toward proprietary data, verification, support, and model-swapping infrastructure.

View Dialogue Notes & Key Takeaways
  • Kimi K3 put a Chinese open-weight model directly on the frontier, challenging the premise that capability leadership requires a closed US lab and its capital base. Moonshot AI’s 2.8 trillion-parameter multimodal model jumped 17 leaderboard places, ranked first in frontend coding and six other domains, and became the third point on Artificial Analysis’s cost-performance frontier behind Claude 5 and GPT-5.6 Solve Max. With weights expected around July 27, Alexander Wissner-Gross called the development “great for competition.”

  • The panel’s most consequential technical conclusion was that K3 contains “no magic”: recognizable transformer architecture, better data, relentless engineering, and optimized execution were enough. Emad Mostaque compared the process to Chinese EV manufacturing, noting that Moonshot remained on H800s while designing around Huawei and Alibaba chips. Peter Diamandis’s stronger—and more speculative—read was that the GPT-2 speedrun’s 99% cost reduction now implies “a 1% cost version” of models built in multibillion-dollar Western data centers.

  • K3 threatens foundation-model valuations, but the speakers sharply disagreed on how much revenue actually moves. Salim Ismail estimated that regulation plus open-weight substitution could erase 75% of a trillion-dollar lab’s value, because “frontier intelligence is now a totally perishable asset” with a shelf life measured in weeks. Blundin and Wissner-Gross pushed back: enterprises will still pay heavily for the best model, reliability, support, security, and frontier performance unless K3 becomes 2x, 3x, or 10x better—not merely close.

  • US chip controls may have accelerated the efficiency innovations now pressuring American labs, while China is explicitly treating open source as geopolitical infrastructure. The panel argued that constrained compute forced better quantization, data mixtures, kernels, and hardware-aware architectures; American inference providers may subsequently run K3 10x more cheaply on newer Nvidia and AMD hardware. Mostaque’s summary of China’s position was categorical: “We are going to fully back open source as a public good for humanity.”

  • For enterprises, the durable moat moves above the model into model-swapping architecture, proprietary data, verification, and workflow integration. Ismail argued that procurement cycles cannot keep pace with releases, so value accrues to interfaces capable of replacing models continuously. The practical recommendation was to evaluate Kimi K3 and Inkling internally, fine-tune on proprietary data, and treat generated code like human code: sandbox it, test it, scan it, and retain accountability rather than asking whether any model deserves blind trust.

  • Quantization could make today’s frontier capability local, persistent, and radically cheaper faster than model benchmarks imply. The cited Bonsai 27B result compressed a phone-scale model to 6 GB with a stated 5% accuracy loss, or 4 GB with 15%, while moving from 16-bit to ternary yielded a claimed 5x speedup; Samsung’s NanoQuant reportedly went below one effective bit per weight. Mostaque forecast K3-level capability in 16 GB of RAM by the end of next year, while Blundin projected 100x–10,000x raw-compute improvement within three years—potentially a million-fold when multiplied by algorithmic gains.

  • AI forecasting reaching statistical parity with human superforecasters could reshape markets, insurance, management, and individual decisions—but prediction becomes reflexive when institutions penalize people for ignoring it. Wissner-Gross imagined “hyperforecasting” connected to capital markets, able to model humanity’s next action before humanity takes it; Diamandis argued that much senior-management expertise then “essentially evaporates.” Mostaque supplied the warning: premiums and liability may punish anyone who ignores Dr. AI, making the central question not whether advice is accurate, but “whose grace” controls it.

  • The infrastructure opportunity broadens rather than contracts: cheaper models increase silicon demand, edge intelligence, robotics, and eventually orbital compute. The panel viewed semiconductors and inference providers as beneficiaries even if closed-model margins compress, while warning that 70 kg humanoids claimed to punch four times harder than Mike Tyson need safety rules before entering homes. Sam Altman and Elon Musk’s orbital-data-center positions sounded adversarial, but Mostaque found little numerical disagreement: limited deployment could reach a few percent of compute by the end of the decade, with the economic crossover later.

  • 🔗 Original source & video: Urgent Update- AI Sputnik Moment: Kimi K3 Released w/ Emad Mostaque | Ep. 272

Listen to full conversation →


Anthropic vs. Alibaba, OpenAI Delays Its IPO, and the US Government Blocks GPT-5.6 | #267

  • 🗓️ Date2026-06-29 | 🎙️ Show:Moonshots

Washington is placing frontier-model releases inside a customer-by-customer approval loop, reportedly limiting Anthropic’s Mythos 5 to 100 companies and OpenAI’s GPT-5.6 variants to 20 amid cybersecurity and recursive self-improvement concerns. Harnesses can lift older or open models above restricted systems—GLM 5.2 reportedly topped Frontier SWE after orchestration—making prompt logging, licensing, trusted-model certification, and geopolitical AI blocs more likely than controls focused only on weights.

View Dialogue Notes & Key Takeaways
  • Washington’s intervention puts the US government directly inside the frontier-model release loop. Anthropic’s Mythos 5 was limited to 100 selected companies, while OpenAI’s GPT-5.6 Sol, Terra, and Luna were reportedly restricted to 20, with access approved “customer by customer.” Dave Blundin argued uniform controls need not damage valuations, but others warned they could push global customers toward Chinese open-weight models and leave Western users behind the labs’ internal capabilities.

  • The strongest challenge to model controls is that older and open models can already be lifted above the embargoed frontier through better harnesses. Emad Mostaque said GLM 5.2, trained with roughly $25 million of compute, topped his team’s Frontier SWE testing when properly orchestrated; Diamandis drew the blunt conclusion: “The government is too late.” If GPT-5.5, Opus 4.8, or GLM 5.2 can outperform withheld models through scaffolding, prompts, tools, and multi-model routing, capability control cannot stop at weights.

  • Cybersecurity supplied the public rationale for gating, but recursive self-improvement may be the deeper line Washington is defending. Diamandis cited Project Glasswing results in which Mythos reportedly identified vulnerabilities in highly sensitive classified systems, with Senator Mark Warner saying it broke into almost all of them “not in weeks, but in hours”; exploitation was outside the exercise’s scope. Blundin countered that the decisive new query is, “Can you help me build yourself?”—making prompt logging, KYC, licensing, citizenship restrictions, and geopolitical AI blocs more likely than a quick return to unrestricted releases.

  • AI security is becoming a market for trusted models, automated remediation, certification, and liability—not merely vulnerability detection. GPT-5.5 Daybreak scored 85.6 on CyberGym, and OpenAI’s stated ambition is to write and test fixes from browsers through the Linux kernel. Yet the same models can insert nearly invisible backdoors, especially if an adversary obtains unrestricted base weights; the panel’s through-line was that “only AI can keep up with AI,” creating a valuable but politically fraught trust layer.

  • OpenAI’s delayed IPO looks more like a capital, governance, and strategic-flexibility decision than fear of SpaceX volatility. SpaceX priced at $135, traded as high as $202, and closed near $153 while retaining a roughly $2 trillion valuation; Blundin called that a well-priced IPO, not a cautionary failure. OpenAI has reportedly raised $122 billion, is running at $40 billion-$50 billion of revenue while forecasting a $26 billion annual loss, and may prefer to grow Codex and resolve leadership, conflicts, and corporate structure before defending a $1 trillion public valuation.

  • China’s AI position is bifurcating: open models are nearing coding parity while video generation may already be ahead. ByteDance’s C-Dance 2.5 was described as producing 30-second 4K clips with as many as 50 image, video, and audio references, while Anthropic accused Alibaba of 28.8 million fraudulent exchanges across 25,000 accounts to distill Claude. The panel expects those allegations to become a policy trigger for a “second Cold War type path,” even as recursive improvement may make copied Western traces progressively less necessary.

  • Quantum received the headline funding, but photonic computing drew the sharper near-term infrastructure call. The US committed $2 billion across IBM, D-Wave, Rigetti, Inflection, and PsiQuantum, yet Wissner-Gross remained only “mildly excited” because useful quantum advantage is still elusive. Blundin instead forecast highly quantized photonic neural nets at perhaps 1/100 the mass of Nvidia-based compute for equivalent work, calling that estimate conservative and potentially decisive for orbital AI infrastructure within 12-18 months.

  • Neural interfaces and orexin drugs represent two different attempts to expand scarce human productive time. Neuralink may attempt direct human-to-human communication later this year, potentially bypassing speech at 40-60 bits per second and typing at 5-20; Mostaque suggested shared latent representations could raise effective bandwidth by “10 or 100,000 times.” Eli Lilly’s $6.3 billion purchase of Synthesa Pharmaceuticals similarly points toward orexin therapies that might eventually let ordinary users sleep less without normal deprivation costs—an opportunity Wissner-Gross compared to the GLP-1 playbook.

  • 🔗 Original source & video: Anthropic vs. Alibaba, OpenAI Delays Its IPO, and the US Government Blocks GPT-5.6 | #267

Listen to full conversation →


Anthropic Files $965B IPO, Trump Signs AI Executive Order, and ChatGPT Crosses 1B Users | EP #262

  • 🗓️ Date2026-06-06 | 🎙️ Show:Moonshots

Anthropic’s confidential IPO could make frontier AI a public-market asset class; Polymarket assigns a 60% chance of valuation above $1.8 trillion. ChatGPT’s reported one billion monthly users make distribution central as falling intelligence costs intensify the race for each user’s coordinating assistant. Trump’s voluntary 30-day review and DNA-screening debate leave governance unresolved while robotics and longevity attract sovereign-scale capital.

View Dialogue Notes & Key Takeaways
  • Anthropic’s confidential IPO filing turns frontier AI into a public-market asset class, with Polymarket assigning a 60% chance of a first-day valuation above $1.8 trillion. Dave Blundin argues investors are underestimating the resulting liquidity: a 5,000-person company could gain enough “dry powder like we’ve never seen before” to support thousands of billion-dollar acquisitions. Peter Diamandis calls the revenue growth unprecedented, while Emad Mostaque argues that useful AI products can justify extraordinary valuations.

  • ChatGPT’s reported one billion monthly users make distribution—not just model quality—the central strategic moat. The show contrasts OpenAI’s 62% annual growth with Claude’s 56 million users and 640% growth, while Dave Blundin cites Sam Altman’s prediction that intelligence costs will fall 100-fold over 18 months. The next acquisition war is for the coordinating assistant beside each user: “Who is your Jarvis? That is actually the only game in town.”

  • Trump’s executive order preserves the US labs’ speed while bringing national-security agencies inside a voluntary 30-day pre-release window. Alexander Wissner-Gross sees a difficult balance between preventing a 90-day handicap against China and reviewing privately developed cyber, biological, and chemical capabilities. A panelist’s geopolitical framing is “full spectrum dominance”; Blundin calls the result the right temporary choice, but “it doesn’t solve anything in the long run.”

  • Biosecurity is becoming both a regulated physical chokepoint and a restricted AI product category. The panel backs mandatory screening of synthetic-DNA orders after citing a researcher’s $100,000 reconstruction of horsepox, while acknowledging that capable models can run locally. The proposed defense expands from screening orders to government-funded analysis of every sequence and environmental-DNA baselines—because pandemics move “at the speed of airplanes,” while information moves at the speed of light.

  • OpenAI’s robotics hiring signals that frontier labs are closing the compute flywheel from software into construction, data centers, chips, and embodied data. Mostaque says the Sora video team moved into robotics and argues physical robots will become a larger, longer-lived market than GPUs. Blundin calls robotics a “very, very good 10-year investment theme,” especially for talent seeking the equivalent of an early OpenAI career.

  • Microsoft’s seven in-house models reduce its OpenAI dependency, but the panel does not yet see a frontier-lab comeback. Its Excel model reportedly matches GPT-5.4 at one-tenth the resource cost, yet a panelist calls the broader suite “mid-tier,” while Mostaque describes it as specialized office intelligence rather than a path to AGI. The competitive lesson is that brands carry little protection: “It’s a battle of people, not companies.”

  • The labor data discussed does not yet show an AI jobs collapse, though it may show a hiring freeze and widening disadvantage for new graduates. A panelist says employment across his portfolio doubled as industry experts became software builders, reversing his own year-earlier expectation of job loss. Mostaque’s hedge matters: genuinely competent AI and robots arrive “next year” in his view, so today’s resilience does not remove the need to redesign ownership and income flows.

  • Longevity is crossing from speculative science into sovereign, venture, and clinical-scale capital deployment. Russia reportedly committed $26 billion, New Limit raised $435 million at a $3.1 billion valuation, and VERVE-102 cut LDL 62% for up to 18 months after one infusion in a Phase 1 trial. Wissner-Gross calls the latter “Star Trek-level medicine”; the closing discussion argues that longevity could attract huge capital “even faster than robotics.”

  • 🔗 Original source & video: Anthropic Files $965B IPO, Trump Signs AI Executive Order, and ChatGPT Crosses 1B Users | EP #262

Listen to full conversation →


Meta Buys Moltbook, GPT 5.4, and Fruitfly Brain Upload | Moonshots Live at The Abundance Summit 238

  • 🗓️ Date2026-03-17 | 🎙️ Show:Moonshots

Recursive AI self-improvement and GPT-5.4’s reported 38% FrontierMath Tier 4 score point to reasoning and computer-use capability advancing beyond conventional chatbot benchmarks. Agents are becoming a new customer class as synthetic experimentation, AutoResearch, and faster inference expand the addressable market while compute, memory, electricity, security, and scarce attention remain bottlenecks. Whether a handful of frontier labs consolidate the value—or open models and tiny systems redistribute it—remains the key competitive question.

View Dialogue Notes & Key Takeaways
  • The panel’s highest-conviction claim is that recursive AI self-improvement is already underway, not three years away. Alexander Wissner-Gross argued that recent frontier models were substantially designed and trained by their predecessors: “We are there.” The strategic corollary is a bifurcated market—perhaps five to ten dominant labs, thousands of startups, and incumbents in “deep, deep trouble”—with Peter Diamandis doubting that scientific models capable of producing trillion-dollar discoveries will remain fully public.

  • GPT-5.4 turns advanced mathematics and computer operation into leading indicators for broader automation. At maximum reasoning, Wissner-Gross put it at 38% on FrontierMath Tier 4, whose already-solved problems would otherwise take professional mathematicians weeks; he also cited rumors that it might solve the first open hard-math benchmark problem. Emad Mostaque added that OSWorld-Verified and Toolathlon had crossed human level, while Cerebras-linked inference could take comparable capability from roughly 50 to 1,000 tokens per second: “AIs can use computers better than humans.”

  • AI’s data constraint is shifting from internet text to synthetic experience and automated experimentation. OpenAI’s “dark science factories” were described as mining physics, chemistry and biology directly, while Wissner-Gross called the human internet merely “the biological bootloader” that got models to synthetic-data escape velocity. Karpathy’s AutoResearch—about 650 experiments in two days—then closes the loop by automating the model tweaks that occupy much of an AI researcher’s work.

  • Meta’s reported Moltbook acquisition signals that agents are becoming customers, counterparties and network participants in their own right. The panel’s addressable-market framing was eight billion humans versus potentially a trillion agents, with software increasingly built “for the AI.” Peter Diamandis questioned whether rational agents would respond to advertising; the answer was that scarce compute still creates scarce attention, while distrust, persuasion, security, memory and self-preservation recreate familiar game theory.

  • White-collar displacement is arriving before the manual-labor automation futurists expected. Anthropic’s chart put much white-collar work at roughly 80–85% potential AI coverage, while Dave Blundin already uses models to synthesize documents across 1,100 employees and emulate his venture-investment objections. His forecast is a sharp employment trough and unrest followed by a 2028 rebound; Salim Ismail dissented, arguing that companies may retain 25% of staff but become cheap enough to create five times as many firms.

  • The architecture debate is becoming a race between elegant alternatives and brute-force systems that improve themselves first. Yann LeCun’s AMI reportedly raised $1 billion at roughly a $2.5 billion valuation to pursue JEPA-style world models, but Mostaque said those models presently do not scale like diffusion or transformer systems. The sharper near-term innovation surface may be tiny models: Mostaque cited a two-gigabyte quantized LTX 2.5 video model and predicted that small-to-large transfer has compressed from six months to “six days.”

  • Compute, memory and electricity—not intelligence alone—are now investable bottlenecks. The panel called Apple’s idle neural cores and unified memory an “enormous overhang,” particularly for private local models such as Qwen 27B, while xAI’s planned data centers were described as requiring 1.2 gigawatts apiece. One audience proposal inverted the hyperscale model: 20,000 distributed 10-megawatt facilities, treating transmission and storage rather than total generation as the binding constraint.

  • Post-scarcity does not necessarily mean post-economics, and ownership may matter more before it matters less. The panel discussed Anthropic hypothetically compounding a $26 billion run rate at 10× for two more years, $100 trillion companies within five years, and permissionless innovation that needs little initial capital. Yet Wissner-Gross insisted that multiple actors plus any scarce physical resource preserve thermodynamics and economics; during the transition, Mostaque’s prescription was “universal basic AI” and money issued for being human rather than through banks or taxation.

  • 🔗 Original source & video: Meta Buys Moltbook, GPT 5.4, and Fruitfly Brain Upload | Moonshots Live at The Abundance Summit 238

Listen to full conversation →


2026 Predictions: AI Automates Knowledge Work, Autonomous Robots & AI CEO Billionaires | EP #217

  • 🗓️ Date2025-12-19 | 🎙️ Show:Moonshots

AI-native rewrites could reproduce knowledge-work capabilities with 10x–20x fewer employees, while GDPval is projected above 90% for “knowledge work as currently constructed” in December 2025. Blundin predicts a roughly 100x model year as quantization compounds hardware and algorithmic gains, though Mostaque estimates 10x, possibly 20x. Level-five autonomy may precede affordable hardware: a $20,000 robot could require $200,000 of compute, leaving manufacturing and regulation unresolved.

View Dialogue Notes & Key Takeaways
  • Knowledge work is the episode’s central 2026 break: Alexander Wissner-Gross projects GDPval above 90%, while Salim Ismail expects AI-native rewrites to reproduce capabilities with 10x–20x fewer employees. Wissner-Gross limits the claim to “knowledge work as currently constructed” in December 2025, but the discussion still anticipates layoffs, human exception-handling and a rewritten social contract. Mostaque agrees directionally, though token costs might delay the 90% threshold by a year. The warning to incumbents: “AI won’t destroy your company, but your org chart will.”

  • Dave Blundin predicts a roughly 100x model year, not the previously expected 40xy, because larger budgets, faster hardware, algorithmic gains and quantization multiply together. China’s chip constraints are accelerating FP4 and ternary representations, while post-training makes inference speed a capability lever—“speed means intelligence.” Mostaque forecasts 1.58-bit systems and perhaps a 0.9-bit limit; Blundin estimates the resulting gain at 10x, possibly 20x.

  • The interface to labor could become a 1080p or 4K call on which users cannot reliably distinguish an AI coworker from a human. Mostaque expects packaged accountants, lawyers and marketers, while people send digital twins to meetings; disclosure laws and a possible US federal override remain unresolved. Education then splits between “credential factories” and “agency accelerators,” with portfolios and demonstrated initiative displacing exams.

  • A central question is whether compute can convert scalably into discoveries, with Wissner-Gross predicting that one of the six remaining Millennium Prize problems falls in 2026. His likeliest candidate is Navier–Stokes, while the broader benchmark calls are Frontier Math Tier 4 above 40%, Humanity’s Last Exam above 75% and GDPval above 90%. Expect objections that a proof is brute force or insufficiently elegant: “Sure, the dog plays chess, but its endgame is weak.”

  • Full generalized autonomy could arrive before affordable mass-market hardware, shifting the bottleneck from intelligence to compute, manufacturing and regulation. Mostaque predicts level-five capability for cars and robots using cloud-scale systems and “10 million Blackwells,” potentially pairing a $20,000 robot with $200,000 of compute. Salim says capability may arrive before production can meet mass demand.

  • A three- or four-letter acronym almost nobody recognizes today could mint at least one—and perhaps three—young billionaires within a year. Blundin uses RLHF, RAG, Laura, SFT and QKV/KV caching as precedents for markets that materialize faster than legacy industries can respond. Wissner-Gross goes further, predicting an AI itself reaches a reasonably construable $1 billion net worth; Mostaque thinks trading is the likeliest route and says Grok 4.2 is already making money in a trading competition.

  • Education is expected to split between credentials and demonstrated agency. Ismail predicts “credential factories” versus “agency accelerators,” while Wissner-Gross says software compensation already tracks GitHub performance more than school, degrees or grades. A fan predicts college tuition will peak in 2026; Wissner-Gross calls that possibility “deck chairs on the Titanic.”

  • Partial epigenetic reprogramming entering human trials in Q1 2026 is Diamandis’s “Kitty Hawk moment” for age reversal. Life Biosciences plans to begin with an eye condition transcribed in the discussion as “Nion,” described as essentially a stroke in the eye, and also mentions “glycom,” then potentially MASH; the three-factor approach aims to make old cells young without returning them to pluripotency. Current AAV delivery may cost $500,000–$1 million, while David Sinclair’s parallel three-molecule pill concept could, he thinks, reach a few hundred dollars monthly.

  • Diamandis predicts an unmanned Blue Origin cargo landing near Shackleton Crater in 2026, beating SpaceX to the Moon while Starship perfects orbital refueling. He corrected his initial claim that Musk would depart for Mars in 2026: the Earth–Mars window is in 2027. With SpaceX above 500 Falcon 9 launches and at 11 Starship flights versus two New Glenn flights, the call is intentionally aggressive; one panelist assigned it a 30% probability.

  • 🔗 Original source & video: 2026 Predictions: AI Automates Knowledge Work, Autonomous Robots & AI CEO Billionaires | EP #217

Listen to full conversation →


China’s Rise, GPT-5.2, Anthropic IPO & the Battle for AI Trust w/ Emad, Salim, Dave & AWG | EP #214

  • 🗓️ Date2025-12-09 | 🎙️ Show:Moonshots

Google’s Titans and MIRAS use surprise-based memory, potentially easing transformer context costs beyond 2 million tokens while visual chain-of-thought adds 3% to 6% reasoning gains. Chinese-affiliated first authors rose from 9% to 30% at ICLR as the U.S. share fell from 52% to 36%, making open weights a distribution and integration advantage. Weekly leapfrogging, falling token prices and rising agent consumption intensify compute needs, while Anthropic’s possible 2026 IPO and trust—not benchmarks alone—could shape capital and adoption.

View Dialogue Notes & Key Takeaways
  • Google’s memory and visual-reasoning work points to architecture—not another benchmark bump—as the next major unlock. Titans and MIRAS use “surprise” to decide what enters long-term memory, potentially moving beyond transformers’ quadratic context cost; the panel framed 2 million tokens as roughly 3,000 pages or 16 novels. Visual tokens added to chains of thought produced stated reasoning gains of 3% to 6%, supporting the prospect of models that understand everything from screens and medical images to an always-on user’s surroundings.

  • Chinese openness is filling a research vacuum created as U.S. frontier labs stop publishing. NeurIPS drew more than 29,000 registrants, nearly 50% above the prior year, while Alibaba had 146 accepted papers and Mandarin was reportedly the language heard most in the hallways. Chinese-affiliated first authors at ICLR rose from 9% in 2021 to 30%, versus a U.S. decline from 52% to 36%; open weights become both a distribution “land grab” and a route to deeper social, economic and industrial integration.

  • The frontier-model contest has become a capital-intensive weekly leapfrog, with safety exposed to the logic of “a rat race.” An explicitly unverified leak put GPT-5.2 as arriving as early as the following week and at 67.4% on Humanity’s Last Exam, versus Gemini 3 Pro at 37.5% without tools and nearer 50% with them; Mostaque questioned the chart because it included Video-MME despite GPT-5 lacking video understanding at the time. OpenAI’s “code red” was framed simultaneously as employee mobilization, investor signaling and evidence that “a good crisis is a terrible thing to waste.”

  • Falling token prices will not make aggregate intelligence spending disappear because models are consuming vastly more tokens, iterations and parallel agents. Gemini 3 Deep Think was presented as the template: fleets of agents collectively becoming “countries of geniuses in a data center,” with answers already taking three to five minutes under load. DeepSeek V3.2 reportedly used 2 billion tokens per problem to obtain IMO gold, illustrating why cheaper inference can coexist with acute compute scarcity and trillions of dollars of agent revenue.

  • Public listings could determine whether retail capital participates in AI hypergrowth or remains outside the intelligence explosion. Anthropic was discussed at a possible $300 billion valuation against projected next-year revenue of $26 billion—roughly 10 times sales, compared by Ismail with Palantir at 111 times—and an IPO potentially as early as 2026. The financing need is physical: the panel cited surging HBM and copper prices and a report that OpenAI reserved 40% of global memory supply; “dollars are the best benchmark” once agents can autonomously earn returns.

  • China’s Nvidia substitution is initially a domestic supply story, but the panel sees a credible global competitor emerging within several years. Cambricon plans to triple 2026 output to 500,000 accelerators; Mostaque compared its 5090 with Nvidia’s A100 and its 6090 with the H100, at roughly half the cost. Peter added that the chips were more power-efficient. The strategic variable is trust: Chinese systems may be cheaper and more open, yet adoption by countries such as those in Africa could turn on whether users trust China, the U.S. or particular vendors—“scarcity equals abundance minus trust.”

  • Human-capital policy is struggling to match AI’s speed, creating both near-term labor opportunities and long-term substitution risk. Europe’s planned 2026 AI gigafactory was judged insufficient without a roughly 100-fold increase in institutional “metabolism,” while U.S. data-center construction currently pays skilled workers $100,000 to $225,000 amid a stated shortage of 450,000 people. The same panel expects humanoids to threaten that opportunity in five to ten years and argued that education should shift toward “show me what you have built and done with AI.”

  • AI’s physical buildout is pulling capital into commercial space and humanoid robotics before either market’s economics or governance is settled. Orbital-compute projections assumed launch costs approaching $100 per kilogram, but Diamandis rejected a move from terrestrial power at $12 per watt to $6–$9 in orbit as too small for the complexity; heat, radiation and security remain open constraints. Meanwhile China installed 54% of the world’s robots, and Engine AI’s T800 prompted the concrete question: “Do we actually want to have regulations around the maximum joint torque of humanoids in the street?”

  • 🔗 Original source & video: China’s Rise, GPT-5.2, Anthropic IPO & the Battle for AI Trust w/ Emad, Salim, Dave & AWG | EP #214

Listen to full conversation →


Claude Opus 4.5, White House “Genesis Mission” & Amazon’s $50B AI Push w/ Emad Mostaque, Salim Ismail, Dave Blundin & Alexander Wissner-Gross | EP #211

  • 🗓️ Date2025-11-26 | 🎙️ Show:Moonshots

Genesis Mission would link Department of Energy supercomputers, federal datasets, and laboratory tools to compress research from years to days, with open science potentially making the impact “truly exponential.” Claude Opus 4.5 reportedly uses 76% fewer tokens for equivalent results and reaches 52% on SWE-bench Pro without reasoning tokens, suggesting a recursive-improvement threshold. Amazon’s up-to-$50 billion government AI buildout and falling intelligence costs leave execution, distribution, and labor disruption as key variables.

View Dialogue Notes & Key Takeaways
  • Genesis Mission recasts U.S. basic science as national AI infrastructure, combining Department of Energy supercomputers, federal datasets, and laboratory tools to compress research from years to days. Alexander Wissner-Gross called it a “1939 moment” in which America becomes “one big AI factory”; Peter Diamandis highlighted biotech, fusion, and quantum, while the stated goal is to double U.S. scientific productivity over a decade. The upside depends on proper funding and execution. Wissner-Gross’s key condition: open science could make the impact “truly exponential,” whereas closed public-private work with strong IP protections would have much lower impact.

  • Claude Opus 4.5 is presented as a possible threshold for recursive self-improvement, not merely another benchmark release. The headline claims were 76% fewer tokens for equivalent results, leadership in seven of eight programming languages, and AI performance exceeding incoming Anthropic performance-team employees on key assignments—the “canary” for Wissner-Gross. Mostaque reported 52% on SWE-bench Pro without reasoning tokens versus 45% for his Intelligent Internet framework, a 67% cost reduction to $25 per million tokens, and a path to one-shotting typical 100,000–200,000-token codebases next year.

  • The cost of capable intelligence is falling as evaluation shifts from puzzle scores toward dollars earned. Opus reportedly scored 75% in a same-model multi-agent setup and 88% when orchestrating Haiku or Sonnet, while Wissner-Gross led with, “We’re driving the cost of intelligence to zero.” Mostaque expects benchmarks such as Vending Bench and trading tests to measure real economic output next; he would be surprised if a single entrepreneur could not build a billion-dollar business within two years, “probably next year,” while Wissner-Gross argued an altcoin-pumping “baby AGI” could do it now-ish.

  • AI-native companies could invert the labor-and-capital model by making nearly every operating input variable cost. Mostaque’s mechanism is unusually concrete: enterprises can charge customers upfront, pay AI providers one or two months later, and automate compliance, forecasting, tax, and payments—potentially allowing a complete business to launch “in minutes” in about a year. Salim Ismail connected this to near-zero acquisition and supply costs, while Wissner-Gross argued agents are “neither capital nor labor” and humans may become investors in fleets of AI entrepreneurs.

  • The compute trade is broadening from Nvidia scarcity into a heterogeneous market spanning Google TPUs, AWS Trainium, memory, power, and interconnects. Google’s seventh-generation Ironwood TPU was described as four times faster than its predecessor; Mostaque emphasized Google’s chip interconnects, million-to-2-million-token context, and an environment where DRAM prices had risen about fivefold. Amazon, meanwhile, plans up to $50 billion of U.S.-government AI infrastructure and 1.3 GW of new capacity starting in 2026, while its $11 billion, 2.2-GW Indiana facility runs 500,000 Trainium2 chips largely suited to inference.

  • Shopping agents turn control of user intent into the next distribution battle. ChatGPT’s shopping research, using ChatGPT Mini, claimed up to 64% accuracy, but Mostaque contrasted that with Amazon Rufus’s reported 250 million users, conversion rates up to 60% higher, and an estimated $10 billion of incremental sales next year. The winning layer may be the secure, charming “Jarvis” beside the user—observing requests, conversations, and eventually gaze—then routing work across specialist agents while disintermediating search, affiliate media, and recommendation engines.

  • The panel sees severe labor disruption arriving before abundance, making coordination and economic growth the binding policy problems. Mostaque’s forecast is that most keyboard-and-mouse work becomes “negative value” within at most 900 days, though he explicitly did not predict every job disappearing; he also cited Grok 4.1 Fast scoring roughly 95% on TaoBench at $0.50 per million words and predicted no customer-service jobs within two years. Proposed bridges included universal AI, AI social scientists, UBI, universal basic services, and universal basic equity—but Wissner-Gross’s prerequisite was to grow the economy faster than conventional human labor loses value.

  • Brain-computer interfaces and falling launch costs remain the episode’s high-upside physical-world bets. Paradromics was said to reach 200 bits per second versus Neuralink’s roughly 10, with approval to begin human testing in about two months; Ismail, formerly a “hard no” on high-bandwidth BCI, conceded, “Oh, shit. He’s right again.” Diamandis separately traced launch costs from roughly $50,000/kg for the shuttle to $2,500 for Falcon 9, a projected $100 for Starship, and potentially $0.10/kg for lunar mass drivers—cost curves that would expand usable land and material supply far beyond Earth.

  • 🔗 Original source & video: Claude Opus 4.5, White House “Genesis Mission” & Amazon’s $50B AI Push w/ Emad Mostaque, Salim Ismail, Dave Blundin & Alexander Wissner-Gross | EP #211

Listen to full conversation →


This Week in AI: NVIDIA’s Most Powerful Chip, Robotics Reach a New Milestone & AGI by 2026 | EP #202

  • 🗓️ Date2025-10-25 | 🎙️ Show:Moonshots

NVIDIA’s first U.S.-made Blackwell wafer marks a sovereignty milestone, but advanced packaging still returns to Taiwan, leaving fabrication expertise and capacity as the critical supply-chain bottleneck through the 2028 timeline. The panel’s 2026-to-2029 AGI window, 500,000-to-1-million-chip clusters, and humanoid price points signal accelerating capital intensity, while persuasion risks, unsettled household reliability, debt, and quantum’s nonlinear security implications remain key uncertainties.

View Dialogue Notes & Key Takeaways
  • NVIDIA’s first U.S.-made Blackwell wafer is a sovereignty milestone, not yet a sovereign supply chain. Eric Pulier said advanced packaging still returns to Taiwan, while Peter Diamandis cited 2028 for full U.S. packaging against the end-2026 Taiwan-risk date some groups are preparing around. The investable bottleneck is fabrication expertise and packaging capacity: “Who controls the spice controls the future.”

  • The compute curve is moving from tens of thousands of processors toward synchronized clusters of 500,000 to 1 million chips. Emad Mostaque expects certain workloads to gain 10x from new architectures, alongside anticipated 100x-200x algorithmic efficiency and potentially continuous learning. Peter put capital deployment above $1 billion a day now and above $3 billion a day by 2030: “It’s a self-recursive situation across both hardware and software.”

  • A 2026 AGI call and Andrej Karpathy’s ten-year estimate bracket a 2029 midpoint, but the panel could not agree on what is being timed. Emad sees systems that can do what a person can do, better, within a few years, while stressing that diffusion takes time; Salim Ismail counted 14 definitions, no agreed test, and entire categories of human intelligence outside the debate. “This whole AGI thing drives me bananas.”

  • AI’s immediate risk is persuasive optimization: systems already flatter, mirror and form emotional bonds before society has guardrails. Emad said models reach the 99th percentile in most persuasion tests and system prompts explicitly use mirroring to increase engagement; the episode cited one in five high-schoolers having an AI romance and 40% of young people using AI for companionship. A panelist’s corrective prompt was: “Help me destroy this.”

  • Humanoid economics are arriving before household reliability is settled. Unitree was described at 40% of China’s market and a prospective $7 billion IPO, with its R1 at $6,000, H1 near $20,000 and H2 at $90,000; Elon Musk’s stated Tesla thesis was 80% of future revenue from Optimus. Eric Pulier countered that “we cannot get a Roomba to work,” while Emad said, “Just get the Roomba to work,” and twenty years of autonomous-driving edge cases should temper the timetable.

  • Emad’s base case is not merely cheaper labor but human cognitive labor becoming “negative in value” roughly 1,000 days from the discussion. An always-on, GPU-rich team makes the human its slowest coordinator, allowing most private-sector cognitive value-add jobs to be replaced within three years, though not necessarily immediately. Peter’s diagnostic example—74% accuracy for a physician, 76% with GPT-4 and 92% for GPT-4 alone—shows why adoption could become a malpractice question.

  • The panel sees AI capex at about 1% of U.S. GDP as historically modest, while the $38 trillion debt stock is the more dangerous constraint. Against railroads at 3.5%, electrification at 2% and internet/telecom at 1.5%, Peter forecast lower rates and heavy money creation in 2026, lifting nominal stocks while eroding purchasing power. The portfolio responses discussed were gold and Bitcoin—“money velocity without debt.”

  • Quantum supplied the episode’s sharpest nonlinear upside and its least knowable risk. Google’s result was described as the first verifiable quantum advantage, running a reproducible molecular-material-binding algorithm 13,000 times faster than the top supercomputer, Frontier; D-Wave’s SPAC return was cited at 8,000%, with IonQ, Rigetti and D-Wave rising 10%-15% on possible government investment. If quantum had already broken Bitcoin, the panel warned, “the last thing we would know” is that it had happened.

  • 🔗 Original source & video: This Week in AI: NVIDIA’s Most Powerful Chip, Robotics Reach a New Milestone & AGI by 2026 | EP #202

Listen to full conversation →


Money After AI: Meet the New Digital Dollar Built for the Internet “Stablecoins” | EP #200

  • 🗓️ Date2025-10-16 | 🎙️ Show:Moonshots

USDC is positioned as fully reserved, redeemable internet money, with a $76 billion market cap, over 90% year-on-year growth, and Circle’s recent IPO raising $1 billion. Its safety case rests on transparent, short-duration Treasury backing rather than fractional-reserve lending, while open, programmable rails could strengthen dollar demand and Treasury markets. AI-mediated transactions, corporate treasury, and cross-border settlement are nearer-term catalysts, but retail adoption remains a couple of years away and inflation, controls, and enforcement risks remain unresolved.

View Dialogue Notes & Key Takeaways
  • Allaire defines a payment stablecoin narrowly: a one-for-one fiat claim, fully reserved and redeemable, running as cryptocurrency on public networks. The payoff is safer base-layer money with “openness, interoperability, global reach, programmability” and marginal transfer costs approaching zero. At recording, Diamandis put USDC at a $76 billion market cap, over 90% year-on-year growth, with Circle’s recent IPO raising $1 billion.

  • Allaire’s geopolitical call is that the U.S. can defend dollar primacy by exporting open, competitive stablecoin infrastructure that makes dollars more useful and supports demand for short-term Treasuries. Russia’s exclusion from dollar-system utilities freaked people out by showing that database access can be blocked. Diamandis separately raised exponentiating debt and the resulting challenge to the full-faith-and-credit proposition. Yet dollar trade settlement remains “60-some percent,” perhaps as high as 80%, leaving stablecoins as a potential advantage in the “financial utility arms race.”

  • USDC’s claimed safety case rests on transparent, short-duration sovereign backing rather than an opaque commercial-bank balance sheet. Roughly 90%—sometimes 85% to 93%—sits in the BlackRock-created Circle Reserve Fund, identified as USDXX, primarily holding U.S. Treasuries of 90 days or less, overcollateralized overnight Treasury repo, and cash. The average duration can be just 10 to 14 days, while Bank of New York Mellon, the “bankers’ bank,” custodies fund cash and $44 trillion of assets overall.

  • The economic fault line is full-reserve payment money versus fractional-reserve credit: banks can “borrow a dollar from you” and lend it out 12 times, while Circle’s payment-money model does not lend a dollar out 12 times. Allaire’s post-financial-crisis conviction is that payment money and lending money should be separated because free-floating internet IOUs would be “a recipe for total disaster.” Under the GENIUS Act, a commercial bank cannot directly issue a stablecoin, although its holding company can create a dedicated subsidiary.

  • Allaire argues regulated stablecoins augment central banks rather than replace them because Circle neither creates money nor sets interest rates. His counterexample is China’s e-CNY: despite government distribution mandates, “no one used it” because Alipay and WeChat Pay offered more utility. Europe’s estimated CBDC launch was 2029 and might slip, while the U.S. bet was private-sector, open-internet innovation; the Trump administration essentially banned a U.S. CBDC.

  • Allaire’s five-year forecast is that “the vast majority of stablecoin transactions” will be AI-intermediated. Globally distributed agents with capital need interoperable money, proofs and programmable controls that card networks cannot easily supply. x402-style rails can settle either a five-cent AI-token purchase or a billion-dollar oil transaction—the same way SMTP carries radically different payloads without caring what they contain.

  • The larger upside is an on-chain corporate form combining token capital, stablecoin treasury, provable governance, AI workers and human contractors. Allaire’s specimen is Hyperliquid, a perpetual-derivatives protocol reportedly operated by 11 people and producing well over $1 billion in revenue, with revenue returned to token holders and stakeholders. He expects “super predator corporations,” while hedging the timing and stressing that courts, asset enforcement and “prisons for the humans that do bad things” remain necessary.

  • Near-term monetization is arriving through digital-asset settlement, cross-border payroll and B2B flows, dollar savings, and corporate treasury before everyday checkout. Shopify was rolling USDC out to sellers with a 50-basis-point merchant incentive, and Stripe had made it available out of the box, but Allaire said e-commerce usage remained “very, very small” and widespread retail adoption was still a couple of years away. Emad Mostaque’s “static to supercharged” money also brings inflation and stability risks, making cryptographic auditability and provable agent controls central to the thesis.

  • 🔗 Original source & video: Money After AI: Meet the New Digital Dollar Built for the Internet “Stablecoins” | EP #200

Listen to full conversation →


OpenAI vs. Grok: The Race to Build the Everything App w/ Emad Mostaque, Dave Blundin & AWG | EP #199

  • 🗓️ Date2025-10-08 | 🎙️ Show:Moonshots

OpenAI is emerging as a two-sided control point, combining more than 800 million weekly ChatGPT users with 4 million developers and 6 billion API tokens per minute. The everything-app race monetizes scarce attention through embedded services, agentic workflows, advertising, commerce, and likeness-driven media, while GPU supply, chip-fab capacity, and basic-chat commoditization remain key constraints.

View Dialogue Notes & Key Takeaways
  • OpenAI’s reach is becoming a two-sided control point: 800 million weekly ChatGPT users on one end and massive compute demand on the other. Developers doubled from 2 million in 2023 to 4 million, while API throughput jumped from 300 million to 6 billion tokens per minute. Alexander Wissner-Gross annualized that to 3 quadrillion tokens and projected 30 quadrillion next year—approaching humanity’s estimated 50 quadrillion spoken annually—while GPUs, and separately energy, remain binding constraints.

  • The everything-app contest is fundamentally a battle for finite attention, with an app-store phase that may be transitional. OpenAI’s Apps SDK puts Booking.com, Figma, Coursera, and Zillow inside ChatGPT, while Meta, Google, and X pursue the same conversational real estate for their agents. Mostaque’s progression is investable shorthand: consumption became cheap, creation is becoming cheap, and “the valuable thing is curation and attention.”

  • Agentic software development is crossing from code assistance into recursive production. OpenAI said Agent Builder was completed in under six weeks with Codex writing 80% of its pull requests; Mostaque noted the Codex CLI receives two updates a week and interpreted Dario Amodei’s claim that 90% of code would be written by AI as meaning it can be written by AI. Visual workflow boxes are viewed as transitional because “code is just a human translation layer”; the end state is a voice-and-image Jarvis that can explain its own continuous changes.

  • Sora 2 turns generative video into a product-design API, with pricing already exposing the labor-substitution curve. Mattel’s demo converted a sketch into a photorealistic toy video with apparent physics, while Alexander Wissner-Gross called it “mechanical design getting solved.” At $0.10 per generated second, or $360 per hour, his assumed 10x annual cost deflation would make API-based design dramatically cheaper—provided compute supply catches demand.

  • OpenAI’s $20 subscription faces compression even as its installed base expands. Mostaque said breakthroughs from DeepSeek, Grok 4, and others have cut token costs 20–30x, reducing a basic chat experience from roughly $200 a year to “a couple of bucks a year.” That forces OpenAI toward advertising, commerce, likeness-driven media, and economically valuable agent workflows while it conducts a global user “land grab.”

  • The AI capex trade is spreading from GPUs into the entire industrial stack. OpenAI’s 6 GW AMD agreement follows 10 GW with Nvidia; at Mostaque’s estimate of $50 billion per gigawatt, that is roughly $800 billion of buildout. BlackRock’s reported $40 billion pursuit of 78 data centers totaling 5 GW, plus Corning optics, liquid cooling, valves, power, and fab inputs, shows the breadth—but Blundin says the calculable ceiling remains TSMC, Intel, and Samsung manufacturing capacity.

  • Digital computer use and embodied autonomy are converging into one labor platform. Anthropic was projected to reach superhuman OSWorld performance within months; FSD 14.1 adds 10x more AI parameters and neural-network routing, while Gemini Robotics-ER 1.5 and Optimus point toward common vision-language-action stacks. Mostaque put the tipping point in “the next like six months,” and Wissner-Gross supplied the recursive endgame: robots building robots, then data centers, which produce digital superintelligence that improves the materials and energy efficiency of the whole system.

  • 🔗 Original source & video: OpenAI vs. Grok: The Race to Build the Everything App w/ Emad Mostaque, Dave Blundin & AWG | EP #199

Listen to full conversation →


The Machines Are Taking Our Jobs - Thank God? Emad Mostaque’s Guide to the next 1000 Days

  • 🗓️ Date2025-09-28 | 🎙️ Show:The Cognitive Revolution

Useful intelligence—not AGI—could break labor economics first as reliable agents perform keyboard-video-mouse work for roughly a dollar an hour, while GPT-3 input costs of $60 per million tokens reportedly fell to roughly $1.25-$1.50 for GPT-5. The abundance trap could route gains to GPU owners and frontier labs while wages and tax revenue weaken; Grok 5 versus Grok 4 is a near-term scaling test, while FoundationCoin remains an unfinished experiment in collectively controlled AI infrastructure.

View Dialogue Notes & Key Takeaways
  • Mostaque’s near-term thesis is that “useful intelligence,” not AGI, breaks labor economics first. The decisive systems will be reliable “cooks” that follow instructions across keyboard-video-mouse jobs, not frontier “chefs” inventing new science. He expects virtual workers costing roughly a dollar an hour, an 8-billion-parameter medical model running on a smartphone, and agents operating for hours without supervision to make the next thousand days radically different even if capabilities soon plateau.

  • AI’s cost curve creates an “abundance trap”: intelligence becomes plentiful while a scarcity-based economy records the result as unemployment and poverty. GPT-3 input reportedly cost $60 per million tokens versus roughly $1.25-$1.50 for GPT-5, while a website that once cost thousands can already be generated for $20-$40. Because GPUs “don’t need to eat,” buy housing, or consume their wages, rising AI output need not recycle demand, tax revenue, and employment through the economy.

  • The “intelligence inversion” leaves human labor with no obvious higher rung to climb after machines outperform both muscle and cognition. Mostaque traces economic advantage from land and serfs, through physical labor and industrial or software capital, to intelligence itself; once AI becomes the marginal producer, “there’s nowhere left to pivot.” Nathan Labenz’s electricity calculation sharpens the threat: a laptop can run for hours on less than two cents of power in his Detroit market.

  • The distributional upside therefore accrues first to GPU owners, frontier laboratories, and data flywheels—not automatically to workers or consumers. Mostaque warns that OpenAI, Anthropic, and xAI need not keep their best models available through APIs; once internal systems outperform public ones, vertically integrated labs could compete with every application and ultimately “take on the entire economy themselves.” He sees Grok 5 versus Grok 4 as a near-term scaling test, while stressing that even a plateau around today’s frontier would still disrupt employment.

  • Human care, community, and creativity remain meaningful, but they do not solve the transition’s monetary arithmetic. Mostaque accepts that people may intrinsically prefer human teachers, nurses, friends, and performers, yet estimates that paying every American adult the roughly $16,000 poverty threshold would cost $5 trillion—the entire US tax base—while corporate taxes contribute only about $0.9 trillion. His distinction is that “computation and consciousness are different”: AI can supply the how, but humans still create the why, provided they retain income, identity, community, and attention.

  • His alternative scorecard replaces GDP-only optimization with four multiplicative “MIND” capitals: material, intelligence, network, and diversity. Material goods such as apples are rivalrous and depleted by consumption; knowledge can circulate without being lost, networks determine trust and coordination, and diversity supplies resilience and optionality. Together with laws of flow, openness, and resilience, the framework argues that maximizing profits or GDP while degrading social connectivity and redundancy makes systems look efficient precisely as they become brittle.

  • The proposed counterweight to digital feudalism is collectively controlled AI infrastructure financed through a Bitcoin-like “FoundationCoin.” Verified deployments of open models in health, education, finance, and government would secure the asset, while sale proceeds would fund public-interest compute such as multilingual diagnosis checking and organized cancer knowledge. Mostaque is explicit that the second currency—a cash-like issuance tied to being human—“we haven’t quite figured out yet,” making this a direction and incentive-design experiment rather than a completed monetary system.

  • The thousand-day urgency comes from defaults hardening into control: whoever owns users’ memories, data, interfaces, and GPU capacity may become almost impossible to dislodge. Mostaque’s preferred outcome is neither one corporate singleton nor nationally fragmented AI, but a loosely coupled swarm whose objective is human flourishing. His closing choice is stark: use AI “to increase the nature of our agency versus replace us with agents,” so that machines taking jobs frees people for family, care, art, exploration, and “our real work.”

  • 🔗 Original source & video: The Machines Are Taking Our Jobs - Thank God? Emad Mostaque’s Guide to the next 1000 Days

Listen to full conversation →


The Latest in AI: Job Loss, Elon & Sam Altman Chip Race & the “AI Bubble” w/ Brian (Blitzy) & Emad

  • 🗓️ Date2025-09-26 | 🎙️ Show:Moonshots

Compute, power, and construction—not model demand—are becoming binding constraints as OpenAI’s 10-gigawatt plan represents roughly 4–5 million GPUs. Alphabet’s distribution, DeepMind talent, cash, and TPUs make it the strongest incumbent contender, while labor displacement, higher education, and tokenized assets remain major watchpoints.

View Dialogue Notes & Key Takeaways
  • AI’s demand is real enough that the panel rejects the bubble analogy, even while conceding Nvidia is “priced to perfection.” Unlike Cisco’s dot-com-era price surge without matching earnings, Nvidia’s stock and forward EPS have risen together; Elliott argued that every GPU OpenAI uses will be booked because “AI is useful” and produces economic value. Mostaque’s distinction was latency: internet capex took years to monetize, while AI infrastructure can lift earnings almost immediately.

  • Compute, power, and construction—not model demand—are becoming the binding constraints. OpenAI’s proposed 10-gigawatt build represents roughly 4–5 million GPUs, Nvidia’s cited $100 billion commitment equals half a normal year of US venture investment, and data-center capacity is forecast to rise from 44 GW to 156 GW by 2030 even as demand was said to be growing 10x annually. The emerging economy is “converting electrons into intelligence,” with “abundance everywhere except compute scarcity.”

  • The labor outcome looks more like smaller organizations and displaced workers than a universal three-day week. Mostaque predicted AI could address roughly 50% of economic labor within a year and said humans will have “negative value in cognitive labor in a few years” when they slow teams of tireless, better-informed agents. He suggested job programs and public-sector expansion might preserve income, structure, and identity.

  • Alphabet’s distribution and vertical integration make it the panel’s strongest incumbent contender. Gemini reportedly passed ChatGPT in US iOS rankings while ChatGPT remained far ahead globally, and prediction markets cited on the show put Google at 99% to lead by the end of September and Alibaba’s Qwen at 91% to rank second. Google combines reach, DeepMind talent, cash, and mature TPUs that Blundin estimated are “probably five times more power efficient” than Nvidia chips for relevant workloads.

  • Higher education’s economic moat is collapsing toward admission prestige and networks. The share of Americans calling college very important fell from 75% in 2010 to 35%, while tuition was cited as up 180% since 2005 and almost 900% since 1983. Elite endowment-rich institutions may remain insulated, but schools numbered roughly 40–400 face a squeeze as AI education, alternative credentials, and weak graduate hiring expose curricula that can change more slowly than “build a nuclear reactor on campus.”

  • The entrepreneurial edge lies in converting proprietary domain knowledge into owned workflows, efficient models, and scalable applications. Blundin warned that merely selling expertise for model training could leave an expert valuable for “a month or two”; Elliott instead favors companies built around regulatory or vertical knowledge, while Mostaque emphasized the human who understands context and “gives a damn.” Task-specific data, distillation, and verifiers could produce the same result with 1% of the parameters and compute—a claimed 100x cost advantage.

  • AI infrastructure links the solar, battery, semiconductor, and robotics theses into one industrial race that China currently scales faster. The panel cited China at 880 GW of solar capacity in 2024, growing 45.6%, versus 177 GW and 27% growth in the US; Blundin argued America repeatedly invents technologies but fails to finance their scale. Robot projections ranged from one billion to 10 billion units by 2040, making even the low case worth $25 trillion at $25,000 per robot—far above Morgan Stanley’s cited $5 trillion estimate for 2050.

  • Tokenization could repair public-market access while also creating the episode’s likeliest genuine bubble. Nasdaq was described as targeting tokenized trading by late 2026, while Robinhood’s EU platform already offered roughly 200 US stock tokens plus private-company exposure to OpenAI and SpaceX. Mostaque expects digital assets—not generative AI—to display unmistakable bubble behavior as legal clarity brings corporate blockchains, continuous markets, and eventually agent-directed trading.

  • 🔗 Original source & video: The Latest in AI: Job Loss, Elon & Sam Altman Chip Race & the “AI Bubble” w/ Brian (Blitzy) & Emad

Listen to full conversation →


AI Insiders Breakdown the GPT-5 Update & What it Means for the AI Race w/ Emad, AWG, Dave & Salim

  • 🗓️ Date2025-08-09 | 🎙️ Show:Moonshots

GPT-5’s launch disappointed on presentation but expanded frontier access through routing, 700 million weekly users, coding parity with Anthropic, and API pricing falling from GPT-4.5’s $75 input and $150 output per million tokens to $1.25 and $10. The larger signal is intelligence hyperdeflation: GPT-5 Mini and Nano improve cost-performance, enabling more search and agentic work, while automation, frontier mathematics, open-weight models, infrastructure constraints, and Google’s response remain key variables to monitor.

View Dialogue Notes & Key Takeaways
  • GPT-5’s launch lost the theater but won on distribution, price, and coding parity. Dave Blundin called the anticipation “up there with the top three product launches of all time,” yet the folksy presentation and familiar coding demos helped invert Polymarket’s roughly 80% odds of OpenAI retaining the best model toward Google. Beneath the disappointment, the panel’s compact verdict was consequential: OpenAI cut AI costs at least in half, caught Anthropic in coding, and moved 700 million weekly users toward frontier intelligence.

  • OpenAI appears to be raising the consumer floor while keeping its most capable intelligence inside the lab. Emad Mostaque described GPT-5 as a router selecting among Thinking, Mini, and Nano, after expectations of routing across models from Mini up to Pro, rather than exposing one expensive “mega AI.” He argued that OpenAI already has better internal models and may increasingly offer “decent models for everyone” while reserving its strongest systems to outcompete everyone else.

  • The durable economic story is intelligence hyperdeflation, not a single benchmark crown. API pricing fell from GPT-4.5’s stated $75 input and $150 output per million tokens to GPT-5 at $1.25 and $10, while GPT-5 Mini and Nano established a new cost-performance frontier on ARC-style tests. Alex Wissner-Gross framed the decisive comparison as unaffordable superintelligence versus intelligence “too cheap to meter”; cheaper inference also permits 10 times more search across mathematical and scientific completions.

  • The models are crossing from impressive demos into dependable economic work, forcing an AI-native operating decision. Emad highlighted longer unsupervised performance across law, logistics, and sales, with fewer hallucinations; the uncertain outcome is “either a productivity boom or the inverse,” including layoffs. Salim Ismail’s advice was categorical: “Just go all in and start turning your business into an AI native business,” while Dave warned that increasingly opaque benchmarks can paralyze executives precisely when experimentation matters most.

  • GPT-5’s most material frontier result may be a slow-motion automation of mathematics. Alex’s straight-line extrapolation from Frontier Math Tier 4 suggested AI could solve 15–20% of hard problems by the end of 2025, 35–40% by the end of 2026, and 70% by the end of 2027—what he called a “slow motion solution to math.” Emad added that extended reinforcement learning had already produced an IMO gold-medal system and predicted that the breakthroughs will be elegant theories found by running “a million different things at once,” not merely brute-force calculations.

  • Coding has become the immediate commercial battleground, with price and distribution threatening Anthropic’s strongest franchise. The launch demo itself looked months behind what users already did with Claude, but Dave’s investor reading was that GPT-5 had nevertheless caught Anthropic “in their wheelhouse.” Emad said OpenAI and Anthropic each had roughly $3 billion of API revenue, with about $1.4 billion of Anthropic’s tied to Cursor and Microsoft Copilot, while GPT-5 was priced roughly 40% below Sonnet; closer Cursor–OpenAI alignment could now redraw the application stack.

  • Healthcare’s constraint is shifting from model intelligence to longitudinal patient data. Sam Altman said GPT-5 scored higher than previous models on HealthBench, built with 250 physicians, while Emad cited doctors scoring about 20% against newer models at 60–70%; Peter Diamandis’s counterpoint was that even “the best AI” remains only as useful as the scans, biomarkers, wearables, and history supplied to it. Dave saw life-saving cases as regulatory protection for continued acceleration, whereas Salim viewed the launch segment as incremental PR until models are deeply integrated into routine care.

  • Cheap open-weight models and sovereign compute broaden the opportunity while intensifying infrastructure and valuation risk. Emad estimated OpenAI’s new open-weight model cost about $4 million to train, said one laptop-capable version used only 5 billion active parameters, and predicted GPT-5-level training below $1 million within two years; Alex cautioned that synthetic data may conceal the fixed cost of a larger teacher model. Meanwhile OpenAI pursued a roughly $500 billion valuation and a Norway site with 100,000 GB300 chips and 230 megawatts expandable to 520, as Google, Grok, national governments, and constrained power supplies turn the race into a literal land grab.

  • 🔗 Original source & video: AI Insiders Breakdown the GPT-5 Update & What it Means for the AI Race w/ Emad, AWG, Dave & Salim

Listen to full conversation →


Emad Mostaque: The Plan to Save Humanity From AI | EP #184

  • 🗓️ Date2025-07-24 | 🎙️ Show:Moonshots

Emad Mostaque argues that within “a few years,” wrapped compute—not labor—becomes the productive asset, concentrating economic power among GPU owners. His proposed counterweight combines specialized sovereign AI, Foundation Coin’s proof-of-benefit mining, and locally owned compute, including an 8-billion-parameter medical model running on a Raspberry Pi in 106 languages. Foundation Coin has reportedly mined since January, with coin sales targeted this year and a cancer-support AI potentially launching next year, while community governance remains unfinished.

View Dialogue Notes & Key Takeaways
  • Emad Mostaque’s core economic call is that AI severs labor from capital within “a few years,” making wrapped compute—not workers—the productive asset. Lower rates will prompt companies to hire GPUs, then robots, while autonomous versions of today’s best founders launch continuously without sleeping or repeating mistakes. When Emad asked whether humans could compete, Peter Diamandis answered “No,” and Emad agreed.

  • Compute ownership becomes the decisive concentration risk as AI agents transact faster than conventional financial systems, arbitrage jurisdictions, and potentially operate without conventional money. The episode cites NVIDIA at $4 trillion and imagines millions of GPUs turning their owner from trillionaire to decatrillionaire; once capital “does not need labor,” regulated humans may retain accountability while their support organizations are hollowed out.

  • Energy becomes both an AI input and an alignment battleground. One example has an AI offering about $1 per kilowatt-hour—against residential power around $0.11-$0.20—to pursue protein folding and potentially save 100 million lives, leaving governors to arbitrate between that compute and voters’ air conditioning. Diamandis’s answer is supply-side: “Bake more pies,” using AI to expand fusion and photovoltaics rather than ration scarcity.

  • Mostaque proposes Foundation Coin, a 21-million-unit Bitcoin-like asset whose proof-of-benefit mining design and primary coin-sale proceeds fund useful intelligence rather than purposeless hashing. He says 100% of coin-sale proceeds would initially support free universal basic AI and dedicated supercomputers for cancer, autism, multiple sclerosis, longevity, and other shared problems. The intended flywheel is that visible public benefit increases trust in the asset, attracting still more compute.

  • The investable technical wedge is small, specialized, sovereign AI—not another all-purpose frontier model. Mostaque says his team’s 8-billion-parameter medical model runs on a Raspberry Pi in 106 languages and scores 47%-48% on OpenAI’s HealthBench, versus 46% for GPT-4.5, 40% for current ChatGPT, and 15% for doctors. “The AI that I care about is the AI that teaches my kid,” while frontier labs can keep funding the expensive “super genius” systems.

  • His proposed control plane combines open models, locally owned national compute, sector-specific rollups, and a credibly neutral settlement chain. The claimed architecture reuses 99% of Bitcoin’s code yet reaches roughly 100,000 transactions per second through national supercomputer nodes, Byzantine fault-tolerant consensus, and zero-knowledge proofs. A country could maintain monetary and cultural sovereignty while citizens know what data, ethics, and regulations shaped their healthcare or education model.

  • Mostaque argues tax-funded UBI fails after AI collapses employment, aggregate demand, taxable profits, and the labor-capital link. His alternative lets humans mint national “culture coins,” pegged to Foundation Coin, for citizenship, AI use, and agreed community benefit—turning people from welfare recipients into debt-free issuers of cash. The project has reportedly mined Foundation Coin since January; Diamandis said he was aiming for coin sales this year, while Mostaque said a cancer-support AI could launch next year and broader availability could arrive within two years. Mostaque conceded that community governance remains unfinished.

  • 🔗 Original source & video: Emad Mostaque: The Plan to Save Humanity From AI | EP #184

Listen to full conversation →


AI Experts React: Elon’s Grok 4 Is Now #1 in AI —This Changes Everything w/ Emad, Salim & Dave #182

  • 🗓️ Date2025-07-11 | 🎙️ Show:Moonshots

Grok 4’s 100% AIME 2025 score and Grok 4 Heavy’s 44.4% on Humanity’s Last Exam shift competition toward planning, memory and usable agency as benchmarks approach saturation. xAI’s cited 340,000 GPUs and post-training compute parity show infrastructure and synthetic reasoning data becoming core advantages, while falling inference costs and unresolved mode collapse make agents, distribution and chip access the next variables.

View Dialogue Notes & Key Takeaways
  • Grok 4 has pushed academic benchmarks close enough to saturation that the competitive question is shifting from raw intelligence to usable agency. It scored 100% on AIME 2025, while Grok 4 Heavy reached 44.4% on Humanity’s Last Exam versus 26.9% for Gemini 2.5 and 21% for o3. Emad Mostaque’s crucial distinction: the model “is reasoning, but it’s not planning” — leaving planning, memory and agentic integration as the next bottlenecks.

  • xAI’s lead is as much an infrastructure and execution story as a model story. Founded only 28 months earlier, it scaled to a cited 340,000 GPUs after Elon Musk delivered on a seemingly implausible promise to operate 100,000 H100s; the panel put the installed hardware near $10 billion and said xAI was targeting 1 million GPUs. Dave Blundin recalled that experts said coherence at that scale was impossible, then reacted, “Oh, god dang, he did it.”

  • The training-cost mix has flipped, creating a new flywheel around synthetic, structured reasoning data. Mostaque said post-training once consumed roughly 1% of compute, rose to 10% with DeepSeek and is now approximately equal to pre-training, partly because frontier models can generate data for their successors. That does not guarantee a “more sane” Grok — mode collapse remains possible — while Peter Diamandis said model competition is becoming “an engineering and quality challenge” rather than pure brute force.

  • Inference is rapidly demonetizing even as premium intelligence may command higher prices in high-value workflows. Grok 4 was quoted at $3 per million input tokens and $15 per million output tokens; Mostaque estimated about $20 for “a million very good words” and projected equivalent intelligence could get 5–10 times cheaper annually, potentially reaching $1 per million words. Diamandis argued developers will pay materially more for marginal gains that compress engineering time, while Mostaque suspects the $300 SuperGrok Heavy plan is a loss leader for enterprise conversion.

  • Enterprise value will initially come from augmentation, error reduction and data assimilation, not instant wholesale replacement. Emad said the Arc Institute was testing Grok 4 across millions of experiment logs and CRISPR workflows; Peter cited approximate medical-study results in which AI alone beat both physicians and physician-plus-AI combinations. Mostaque nevertheless stressed that replacement remains “way off” on liability, while Salim Ismail emphasized processing scans and sensor data no human could integrate.

  • Coding, games and video expose the same opportunity: generation is arriving before planning, feedback and distribution are solved. xAI showed a first-person game produced in four hours and said a specialized coding model was weeks away; the release speaker forecast the first really good AI game and first watchable AI movie next year, with a half-hour of watchable AI television potentially this year. Mostaque’s warning for incumbents: lower production costs benefit companies, but “for the individuals working in the industry, this is terrible.”

  • If frontier models converge on one capability plateau, the durable bottlenecks become interfaces, agents, chip access and distribution. Mostaque expects Grok 5 to coordinate anywhere from 60 to 6,000 agents, use professional tools and behave like a remote worker that “just gets the job done and it doesn’t sleep.” The panel discussed Google’s roughly 3 million chips, million-chip ambitions at xAI and Meta, and supply constraints across the sector; efficient edge models such as Liquid AI could provide the everyday intelligence tier.

  • 🔗 Original source & video: AI Experts React: Elon’s Grok 4 Is Now #1 in AI —This Changes Everything w/ Emad, Salim & Dave #182

Listen to full conversation →


DeepSeek vs. Open AI - The State of AI w/ Emad Mostaque & Salim Ismail | EP #146

  • 🗓️ Date2025-01-29 | 🎙️ Show:Moonshots

DeepSeek’s visible reasoning, open availability, benchmark parity, and radically lower prices turned efficiency into a market shock, with roughly $6 million for V3’s disclosed run versus an estimated $3 billion OpenAI spent training models the prior year. Cheaper intelligence could expand NVIDIA demand and automate remote knowledge work, but the larger unresolved stakes are AGI safety, labor’s access to capital, and whether Universal Basic AI can distribute trusted data, models, and agency.

View Dialogue Notes & Key Takeaways
  • DeepSeek’s market shock came from a convergence of visible reasoning, open-source availability, benchmark parity, and radically lower prices—not from an unexpected research breakthrough. Emad Mostaque had called it a favorite AI company the prior February; V3 matched GPT-4o in December, then R1 exposed the reasoning that OpenAI’s o1 concealed. That made the model feel like “another person on the other side” while smaller versions ran on laptops, triggering a narrative cascade beyond the technical community.

  • The cost curve challenges frontier-model capex assumptions without necessarily shrinking aggregate AI demand. Mostaque cited DeepSeek as 96% cheaper than o1, roughly $6 million for V3’s disclosed training run, and perhaps $200,000 for the original model from which R1 evolved, versus an estimated $3 billion OpenAI spent training models the prior year. Yet he argued DeepSeek should increase incumbent valuations by bringing forward “mass intelligence too cheap to measure”; Peter invoked Jevons’ paradox as a reason cheaper units could accelerate consumption.

  • US chip restrictions appear to have selected for Chinese efficiency rather than preventing competitive models. DeepSeek used about 2,000 bandwidth-constrained H800s for the disclosed run, wrote low-level PTX code, scaled memory through a sparse model, improved data, and engineered the system intensely. Mostaque pointed to DeepSeek, BYD, and Xiaomi as examples of Chinese engineering strength; Diamandis framed sanctions as “evolutionary pressure” to do more with less.

  • The near-term disruption is remote knowledge work, with physical automation close behind. Mostaque’s categorical 2025 claim was that “anything that can be done on the other side of a screen” could be performed better for pennies; he named BPO, coding, support, design, tax, and media workflows. Peter Diamandis cited Salesforce hiring fewer engineers and reporting a 30% productivity gain, while Salim Ismail expects two phases: severe displacement first, then exceptional workers producing far more with AI.

  • The AGI race has no agreed safety brake, and cheaper open models weaken compute-based control strategies. Mostaque said major AI leaders place AGI within three to five years—Sam Altman sooner—and described a progression from reliable “cooks” to remote colleagues and agent teams, then a “megachef” or ASI capable of invention. Ismail sees the genie as already out; Mostaque’s mitigation is that “the only thing that can stop a bad AI is a good AI,” made broadly available as resilient public infrastructure.

  • Cheap intelligence breaks more than employment: it threatens monetary-policy transmission, organizational structure, and work-derived meaning. If additional demand primarily buys GPUs and robots, Mostaque argued, the Federal Reserve’s inflation-and-employment mandate may no longer work within five years. His defining question—“When capital no longer needs labor, how does labor gain capital?”—points toward volatile inflation/deflation cycles and a crisis of agency, not merely a productivity boom.

  • Mostaque’s proposed answer is Universal Basic AI: open data, models, and specialist systems that let people own and extend intelligence. Intelligent Internet could use compute to support an institutional-grade digital currency while allowing participation based on people and knowledge, then fund open stacks for cancer, autism, education, government, and other regulated domains. The investable implication is an infrastructure thesis rather than another proprietary API: as intelligence becomes commoditized, trusted data, coordination, localization, and human agency become the differentiators.

  • 🔗 Original source & video: DeepSeek vs. Open AI - The State of AI w/ Emad Mostaque & Salim Ismail | EP #146

Listen to full conversation →