
Joey Brookhart
Frontier Insights
Core Frontier Thesis: Compute economics are shifting from raw infrastructure rental to high-margin token monetization. While Vera Rubin and AMD’s MI455X center hardware supremacy on memory bandwidth, hyperscalers are aggressively peeling away inference workloads via custom silicon (Trainium, Maia).
Strategic Moves: Hyperscalers like AWS are unlocking massive operating leverage—converting raw gigawatts into Bedrock token revenue, heavily propelled by Anthropic’s API-driven enterprise scale, while OpenAI pivots toward disciplined pricing to recover margins.
Risks & Warnings: Memory cost inflation, delivery execution (AMD/Helios), custom silicon software bottlenecks, and unsustainable subsidization of coding tokens pose immediate margin threats.
Key Views & Dialogues
Ep. 020 - Anthropic vs OpenAI Usage, Margins, Meta Compute, Future of MSL (Tokenomics)
- 🗓️ Date:
2026-07-18| 🎙️ Show:SemiAnalysis
Anthropic’s enterprise/API mix is producing operating leverage: over 80% of ARR is API-based, Q2 operating profit was positive, and Q3 could exceed $1 billion. OpenAI’s free-user base weighs on margins, but 5.5 and 5.6 have reportedly restored a two-horse race, while subsidized coding plans and RL environments leave unit economics and capability scaling unresolved.
View Dialogue Notes & Key Takeaways
Enterprise token austerity is aimed at the wrong workloads: coding consumes the budget, while premium-model emails are effectively free. Crystal says most employees never approach their caps, so pooled company budgets make more sense than per-person ceilings. Max calls engineering-only access a “caste system,” while Joey says the top 1% of companies already spend roughly $100,000 per employee annually on AI and continue when the ROI survives scrutiny.
The $200 coding subscriptions are heavily subsidized and may be loss-making at power-user utilization. The team estimates Codex can provide about $12,000 of API-equivalent usage and Anthropic about $8,000; break-even utilization falls to 10% for Claude Max 20x and 5.7% for OpenAI Pro 20x. At likely utilization levels, Max says these plans “might just be negative margin.”
Anthropic’s enterprise/API-heavy mix is already producing operating leverage that OpenAI’s free consumer base cannot match. More than 80% of Anthropic ARR is API-based and historically more than 90% enterprise; Joey says some profitability is on a non-GAAP basis excluding stock compensation, while operating profit was positive in Q2 and could reach $1 billion-plus in Q3. OpenAI’s roughly 950 million weekly users convert only about 6% to paid plans, and servicing the free population lowers blended gross margin by around 20 points.
OpenAI has nevertheless returned to a genuine two-horse race because model quality, not incumbency, remains decisive. After likely flat ARR growth in March and April, 5.5 became the predicted inflection point; some people now judge 5.6 “as good as Opus 4.8” at half the price and likely with a smaller model. Net-new monthly ARR is reportedly comparable with Anthropic, prompting Dylan’s reversal: “I am nothing if not flexible.”
Compute optionality may matter as much as compute ownership for xAI and Meta. Max’s revised view is that surplus capacity can be rented at 3x-4x market rates while teams retain enough to prove they can reach the frontier, protected by 90-day clawback clauses. Meta compute could similarly justify heavier 2027-28 capex and become a profitable fallback if MSL disappoints; by contrast, Max reads Google’s non-recallable, long-term TPU commitments as evidence of weak conviction in building RSI.
Hyperscalers win token-as-a-service distribution through existing enterprise relationships, while independent inference providers can prosper without winning the market. Anthropic’s indirect token volume may have risen from roughly 5% to 20% of its business in six months, and AWS may collect a 20%-30% revenue share largely because customers already buy through Bedrock. Together, Fireworks and Baseten can still grow rapidly because inference could be “the largest market ever,” even if open-source volume grows more slowly than frontier-model volume.
RL environments are becoming both the next capability bottleneck and an unusually lucrative engineering market. Frontier-lab data budgets could exceed $10 billion this year — about 10x last year — and might increase another 10x in 2027; top coding tasks command well over five figures each, while people who are very good at creating RL tasks can earn seven figures or more annually. The episode makes its exponential thesis falsifiable with a $400 billion Anthropic ARR over-under for end-2027: Max and Dylan take the over, Crystal, Joey and Jeremy the under, and the loser must explain publicly why they were wrong.
🔗 Original source & video: Ep. 020 - Anthropic vs OpenAI Usage, Margins, Meta Compute, Future of MSL (Tokenomics)
Ep. 013 - AWS Margins Jump 10% While Azure and GCP Flatline (Tokenomics) | Jordan Nanos, Jeremie Eliahou Ontiveros, Joey Brookhart, Crystal Huang
- 🗓️ Date:
2026-06-01| 🎙️ Show:SemiAnalysis
AWS’s margins are rising while it adds more than a gigawatt of capacity per quarter, signaling a shift toward higher-margin Bedrock token sales. Claude represents an estimated 80%-93% of Bedrock usage, while Anthropic’s API-led growth reaches $47 billion ARR. Whether this operating leverage survives stabilized utilization, coding-token cuts, and the SpaceX/xAI capacity deal remains the key risk.
View Dialogue Notes & Key Takeaways
AWS’s improving margins versus Azure and GCP come from selling Claude through Bedrock as higher-margin tokens, not simply renting accelerators. AWS is adding more than a gigawatt of capacity per quarter while margins rise; Azure and GCP remain far more exposed to lower-margin infrastructure-as-a-service. Joey’s framing: token sales retain more upside than five-year take-or-pay contracts.
The capacity ramp normally crushes near-term cloud margins before clusters reach “stabilized” utilization. CoreWeave-style providers pay depreciation, leases, and labor while complex systems such as GB200 wait months for activation and produce no revenue. AWS’s ability to absorb the same costs while expanding margins is therefore “a pretty good sign” for its eventual return on capital.
Claude’s API-heavy growth is giving both Anthropic and AWS unusually strong operating leverage. Joey cites Anthropic at $47 billion of ARR, with probably $10 billion of net-new ARR per month across March, April, and May and roughly 80% of the increase coming from APIs. Amazon was “in the right place at the right time”: Claude represents an estimated 80%-93% of Bedrock usage.
Anthropic’s $65 billion Series H at a $965 billion post-money valuation looks less extreme against its growth and profitability. Joey compares the roughly 20x ARR multiple with the 80x levels reached by software names in 2021 and says Anthropic is profitable excluding stock-based compensation. The caveat is significant operating deleverage if enterprises curb coding-token consumption, but “there’s no train that’s slowing right now.”
The SpaceX/xAI compute deal produced the episode’s sharpest disagreement over AI demand. Jeremie sees a former compute buyer becoming a supplier and possibly “giving up on the frontier race”; Jordan sees overwhelming Anthropic demand, valuable GPU-recall optionality, and a rational way for xAI to earn revenue until its own research and distribution can use the capacity.
Whether AI becomes winner-take-all depends on whether spending concentrates in open-ended tasks where “good enough” never arrives. Crystal argues a third- or fifth-ranked model could still replace substantial labor; Jeremie counters that legal work, science, healthcare, and analysis reward continually buying more intelligence. Jordan’s formulation is even broader: “coding is not coding, it’s computer use.”
The durable winners may be hyperscalers combining frontier-model access, enterprise distribution, and custom silicon. Bedrock could become the majority of AWS’s AI business by year-end, while Azure and GCP remain 80%-90% infrastructure-as-a-service in the panel’s model. Trainium and TPUs gain another advantage when token buyers never need to know which accelerator served them: “Winners win, losers lose.”
🔗 Original source & video: Ep. 013 - AWS Margins Jump 10% While Azure and GCP Flatline (Tokenomics) | Jordan Nanos, Jeremie Eliahou Ontiveros, Joey Brookhart, Crystal Huang
AI Chip & Silicon Round-up 2026
- 🗓️ Date:
2026-03-13| 🎙️ Show:SemiAnalysis
NVIDIA’s Vera Rubin VR200 targets 35 petaflops and 288 GB of HBM4, while AMD’s MI455X could gain a memory edge if delivered on time. Inference is increasingly shifting to custom silicon, with Trainium 3, Maia 200, and Ironwood offering credible alternatives; external access and execution remain key uncertainties.
View Dialogue Notes & Key Takeaways
NVIDIA’s Vera Rubin VR200 is called the most anticipated release of 2026, and the narrator expects it to “very likely once again top the charts.” Built on TSMC’s N3B with HBM4, a single package targets 35 petaflops of FP4 and 288 GB at 22 TB/s, scaled into the NVL72 rack via NVIDIA’s NVLink scale-up network—with the caveat that 22 TB/s is “the speed NVIDIA is targeting.”
AMD’s 2026 could be a turning point—conditional on execution: “if MI455X and the Helios Rack are on time.” CDNA 5, 320 billion transistors in a mix of 12 2-nanometer and 3-nanometer logic chiplets connected via advanced 3.5D packaging, and the headline edge is memory: 432 GB of next-generation HBM4 at nearly 20 TB/s, an advantage that holds “at least until Rubin Ultra, which will come with a full terabyte of HBM.”
The structural call of the episode: “More and more inference is moving away from NVIDIA to custom silicon.” Microsoft’s Maia 200 (140B transistors, 216 GB HBM3E, over 5 and 10 petaflops at FP8 and FP4, respectively) will run future ChatGPT models; Meta’s MTIA v3 is very likely to use HBM and “offers good margins,” so external hardware can handle training while internal inference goes in-house.
Trainium 3 is flagged as a top contender for future large-scale deployment, based on Trainium 2’s footprint. Hundreds of thousands of Trainium 2 chips already sit in AWS’s Canton and New Carlisle data centers; Anthropic’s Claude Code was trained on and runs on Trainium, and “OpenAI will use 2 GW of Trainium compute starting this year.”
Google’s Ironwood (TPU v7) is the efficiency heavyweight if external customers can use it as well as Google can. 192 GB HBM3E, likely 100B+ transistors, and optical circuit switches—“tiny physical mirrors”—linking superpods of up to 9,216 TPUs.
Qualcomm’s LPDDR bet might have been a great idea a year ago, but memory prices are “skyrocketing across the board,” including LPDDR5X. The AI 200 (768 GB LPDDR5X) is unlikely to make waves, but the AI 250 is supposed to bring a compute-near-memory architecture that, according to Qualcomm, delivers a 10x effective-bandwidth gain with LPDDR6—“don’t put on your party hats just yet, but do keep Qualcomm in mind.”
The SRAM outliers get respect, not conviction: NVIDIA thought Groq’s deterministic LPU “was worth about 20 billion dollars,” while Cerebras’s WSE3 with 44 GB of on-wafer SRAM “isn’t really cutting it anymore,” even if it is SRAM. Intel’s Jaguar Shores (18A, 288 GB HBM4) looks competitive on paper, but it does not seem targeted for 2026 and might be a 2027 product.
🔗 Original source & video: AI Chip & Silicon Round-up 2026