Pioneers Insight Method Research Author
The Future of Frontier Model Architectures with Walter Goodwin, Founder & CEO of Fractile
Back to Episodes

The Future of Frontier Model Architectures with Walter Goodwin, Founder & CEO of Fractile

Summary

  • Fractile’s core bet is that frontier inference will be constrained by memory bandwidth and memory cost, not merely compute. Goodwin wants to combine DRAM’s capacity and economics with the “many thousands of tokens per second” associated with SRAM-based systems, enabling long-context agents and models with many trillions of parameters.

  • Goodwin sees much of today’s AI-chip variety as architecturally similar beneath the branding. NVIDIA and AMD GPUs, hyperscaler ASICs, and TPUs commonly rely on HBM, tensor cores, and TSMC advanced packaging. Fractile’s differentiation is attempting innovation across the entire stack rather than outsourcing physical implementation after front-end design.

  • Fractile abandoned an initial SRAM architecture after concluding that model size and context length would outgrow it. SRAM delivers extraordinary bandwidth but insufficient cost-effective capacity; even fast inference chips may fall back to GPUs for long-context processing. The company instead pursued “aggressively high bandwidth from the world’s cheapest memory,” DRAM, with a platform expected to be fully operational in the second half of next year.

  • The company’s roughly 150-person, vertically integrated team is designed to shorten the feedback loop between workload research, architecture, physical design, packaging, and foundry interaction. Goodwin argues that a durable three-to-six-month lead can decide deployments: “If you find a way to structurally secure a three- to six-month lead, you win all those implementations.”

  • AI may radically compress chip development, but it cannot abolish fabrication time or silicon economics. Foundry turnaround still takes three to five months, production ramps take 12-18 months, and a financially viable chip needs a three-to-five-year amortization window. Faster design therefore means maintaining more “irons in the fire,” not shipping a fundamentally new processor every few weeks.

  • Fractile claims 25 times more bandwidth per chip than an HBM-based design, potentially changing which model architectures are economical. Goodwin’s example is mixture-of-experts sparsity: moving from 1-to-16 toward 1-to-128 or 1-to-256 could save substantial computation, but current accelerators become bandwidth-bound. Compute has scaled roughly one million times in 20 years while memory bandwidth rose only about 40 times.

  • Goodwin expects frontier labs to keep buying from multiple hardware suppliers even as they develop proprietary silicon. In-house chips provide bargaining power, supply diversity, and control, but exclusive dependence is dangerous: a rival’s fivefold algorithmic efficiency breakthrough could leave a lab stranded for nine months while it redesigns hardware. Independent platforms remain valuable because they let labs compete at the model layer without betting survival on one architectural path.

Deep dive

1. Full-stack integration is Fractile’s answer to ASIC sameness

  • Founded in summer 2022, Fractile began with two linked convictions: foundation models would generalize broadly, and additional inference-time compute could bring AlphaGo-like scaling to language. Its mission became running the world’s largest models “much, much faster,” including at extreme context lengths.

  • Goodwin’s industry framing: the apparent “luxury of choice” conceals common foundations. Google’s TPU, Meta’s MTIA, Microsoft’s Maia, OpenAI’s Jalapeño, NVIDIA and AMD GPUs, and other ASICs are ultimately often developed and shipped with a small set of ASIC companies. Beneath the branding, they generally share HBM, tensor-oriented matrix multiplication, and advanced TSMC packaging.

  • The conventional handoff starts with workload-aware architects, proceeds through RTL and front-end design, and ends with specialists synthesizing that logic into GDSII—the transistor-and-metal-layer layout sent to a foundry. Development and supply companies such as Broadcom—which Goodwin identifies as the largest, with a $2 trillion market capitalization—help implement these projects, including analog connectivity IP and node-specific physical placement.

  • Fractile keeps workload analysis, front-end design, physical design, backend implementation, and advanced packaging inside a team of about 150. The advantage is a “much more flexible closed loop” for choosing architectures that may not reach volume production for another one or two years.

2. Fractile traded SRAM purity for scalable DRAM bandwidth

  • Goodwin says the first decisive bet was inference itself: some chips that recently reached the inference market had been training chips just 18 months earlier. Deployment cost may sound marginal, but it is incurred “every time you deploy these models,” making inference economics structurally important.

  • For its first two years, Fractile pursued an SRAM design resembling Groq or Cerebras. Because SRAM sits beside the logic, it supplies enough bandwidth to move weights or KV cache rapidly and produce thousands of tokens per second.

  • By late 2023 and through 2024, increasing parameter counts and context lengths changed Goodwin’s view. SRAM’s capacity became the limiting factor; fast-output systems could not handle long-context processing and therefore still relied on GPUs for that work.

  • The revised architecture pursues high-capacity, lower-cost DRAM with exceptionally high bandwidth. Goodwin’s “faster horses” distinction matters: a faster chatbot is incremental, while running models with many trillions of parameters at thousands of tokens per second could enable long-lasting agents and radically speed their performance.

3. Faster design creates more bets, not disposable chips

  • Sarah Guo’s pushback — worth keeping: chip generations already arrive roughly annually, major architectural changes move more slowly, and physical supply chains impose unavoidable constraints. How, then, can Fractile credibly claim a software-like increase in architectural velocity?

  • Goodwin concedes the tension. New models appear roughly every two weeks; attention mechanisms, MoE sparsity, and attention sparsity can shift every few weeks, especially among advanced Chinese open-source systems. Yet common requirements persist: high memory bandwidth and autoregressive generation at very small batch sizes, creating a trade-off among bandwidth efficiency, cost, and service speed.

  • Some delays are fundamental: a rush foundry cycle still takes three to five months, scaling a flagship platform takes 12-18 months, and the chip must have a useful life of more than three years to justify a three-to-five-year amortization window. “I don’t buy” fundamentally new chips every few weeks, Goodwin says.

  • Compression instead buys optionality. Fractile wants a dynamic set of aligned bets ready to run, so it can select a durable flagship and scale it—while competing against NVIDIA systems that combine six to nine custom NVIDIA chips. The prize is a repeatable three-to-six-month lead over competitors.

4. AI moves architecture faster than physics

  • Guo cites a semiconductor CEO’s private estimate that moving directly from a chief architect’s idea to completed GDSII could take ten years. Goodwin’s heuristic is to challenge the assumptions, divide that timeframe by four, and subtract a little more.

  • Intelligence is only one part of the loop. Placement and routing attack NP-hard problems with algorithms that can run for days, while finite-element, thermal, and other simulations create their own bottlenecks—an Amdahl’s-law problem analogous to AI-guided model research being limited by experiment time.

  • Goodwin expects approximations and surrogate-model-like tools to accelerate iteration, but not soon replace final sign-off from Cadence and Synopsys, which have worked with TSMC and other foundries for decades to verify DRC/LVS compliance with foundry rules. His deeper prescription is to spend far more reasoning before each costly experiment: potentially “a hundred years” of human-equivalent thought before and after a real-world test.

5. Bandwidth can reshape models and preserve an independent chip market

  • Fractile must serve today’s “hardware lottery”—models trained for GPU- or XPU-like systems—while using its claimed 25-times bandwidth advantage to exert “gravitational forces” on future architectures. Goodwin calls the opportunity a bandwidth scaling law alongside familiar FLOP scaling laws.

  • Mixture-of-experts models carry the clearest example. Increasing sparsity from 1-to-16 toward 1-to-128 or 1-to-256 could reduce computation at equal intelligence, but HBM systems struggle to serve those models efficiently. Similarly, some attention mechanisms are less demanding on bandwidth but require more compute for a given level of intelligence.

  • The imbalance is historical: over 20 years, compute increased roughly one million times while memory bandwidth rose about 40 times. Raising the bandwidth frontier could reduce the number of operations required for a given level of intelligence rather than merely execute the existing workload faster—a multiplier on overall model performance.

  • Goodwin expects large AI players and hyperscalers to retain multiple suppliers for existential supply resilience and NVIDIA pricing leverage. He sees a premium segment optimized for speed above all else, especially under open-source pressure—otherwise, he says, everyone would use Kimi models all the time. Proprietary hardware alone is an asymmetric risk: if a rival achieves fivefold efficiency on an incompatible design, a lab might “disappear for those nine months” while catching up. Shared independent platforms let frontier labs keep competing through model quality and inference speed.