Pioneers Insight Method Research Author
Positron’s Sohmers: Data-Centre Backlash and AI’s Real Bottleneck
Back to Episodes

Positron’s Sohmers: Data-Centre Backlash and AI’s Real Bottleneck

Summary

  • Sohmers’s core hardware call is that inference is becoming a memory problem even though training remains compute-bound. From 2014 to 2024, GPU FLOPS improved roughly 120x while memory bandwidth rose only 17x; Positron’s base case is to perform a workload requiring 500 MW of NVIDIA equipment in 100 MW. Operators will still maximize the facility, however, because efficiency means “more tokens, more intelligence per joule,” not less demand.

  • Cached tokens are the model providers’ hidden profit engine. Sohmers estimates processing a cached token costs roughly one-thousandth as much as recomputing it, making cache reads an “obscene margin” product and helping explain Anthropic’s reported 80-point API gross margin. OpenAI and Anthropic could become “massively profitable overnight” by stopping training, he argues; frontier expenditure, not inference economics, drives the cash burn.

  • A unilateral Western slowdown would concentrate AI power without reliably slowing China. Sohmers calls restrictions on who may perform matrix multiplication a modern “road to serfdom”: most tokens may come from large providers, but restricting the underlying capability turns them into “new lords and kings.” Harry’s pushback lands—“you can’t pace the frontier unless the global AI community paces the frontier”—and Sohmers agrees export controls cannot preserve a permanent lead.

  • Sohmers sees anti-data-centre politics as an almost entirely “Chinese psyop,” because false resource claims could hobble Western capacity while China keeps building. He claims one In-N-Out uses more water than the largest US data centres and notes closed-loop cooling, while arguing new facilities bring generation sufficient for their own use and beyond. He instead points to permitting, grid structure and economics—not an absolute lack of land or generation technology—as the binding constraints.

  • Most planned hyperscale capacity should still be built, although rejected projects will move jurisdictions. The US retains large tracts of remote federal land with geothermal, solar and potential nuclear resources; Sohmers is therefore not worried about an existential capacity shortfall. Ocean-based pumped-hydro facilities such as Pantala and, longer term, space data centres provide additional options—though he concedes terrestrial construction remains cheaper and easier today.

  • Small and on-device models could increase rather than cannibalize frontier-model demand. Roughly 80-85% of tokens currently come from the top four model companies, Sohmers estimates, with perhaps 5% eventually running on-premises; local agents continuously reading email, calendars and messages could autonomously escalate difficult work to cloud models. “The bottleneck is actually a human making some sort of decision,” so trusted local orchestration could unlock orders-of-magnitude more frontier tokens.

  • KV caching is economically essential because valuable agentic workloads are overwhelmingly repetitive, but it turns inference into a storage-management problem. SemiAnalysis’s Agent X traces reportedly find about 96% of tokens cached; meanwhile a quantized 1.8-trillion-parameter GPT-4 would occupy roughly 900 GB, a hypothetical 10-trillion-parameter Claude Fable about 5 TB, and one long-context user session could approach 100 GB. At only 50 users, context can outweigh the model itself.

  • Falling token prices understate how quickly the economic value of intelligence is rising. The index Harry cites fell from $60 to below $1 per million tokens in five years, but Sohmers estimates today’s tokens may be 100x more useful and the value per unit of intelligence closer to 1,000x higher. He calls GPT-6 Astra AGI, citing its ability to compress a two-to-three-week chip-design task into roughly 50 hours; future frontier agents may therefore be sold as million-dollar annual workers rather than metered purely by tokens.

Deep dive

1. Inference moves AI’s bottleneck from arithmetic to memory

  • Positron spans chips, low-level software, complete systems and rack-scale deployments for generative-AI inference. Harry introduces the company as having raised an $875 million Series C at a $5 billion valuation.

  • Sohmers’s distinction is causal: training already has its corpus, so inputs can be massively parallelized and “crushed through with a bunch of compute.” Inference generates each token autoregressively, without knowing the fifth word ahead, and must reread the model’s weights for every output token.

  • The memory wall predates generative AI. Between 2014 and 2024, Sohmers says, FLOPS per NVIDIA GPU improved roughly 120x but memory bandwidth only 17x, leaving every memory-bound workload progressively further behind available compute.

  • SRAM cells have shrunk much more slowly than general-purpose transistor groupings, while architectures such as AlexNet and ResNet rewarded additional compute. Transformers changed the incentive, but Sohmers says the market only truly noticed after GPT-3 scaled to 175 billion parameters and ChatGPT arrived in late 2022.

2. Cache economics explain the API providers’ margins

  • Sohmers’s underappreciated token-economics call: “You make all of your money on selling cached input and output tokens.” Providers charge to establish a cache, then offer cheaper reads whose actual processing cost may be roughly one-thousandth of fresh computation.

  • Anthropic’s reported 80-point API gross margin does not surprise him, although “I’m impressed by 80 points of margin in basically any industry.” Competition should compress it, which he welcomes as healthier for the ecosystem even if Positron’s customers consequently retain less margin.

  • The popular image of OpenAI and Anthropic as structurally loss-making is “absurd” to Sohmers: stop frontier training and they could become “massively profitable overnight.” Harry notes that pacing would reduce their largest cost; Sohmers jokes it could be useful “ahead of IPO,” while explicitly saying he does not believe that is the actual intention.

3. Pacing the frontier risks becoming a road to technological serfdom

  • Sohmers takes safety seriously—he wants humanity to survive “to the stars and beyond”—but says he believes much more in AI’s positive potential. A pause could validate broader Luddite efforts to stop the technology altogether rather than narrowly mitigating catastrophic risk.

  • His personal “P doom” is concentrated capability. If only selected companies or governments may legally deploy advanced computation, those institutions become “the new lords and kings and everyone else is back to serfs”; regulating matrix multiplication would attack classical liberal freedom at a foundational level.

  • Harry’s cynical reading is that incumbents gain strategic cover: compliance burdens Meta, pacing improves IPO economics, and Elon Musk gets time to catch up. Sohmers largely agrees for everyone except Dario Amodei and Anthropic, whom he describes as sincere believers in both AI’s promise and its risks.

  • Sohmers’s sharper warning is that executives inviting regulation assume they will remain in control. Dario may expect to become the regulator, but could instead discover that “it just becomes a bureaucracy that halts all progress,” with existing capability captured and squandered by officials.

4. A Western pause cannot bind China

  • Harry’s central objection—“you can’t pace the frontier unless the global AI community paces the frontier”—draws an immediate “agreed.” Sohmers thinks Western leaders overestimate the durability of their lead, much as militaries can remain optimized for “the last war” while cheap drones reshape conflict.

  • A CCP-controlled superintelligence could impose the same technological serfdom he fears domestically, with worse consequences for some people. China currently benefits from proliferating open technology, but Sohmers expects that “as soon as they get into pole position, the ladder gets pulled up.”

  • He doubts the CCP would allow the roughly one billion Chinese people who are not party members to benefit equally. The present open-source strategy should therefore not be confused with a lasting commitment to broadly distributed capability.

  • Sohmers favors free trade and open exchange, but makes an exception for totalitarian regimes that enjoy the liberal system’s benefits while remaining closed themselves. Exporting capabilities that strengthen those regimes is, in his framing, not reciprocal free trade.

5. Anti-data-centre politics may relocate capacity, not eliminate it

  • Sohmers calls the left-right convergence against data centres “almost entirely a Chinese psyop.” Harry supplies the strongest countercase: communities see higher electricity and water prices plus “ugly data centers in their backyard,” even if the facilities also move regional economies.

  • On aesthetics, Sohmers proposes turning the largest sites into civic monuments—future generations should see them as technological “Great Pyramids.” On resources, he says many claims are “patently false”: one In-N-Out allegedly uses more water than the largest US data centres, while golf courses use orders of magnitude more.

  • China, meanwhile, adds gigawatts of generation, builds enormous facilities and can bulldoze homes or impose rolling blackouts. Sohmers values Western freedom to object, but sees a strategic disadvantage when public criticism rests on misinformation while China can compel construction.

  • Harry argues that Europe can barely permit “a paper airplane” and the US is drifting the same way. Sohmers’s counterweight is geography: he estimates, while acknowledging that he does not know the exact number, that more than 90% of federal land is open Western desert, including Nevada sites suitable for geothermal, solar and potentially nuclear generation far from populations.

6. Energy is universal, but capital and market design bind first

  • Sohmers expects major providers’ planned capacity largely to materialize, though not necessarily where originally proposed. Communities may stop individual facilities, but projects can move, and operators are investing more in local education after underestimating the political resistance. Although terrestrial construction remains cheaper and easier near term, he says he would never bet against Elon and sees space data centres as a long-term possibility.

  • His categorical power-market claim is that data centres cannot draw electricity already allocated to homes. He says new sites bring generation covering their own consumption “and beyond,” but grid rules can prevent excess capacity from connecting while incumbent utilities resist supply that would lower consumer prices.

  • Positron’s base case converts what would require 500 MW of NVIDIA equipment into 100 MW. That will not produce smaller facilities: operators will still deploy the maximum available power, receiving “more tokens, more intelligence per joule.”

  • Over centuries, “all progress is gated by energy,” from fire through nuclear power. Nearer term, Sohmers sees economics as the tighter constraint: how much debt will the world accept to finance generation and infrastructure? He ties his sovereign-debt concern to governments’ ability to print currency and rising perceived risk in Treasuries and bond markets. He is not worried about AI-stack companies missing revenue targets; when discussing corporate debt, he says he trusts Oracle’s business model and execution more than the US government.

7. KV caching saves compute by creating a vast memory hierarchy

  • A token averages roughly half to 75% of an English word; a sequence is the ordered accumulation of prompt and generated tokens. Early transformers recomputed prior tokens repeatedly, while the KV cache stores the attention mechanism’s key and value matrices so that work need not be repeated.

  • The trade is attractive because attention compute grows quadratically with sequence length while stored K and V data grows linearly. The cost is a unique, expanding cache for every user and hard operational decisions about how long each session deserves expensive memory.

  • Quantization compresses FP32 to FP16 or BF16, then FP8 and FP4. A naive 16-bit-to-4-bit conversion saves 75% of space but can degrade benchmarks 20-30%; advanced shared biases and multipliers reach roughly 4.5 bits per value while keeping damage within about 1%.

  • SemiAnalysis’s Agent X traces of agentic coding sessions reportedly show approximately 96% of tokens are cached. Operators therefore retain active sessions in accelerator memory, recent ones in host memory—typically 4-10x larger—and older contexts in local or networked NVMe before eventually relegating them to slower storage.

8. Edge intelligence should create more frontier-model consumption

  • Sohmers still expects frontier models to grow from one trillion toward five, 10, 50 or 100 trillion parameters. He concedes this is “sort of on vibes”: scaling is called a law because capability gains have been observed, not because continued improvement has been mathematically proved.

  • For high-value internal Positron work, he would “very gladly pay ten times more” for ten times the output. Smaller proprietary models make sense for privacy, on-premises deployment and cost, but most enterprises lack the resources to push the capability frontier themselves.

  • His market split is highly concentrated: roughly 80-85% of tokens come from the top four model companies, with another 5-10% from the next three or four. He can believe on-premises systems eventually account for 5%, but Positron is designed around the high-volume tier.

  • The apparent substitution from phones and laptops is actually orchestration. A local model monitoring email, calendars and messages can decide continuously which tasks need a smarter cloud model, removing the human prompt as bottleneck and generating far more frontier requests than sporadic manual use.

9. GPT-6 Astra made AGI tangible through real engineering work

  • Sohmers says his first 24 hours with GPT-6 Astra felt as magical as GPT-3.5 in November 2022, when he spent four or five hours prompting ChatGPT after its low-key NeurIPS launch. Harry pushes back that Astra feels “kind of the same as before,” forcing the concrete case.

  • In software, Astra solved hard problems other models circled unsuccessfully, found bugs and performance improvements, and identified unexplored areas of Positron’s codebase. Sohmers calls that a step-function improvement, though “not mind-bogglingly” so.

  • Computer use supplied the bigger shock: Blender animation was basically impossible with GPT-5.6, while Astra handled it with high fidelity; he also had Astra design his home interior from a couple of pictures. That breadth led Sohmers to say plainly, “GPT-6 Astra, I do think is AGI.”

  • The decisive Positron test took a Kecak encryption block from specification through Verilog and the full RTL-to-GDS flow using TSMC N3 PDKs. Astra met timing above 1 GHz in a little over 50 hours, compressing what Sohmers estimates would take a newly assigned engineer two to three weeks.

10. Context quality and intelligence value matter more than sticker price

  • Positron’s upcoming generation is intended to offer eight times the memory of NVIDIA’s highest-memory SKU. Yet moving conventional attention from one million tokens to 10 million remains “really, really hard”; useful agents need enough context for entire codebases, not merely larger parameter counts.

  • Chinese labs have responded to hardware constraints algorithmically. DeepSeek V3’s multi-head latent attention reduced KV-cache size by spending more FLOPS, while gated DeltaNet derivatives can cut attention time roughly 75%; Sohmers cautions that “there’s no such thing as a free lunch” and says—based on rumors plus what he calls good information and belief—that major US labs have not adopted MLA.

  • Advertised length is not usable recall. GPT-5.6 found a hidden value in long-context “needle in a haystack” tests only about 70% of the time, Sohmers says, versus more than 95% for GPT-6 Astra—evidence that effective context still has substantial headroom.

  • Harry’s cited token index fell from $60 to below $1 per million, but Sohmers says the old token would now be worthless: current tokens are conservatively 100x more valuable, making the value per unit of intelligence perhaps 1,000x higher. Pricing may shift toward useful results or a hypothetical flat annual fee—he floats $1 million as a random example—for unlimited use of a virtual agent worker, while the deeper alignment task is aligning human incentives around regulation, energy and economics.