
Thomas Sohmers
Key Views & Dialogues
20VC: “Anti-Data Centres is a Chinese Psyop” | How Many Planned Data Centers Will Actually Get Built? | Is Energy AI’s Biggest Bottleneck? With Thomas Sohmers, Co-Founder @ Positron
- 🗓️ Date:
2026-09-19| 🎙️ Show:20VC
Positron’s Thomas Sohmers says inference is becoming memory-bound: GPU FLOPS rose roughly 120x from 2014 to 2024, versus 17x for memory bandwidth, potentially compressing a 500 MW NVIDIA workload into 100 MW. KV reads may cost one-thousandth of recomputation, helping explain Anthropic’s reported 80-point API gross margin, while permitting, grid rules, financing and unilateral export controls remain unresolved constraints on capacity and Western AI’s lead.
View Dialogue Notes & Key Takeaways
Sohmers’s core hardware call is that inference is becoming a memory problem even though training remains compute-bound. From 2014 to 2024, GPU FLOPS improved roughly 120x while memory bandwidth rose only 17x; Positron’s base case is to perform a workload requiring 500 MW of NVIDIA equipment in 100 MW. Operators will still maximize the facility, however, because efficiency means “more tokens, more intelligence per joule,” not less demand.
Cached tokens are the model providers’ hidden profit engine. Sohmers estimates processing a cached token costs roughly one-thousandth as much as recomputing it, making cache reads an “obscene margin” product and helping explain Anthropic’s reported 80-point API gross margin. OpenAI and Anthropic could become “massively profitable overnight” by stopping training, he argues; frontier expenditure, not inference economics, drives the cash burn.
A unilateral Western slowdown would concentrate AI power without reliably slowing China. Sohmers calls restrictions on who may perform matrix multiplication a modern “road to serfdom”: most tokens may come from large providers, but restricting the underlying capability turns them into “new lords and kings.” Harry’s pushback lands—“you can’t pace the frontier unless the global AI community paces the frontier”—and Sohmers agrees export controls cannot preserve a permanent lead.
Sohmers sees anti-data-centre politics as an almost entirely “Chinese psyop,” because false resource claims could hobble Western capacity while China keeps building. He claims one In-N-Out uses more water than the largest US data centres and notes closed-loop cooling, while arguing new facilities bring generation sufficient for their own use and beyond. He instead points to permitting, grid structure and economics—not an absolute lack of land or generation technology—as the binding constraints.
Most planned hyperscale capacity should still be built, although rejected projects will move jurisdictions. The US retains large tracts of remote federal land with geothermal, solar and potential nuclear resources; Sohmers is therefore not worried about an existential capacity shortfall. Ocean-based pumped-hydro facilities such as Pantala and, longer term, space data centres provide additional options—though he concedes terrestrial construction remains cheaper and easier today.
Small and on-device models could increase rather than cannibalize frontier-model demand. Roughly 80-85% of tokens currently come from the top four model companies, Sohmers estimates, with perhaps 5% eventually running on-premises; local agents continuously reading email, calendars and messages could autonomously escalate difficult work to cloud models. “The bottleneck is actually a human making some sort of decision,” so trusted local orchestration could unlock orders-of-magnitude more frontier tokens.
KV caching is economically essential because valuable agentic workloads are overwhelmingly repetitive, but it turns inference into a storage-management problem. SemiAnalysis’s Agent X traces reportedly find about 96% of tokens cached; meanwhile a quantized 1.8-trillion-parameter GPT-4 would occupy roughly 900 GB, a hypothetical 10-trillion-parameter Claude Fable about 5 TB, and one long-context user session could approach 100 GB. At only 50 users, context can outweigh the model itself.
Falling token prices understate how quickly the economic value of intelligence is rising. The index Harry cites fell from $60 to below $1 per million tokens in five years, but Sohmers estimates today’s tokens may be 100x more useful and the value per unit of intelligence closer to 1,000x higher. He calls GPT-6 Astra AGI, citing its ability to compress a two-to-three-week chip-design task into roughly 50 hours; future frontier agents may therefore be sold as million-dollar annual workers rather than metered purely by tokens.
🔗 Original source & video: 20VC: “Anti-Data Centres is a Chinese Psyop” | How Many Planned Data Centers Will Actually Get Built? | Is Energy AI’s Biggest Bottleneck? With Thomas Sohmers, Co-Founder @ Positron