Pioneers Insight Method Research Author
159: Musk’s Terafab Space Compute, Nvidia Reclaims the CPU | Talking New AI Compute Trends with Zhang Lu of Fusion Fund
Back to Episodes

159: Musk’s Terafab Space Compute, Nvidia Reclaims the CPU | Talking New AI Compute Trends with Zhang Lu of Fusion Fund

Summary

  • Terafab is not really competing for a chip fab; it is competing for compute sovereignty across Musk’s ecosystem. The plan is to integrate Tesla, SpaceX and xAI from chip design and manufacturing through deployment and applications, with an ultimate target of producing 1 TW of AI compute annually and deploying 80%-90% of it in space. Zhang Lu described it as “a giant cross-company industrial operating system,” but expects both the rollout timeline and funding needs to far exceed Musk’s own expectations.
  • Orbital data centers are not a mature startup bet over the next 2-3 years. Cosmic radiation, chip packaging, launch, maintenance and repairs all present major technical and cost challenges; serving Earth-based applications would also introduce latency. Zhang’s estimated maturity window is seven or even ten years. “If everyone wants to explore building data centers in space, why not build them in nearby Canada?”
  • SpaceX needs the entire space economy—not just rockets and Starlink—to support a trillion-dollar IPO narrative. Zhang believes orbital deployment also reflects Musk’s deeper goal: “He doesn’t want any government to regulate him.” With jurisdiction over space, orbital enforcement and lunar resource ownership still unclear, SpaceX could become both the main enabler of the space economy and a future rule-setter.
  • The nearer-term opportunity than orbital data centers is serving a space industry that is “AI native, robotics native.” Microgravity factories could explore crystals, materials and protein structures difficult to produce on Earth; AI could manage satellite traffic and data trading, while robots extract water, hydrogen and oxygen from lunar regolith to build lunar “gas stations.” Genuine demand for space-based compute will emerge only after these local applications scale.
  • The more investable opportunities today are the efficiency layer of terrestrial AI infrastructure and other space infrastructure. Interconnect, optical switches, power consumption, memory, security, inference costs and system integration all present clear bottlenecks; startups must also decide whether they will serve future hyperscale clusters or build capital-intensive clusters themselves. Zhang’s direct conclusion: “I don’t think it’s a good time. I think it’s still a little too early.”
  • Nvidia is rewriting itself from a GPU company into an “AI factory” that produces tokens, while inference turns revenue from one-off demand into recurring cash flow. The training-to-inference compute mix has shifted from roughly 70%-80% versus 10%-20% toward an even split, and could eventually invert to 20%-30% versus 60%-70%; Jensen Huang therefore expects Blackwell and Vera Rubin to support more than $1T in data-center revenue by 2027.
  • The Agent era will lift demand for heterogeneous compute across CPUs, GPUs, LPUs and NPUs, making platform control more important than winning on any single chip. Nvidia is using Vera CPU to cover tool calls, code execution, multi-Agent workloads, reinforcement learning and simulation, while adding Groq’s low-latency, high-throughput inference capabilities to Vera Rubin; even if its CPUs do not always beat AMD’s, control of the full system, rack and one-stop purchasing creates a major integration advantage.
  • Enterprise AI has entered a phase in which budgets, deployments and exits are all accelerating. Regulated industries favor private-cloud or on-premises deployments of small language models to balance accuracy, privacy, latency and cost; companies in the Fusion Fund network have reached AI budgets of up to $12B, while some sub-10-person companies grew annual revenue from zero to $20M. Finance, healthcare, insurance, industrials and supply chains could be among the fastest adopters, and some portfolio companies founded less than two years ago were acquired last year at 10-20x returns.

Deep dive

1. Terafab aims to turn Musk’s ecosystem into an industrial operating system

  • Manqi summarized the plan as follows: on March 21, Musk proposed bringing Tesla, SpaceX and xAI together to build in-house full-stack capacity from chip design and manufacturing through application deployment, with an ultimate target of producing 1 TW of AI compute annually, 80%-90% of it in space.
  • Zhang believes this is not a sudden idea. Musk has long wanted chips, infrastructure, models and end products under one system; the core motivation is to “avoid being bottlenecked by others as his own compute needs evolve.”
  • She pointed to Google Gemini’s rapid improvement as a reference case: owning the TPU, compute, models and data makes system-level and cost optimization easier. Musk wants to extend that full-stack capability into cars, robots, satellites and the physical world, creating “a giant cross-company industrial operating system.”

2. Orbital data centers first run into radiation, maintenance and latency

  • Zhang’s first reservation is cosmic radiation: its impact on chip performance is “far greater than people think,” and practical questions remain around how chips including TPU should be packaged and protected.
  • Even if SpaceX continues to drive down launch costs, launch is only one part of the total cost. When hardware fails in a large cluster, routine operations and repairs are extremely expensive; “maintenance costs are actually higher.”
  • The customer for the compute matters just as much. If orbital systems support future satellites, robots and other space AI, local compute can reduce latency; if the main near-term use is sending results back to Earth, the distance will weaken the economics.
  • Zhang used a deliberately simple counterexample to test the energy and cooling narrative: Canada has abundant water and power, a cold climate and plenty of open land. “Why not build a data center in nearby Canada” instead of jumping directly to space?

3. A trillion-dollar SpaceX needs a “space economy” valuation framework

  • Zhang explicitly links Terafab to a potential SpaceX IPO: market expectations have reached the trillion-dollar range, but “no company has ever gone public at such a high price,” and rockets plus Starlink alone are unlikely to support the full story.
  • Musk wants SpaceX to represent the space economy—the ecosystem, infrastructure and full potential market value of space. Terafab, lunar bases, orbital services and edge compute are therefore both long-term plans and assets that support the valuation upside.
  • She also emphasized that Musk has been consistent: an IPO narrative does not mean the vision is fake. If space AI applications eventually generate massive demand, building data centers close to those workloads can work; the technology and demand still need time to mature.

4. A regulatory vacuum may explain orbital deployment better than solar power

  • Zhang’s sharper read on Musk’s motivation is that “a core reason he wants to build data centers in space is that he doesn’t want any government to regulate him.” A 1 TW buildout on the territory of a particular country would require extensive regulatory approvals and political coordination.
  • Low Earth orbit is subject to some international coordination, but it lacks a clear jurisdictional framework comparable to territory or airspace. Rules may exist around who launches satellites and who is responsible for removing them, but there is no clear, continuous enforcer.
  • That gives orbital slots and lunar resources a first-mover character: launching more satellites first could mean controlling scarce orbital capacity. Musk could become not only a major supporter of the space economy, but also “a rule-setter for the entire future space economy.”

5. Space factories are the nearer-term AI and robotics application

  • Zhang sees space factories as one of the most compelling near-term use cases. Microgravity or zero-gravity environments could produce new crystals, materials and protein structures unavailable on Earth, or create more perfect and symmetrical crystals, offering new options for industries constrained by material bottlenecks.
  • Her core definition is that space is naturally “AI native, robotics native.” Sending humans up requires an ecosystem to sustain life, health and safety; robots require far less support, so automation is not replacing labor but providing the precondition for the industry to exist.
  • This also points to the more credible long-term path for orbital data centers: build space factories, satellite networks and robotic systems first, then use local clusters to support local AI rather than moving terrestrial workloads into orbit from the outset.

6. Satellite traffic and lunar gas stations sketch out demand for local compute

  • One Fusion Fund portfolio company uses AI to manage satellite traffic, reducing collisions, asset losses and space debris. The process can also facilitate satellite data trading, applying high-quality data to mineral detection, wildfire alerts and weather services on Earth.
  • Another company is building a space “gas station”: robots extract water from lunar regolith and separate it into hydrogen and oxygen, all of which can support rockets or spaceflight. If the model works, spacecraft would not need to carry all their fuel from Earth and could refuel on the Moon.
  • Zhang envisions every satellite becoming a small edge device. Only once satellite data, automated facilities and space AI applications reach sufficient scale will local data centers have both a demand base and a latency advantage.

7. The startup window for orbital data centers may still be seven to ten years away

  • Manqi mentioned Starcloud, a US company already working on space computing services. Zhang is familiar with the company and said it is discussing a partnership with the lunar gas-station project described above. Musk’s vision has clearly accelerated investor attention.
  • But she does not view this as a mature opportunity over the next 2-3 years. Orbital data centers may require a seven- or even ten-year technology cycle, allowing packaging, launch and maintenance costs to mature while the space economy generates enough local workloads.
  • Her answer to founders left no room for ambiguity: “I don’t think it’s a good time. I think it’s still a little too early.” The more rational observation period is the next 3-5 years, after which entry should be judged against the actual growth rate of the space ecosystem.
  • Capital structure is another constraint. Data centers are capital-intensive businesses, so startups must decide whether to own the clusters themselves or provide the hardware, software and space infrastructure that make clusters more efficient; the latter is better suited to startups today.

8. Terafab’s nearer-term opportunities are concentrated in cluster efficiency

  • Zhang broadens the opportunity from “building another chip” to full AI infrastructure: data-center costs, power consumption, networking, storage and system coordination all determine real-world performance, and infrastructure itself will form moats like energy and transportation.
  • Her specific examples include next-generation interconnects and optical switches. AI power consumption does not come only from training; a large amount is spent on communication and data transfer. Faster, lower-power interconnects are prerequisites for large-scale deployment.
  • Capital is shifting as well. Large VCs that previously invested only in software are moving into deep tech, looking for hardware and software that solve AI infrastructure bottlenecks. Terafab is better understood as a demand signal than as a mandate for every company to replicate its capex.

9. Nvidia is remaking itself as an AI factory that produces tokens

  • At GTC, Zhang observed that Jensen Huang no longer defines Nvidia as a chip or GPU company, but as a “full-stack artificial intelligence infrastructure company” built to support the next phase of the token economy.
  • The product unit has also shifted from a single chip to an ecosystem. Zhang recalled that when she was in school, a company releasing 1-2 chips a year was considered fast; at this GTC, Nvidia launched 7 chips at once, alongside interconnects, inference infrastructure and software platforms.
  • Nvidia is selling “more than a card or a chip”: a complete system spanning GPUs, CPUs, networking, storage, CUDA, Agentic AI and inference deployment. The customer’s procurement question is shifting from which chip to choose to whom to buy the entire AI factory from.

10. Inference turns compute demand into recurring cash flow

  • Zhang describes training as “one-off cash flow”: a model consumes large amounts of GPU capacity in a concentrated phase. Inference continues as long as the product is used, and becomes “recurring cash flow” once Agents stay online and repeatedly call tools and data.
  • She recalled that early discussions put the compute mix at roughly 70%-80% for training and 10%-20% for inference. Today it is close to an even split; in the future it could invert to 20%-30% training and 60%-70% inference.
  • Jensen’s outlook for Blackwell and Vera Rubin rests on the same assumption: by 2027, the associated data-center revenue could exceed $1T, provided token consumption and inference workloads far outstrip the training demand seen in the early phase.

11. Agents put CPUs back at the center of AI compute

  • Previous industry research by Fusion Fund found that some new model architectures are more efficient on CPUs than GPUs. Zhang therefore expects future AI training and deployment to rely on multiple compute architectures rather than a single one.
  • Agents do more than generate tokens: they continuously call tools, execute code and coordinate multiple Agents. Reinforcement learning and simulation also rely heavily on CPUs. As Agents shift from occasional Q&A to continuous operation, CPU demand will rise in parallel.
  • Nvidia launched Vera CPU this time as a processor for Agentic AI and reinforcement learning, with claimed efficiency gains of 2x over traditional CPUs. Initial partners include Alibaba, ByteDance, Oracle and Meta.
  • CPU control also has platform significance. If the CPU remains in someone else’s hands, Nvidia cannot fully define system- and rack-level performance. Even if AMD eventually produces the stronger individual CPU, Nvidia’s one-stop procurement and integration advantages could remain substantial.

12. Groq and 2 early acquisitions round out Nvidia’s platform

  • Fusion Fund had 5 companies acquired last year, 2 of them by Nvidia: Lepton AI, founded by 贾扬清, and Nexusflow, founded by 焦建涛. Both companies were less than 2 years old, and Nvidia moved quickly from initial contact at the start of the year through acquisition and internal integration.
  • Zhang said Lepton subsequently entered Nvidia’s DGX platform as part of its GPU cloud strategy. She could not discuss the integration details, but rapid website redirects and product launches showed Nvidia’s ability to “pull capabilities into the ecosystem the moment it sees them.”
  • Groq was founded in 2016 by former core members of Google’s TPU project. Rather than patching the GPU, it redesigned the inference path for low latency and high token throughput. Its advantage is limited by specific model sizes, but it complements the general-purpose training and inference capabilities of GPU plus CUDA.
  • Reports cited by the program put the transaction value at roughly $20B, with a special structure combining non-exclusive technology licensing and talent acquisition. Zhang said Groq’s chip alone “was not worth the price Nvidia paid”; its value was completing Vera Rubin’s story across CPU, GPU, inference accelerators, networking and storage.

13. Full-stack competition will ultimately extend to world models and real 3D data

  • Zhang sees Google, Nvidia and Apple as the companies with relatively complete full-stack capabilities today. Apple lacks its own AI models, but has an integrated stack of chips, products and endpoints. Meta and others are investing and acquiring to build up their chip capabilities.
  • She summarizes model evolution as language models, multimodal models, Agents and then world models. The final stage requires not only models but also high-quality 3D data from the real world.
  • That is why she is bullish on Musk’s ecosystem: Tesla provides transportation, factories and real-world visual data; SpaceX has engineering, satellite and space data; humanoid-robot data may be added in the future. Other companies may have only video, while Musk has diverse “real 3D data of the world.”

14. xAI’s personnel turmoil reflects Musk’s speed—and its cost

  • Zhang believes the steady departures of co-founders may indicate that model capability has improved more slowly than Musk expected. His goal is not to catch up but to surpass, so he adjusts rapidly once he believes the direction is wrong. She uses “done is better than perfect” to describe the adjustment logic of startups.
  • He is remarkably patient when recruiting: one co-founder known to Zhang was persuaded by Musk for 3 consecutive years. After joining, the team often worked until 5 or 6 a.m., because Musk might arrive at the office at 1 or 2 a.m. to continue brainstorming.
  • Her condensed assessment is “a charismatic tyrant”: his vision, drive and recruiting ability are all exceptionally strong, but he places the goal above personal relationships, gains and sacrifices. “If he thinks it isn’t working, he adjusts immediately.”
  • Zhang is not bearish on xAI as a result. After being folded into SpaceX, it can draw on talent and resources from Tesla and SpaceX, as well as real-world, satellite and space data; it is no longer an ordinary startup. If SpaceX is eventually lifted by the space economy, xAI could accelerate with it.

15. Google TPU’s advantage is mainly inside Google’s full stack

  • Google has invested in TPU for more than 10 years, showing that it recognized the importance of inference early. TPU is jointly optimized with Google’s models, data, cash flow and application feedback, forming an important foundation for Gemini’s rapid improvement.
  • Zhang recalled that Google’s internal use of TPU put training costs at roughly one-third of OpenAI ChatGPT’s. The key was not an isolated chip victory, but system-level performance and cost optimization.
  • Third parties lack Google’s complete technology stack, so TPU performance and costs both suffer outside it. Manqi added that Google could gradually expand the developer base familiar with the architecture through software systems such as JAX and support for Google alumni launching startups.
  • Zhang still does not see TPU posing an effective short-term threat to the GPU. AI penetration in finance, healthcare and insurance may still be below 1%, leaving enough total demand for multiple architectures; Nvidia also uses CUDA and a more complete platform to raise switching costs.

16. Enterprise AI is opening budget, revenue and exit channels at the same time

  • Zhang sees a consensus that highly regulated industries including finance, healthcare and insurance need vertical small language models deployed on-premises or in private clouds. They prioritize low latency, fewer hallucinations, data privacy and cost optimization over handing all data to a general-purpose frontier model.
  • Data alone is not yet an asset. Enterprises must first perform data curation, then add security and privacy middleware before connecting to AI applications. Fusion Fund’s CXO network includes roughly 45 CTOs from Global 1000 companies, with the largest enterprise AI budget reaching $12B; some financial and insurance companies can integrate new AI technology in 3-4 months.
  • Healthcare is one of Zhang’s most favored verticals. US healthcare accounts for roughly 20% of GDP, Eli Lilly and Nvidia have reached a multibillion-dollar partnership, and ChatGPT and Claude have launched healthcare-specific products. High-quality data, small models and federated learning can reduce institutions’ concerns about data privacy.
  • Finance, healthcare and insurance may integrate AI fastest, while logistics, supply chains and industrial Physical AI are also expanding. Some B2B teams with fewer than 10 employees grew annual revenue from zero to $20M because traditional enterprises prefer working with startups rather than handing core data to large technology platforms.

17. The US market amplifies small-team value through rapid procurement and M&A

  • Zhang attributes the concentration of global AI talent in the US to commercialization efficiency. Traditional banks such as JPMorgan Chase are integrating AI quickly, while Koch Disruptive Technologies can bring technology into its portfolio companies, sometimes in cycles of only 1-2 months.
  • Fusion Fund had 5 companies acquired last year, 4 of them founded less than 2 years earlier, with some deals generating 10-20x returns. Zhang would rather see companies scale over the long term, but this “short, fast” exit cycle recirculates capital, talent and serial founders.
  • Manqi closed by describing organizational forms at 2 extremes: Tesla, SpaceX and Nvidia are moving toward extreme vertical integration. She also cited individual developer Peter Steinberg’s OpenClaw, arguing that an individual or tiny team building a product and then being absorbed by a major platform is one of the more predictable startup opportunities; the episode’s postscript said Peter joined OpenAI soon after OpenClaw took off.
  • She also used Groq’s roughly 200 employees and $20B transaction to estimate that “each employee was effectively worth $100M.”
  • Zhang will continue investing through Fusion Fund’s fourth fund: after backing 10 companies last year, she plans to invest in 7-10 more this year and support 3 companies toward IPOs. Her personal focus remains inference costs, energy, memory, security and system integration, alongside medical robots and micro- or nanorobots; the space economy remains a key area to monitor and invest in over the next 2-3 years.