Huawei's AI Supernode: Market Impact with Magic Compute's Xu Lingjie
Summary
The significance of Huawei Cloud Matrix 384 is not that a single Ascend 910 catches up with the most powerful GPU, but that 384 chips and systems engineering push aggregate capability into NVL72’s competitive range. The 12-rack system delivers roughly 300P of compute. Xu Lingjie estimates single-chip FP16/BF16 performance at about 780T FLOPS, or roughly 80% of H200’s static performance. A post-recording update said more than 6,000 Ascend chips had also completed long-term, stable training of a 718B MoE model. “Each component is not that powerful on its own, but stronger cluster capabilities build the overall result.”
Strictly speaking, Cloud Matrix 384 is closer to a scale-out cluster, while NVL72 is closer to a scale-up supernode that fully interconnects 72 GPUs within a single rack. Huawei’s solution uses 12 racks, 6,912 400G optical modules and 3,168 fiber strands to achieve scale; overseas media estimate potential power consumption at roughly 4x NVL72’s. Post-recording reports put the full system at about $8M, versus roughly $3M for NVL72, though the former is not an official Huawei quote. China’s electricity prices and cost structure could still make this high-power route viable: “This is a path US vendors have not taken, but Chinese vendors can.”
The core bottleneck in AI compute is shifting from individual chips to memory, interconnects, and cooling, power delivery and coordination efficiency at the data-center level. Storing DeepSeek’s 671B parameters in 8-bit format alone requires 671GB; with KV cache and other overhead, usage often exceeds 1,000GB, above the 640GB available in an 8-GPU H100 server. Longer contexts and deeper reasoning will amplify the pressure. “Bigger problems need bigger systems,” making NVLink, NVSwitch, HBM, liquid cooling and power delivery increasingly important.
Inference economics increasingly depend on large-scale MoE expert parallelism, rather than merely squeezing a model into a few cards. DeepSeek spreads experts across hundreds of cards and raises utilization through high concurrency. Xu says his team has made the full-size DeepSeek profitable as an API business on hundreds of H20s, but did not endorse the host’s cited 545% margin. His estimates even show H800 compute at roughly 5—6x H20’s, for only 2—3x the price, potentially making unit compute cheaper in the cloud.
Huawei’s supernode will have limited near-term impact on Nvidia overall, but it will first reorder the淘汰顺序 of domestic AI-chip vendors. Nvidia’s problem in China remains one of supply and policy—“the product can’t get in.” Domestic vendors, meanwhile, will be compared by customers on hundreds-of-card scaling, throughput and unit token cost. “NVIDIA may not be that scary; Huawei is scarier.” Vendors without supernode reserves or architectures suited to scaling could be eliminated.
Nvidia’s moat over the next two to three years will remain the closed loop formed by CUDA, NVLink/NVSwitch and the world’s most advanced supply chain. TSMC, the latest HBM, Supermicro and Foxconn allow it to design continuously “right up against the global technology limit.” The real industry threat comes from customer concentration: Google, Amazon, Microsoft and others provide massive purchases while possessing the scale to develop TPU, Trainium and Maia in-house. Nvidia’s semi-custom division may also be intended to keep these customers tied to it beyond standard products.
The next major industry variable may emerge first in energy, optical interconnects and new computing paradigms—not simply in another faster GPU. Scale-up could reduce cross-rack optical-module demand while creating opportunities in CPO, liquid cooling, high-voltage DC power and waste-heat recovery. Supernodes may lower unit token costs and energy use, but total electricity consumption is still rising. Xu calls human sustainability Nvidia’s biggest soft spot and warns against ruling out near-memory computing, neuromorphic computing and spiking neural networks too early: “When it disrupts you one day, it will be too late.”
Deep dive
1. Cloud Matrix 384 Uses System Scale to Close the Single-Chip Gap
Cloud Matrix 384 comprises 12 racks, with four 8-GPU servers per rack, for a total of 384 Ascend AI chips. Six racks sit on each side, with switching equipment in the middle. The system is interconnected through 6,912 400G optical modules and 3,168 fiber strands.
Xu Lingjie’s detailed figures put total compute at roughly 300P, 67% above NVL72, with approximately 2x the network interconnect and aggregate memory bandwidth. Back-calculating from the cluster data implies about 780T FP16/BF16 FLOPS and roughly 128GB of HBM per chip, along with eight 400G interconnect ports.
Single-chip static metrics are about 80% of the previous-generation H200, so the point is not to match it card for card, but to “trade quantity for aggregate compute.” Xu praised Huawei for “blazing a new trail for China” under restrictions, then immediately added that it may not yet be a supernode in the strict sense.
The episode introduction added two post-recording developments: reports put the complete system’s price at about $8M, versus roughly $3M for NVL72, though the figure is not an official Huawei price; Huawei’s Pangu team also said it had completed long-term, stable training of a 718B MoE model on a cluster of more than 6,000 Ascend chips.
2. The Scale-Out/Scale-Up Distinction Determines What “Supernode” Really Means
Xu uses a beverage shop as an analogy: replacing a small blender with a larger, more powerful machine is scale-up; buying several identical machines to handle rising traffic is scale-out. Both expand compute, but their internal coordination efficiency is entirely different.
NVL72 places 72 chips and roughly 18 servers in a single rack, using high-speed switching to create full point-to-point interconnectivity. It represents a shift from the chip and server levels to the rack level. Cloud Matrix 384 spans 12 racks and is closer to a high-density scale-out cluster.
Cheng Manqi retained Huawei’s product name while accepting Xu’s technical distinction: Huawei officially calls it a supernode, but if a “single large node” means shared high-speed interconnect and memory capabilities, NVL72 has the stronger scale-up attributes.
3. The Data Center Is Becoming a Computer
Xu’s broader framework is “data center as a computer,” or even “data center as a GPU”: once a task can no longer be handled by a single chip or server, switching, power, cooling, heat dissipation and software optimization become part of the computing system itself.
Stacking 72 cards with ordinary network adapters produces a very different result from connecting 72 cards through a dedicated high-speed network. GPUs remain the most expensive and critical components, but they can no longer determine cluster performance on their own.
Xu recalled that when he joined Nvidia in 2008, the company was still primarily a PC graphics business and a “nobody” in high-performance computing and CUDA. In his memory, early NVLink appeared around the V100 generation in 2017. After acquiring Mellanox, Nvidia further strengthened its switching capabilities and gradually evolved from a GPU company into a data-center company.
4. The First Wall Large Models Hit Is Memory, Not Just FLOPS
Storing DeepSeek’s 671B parameters in 8-bit format consumes roughly 671GB by itself. Add KV cache and working space, and actual memory usage often exceeds 1,000GB. An 8-GPU H100 server has only 8×80GB, or 640GB—at least two servers are needed just to hold the parameters.
Longer contexts expand KV cache, while deep reasoning increases context requirements. Xu condensed the industry’s central problem into one sentence: “Bigger problems need bigger systems.”
A post-episode note used another set of figures to underscore the imbalance between compute and memory: from the A100 in 2021 to the B200 in 2025, GPU compute rose 64x, while memory capacity increased only 1.2x. The conclusion was that memory is becoming the harder constraint.
5. NVSwitch Turns Communication Links into Part of the Compute
Cheng Manqi asked about the relationship between the “bus,” NVLink and NVSwitch. Xu said there was no need to get trapped in terminology: whether called a bus or a high-speed interconnect protocol, the goal is to let chips communicate, synchronize and share data at high speed.
NVLink provides high-speed point-to-point connections, while NVSwitch organizes multiple links into a larger switching network. Based on Nvidia’s roadmap at the time, at least 144 GPUs could eventually be connected into a fully interconnected system.
The more important shift is in-network computing. Results distributed across GPUs need not be sent back layer by layer to a single GPU for aggregation; they can be merged and computed on the central switch chip as they move through the network. “In the current environment, compute and interconnect are becoming increasingly inseparable.”
6. The DPU Story Is Turning into an RDMA-NIC Opportunity
Xu helped Biren invest in DPU company Yunmai Xinglian in early 2021. What began as a strategic investment later looked more like a financial investment and exited profitably. He says the company should already have RDMA-NIC products on the market.
His view is not that the DPU concept has disappeared, but that demand has become more specific. As distributed training and inference drive up communication volumes, the shift from DPUs toward RDMA NICs and high-speed interconnect products could still represent a “very large opportunity.”
7. Broad AI Infra Leaves Room for Hardware-Software Startups
Narrowly defined, AI infra usually means software optimization to run models faster and more cheaply. Xu’s broader definition covers the middle layer between chips and applications, including servers, racks, supernodes and hardware-software coordination.
Magic Compute does not develop its own chips. It plans to start with overseas chips, building non-reference, high-density integrated systems, and then work with domestic chips on large-scale clusters. Xu handles chips and servers, while co-founder Jin Chen focuses on software, systems and algorithms.
The opportunity Xu saw at the end of 2023 was that as China’s chip competition shifted from a more market-driven model toward one shaped more by resources and policy, differentiation remained possible across the supply chain. The emergence of NVL72 and DeepSeek V2 in 2024 convinced the team that “large systems” could be a durable entry point, leading it to begin building prototypes.
8. Training Clusters Sell Certainty; Inference Clusters Sell Unit Token Cost
A cluster cannot be judged by peak compute alone. Memory capacity and bandwidth, supported precisions, process technology, power consumption, heat generation and cooling all matter. Xu compares it to cars: a sports car optimizes for top speed, while a family buyer may care more about space, price and features. The product must match the use case.
Training prioritizes stability, scale and predictable outcomes. By Xu’s recollection, Llama 3 was trained on more than 15,000 H100s, with hundreds of incidents over two months—on average, a checkpoint recovery could be required every two to three hours.
The larger the cluster, the greater the probability that a card, network or power failure interrupts work. Reliability is therefore no longer a secondary metric. “Reaching theoretical peak” and “completing training continuously for weeks” are two different capabilities.
Inference is an operating expense, and customers ultimately compare TCO: after machines, electricity and operations are combined, how low is the cost per token? Xu categorizes training as capex and inference as opex, explaining why the same hardware may not be equally competitive in both businesses.
9. DeepSeek Turns MoE Inference into a Cluster-Utilization Contest
Cheng Manqi cited DeepSeek Open Source Week’s disclosed 545% cost margin and asked how much more third parties could optimize. Xu did not verify the figure. He only said his team had achieved good results with the full-size DeepSeek on H20s and made the API business profitable.
DeepSeek’s key insight is not fitting the model onto a few cards, but running large-scale expert parallelism across hundreds of cards, particularly during decode. Experts are spread out, with each card holding only a subset, while high concurrency keeps every card working continuously.
“If you only barely get it running, cluster utilization will be quite low.” Xu sees high utilization as the source of DeepSeek’s profitability and says the team is continuing to tune the system on dozens of machines and hundreds of H20s. “The result is improving every day.”
10. MoE Methods Are Reusable, but Static Configurations Are Not
Cheng Manqi used two models to frame the question: Qwen3’s flagship open-source model has 235B total parameters and 22B active parameters; DeepSeek has 671B total parameters and 37B active parameters. With different sizes and expert configurations, does every model require a new inference design?
Xu’s answer is that the direction is the same: distribute the experts and raise per-card utilization through high concurrency. Different models change the operators, number and size of experts, and allocation ratios, but theoretical analysis can first identify a strong starting point, followed by rapid tuning through experiments and experience.
Unified training and inference broadly works for large-scale cloud businesses, but it is not every customer’s answer. An offline deployment with a budget of only RMB2M—3M, enough for one or two servers and with low concurrency, will prioritize capacity, procurement thresholds and local requirements differently.
More counterintuitively, Xu estimates that H800 inference economics may be better than H20’s. Once experts are distributed, the bottleneck shifts from capacity to compute: H800 compute is roughly 5—6x higher, while its price is only 2—3x higher, making unit compute cheaper.
11. Google Turned AI Chips into Pods Before Nvidia Did
Google’s release of its first TPU in 2016 made a strong impression on Xu: AI was no longer a “toy” in the eyes of Silicon Valley engineers; it had to “enter the mainstream and produce real productivity.” That pushed him to move from Samsung to Alibaba Cloud, shifting from chip design into cloud infrastructure.
TPU V2 appeared as a 256-chip pod, V3 expanded to 1,024 chips, and V4 was, in his memory, 4,096 chips. V4 and V5 also used OCS (Optical Circuit Switch) to configure the network dynamically. Google recognized early that chips, networks and clusters were a single product.
Nvidia later launched SuperPOD, but during the generative-AI wave, the product that truly focused industry attention on scale-up was NVL72, released in 2024. “From a single chip to a server, and then to an entire rack” captures the central hardware paradigm shift of this cycle.
12. HBM Has Gone from Supporting Cast to GPU Cost Driver
Memory capacity expanded from the A100’s 40GB and 80GB to the 144GB Xu cites for H200, and is moving toward 192GB and 288GB. The larger and more expensive the memory, the more the system needs high-speed networking to pool and fully utilize it.
Memory accounted for roughly 40%—50% of chip cost in the H100 generation and may rise to 50%—60% with B200. The cost structure is shifting from compute wafers to memory. That is why supernodes pursue not only compute, but also aggregate memory and bandwidth.
Global HBM supply is dominated by SK Hynix, Samsung and Micron. Some domestic vendors can already substitute for products such as LPDDR, but HBM remains an area requiring catch-up and affected by export controls. Asked whether Huawei can produce HBM itself, Xu answered directly: “I can’t comment on that.”
13. Huawei Trades Higher Power Consumption for Aggregate Capability
By Xu’s comparison, Cloud Matrix 384 delivers roughly 300P of compute, 67% above NVL72, with more than 3x the memory capacity and more than 2x the bandwidth. The cost, according to overseas media estimates, could be potential power consumption roughly 4x higher.
This is a choice shaped by regional cost structures. Based on Nvidia GPU economics, electricity accounts for about 10% of annual opex in China and can be lower in some western cities. Overseas data centers often face constraints from both electricity prices and grid capacity.
Cheng Manqi summarized the approach as “adapting to local conditions.” Xu said it reflects the cost structure: if hundreds of cards can run highly parallel workloads and the system is fast enough, lower electricity prices can offset a meaningful share of process and design disadvantages.
His description is “stacking building blocks and piling up components.” Individual components may be weaker, but clusters and optimization can create overall competitiveness. This path may not replicate the US model, but could better fit China’s current manufacturing, energy and supply conditions.
14. Building a 384-Card Scale-Out Cluster Is Not Rare; Full-Interconnect Scale-Up Is Hard
Xu believes that as long as a chip has enough switching ports and a company can find partners for NICs, switches and integration, many domestic chip vendors could assemble a Cloud Matrix 384-like scale-out cluster.
He knows several vendors working on 352-card or 384-card clusters and tuning performance. The real differences lie in network speed, switching capability and topology. Without full interconnectivity, efficiency in some applications will still “take a hit.”
Huawei’s current solution therefore both proves that scale is feasible and raises customer expectations. Vendors can no longer show only single-card benchmarks; they must demonstrate stable coordination across hundreds of cards, along with throughput and unit cost.
15. Huawei Had a “Mellanox on the Network Side” Before Adding Compute
Huawei’s most relevant accumulation for supernodes is not just traditional communications equipment, but more than a decade of work and talent in network processing units and switching. Xu’s analogy is: “Before Huawei became Nvidia, it was already a Mellanox.”
NVSwitch’s core challenges include high-speed SerDes interface IP, advanced process technology and the organizational ability to co-design compute and switching chips. Xu calls NVLink/NVSwitch Nvidia’s second-most important ecosystem moat after CUDA.
He believes domestic vendors had generally not yet built out comparable NVSwitch products at the time and might not launch them over the next two or three years. Huawei theoretically can—and “may already be doing so.” Another route would be to partner with companies that control high-speed interfaces and switching, as Google did with Broadcom.
16. Huawei’s Supernode Will Hit Domestic Peers First; Nvidia Is a Separate Battle
Xu expects Cloud Matrix 384 to have limited short-term impact on Nvidia: “You fight your fight, I fight mine.” The issue in China is not a lack of Nvidia demand, but whether products can enter the market given supply and policy constraints.
The impact on domestic chips is more direct. DeepSeek has accustomed customers to freely trying a full-size 671B model. Similarly, once customers see that a 384-card cluster can deliver several times—or even an order of magnitude—the cost-performance of a small cluster or single machine, it will be difficult to sell products limited to single-card capability.
His conclusion is deliberately sharp: “NVIDIA may not be that scary; Huawei is scarier.” Once Huawei turns supernodes into a procurement standard, domestic vendors whose architectures do not scale and who lack system-level reserves may be eliminated first.
17. The Direct Cost of the H20 Ban Is $5.5B; Reworking Inventory Depends on Fuses
At the time of recording, new H20 restrictions had forced Nvidia to provision $5.5B, and the stock had reacted. The H20 is mainly intended for China; overseas customers that can buy H200 generally have no reason to choose the restricted version.
Xu characterizes H100, H800, H200 and H20 as different configurations built around similar chips. H20 compute is roughly one-sixth of H200’s, with 8-bit compute held below about 300T to stay under export thresholds, alongside the 2,400-bit-ops restriction criterion.
If some compute units are merely configured off rather than permanently blown, Nvidia may be able to use board-level settings and fuses to convert inventory into another overseas product. Chinese customers would find such “modifications” difficult themselves, especially because they would also need to break the encryption and configuration system involving the on-board RISC-V core. “The technical threshold is extremely high.”
18. Zero-Day Adaptation Only Starts the Car; Throughput and Precision Determine Usability
Huawei announced “out-of-the-box, zero-day adaptation” hours after Qwen3’s release. Xu believes getting a comparable model to run is not difficult. The real competition is how much concurrency can be supported while maintaining tokens per second per user, and the total throughput delivered under stable service.
His analogy is: “Getting the car started isn’t hard; accelerating smoothly on the highway is.” The latter determines inference economics.
DeepSeek’s official full-size model uses FP8. Some domestic chips lack native support and must fall back to FP16/BF16, requiring more memory and bandwidth, or quantize to INT8 and accept potential precision loss. So systems may all be called DeepSeek while producing different answer quality in some environments.
The difference between FP8 and INT8 is both a matter of format design and ecosystem choice. Nvidia can push FP6, FP4 and other lower-precision formats into applications, while followers often arrive half a step late. Export controls may force China to develop a separate format ecosystem, creating risks as well as new opportunities.
19. Nvidia’s Lead Is Built on the Global Supply Chain
Xu believes the probability of Nvidia’s global position being technically weakened over the next two or three years is low. Its chip roadmap is strong, CUDA is mature, and NVLink/NVSwitch closes the loop across compute, NICs and switching.
TSMC, SK Hynix and Samsung can prioritize the latest process technologies and custom HBM, while Supermicro and Foxconn handle complex systems manufacturing. Nvidia can consistently obtain technology that is “faster, earlier and better,” designing products right up against the global “technology limit.”
This systems capability also underscores the importance of full-rack delivery. Xu noted that AMD spent roughly $4.9B to acquire ZT Systems last August, adding rack-scale and supernode delivery capabilities.
20. The Customers That Could Truly Weaken Nvidia Are Its Richest and Most Concentrated
North American hyperscalers and Chinese internet giants before the restrictions may have accounted for more than 60% of Nvidia’s purchases. Nvidia’s data-center gross margin often reaches 70%—80%, so even massive investment in in-house chips can make economic sense at that procurement scale.
Google has TPU, Amazon has Trainium and Microsoft has Maia, with help from Broadcom’s design and high-speed SerDes capabilities. In-house development not only cuts costs, but also avoids the long lead times and supply uncertainty that emerged when ChatGPT first exploded.
Cheng Manqi asked: if these chips have been in development for years, why have they still not replaced Nvidia? Xu’s explanation is that TPUs and similar products primarily serve internal needs and were not organized as independent chip companies. As cloud providers, they must also be “customer first”: if OpenAI wants Nvidia, Microsoft still has to buy Nvidia.
Nvidia has created a semi-custom division that may offer major customers portions of its IP and design capabilities. The host worried this could create internal cannibalization. Xu sees it more as a defensive bargaining tool for key accounts, while acknowledging that there is no clear success case yet: “Maybe we filled in too much of the picture ourselves.”
21. Scale-Up Will Reallocate Value Across Optical Modules, Liquid Cooling and CPO
Xu offered an “immature speculation”: if more inference workloads can be handled inside single-rack scale-up systems, cross-rack communication may decline, reducing fiber and optical-module consumption. Cloud Matrix 384 still relies on large numbers of optical modules and fiber strands, reflecting its different path from NVSwitch-style full interconnect within one rack.
Cheng Manqi mentioned the market saying that “Nvidia has Chinese vendors by the throat, while Zhongji Innolight has Nvidia by the throat.” Xu did not comment on specific companies, suggesting instead that the variable be tracked to see how demand shifts as system topologies change.
As interconnect requirements continue to rise, opportunities may move from pluggable optical modules toward faster, more integrated CPO. At GTC 2025, Nvidia introduced a Co-Packaged Optics direction based on CPO. The more chips are consolidated into a rack, the more attention will go to co-packaged optics, liquid cooling and high-density power delivery.
22. Nvidia Is Moving Up from Selling Chips toward the Inference-Service Layer
Asked about Nvidia’s acquisition of Lepton AI, Xu declined to make factual judgments because its founder was in a quiet period, offering only the “bold guess” that it might be related to Dynamo. Dynamo is an open-source modular framework for distributed multi-GPU inference, positioned above the TensorRT inference framework.
One of Dynamo’s key features is PD disaggregation: placing prefill and decode in separate clusters, each optimized independently. If Nvidia controls this scheduling layer, it will not merely sell GPUs; it will also directly influence customers’ service throughput and unit token costs.
DGX Cloud is not simply a copy of public cloud. In one example, Oracle buys a batch of Nvidia GPUs and leases them back to Nvidia; Oracle handles the low-margin machine operations, while Nvidia provides higher-value services on top of the cluster. Nvidia wants to move closer to end customers while managing its “frenemy” relationship with hyperscalers.
23. Nvidia’s Success Was Not Written in Stone in 2010
Xu left Nvidia for AMD at the end of 2010 because he had the chance to follow chief architect Mike Mantor in designing a new generation of Compute Units. At the time, the two companies’ GPU shares were roughly 50/50 or 60/40, and AMD still held clear advantages in die area and power consumption.
He does not believe Jensen Huang foresaw today’s outcome as early as 2012—2014. Xu recalls that Huang had pledged roughly 5% of the company’s shares as convertible-debt collateral to Goldman Sachs and Wells Fargo, with an exercise price, if memory serves, slightly above $20 before the split. “There is no certainty in life. Today’s Jensen is not the same Jensen as ten years ago.”
But long-termism and crisis leadership were present early. People continued to follow Huang after Fermi failed tape-out twice. During the 2008 financial crisis, Nvidia cut headcount by 5% and salaries across the company by 10%, while Huang took only $1. “It is not just saying those things; it is being able to lead everyone out of a crisis.”
24. Energy Constraints May Rewrite Computing Paradigms Before New Chips Do
Xu calls human sustainability Nvidia’s biggest soft spot. Data centers already account for roughly 2%—3% of global electricity use, and AI adoption will continue to push the total higher. Supernodes can lower the cost per token and per unit of output, but “lower unit cost” does not mean lower total electricity consumption.
He sees supernodes as the “only possible path” for extending Moore’s Law at the rack and data-center levels. The supporting infrastructure must be redesigned end to end: 800V and 1,600V high-voltage DC power can reduce AC/DC conversions; lower current reduces line losses; phase-change liquid cooling can centralize waste-heat recovery.
Further out lie near-memory computing, neuromorphic computing and spiking neural networks. Cheng Manqi relayed Yann LeCun’s cautious view that Bell Labs had been researching neural-network simulation hardware since the 1980s, but it had not proved particularly effective over the long term.
Xu does not deny the historical failures, but refuses to use them to rule out the future: “History keeps repeating itself, only in different forms.” Reinforcement learning, neural networks and DeepSeek all underwent sudden breakthroughs from old lines of research. With new paradigms, “when it disrupts you one day, it will be too late.”