Pioneers Insight Method Research Author
29. Inside 2025 GTC (Part II) | The Weight of the Crown: How 黄仁勋 Answered Every Challenge
Back to Episodes

29. Inside 2025 GTC (Part II) | The Weight of the Crown: How 黄仁勋 Answered Every Challenge

Summary

  • At 2025 GTC, Nvidia shifted its core growth narrative from training to inference. 姚欣 said the industry had recognized by June-July 2024 that the marginal returns from training-side scaling laws were weakening. DeepSeek turned that view into mainstream consensus: the next battleground would be cost-performance, applications and token costs. 黄仁勋 barely mentioned training, instead repeating “token token token” and arguing that inference still requires massive compute—roughly 20x the tokens for a single answer, and potentially 100x the compute.
  • Behind Blackwell Ultra, Rubin and the 3-year roadmap was the capital market’s demand that Nvidia prove shipments and growth immediately. Early GB200 products had reportedly raised thermal and system-stability concerns, while DeepSeek triggered questions about whether the architecture itself was changing. Robots therefore gave way to nearer-term delivery commitments: major cloud providers bought roughly 1.2 million Hopper GPUs last year and have now ordered about 3.6 million Blackwell GPUs. “Tell me how many units you can ship this year, and how many next year” is the pressure point. 姚欣’s summary: “Nvidia is in a rush”(英伟达急了).
  • The key impact of DeepSeek is not the disappearance of compute demand, but whether Nvidia can replicate its training-era monopoly in inference. 卫诗婕 cited 袁金辉’s metaphor: “The compute required for training is a swimming pool; the compute required for inference may be a river.” Compute in that river is more distributed and scheduling-intensive, creating room for cost-effective challengers such as AMD. The post-GTC share-price decline shows the market has yet to accept the full “buy more, save more” argument.
  • Dynamo is Nvidia’s most concrete software answer to the inference era—and may also be a rapid catch-up effort prompted by China’s open-source practice. It packages the data center as an “AI factory” that produces tokens, with its introduction featuring terms such as PD disaggregation and expert parallelism, while managing clusters and optimizing deployment and utilization. 姚欣 sees its approach as broadly parallel to the components DeepSeek presented during its February open-source week. 卫诗婕 also speculated that parts of the architecture resemble Kimi’s Mooncake; 姚欣 said open-source projects may borrow from one another. NVL72 sells a cluster of 72 GPUs, while Dynamo adds deployment, scheduling and utilization optimization.
  • What DeepSeek demonstrated was the ability to extract gains through systems engineering, not a way around Nvidia. High-Flyer’s quant background led the team to focus on low-level efficiency for years. By calling PTX, an assembly-like language inside CUDA, it pushed Nvidia hardware harder—but may also have made the code more “Nvidia-only.” 姚欣 rejected the idea that only a handful of people globally can do this, arguing that many teams with more than 10 years of infrastructure experience have similar capabilities.
  • CUDA’s ecosystem moat remains formidable, but the consolidation of foundation models is lowering the cost of porting across hardware. Nvidia once cultivated developers by giving away cards, placing engineers on-site and building out cuDNN, creating a default path of “get it working on CUDA first, then port it elsewhere.” If foundation models converge around a handful of systems, vendors such as AMD will find it much easier to support them one by one. 姚欣 therefore expects the software moat to erode over a 10-year horizon; 80-90% gross margins and 25-30x P/E multiples will be difficult to treat as permanent.
  • The four stages—perception, generative, Agentic and physical AI—keep pushing Nvidia’s long-term TAM into the physical world. 姚欣 compares them with biological evolution: first senses, then the brain, then digital limbs and finally physical limbs. 卫诗婕 sees a complete reflex chain of perception, thought, agency and action. If physical AI is to simulate friction, bouncing and every physical law, the ultimate constraint is energy rather than imagination: “Nvidia’s ultimate adversary is the laws of physics.”
  • China is both a major market Nvidia has lost and a source of technological pressure forcing Silicon Valley to iterate faster. Based on his reading of financial filings, 姚欣 said China once accounted for roughly 24% of Nvidia’s global chip sales; 卫诗婕 added that the share fell below 10% after the sales restrictions, which 姚欣 attributed to the large number of chips barred from sale. Apart from Lenovo, there were almost no Chinese booths at this GTC, and China AI Day moved online. Silicon Valley is nevertheless scrambling for Manus invitations and beginning to acknowledge China’s advantages in cost-performance, applications and data: “Although nobody talks about China, China has shaken Silicon Valley”(“虽然大家不讲中国,但是中国对于整个硅谷产生了冲击”).

Deep dive

1. Nvidia’s core franchise is betting early on where massive compute will be needed

  • 姚欣 traces the story back to around 2000. GPU originally stood for Graphics Processing Unit, but scholars including Bill Dally proposed GPGPU, applying parallel architectures to video, scientific computing and biological simulation. 黄仁勋 recruited Dally as chief scientist, taking Nvidia from graphics cards into general-purpose computing.

  • Roughly 15 years ago, Nvidia tried putting an entire chip into mobile devices, but power consumption forced it toward handheld gaming. The technology later found uses in Switch, automotive applications and early autonomous-driving simulation. In 2016, 黄仁勋 also donated AI servers to then-startup OpenAI, with Elon Musk taking delivery—a snapshot, 姚欣 says, of a company that keeps betting on the future.

  • GTC is therefore less a product launch than a technology outlook conference, and is sometimes mocked as a “futures launch”: announced today, perhaps visible next year. 黄仁勋 is not selling a single chip so much as the next set of use cases that will consume enormous amounts of compute.

  • 姚欣 uses the phrase “what splits apart eventually comes together, and what comes together eventually splits apart” to describe the underlying cycle. IBM mainframes centralized computing; PC clusters dispersed it; Google Cloud centralized millions of servers again. Autonomous driving, IoT, robotics and on-device models may push compute back toward the edge. Nvidia has captured the 20-year through-line of parallel, distributed and accelerated computing.

2. Four kinds of AI push the growth endpoint into the physical world, but energy sets the ceiling

  • 黄仁勋’s sequence is perception AI, generative AI, Agentic AI and physical AI: from recognizing faces and voices, to ChatGPT-style generation, to agents that complete tasks in the digital world, and finally to agents entering the physical world through robots.

  • 姚欣’s personal interpretation is “four stages in the evolution of a living organism”: first senses, then a brain and nervous system, then limbs in the digital world, and finally physical limbs. 卫诗婕 offers another framing—receiving information, computing and thinking, forming an agent that acts, and ultimately changing the physical world.

  • The ambition of physical AI is to simulate every physical law, including bouncing, reflexes and friction. That leads to the pair’s ultimate question: “Is what we call reality actually real, or is it a simulation?” The commercial constraint is more immediate. Cooling alone can account for roughly 40% of the energy consumed by a large data center, so Nvidia must increase token output by 10x or 100x within a fixed energy budget.

3. Robots give way to a shipment roadmap because 黄仁勋 has to answer immediate doubts

  • The pair had expected the conference to devote substantial time to humanoid robots, but robotics received less attention than anticipated. 姚欣 says Nvidia has not abandoned physical AI; investors simply do not want a 10-year story right now: “Tell me how many units you can ship this year, and how many next year.”

  • GB200, previewed a year earlier, has begun shipping, but initial units reportedly raised thermal and system-stability concerns. DeepSeek also prompted the market to question whether the compute architecture itself was changing. 黄仁勋 responded by concentrating on Blackwell Ultra and formally naming the next generation Rubin, previously referred to as X100.

  • More unusually, Nvidia laid out products and delivery timelines for the next 3 years in a single presentation. 姚欣 repeatedly stressed that this was a sign of pressure: “Nvidia is in a rush.” The roadmap was also a written answer for capital markets on the US East Coast.

  • The atmosphere shifted from last year’s “rock ’n’ roll tailwind” to a pragmatic business meeting. 黄仁勋 said he had no teleprompter, and his delivery featured more pauses and verbal slips. 姚欣 felt that some product lines were being “rushed to market,” reflecting an AI industry whose 3-month cycle is forcing giants to adjust rapidly.

4. “1/27” suddenly transmitted industry consensus to capital markets

  • 卫诗婕 calls Nvidia’s post-DeepSeek R1 collapse the “1/27 disaster” in market parlance. She also cited a report described as “the most successful short report in history,” which questioned whether low-cost models could break Nvidia’s high market share and high profitability.

  • 姚欣 says AI practitioners and outside investors were operating on different clocks. In 2023 and 2024, the industry was still building AI data centers and competing to train larger models. But by June or July 2024, insiders were already discussing weaker scaling-law returns on the training side. By late November, OpenAI’s Ilya had also expressed the view that the industry was nearly out of data.

  • The real migration was from model training to model inference and applications; capital markets simply saw it later. DeepSeek’s breakout made corporate decision-makers realize that open-source models were already good enough. The next contest would be over cost-performance, usage costs and deployment—not just higher benchmark scores.

  • 卫诗婕 quoted 袁金辉 of SiliconFlow: “The compute required for training is a swimming pool; the compute required for inference may be a river.” Total inference demand could be larger, but it no longer follows exactly the same centralized architecture as training. That is where Nvidia’s opportunity and risk expand together.

5. 黄仁勋 uses orders and tokens to answer DeepSeek; the market still questions the monopoly

  • The presentation opened with commercial figures. According to 黄仁勋, major cloud providers bought roughly 1.2 million Hopper GPUs last year and have now ordered about 3.6 million Blackwell GPUs. The subtext was clear: customers have not stopped spending; orders are still rising sharply.

  • 黄仁勋 then recast reasoning models as a bullish argument for compute. One comparison said DeepSeek-like models might use roughly 20x more tokens and 100x more compute for the same answer; elsewhere he emphasized token throughput at 100x scale per unit of time. He reduced the conclusion to “buy more, save more.”

  • Nvidia shares still fell after GTC, showing that the question remains open: Nvidia effectively defined the training market, but in a distributed and cost-sensitive inference market, will customers still have no choice but Nvidia? 姚欣 believes that question—not whether inference requires compute—is at the center of the valuation volatility.

6. Dynamo turns the data center into an “AI factory” measured in tokens

  • 黄仁勋 upgraded IDC and AIDC into an “AI factory.” A traditional factory takes in raw materials and produces pots, bowls and pans; a digital factory takes in data and produces text, audio, video and every kind of token. Token thus becomes the smallest unit for pricing digital goods.

  • The key factory metric is not how many GPUs are installed, but how many tokens the system can produce at a given power budget and within a given time—and whether the equipment can run 24/7. 姚欣 sees this seemingly unglamorous language as a shift from myth-making to production efficiency and unit economics.

  • Dynamo can be understood as the operating system for an AI factory. It manages large numbers of GPUs and optimizes data-center deployment, scheduling and utilization. Its introduction features terms including PD disaggregation and expert parallelism.

  • 姚欣 explicitly labels “prompted by DeepSeek” as his own interpretation. He believes Dynamo and the components DeepSeek showcased during its February open-source week are “different expressions of the same idea.” 卫诗婕 also speculated that parts of its architecture resemble Kimi’s Mooncake; 姚欣 responded that all of these projects are open source, so borrowing and cross-reference are possible. Because these features had not appeared in conversations six months earlier or at CES at the start of the year, he suspects Nvidia rapidly absorbed industry advances from the past month or two.

7. Distributed inference splits large models into experts—and turns a single-machine business into a cluster business

  • 姚欣 uses DeepSeek’s full-size 671B model to explain MoE. It can be viewed as hundreds of smaller experts—some 7B, some 20B—with different strengths, such as content organization or logical reasoning, rather than invoking every parameter for every request.

  • If 100 experts are distributed evenly across 100 machines, popular experts become congested while less-used experts sit idle. DeepSeek’s approach reallocates them according to call rates: concentrate low-frequency experts and distribute high-frequency experts, allowing 10 or even 100 machines to serve requests together. 卫诗婕 calls this a “sharing economy,” and 姚欣 agrees that it is a “sharing economy for compute.”

  • NVL72 sells 72 GPUs as a cluster, while Dynamo adds automated deployment and scheduling. Nvidia is therefore selling not only individual cards but also clusters and their software management layer—a route “clearly built for inference,” with technical priorities distinct from training large models.

8. Cooling AI expectations is not the endgame; applications are taking over the industry chain

  • 姚欣 uses Gartner’s Hype Cycle to explain GTC’s darker mood. When a new technology emerges, expectations rise faster than actual progress, followed by disappointment and a slide down the curve. “A bubble bursting is not inherently a bad thing”: it strips out excess and lets the technology mature under more reasonable expectations.

  • The shift from scaling laws and ultra-large models toward AI factories, costs and token output may look less glamorous, but it means capabilities have been validated and prices are low enough. 姚欣 expects AI over the next 1-2 years to become “quietly pervasive,” like the internet and mobile internet, entering a wide range of everyday situations.

  • After DeepSeek, Chinese government, enterprise, education and healthcare sectors that had previously moved slowly also began embracing AI. Infrastructure providers must make tokens cheaper while flexibly supporting 10 million or even 100 million users. The beneficiaries will expand beyond GPU sellers to include model services, compute scheduling, vertical models and application software vendors.

9. CUDA built its wall on developer habits; model consolidation is lowering it

  • 卫诗婕 cited an early counterexample. Around 2011-2012, Jeffrey Hinton reportedly wrote to say that if Nvidia gave him GPUs, he would recommend them at academic conferences. Nvidia declined. 姚欣 believes the timing was simply too early; deep learning had not yet entered the industry’s field of view.

  • Once the deep-learning wave gained acceptance, Nvidia moved quickly. Its engineers went into startups to help them use CUDA and cuDNN, while universities could apply for free 1080 GPUs and others. Nvidia “won over a huge number of developers,” creating the default habit of getting a system working on CUDA first and porting it to other platforms later.

  • There were hundreds of models over the past 2 years, and the cost of adapting each one to different GPUs was extremely high, giving CUDA a major advantage. But once the model wars end, foundation models may consolidate into a handful, much like iOS, Android, Windows and Linux. Hardware vendors will then find one-by-one compatibility far easier, while agent-based programming further obscures the underlying layer.

  • 姚欣 expects cost-effective solutions such as AMD to gain ground as standards emerge, with the CUDA moat “slowly eroding” over a 10-year horizon. Sustaining Nvidia’s current 80-90% gross margins and 25-30x P/E for 10 or 20 years would be a very high bar. 卫诗婕 cited Kevin Kelly: “No technology giant will remain in a monopoly position indefinitely.”

10. DeepSeek’s edge comes from quant-style systems engineering, but PTX did not bypass Nvidia

  • 姚欣 says DeepSeek is not a conventional team of pure algorithm scientists. Its parent, High-Flyer, came from high-frequency quant trading, where being 0.1 microsecond or a few tenths of a millisecond faster can generate returns. That background led it to build its own data centers, optimize the full stack and accumulate large numbers of A100s before the large-model boom. The market had even circulated a figure of around 10,000 GPUs.

  • 姚欣 explicitly rejects the claim that “only a handful of people in the world can handle low-level optimization.” He points to his experience founding PPTV in 2004: with no public cloud available, the team built from servers and operating systems all the way up to cloud architecture. Many companies with more than 10 years of infrastructure experience have similar low-level optimization capabilities.

  • Technically, DeepSeek did not escape CUDA. It called PTX, an assembly-like language within CUDA, to control Nvidia hardware more directly. That improved performance but may have deepened the lock-in. The broader industry lesson is that system architecture and deep engineering still offer enormous room for optimization; scaling model size is not the only path forward.

11. China has faded from the GTC stage but entered Silicon Valley’s competitive assumptions

  • Citing his read of Nvidia’s filings, 姚欣 says China accounted for roughly 24% of the company’s global chip sales before export restrictions. 卫诗婕 adds that the share later fell below 10%, which 姚欣 attributes to the large number of chips barred from sale. 黄仁勋’s recent appearance in a floral padded jacket at Nvidia’s annual meeting also underscored that China remains an indispensable consumer and developer market.

  • Geopolitics was visible on the show floor. China AI Day was scheduled for Beijing time and held online in Chinese. Apart from Lenovo, which still sells servers, GTC had almost no booths from Chinese model or application companies.

  • “Although nobody talks about China, China has shaken Silicon Valley”(“虽然大家不讲中国,但是中国对于整个硅谷产生了冲击”). 姚欣 says Silicon Valley industry participants are hunting for Manus invitations and beginning to seriously compare Chinese companies on cost-performance, products and services—a sharp contrast with some of the negative discussion inside China.

  • 姚欣’s final judgment is that only the US and China currently possess comprehensive AI competitiveness. China is still catching up, but the gap is narrowing. As the contest moves into applications, specialized scenarios and data advantages, Chinese companies may gain relative strength: “We also need to have more confidence in ourselves.”