Pioneers Insight Method Research Author
Silicon Valley Coordinates x CASPA Vice President Jimmy Cheng: GTC Recap—Nvidia's Moat in the AI Era
Back to Episodes

Silicon Valley Coordinates x CASPA Vice President Jimmy Cheng: GTC Recap—Nvidia's Moat in the AI Era

Summary

  • Jimmy Cheng’s central takeaway from GTC: OpenClaw is being positioned as the operating system for agentic AI. The opening remarks called it “the most popular open-source project ever” and said, “It exceeded what Linux did in 30 years”; Jimmy relayed Jensen Huang’s view that it is one of the greatest pieces of software ever built, growing faster than Linux and Kubernetes. Huang also said that “every CEO in the future will have an OpenClaw strategy, just as every CEO used to have a Linux strategy”; in response to security and privacy concerns, Nvidia released NemoClaw with enterprise security built in.
  • Nvidia’s moat is thickening through systemization: the bar has shifted from a chip that is “10x faster” to a two-part test of TCO and software compatibility. Jimmy breaks the AI factory into five layers—energy, infrastructure, chip, model, and application; OEMs must follow the DGX blueprint, while suppliers face double-sourcing risk, leaving Nvidia with most of the industry’s profits.
  • Nvidia’s push into open-source models serves three purposes and will also accelerate custom-chip development. First, it adds competitive pressure among cloud companies, Anthropic, and OpenAI; second, it lets companies such as Salesforce and ServiceNow buy servers rather than tokens; third, it strengthens the flywheel around CUDA’s 7 million developers. Jimmy sees visible model weights lowering the barrier to building accelerators, with CapEx pressure pushing hyperscalers toward in-house chips and a larger market drawing more startups into custom silicon.
  • The hardware grammar of the inference era is SRAM-centric computing and 3D ICs. An LPU is like a taxi—“fast, but with limited seating”—while a GPU is like a bus—“it carries 100 people, but moves more slowly.” Cao Qingyun said Nvidia acquired Groq for $20B in December 2025, integrated it into its systems this March, and launched the LPU; LPX is the full system, with Nvidia selling 192 LPUs, each equipped with 500MB of SRAM. Cerebras won a $10B order from OpenAI; inference chips attracted $3B in VC investment in the first 3 months of this year, half of last year’s full-year $6B.
  • Gemini 3.0 used TPU for all inference and training, with no GPU involvement—a strong signal to in-house chip developers that Nvidia’s chips can be beaten. Amazon/AWS, Google, Meta, and Microsoft have CapEx budgets of $200B, $185B, $135B, and $100B, respectively, totaling $650B, with at least 50% going to chips. Jimmy sees custom silicon as one path to lower costs, while hardware-software co-design will push the industry toward a more diversified structure; UALink is an open format challenging the closed NVLink.
  • Synopsys’ moat is stronger in the AI era: IP’s core value is being silicon proven. Good IP only counts after 300 or 500 tape-outs; it is not a few lines of RTL code, nor a barrier AI can erase in 1 or 2 days. Synopsys released L4 Agentic Engineer, one level below L5 full autonomous, using AI to design AI; Jimmy cited Huang’s view that large numbers of agent engineers will use these tools and that the tools will grow “explosively.”
  • Hardware engineers will not face the same replacement pressure as software engineers in the near term: zero tolerance meets a talent shortage. A chip bug means the chip does not work, and R&D costs can reach $10M; human experience remains irreplaceable. The transition from Hopper to Blackwell took 2 years, while Blackwell to Rubin took 1 year, meaning the same people must do twice the work. Jimmy says the most overlooked issue is physical AI’s sim-to-real gap: real-time physics engines, deterministic sensor processing, and digital twins are all critical—“one second late, and the cup is already on the floor.”

Deep dive

1. GTC’s Main Theme: OpenClaw Is the Operating System for Agentic AI, and Tokens Are the New Currency

  • Jimmy likened this year’s GTC to the Spring Festival Gala. After Jensen Huang’s 2-hour keynote opening, his first takeaway was OpenClaw: Huang called it one of the greatest pieces of software ever built, growing faster than Linux and Kubernetes, and positioned it as the “operating system of agentic AI.” The opening remarks also called it “the most popular open-source project ever” and said, “It exceeded what Linux did in 30 years.” In response to the biggest concerns—security and privacy—Huang announced NemoClaw, embedding enterprise security, safety, and privacy into OpenClaw for corporate use. “Every CEO in the future will have an OpenClaw strategy, just as every CEO used to have a Linux strategy.”
  • Jimmy’s second point was inference, which he said Huang mentioned no fewer than 36 times, along with the launch of Vera Rubin, designed for inference; he also noted that an LPU from “Crock Three” (per the spoken transcript) can be used alongside Vera Rubin.
  • The third point was tokenomics: tokens are the “currency of agentic AI.” An AI factory must compete with other AI factories, and the contest is not simply about FLOPS or TOPS but token per watt.

2. From Selling Chips to Building AI Factories: How Systemization Rebuilds the Moat and Redistributes Profits

  • Jimmy breaks the AI factory into five layers: energy, including power delivery, heat exchange, and thermal management; infrastructure, including InfiniBand; chips, including GPU, CPU, and LPU; models, including Nemotron, Anthropic, and ChatGPT; and applications, including Omniverse. Nvidia has products or positioning at every layer.
  • He highlights two opposing trends. System companies are making their own chips: Apple’s early iPhone chips were developed by Samsung before Apple gradually moved in-house over several generations, while Tesla’s autonomous-driving chips moved from Mobileye to Nvidia and then to internal development. Nvidia has taken the opposite route, moving from chips into systems—developing its first GeForce graphics cards in 1999, launching CUDA in 2006, acquiring Mellanox in 2020, building Omniverse in-house, and launching the Nemotron large model in 2024.
  • The moat can be quantified this way: “Before, a chip that was 10x faster was enough. Today, to compete with Nvidia… you have to beat it on TCO, and you have to be compatible with its software.”
  • Profit pools have shifted accordingly. Nvidia now has 7 chips, while NVIDIA Switch functions as networking, taking the company into areas previously supplied by Cisco, HP, and Dell. OEMs must follow the DGX blueprint; they used to be able to design their own motherboards, but no longer have that discretion. Suppliers may earn attractive margins, but remain exposed to Nvidia’s double-sourcing risk. Nvidia captures most of the profit across the supply chain.

3. Why Nvidia Is Making Its Own Open-Source Models—and How That Drives Custom Chips

  • Cao Qingyun said Huang views open-source models as the second-largest ecosystem player after OpenAI and personally hosted the open-source-model panel. Huang has also entered the field with Nemotron and built OpenClaw on top of it. Jimmy identifies three motives: first, to add competitive pressure among cloud companies, Anthropic, and OpenAI—every new model drives a hardware refresh, and “the last thing he wants to see is model companies stop releasing new models or slow down,” so the cycle must become shorter and faster; second, to democratize generative AI—when Salesforce and ServiceNow use open models, their main purchase is Nvidia servers rather than tokens; and third, to strengthen the developer flywheel around CUDA’s 7 million developers.
  • Asked by Cao whether advances in open-source models would make chips more general-purpose or more customized, Jimmy’s answer was customization. Once model weights are visible, different accelerators can be designed around them, lowering the technical barrier. As usage at companies such as ServiceNow and Salesforce grows, CapEx pressure will push them toward in-house chips. Open source also expands the market: the larger the inference market, the more startups will enter with custom silicon.

4. The Hardware Grammar of the Inference Era: SRAM-Centric Design, LPU, and 3D ICs

  • Agents mark “the beginning of the inference era,” bringing an SRAM-centric design philosophy to the fore. HBM is in high demand, and HBM manufacturers cannot keep up. SRAM is exceptionally fast but limited in capacity; inference does not require processing huge volumes of data, but it does require very fast processing, making SRAM-centric computing a better fit. Cerebras won a $10B order from OpenAI; its chip is roughly the size of an iPad and contains a large amount of SRAM.
  • Jimmy’s analogy is: “An LPU is like a taxi—the taxi is fast, but it carries few passengers; a GPU is like a bus—it carries 100 people, but drives more slowly.” GPUs are stronger on throughput and weaker on latency; LPUs are the reverse. They are not substitutes for one another, but fit different use cases. Cao said Nvidia acquired Groq for $20B in December 2025, integrated Groq into its systems this March, and launched the LPU. LPX is the entire system; Nvidia sells 192 LPUs, with 500MB of SRAM attached to each chip.
  • The underlying trend is 3D ICs. Hopper was still a single chip, while Blackwell had grown large enough to be split into 2 chips. Stacking SRAM onto GPUs will be another likely direction. Feynman is planned for 2028 on TSMC’s 1.6nm process, designed for physical AI, or robotics. Real-time inference has demanding requirements, making SRAM even more valuable; 3D stacking may be needed to place SRAM on separate chips.
  • The funding data supports the thesis: inference-chip companies have raised $3B since the start of this year. Investment in the first 3 months alone reached half of last year’s full-year $6B, a 100% growth rate, signaling strong VC conviction in the infrastructure market.

5. The Competitive Landscape: TPU Shows That Nvidia Can Be Beaten, Pointing to a More Diverse Future

  • Google’s TPU is one of the more successful chips in the broader chip ecosystem, custom-designed for TensorFlow. All of Gemini 3.0’s inference and training run on Google’s own chips, with no GPU use at all. Jimmy called this a strong signal to every in-house chip developer: “Nvidia’s chips can be beaten.” AWS has its own inference and training chips, with Inferentia performing reasonably well; Meta’s MTIA remains in the early stages of development.
  • The CapEx logic behind in-house chips is strengthening. The first question analysts ask on earnings calls is no longer revenue growth or EPS, but CapEx. Amazon/AWS has $200B, Google $185B, Meta $135B, and Microsoft $100B, for a combined $650B, with at least 50% going to chips. Upfront R&D costs for proprietary chips may be high, but after mass production, the cost of each CPU or GPU may be lower than buying Nvidia GPUs. The hyperscalers also have their own models and software stacks, allowing hardware-software co-design for greater efficiency.
  • Jimmy favors a more diversified environment: models, chip companies, and data centers could all diversify. UALink is an open format developed in response to closed NVLink, and Google and Nvidia will each have their own role.

6. The Synopsys View: Silicon-Proven Moats, Zero-Tolerance Engineers, and the Sim-to-Real Gap

  • The host presented two competing views: AI tools lowering the barrier to IP creation could be a negative for Synopsys, while surging agent usage and demand for licenses could be positive. Jimmy described Synopsys as the world’s No. 12 software company, the No. 1 EDA company, the No. 2 IP company, and the No. 1 physical-simulation company; its IP business ranks No. 2, behind only Arm. His response was that IP’s real value is not whether the RTL code works, but whether it is silicon proven. Much of the IP has been taped out 300 or 500 times, a barrier AI cannot erase in 1 or 2 days. In his view, Synopsys’ moat has strengthened.
  • At Synopsys Converge, the company announced L4 Agentic Engineer. L4 is one level below the top-tier L5 full autonomous. The system uses AI to design AI and combines Synopsys’ AI agents with Nvidia AI GPUs to help industries design AI chips. Jimmy quoted Huang as saying that there will be many agent engineers in the future, every engineer will use these tools, and the tools will grow “explosively.”
  • On whether hardware engineers will be replaced like programmers, Jimmy said the two fields are fundamentally different. A software bug can be patched the next day and deployed to the cloud; hardware is a zero-tolerance industry, where a chip bug means the chip does not work and R&D costs can reach $10M. Human experience remains irreplaceable. The industry is also facing a talent shortage while iteration cycles accelerate: Hopper to Blackwell took 2 years, while Blackwell to Rubin took 1 year, requiring the same people to do twice the work. Over the long term, AI will replace some repetitive, manual work, but roles may shift toward higher-value tasks.
  • On the limits of “fully automated chip design,” Jimmy said AI cannot today design a 2nm chip while specifying the area of its PLL, and that this will not happen at least in the near term. AI has been integrated into chip design for roughly 7 or 8 years, with applications in architecture design, RTL verification, synthesis, and backend sign-off; PPA—power, performance, and area—can already improve by 10% or 15%. The industry is now entering the agent phase. But “black boxes may work for some industries; for chip design, they are absolutely unacceptable,” because verification is critical.
  • The most overlooked link is physical AI’s sim-to-real gap. A robot system includes chips, models, sensors, and actuators. A robot may use a given force and angle to pick up a cup in simulation but fail to lift it in the real world—that is the sim-to-real gap. The three challenges are a real-time physics engine that understands gravity; low-latency, deterministic real-time sensor processing—“if you say 3 seconds, it must be 3 seconds, no more and no less,” because 1 second late and the cup may already be on the floor; and a real-time digital twin capable of running 100 simulations before the cup falls. Many of these challenges require hardware-software co-design. Jimmy believes that if the industry keeps working to close the sim-to-real gap, the future is already within reach.