66. Live from GTC: Taking Stock of Nvidia’s Ambitions with 姚欣 and 季宇, and the 2026 AI Q1 Earnings Reports
Summary
- 姚欣 defines 2026 Q1 as the “post-inflection-point moment”: AI has crossed the demand inflection point, while inference and Agents are sending token consumption into a parabolic surge. He points to Google’s token usage growing more than 13x in 6 months and some model and compute-service companies generating more revenue in 1–2 months than in the previous 18 months. Cheaper models, in his view, have not reduced compute demand; through the Jevons effect, they have instead unlocked 100x-scale growth in total volume. Scaling has also moved from pre-train to test-time and inference, and now further downstream to Agent networks.
- Nvidia’s growth narrative has evolved from selling GPUs to delivering the entire “AI factory,” with the company seeking control of the chip, interconnect, cloud infrastructure and model-service layers above energy in the 5-layer stack. The Rubin era may no longer feature standalone accelerator cards, but instead deliver systems such as NVL72 supernodes. 姚欣 expects this to drive more NCPs, sovereign AI projects and a global scramble for energy—an ecosystem position he sums up as: “The cake is baked, the icing is on; you bring the cherries.”
- Optical interconnect is 姚欣’s highest-conviction industry call ahead of GTC, as the marginal gains from process nodes, stacking and rack density narrow while architecture and interconnect still offer room for expansion. Moving from 2nm toward 1.6nm may deliver only about 20% improvement, while shifting NVLink and NVSwitch from copper to optics could reduce power consumption and deployment complexity and virtualize distributed GPUs into larger pools of compute and memory. He would not hazard a call on Nvidia’s near-term stock direction, but said optical-related companies would definitely rise.
- Energy remains the hardest ceiling on AI infrastructure, and the industry’s core metric is shifting from card counts to how many tokens—or even how much human work—one kilowatt-hour can produce. 姚欣 says US data centers above 30MW are already difficult to find, while a conventional training center requires at least roughly 200MW and new 100MW-scale facilities take 18–24 months to build. North American domestic supply may not meet demand over the next 2–3 years, pushing compute sales and deployment toward energy-rich regions.
- Groq’s LPU is the most controversial new addition at this year’s GTC: 姚欣 sees it as a high-end specialized weapon in Nvidia’s inference arsenal, while 季宇 sees an obvious mismatch in the current system design. The LPU trades on-chip SRAM for extremely high bandwidth at the cost of capacity. Nvidia assigns it decode FFN/MoE, leaving the more bandwidth-intensive attention workload—with its dynamically expanding KV Cache—to GPUs. 季宇’s conclusion is that this may be a barely workable configuration designed “to find a use for the LPU,” while weak inter-rack links and massive concurrency could amplify the problems.
- 姚欣 believes inference has already settled into a “one superpower, many contenders” structure. AMD, Google TPU, Broadcom custom chips, Cerebras and Chinese GPUs are all expanding, with the key question being whether Feynman can preserve Nvidia’s lead through 2028. 季宇 still sees Nvidia as a “hexagonal warrior” with no obvious weakness; the real short thesis, he argues, is not who can build a more powerful supercomputer, but whether Nvidia is recreating IBM’s closed, high-barrier mainframe model.
- 姚欣 sees OpenCloud’s “Little Lobster” as a prototype of the AI-era OS, unifying interchangeable models, long-term memory, MCP, skills and computer use into an application-oriented orchestration layer. He expects a two-OS market to emerge over the next 2–3 years, with 1 open-source and 1 closed-source system. He relays the reaction from early-stage Silicon Valley investors—“The past 2 years were a waste”—as many prepare to rebuild their projects around the OpenCloud ecosystem. Agent organizations are also moving from individual productivity gains toward “super-individuals” directing dozens or even hundreds of AI employees.
- The central disagreement is not whether AI demand will grow exponentially, but whether that demand should ultimately be supplied by a handful of giant AI factories or by cheaper, localized personal compute. 卫诗婕 summarizes 4 directions for Agents: local, personal, private and customized. 姚欣 sees another 3–5 years of infrastructure exuberance for Nvidia, driven by the expansion of global token consumption and the AI labor market. 季宇 invokes IBM and Intel to warn that “the world only needs a few AI factories” sounds much like “the world only needs 5 computers”; OpenCloud, he argues, shows that users want AI that is their own, private and personalized.
Deep dive
1. The same episode split Nvidia’s story between pre-show calls and post-show calibration
卫诗婕 recorded the first part with PPIO founder 姚欣 about 1.5 hours before GTC opened on March 17. The reported 2-hour queue to exchange credentials, along with elevated expectations from markets and attendees, was already a marked departure from the tense atmosphere after DeepSeek’s shock 1 year earlier.
姚欣’s pre-show call was that the timeline had formally reached Agentic AI. 黄仁勋 was no longer making last year’s hurried pivot from training to inference, but was “calmly taking control” of the Agent AI and Physical AI roadmaps he had laid out. The real question was what products he would release for Agents.
After 黄仁勋’s keynote, 卫诗婕 brought in domestic GPU entrepreneur 季宇 for a postmortem. 姚欣 read the demand breakout through cloud, model services and industry capital; 季宇 challenged the designs through chip architecture and system economics. Both were bullish on AI growth, but assigned sharply different valuation implications to the “AI factory.”
2. 姚欣 says AI has crossed the “post-inflection-point moment”
姚欣 says the market has lived through “a full spring, summer, autumn and winter” in just 1 year. Since Q4 last year, not only GPUs but also storage, hard drives, memory and other legacy components have entered a price-hike cycle: those industries had been adding capacity at roughly 3%–5% per year, nowhere near enough to keep up with AI demand.
Reality has answered the post-DeepSeek concern over the Jevons effect: inference is becoming cheaper per unit, but total inference volume could grow 100x or more. His most striking example is Google, the giant among giants, where token usage “grew more than 13x in 6 months.”
He also sees a sharper inflection in his own business and among model companies: some businesses are generating more revenue in 1–2 months than they did in the previous 18 months. “We have crossed the base,” he says. Demand is no longer a long-term hypothesis; it is beginning to show up in short-cycle results.
That has also changed risk appetite across the supply chain. The circular investment criticized in the past has not disappeared; it is becoming more common, because the giants would “rather lever up, rather increase investment, even take risks to increase investment,” in an effort to lock in future compute ahead of time.
3. Scaling has moved from pre-training all the way to Agent networks
姚欣 reprises his metaphor from the previous episode: if training is “a pool of water,” inference has become “an ocean.” Better model efficiency has not capped compute demand, because scaling first moved from pre-train to test-time and inference, and has now shifted again to the Agent layer.
An Agent is not a single model call. It observes, calls tools, checks results and continues acting over long periods and multiple steps. Reliably reproducing the full workflow also requires substantial post-training and reinforcement learning, so Agent capability itself rests on an inference-compute base.
His view is that the industry is not merely advancing exponentially, but experiencing “accelerated exponential progress.” Agentic AI, which last year’s 黄仁勋 roadmap appeared to place 2–3 years out, has already broken out this year. “The future is here” is no longer just a slogan.
4. The 2026 model race is centered on long-horizon tasks, personalization and full modality
The first theme is long-horizon tasks. Anthropic’s Claude Cowork, desktop control and Excel workflows, along with OpenCloud’s “Little Lobster” independently completing tasks for ordinary users, are pushing Agents beyond the chat interface into long-running execution engines—and fueling the “SaaS apocalypse” thesis.
The second is personal intelligence: models will read an individual’s email, conversations and historical context and return a different result for each person. 姚欣 points to Gmail’s personalized capability launched in January, arguing that it increases stickiness while directly supporting advertising and other monetization efforts.
The third is multimodality, moving toward full modality across text, images, audio and video. He cites the new wave triggered by ByteDance’s Seedance and expects mixed-modality reasoning to produce a crop of killer applications this year. All 3 themes point to longer context, greater memory pressure and more tokens.
5. Inference is not 1 market, but a set of markets with radically different price-performance curves
姚欣 uses 2 use cases for the same model to illustrate the split. Ordinary chat can tolerate latency in the hundreds of milliseconds and mainly demands low cost; AI coding requires extreme reliability, speed and throughput, and users may accept prices 10x or even 100x higher.
GPU generality is therefore both an advantage and a source of redundancy. The GP in GPGPU stands for general purpose, and CUDA must cover dozens of industries. Specialized architectures for extreme inference cases can spend more to move performance from 85 or 90 points to 99.
That is how 姚欣 interprets Nvidia’s acquisition of Groq: not as an abandonment of the general-purpose GPU, but as an additional SRAM-and-LPU weapon for ultra-low-latency, high-stability, high-price workloads. Nvidia can also interconnect it with ASICs and other custom chips, positioning itself to capture an inference market where “a hundred flowers bloom.”
6. Rubin pushes standalone GPU cards toward full-rack supernodes
姚欣 describes the progression from H to B as the beginning of stacking. B200 can be understood as combining 2 units of capability at the board level; as density, power and interconnect requirements continue rising, the Rubin era may stop emphasizing standalone AI accelerator cards and sell NVL72-style supernodes directly.
In the past, an IDC or AIDC operator acted like an integrator, separately sourcing power, networking, racks and servers before assembling them. Nvidia’s AI Factory pre-integrates the rack servers, interconnect, debugging and part of the base software, so delivery is close to “plug it in and run.”
That threatens the existing division of labor among IDCs and channels while allowing Nvidia to capture more system profit. 姚欣 is particularly interested in asking NCPs at GTC how cloud and IDC operators—previously responsible for stocking and shipping Nvidia products—will reposition themselves once technical integration becomes easier and securing energy becomes the core capability.
7. The AI factory turns NCPs and sovereign AI into energy proxies
姚欣 expects the number of NCPs to rise, but their core competitive advantage will shift from system integration to securing electricity locally. Energy is inherently distributed around the world, and companies with the deepest knowledge of local energy infrastructure could become the builders of the next generation of sovereign AI.
黄仁勋 has compared compute to oil and argued that every country should have its own “extraction plant.” If Rubin and eventually Feynman are delivered as complete systems, local operators may only need to build the data center and connect the power; Nvidia can package everything above energy in 1 stack.
姚欣 therefore no longer views Nvidia as a GPU company, but as an “equipment exporter for AI factories.” It is not selling 1 machine tool inside a factory; it is selling the entire production line, with the output moving beyond compute toward tokens and intelligence.
8. Optical interconnect is the next growth curve as process and stacking gains slow
姚欣’s only semiconductor keyword before GTC was “optics.” In his estimate, moving from 2nm toward 1.6nm offers only about 20% improvement; continuing to stack from 56 or 72 to higher densities is also yielding fewer of the 2x gains seen in the past.
The real bottleneck in expanding racks is moving data between GPUs. Ordinary local networking is separated from chip-level interconnect by several orders of magnitude. Nvidia has spent years advancing IB, NVLink and NVSwitch through acquisitions and internal R&D, seeking to bring multi-machine communication closer to local GPU-to-GPU links.
Copper interconnects generate heat; fiber is lighter and thinner and easier to deploy. Cabling across clusters of tens of thousands of cards can become extraordinarily long, so optics could reduce power, weight and wiring costs. Optical channels can also continue scaling through upgrades to optical modules and architecture.
The end goal is to connect GPUs across different machines into 1 virtual giant GPU with 1 giant memory pool. MoE models distribute large, medium and small experts across multiple machines, while inference is moving from single machines and 8-card or 16-card configurations to parallelism across hundreds of cards. That directly expands demand for optical modules, optical communications and high-speed interconnect.
9. “Intelligence output per kilowatt-hour” replaces card counts as the ultimate metric
姚欣 says Nvidia cannot solve the energy crisis directly; it can only raise output from the same energy input. More advanced process nodes, stacking, optical interconnect and software optimization may each offer limited gains, but together they can improve compute and token output per kilowatt-hour.
The simplest AI-factory calculation is how many tokens 1 kilowatt-hour can generate. Going further, institutions such as Sequoia have begun asking how much actual human work 1 kilowatt-hour can complete. Chips, model compression and service orchestration are all aimed at converting energy into more sellable intelligence.
Demand is more aggressive than human software usage. An Agent can work 24 hours a day without eating, drinking or sleeping. Once machines call one another, consumption and iteration run faster than when humans use tokens directly, so energy pressure will rise alongside the AI labor force.
While searching for compute in the US, 姚欣 found facilities above 30MW extremely scarce, while a conventional training center requires at least roughly 200MW. New 100MW-scale facilities often take 18–24 months. Even 马斯克, despite emphasizing clean energy, has powered xAI training clusters with diesel generators—a snapshot of the supply-demand imbalance.
10. Geopolitical conflict is turning compute scheduling into a global reliability problem
姚欣 says some nodes in the Middle East temporarily went offline because of airstrikes, damage to energy infrastructure or unstable power, forcing their workloads to shift elsewhere. With data centers worldwide already close to full utilization, those migrations can create inference latency and network volatility.
He explains the amplification effect with 5% of compute going offline. If the receiving region was already carrying 10% of total load, adding another 5 percentage points means a 50% short-term increase. Idle resources generally cannot be brought online immediately. He also interprets Gemini’s previous outage as a chain reaction of this kind of regional shock.
One model service has even tried encouraging programmers to use it during the early morning in North America, moving peak demand to nighttime capacity in the Eastern Hemisphere. “Compute west, inference east” means global inference has become 1 network: energy, war and regional node failures all propagate to users everywhere.
11. The 5-layer cake is Nvidia’s alliance strategy for absorbing the upstream and downstream
姚欣’s blunt read is that “Nvidia is no longer a GPU company; it wants to swallow its upstream and downstream.” It invests in new clouds, NCPs, models and applications, but avoids directly replicating AWS or Google Cloud, instead using capital and ecosystem partners to build its own cloud alliance.
He cites Nvidia’s additional $2B investment in Nebius and says multiple Silicon Valley compute-service competitors have received Nvidia funding. The chip company supplies capital to cloud and model companies, which then buy chips and services; related-party transactions and circular investment become tools for locking in the ecosystem.
Nvidia is also using open-source model partnerships to embed models into a full service stack, gradually moving from hardware delivery to token delivery. 卫诗婕’s summary is apt: “The cake is baked, the icing is on; you bring the cherries.” Applications can bloom across the top layer, while Nvidia seeks influence over the 3 layers below.
12. The infrastructure triangle may ultimately invert toward applications and labor
姚欣 continues to use the “upright triangle, inverted triangle” framework. As application demand keeps exploding, today’s upright triangle could eventually become inverted, with the mature application-side market representing the largest slice and the base relatively smaller.
Even after the inversion, the AI foundation could exceed the total IT investment in human history. The truly larger upper layer is not just SaaS or consumer internet, but the “tool-use market” described by 黄仁勋—its substitute is labor, not the roughly $600B software market already in existence.
He assigns Nvidia’s infrastructure boom a 3–5-year window and compares it with the 2010–2015 mobile-hardware cycle. Once hardware growth stabilizes, the next Douyin or TikTok becomes the focus. OpenCloud and digital employees may be early signals of the AI application cycle.
13. Long-horizon tasks stretch token time from 15 minutes to days
姚欣 believes the combination of models, deployment and Agents can still deliver roughly 10x annual improvement in final results over the next 3–5 years. The gains will not come only from a single model, but from multi-model calls, multi-agent collaboration and longer execution times.
Last year, Deep Research- or Manus-style tasks might run for 15 minutes. Some tasks now run for 1 hour, and he expects tasks taking 3 or 5 days to deliver a result by year-end. Token consumption could rise 100x or 1,000x as a result, while reliability may become sufficient for users to adopt the output directly.
That is when AI labor begins to become real: the model no longer answers a single question, but takes over an entire work cycle. 姚欣 even offers a “cocksure assertion” that some CEO work may begin to be replaced this year, citing experiments with “zero-person companies” planned, organized and monetized by AI.
卫诗婕 does not accept the replacement narrative wholesale as a source of anxiety. The discussion ultimately returns to a new human capability: directing a matrix of 5, 10 or 100 AI employees may matter more than personally completing every task.
14. Agents extend the time before circular investment is disproved
姚欣 expects circular investment among Nvidia, cloud companies and model companies to increase again this year. The OpenAI–Nvidia–Oracle-style “Stargate” alliance came under pressure from massive losses, but OpenAI’s $110B fundraising and Anthropic’s rising valuation have injected fresh capital into the chain.
Investors originally worried that model companies would spend $100B per year on infrastructure and eventually need $100B or even trillions of dollars in revenue, while global software revenue was only about $600B. On a software-company valuation framework alone, the math was difficult to close.
Agents expand the addressable market from software budgets to labor spending and cover work scenarios that software could not previously serve. Trillion-dollar revenue then becomes imaginable again. 姚欣’s analogy is that the AI series is not moving from episode 30 to its finale; it has suddenly expanded to 50 episodes. “The original imagination was not the finale at all.”
15. “There is no AI bubble” and “valuations will correct” can both be true
姚欣 does not deny the bubble. He defines a bubble as temporarily elevated valuation, not nonexistent demand. Each time technology arrives 1–2 years early, investors steepen their forecast curves; when the future temporarily disappears from view, they revisit whether stocks have become too expensive.
The private market is more willing to believe the logic first and wait for financial delivery, tolerating 10 years of cloud- or SaaS-style losses. Public markets “see it first, then believe in the future,” relying more heavily on revenue, profit and company-level evidence. Chinese and US public markets therefore may not share Silicon Valley’s optimism.
VCs are also investing in multiple leading model companies at once, betting on the category rather than 1 winner. If the overall model and Agent market takes hold, the portfolio can still generate returns even if 1 company fails. That is why private markets remain willing to support high spending and circular transactions.
16. The most tradable pre-show calls were optics, storage and legacy components
姚欣 refuses to make a definitive call on Nvidia’s near-term stock price. The long-term trend could still be upward, but the base is already high and any miss—such as weaker growth in China—could trigger a shock. The future may look like slower straight-line gains with steadily higher volatility.
He is more certain about optical interconnect. Once Rubin and Rubin Ultra disclose the optical components required per rack, investors will quickly be able to draw a 3-year demand curve from shipment volumes. Earlier disclosures at CES about hard-drive requirements for new systems had directly lifted NAND-related names such as SanDisk.
He extends the opportunity set over the next 1–2 years to power, storage, hard drives, memory and a broad range of semiconductor components. Those industries had historically expanded capacity at low-single-digit rates, while GPUs doubled on the AI cycle. The entire IT supply chain needs to catch up, and in his view its cycle visibility may be even better than Nvidia’s.
17. OpenCloud gives the “AI-era OS” a concrete shape for the first time
姚欣 has long argued that no single foundation model can become the OS. Models are replaceable and iterate quickly, with each vendor leading in particular capabilities. A real OS should orchestrate heterogeneous underlying resources so upper-layer developers do not have to adapt separately to every hardware and model stack.
OpenCloud’s first OS feature is model choice. The second is personalization and long-term memory, which goes beyond a traditional database. The third is a tool-calling layer made up of MCP, skills and computer use. Prompts and model switching are hidden underneath; users care only about the capability they want.
Skills in particular resemble the App Store in the mobile-internet era. Install different skills and the same Agent can become a financial analyst, meeting assistant or something else. “How smart a smartphone is depends on which apps you install”; an Agent’s capabilities now depend on which models and skills it can orchestrate.
A unified layer lowers the development barrier, much as iOS and Android ended the Symbian era of device-by-device adaptation. 姚欣 has worked on Symbian ports himself and stresses that only once the OS matures can application developers move from manual integration to large-scale commercial breakout.
18. The AI OS will likely reproduce a dual-track market: 1 open-source, 1 closed-source
Looking back at Windows and Linux, then iOS and Android, 姚欣 sees a recurring pattern in which each generation has a commercially led closed-source OS alongside an open-source OS that maximizes ecosystem participation. OpenClaw is described as the fastest-growing system-level application on GitHub and already has a large ecosystem, making it look more like the open-source foundation for now.
Governments, industrial users and large enterprises, however, impose security, audit and compliance requirements on OpenCloud, leaving room for a closed enterprise OS. Anthropic’s Claude Cowork is gradually adding desktop access, task execution and controlled auditing, while Genspark Cloud and Manus are also competing for the position.
He gives the structure 2–3 years to evolve. The end state may still be 1 open-source and 1 closed-source system, but the winners are undecided; neither OpenCloud nor Anthropic has secured a seat. China will also develop its own open-versus-closed competition around local internet habits and compliance requirements.
19. Silicon Valley founders are already rebuilding around the Agent OS
姚欣 observes that OpenCloud’s popularity among early-stage Silicon Valley founders, YC and VCs is no lower than in China. One investor directly said, “The past 2 years were a waste,” because old projects often vertically integrated their own models and services and are now preparing to redesign around the OpenCloud ecosystem.
Big companies may respond more slowly than small teams. Some domestic giants previously worried about Doubao reading private-domain data and initially moved to block access; when faced with the open-source “Little Lobster,” they quickly launched compatibility with WeChat, Feishu and other platforms, showing how open-source ecosystems reduce defensive instincts.
AI-native organizations are splitting into 2 camps. Ordinary users can raise output by 3–5x with tools, while a small number of super-individuals can achieve the output of dozens or hundreds of people because they are not treating OpenCloud as 1 employee; they are building an entire company and a multi-Agent collaboration system.
YC partner Gary Tan publishing skills himself is, in 姚欣’s eyes, a symbolic example: “Even investors are writing their own skills now.” The scarce capability in this startup cycle is shifting from hiring more people to organizing an Agent matrix.
20. The “1 superpower, many contenders” inference market has moved from forecast to reality
姚欣’s list of contenders includes AMD, Broadcom custom chips, Google TPU, Cerebras and Chinese GPUs. Cerebras is pursuing a distinct route through wafer-scale integration; AMD is gaining adoption among major technology companies; and TPU, after nearly 10 years of iteration, now covers training, inference and model optimization.
Google is becoming more aggressive in opening TPU services, while model companies such as OpenAI also want to embrace TPU because inference costs may be lower than GPU costs under comparable conditions. TPU and GPU both use relatively conventional external-memory architectures, making them less divergent from each other than either is from an on-chip-SRAM LPU.
Nvidia remains the “1,” but the question has shifted from whether there are multiple contenders to whether there will be a second “superpower.” 姚欣 says it depends on whether Feynman can extend the lead through 2028; 季宇 emphasizes that ecosystem, generality and all-around capability still give GPUs a major advantage.
21. After GTC, 季宇 remembers “racks of every shape and description”
季宇’s first post-show impression was not a particular GPU, but the fact that CPUs, GPUs, networks and different combinations had all been turned into high-density racks. Racks used to be built around GPUs; now CPU, GPU, CPU+GPU and various GPU combinations are each standalone products, leaving him “dazzled.”
He agrees that the 5-layer cake reflects Nvidia’s need to find new growth. GPU is already close to an absolute monopoly; if Nvidia wants to keep discussing revenue in “T”-scale units, it must capture profit beyond GPUs—from CPUs, interconnect, storage, infrastructure, data centers and energy.
But 季宇 compares the scene with IBM mainframes in the 1980s. Operating systems, compute, storage and even printing were all delivered through expensive specialized machines: technically supreme and awe-inspiring, but closed and difficult to democratize. The more successfully Nvidia builds AI factories, the stronger that analogy becomes.
22. The LPU trades SRAM for bandwidth and writes the capacity problem into the architecture
季宇 explains that conventional GPUs use external DRAM or HBM to store weights and working data. On-chip SRAM is integrated with the logic circuits and offers extremely high access bandwidth, but capacity is generally only in the hundreds-of-megabytes range and unit cost is high. He also mentions Nvidia’s $20B acquisition of Groq last year.
Groq’s LPU takes an extremely aggressive path: it abandons conventional large memory, uses on-chip SRAM to eliminate the narrow channel between compute and storage, then combines hundreds of LPUs into a customized system to assemble enough capacity. The benefit is extremely fast inference; the cost is system scale and static constraints.
Large models have not only hundreds of billions or trillions of parameters, but also dynamic storage generated by each user’s context. The more requests and the longer the conversations, the larger the KV Cache. LPU capacity is difficult to expand on the fly, so a performance issue can quickly become a situation where the data “simply does not fit.”
Reducing users to control context creates another economic problem: a system of hundreds of expensive chips serving only a small number of people cannot achieve viable utilization or returns. 季宇 sums up the dilemma as “full of contradictions.”
23. Attention, MoE, prefill and decode require 4 different resource mixes
Attention processes relationships within the context. Every generated token must read the prior context and form the KV Cache, making attention highly memory-bandwidth intensive. When the raw material cannot be delivered, additional compute units can only finish a small amount of work instantly and then wait.
FFN later specialized into MoE, splitting 1 universal layer into many experts. A single request activates only a few experts, but when hundreds, thousands or even tens of thousands of users run concurrently, requests line up behind every expert and compute can be fully utilized. MoE then becomes a more compute-intensive workload.
Prefill handles one-time input—for example, reading 10,000 tokens at once—and can execute in parallel. Because weights are loaded once to process a large amount of content, it is more compute-intensive. Decode generates 1 token at a time, requiring both attention to read the KV Cache and scheduling of MoE experts, making system behavior more complex.
The difference in input and output costs also comes from this split: input tokens are generally cheaper, while output tokens are more expensive. Agents change the ratio again by continuously reading webpages, documents and tool results while outputting only short tool-call instructions. The workload no longer resembles a chatbot, where input and output are roughly balanced.
24. Nvidia’s decode-FFN assignment for the LPU misses its strongest use case
The GTC diagrams assign GPU racks to prefill and decode attention, and LPU racks to decode FFN—the MoE expert layer. 季宇 sees a direct mismatch: the LPU’s greatest advantage is bandwidth, while MoE, once optimized for high concurrency, is more compute-intensive.
He cites the diagram’s roughly 300P of compute and 40P of bandwidth as an intuitive illustration. Each unit of data may require very little computation before consuming the available bandwidth; a compute-intensive task might perform hundreds of calculations per unit of data and does not need such expensive bandwidth.
The true bandwidth-heavy workload is attention, but its KV Cache varies dynamically with user count and context length—precisely the weakness of the LPU’s limited capacity. Put attention on the LPU and the system may not merely run inefficiently; it may suddenly run out of space in production.
季宇 speculates that fixed model weights can at least be preloaded into the LPU rack, eliminating the risk of dynamic overflow, which is why Nvidia assigns it the expert layer. It is a static design that “barely works,” but looks more like an effort to “force a use for the LPU” than an optimal resource match.
25. Inter-rack links and massive concurrency further weaken the LPU thesis
Nvidia’s strength is in-rack NVLink, but 2 racks do not have an interconnect of comparable strength. If 1 model is split between GPUs and LPUs, every generated token could travel between the 2 racks dozens of times, turning the inter-rack link into a new system-wide bottleneck.
Before its acquisition, Groq used cloud services to demonstrate extremely high output tokens per second, creating a powerful visual effect. 季宇 stresses that single-request speed does not determine whether a system can scale commercially; massive users, dynamic context and cross-rack traffic do.
“Raising lobsters” also does not give the LPU a personal-use case. Even with just 1 local user, 1 LPU cannot hold the model; hundreds of chips must be combined into a system costing millions or even tens of millions of yuan. Building that supercomputer for 1 person is economically untenable from the start.
26. Agents push chip demand back from output reasoning toward massive input
季宇 warns that changing workloads are the biggest risk for specialized chips. Early dense models woke all parameters for every generated token and were relatively bandwidth-intensive; MoE made the workload sparse, so the requirements changed. An extreme architecture customized for the previous model generation can quickly become mismatched.
In the chatbot era, a user asked 1 question and the model gave 1 answer, leaving input and output roughly balanced. O1-style reasoning caused output tokens to surge because internal reasoning also counts as output. Agents, by constantly retrieving documents, webpages and tool results, push input tokens back into a dominant position.
季宇’s own Claude Code usage shows input tokens running about 150–200x output tokens. An Agent may output only a few characters for a tool call but then read pages of returned content back into the system, requiring inference infrastructure to remain “steady across the board” rather than optimize a single token path.
GPUs and TPUs therefore remain the safer general-purpose solutions. The LPU can create value in extreme low-latency scenarios, but Nvidia has yet to explain away its current capacity, interconnect and concurrency issues as a foundation for mass Agent inference.
27. “AI factories everywhere” and “AI for everyone” are mutually exclusive end states
黄仁勋’s roadmap is to build large AI factories around the world and continuously produce tokens for phones, watches, headphones, cars and robots. 季宇 accepts that demand could grow exponentially, but does not agree that only giant factories can meet it economically.
His historical comparison is IBM and Intel. IBM pushed computing technology to its peak, while Intel used cheap x86 personal computers and servers to put that progress in ordinary people’s hands, enabling early Google and others to build internet services on low-barrier hardware.
“The world might need only 5 computers” did not become the end state of computing. 季宇 gets the same feeling from “the world only needs a few AI factories.” True democratization is not making a handful of supercomputers more powerful; it is making high-quality models affordable and deployable for more individuals, universities and companies.
OpenCloud offers a full picture of personal AI. Mac Mini-style devices still depend on cloud tokens today, but users are already signaling a desire for local, private and personalized AI. If high-quality models can move directly into consumer devices, that would not supplement the AI factory; it would represent a different industry structure.
28. Nvidia’s premium strategy improves unit economics while raising the absolute barrier
季宇 acknowledges that the unit economics of Nvidia’s machines improve with every generation: prices rise, but performance rises more. The company pursues maximum compute, bandwidth and interconnect with little regard for cost, making it a “hexagonal warrior” with almost no weak spot.
The problem is that absolute investment rises with every generation. 10 years ago, individuals, laboratories and ordinary companies could still buy consumer GPUs or compute cards and participate in deep learning. Today, much of what GTC displays in rack form “has nothing to do with me anymore”; the core customer is now the hyperscaler.
He has watched GTC for years partly to see whether 黄仁勋 would pivot toward democratizing products, but the signals instead keep reinforcing the AI factory. Mainframes and personal computers cannot both be maximized as markets: once cheap standardized equipment becomes good enough, customers stop paying for closed mainframes, and Nvidia is unlikely to “revolutionize itself.”
29. The disappearance of CPX and Vera’s bundling expose the commercial choices behind the roadmap
CPX originally sought to replace HBM with cheaper GDDR, combining high compute with low-cost memory for better price-performance. But after memory prices rose across the board following September last year, the gap between GDDR and other memory narrowed, potentially erasing the economic advantage of the heterogeneous design. It has therefore faded from this year’s roadmap.
季宇 uses the example to stress that a GTC presentation does not mean a product is finalized. Nvidia releases a design, watches the market and supply chain, then adjusts the form factor. The LPU could receive a more coherent story next time, or gradually disappear like CPX.
Vera, the CPU after Grace, also serves system bundling. A pure Vera server does not use conventional pluggable memory modules; instead, nonstandard memory dies are mounted flat on the board. 季宇 believes the CPU itself is not the main profit pool. The more important commercial objective is using the CPU to bind customers to nonstandard memory.
He summarizes Nvidia’s sales logic as “carrot and stick.” The carrot is the incremental performance of a nonstandard system; the stick is that customers who want GPUs must also buy the CPU, interconnect and memory. If feedback on GB200 bundling is weak, it could force the company back toward forms such as B200 and B300 that do not bundle Grace.
30. Feynman is still a futures contract; the real debate is whether Nvidia becomes the next IBM
Feynman comes after Hopper, Blackwell and Rubin, and even the Rubin product has not yet been fully seen; this GTC offered limited disclosure. 季宇 says it will naturally be called “built for Agents,” because Nvidia always aligns a product with the hottest demand at launch. GB200 was once described as “built for Scaling Law training” and later repositioned as “built for inference.”
姚欣 believes Nvidia still has a 3–5-year runway. Feynman, optical interconnect, Agent scaling and global energy expansion are enough to sustain an AI-infrastructure boom. 季宇 sees the real bear case not as insufficient Nvidia technology, but as the fact that “its thinking is too much like IBM’s.”
Their greatest point of agreement is that AI demand is nowhere near a peak; their greatest disagreement is the structure of supply. 姚欣 sees AI factories turning electricity into tokens. 季宇 sees an “AI industry trapped inside mainframes.” Behind 黄仁勋’s confidence as the “token king” lie both Nvidia’s ambition and the system-level economics that markets will need to validate over the next several years.