周默 on U.S. Tech Q2 Earnings: CapEx Growth Outpaces AI Revenue Again
周默 on U.S. Tech Q2 Earnings: CapEx Growth Outpaces AI Revenue Again
Summary
- 周默称过去一个月“可能是我们有股票以来,人类历史上回调最猛烈的一次”。 七月半导体板块蒸发超$1T市值,大部分半导体股一个月腰斩,而Magnificent Seven仍在创历史新高。下跌链条由仓位触发:五月此前平均约10%–15%的半导体配置被部分美国基金提到30%–35%,六月对冲基金开始加空头,散户仓位处于历史最高;七月底又叠加一家知名基金降杠杆、接近爆仓,其覆盖整个美股的高gross仓位可能影响数百亿美元半导体仓位,“那一周可能跌了30%到40%”,最终以流动性危机收尾。
- ROI叙事在Q2出现二次反转:上半年AI ARR增速快于CapEx,但供应链通胀让CapEx增速重新超过ARR。 Anthropic较去年底约增长8倍,单GW数据中心年收入从$10B升到$30B;但单GW数据中心成本从去年底约$35B涨到明年底的$55B–$60B,涨幅约50%。2027年30GW对应的CapEx从$1T变成$1.7T——需要的AI ARR也从$300B–$500B推高到$600B–$800B,仅计算当前闭源模型收入很难达到。
- 微软是本季逻辑反转最猛的一家,周默在财报前转变态度:这是他“过去一年唯一一次看到Copilot加速增长”的季度。 一家美国大型金融客户从几千个seat试用直接采购100K个seat,付费席位从20M升破30M。Azure增速43%超指引、下季指引45%,且微软是“这几家CSP和Meta里面唯一一个提到明年现金流为正的公司”;Satya首次以业绩会形式给出开源模型合作ROI约30%,即3年回本。核心逻辑是“开源模型的定价本质上是CSP定价,不是开源模型定价”。
- Meta被惩罚的不是本季数字,而是明年的问题:约$220B CapEx相当于收入的70%、利润的两倍,问题已从“花光现金流”变成“能不能再造一个Meta”。 本季广告收入增长27%,剔除一次性费用后利润率仍在增长;但CapEx将对明年毛利率造成6–7个百分点拖累,明年利润增速可能只有低个位数。卖算力也没那么快——Meta缺SLA级服务能力,想租H卡但不想租B卡,而买方希望两者都要,最早年底或明年初才可能有算力租赁收入,市场却希望本季度就看到AI收入路径。
- AWS增速36.7%、连续第五个季度加速,单季环比新增$4.6B收入比历史最大增量高80%,最大driver是Anthropic。 其API约70%经CSP销售,其中接近90%通过AWS后台,AWS在该模式中抽走40%;Anthropic二季度ARR环比又增约100%–120%。Blackwell从出货到确认收入有约一个季度时滞,Q1、Q2的拉货都将确认到Q3,“到了三季度,我们觉得AWS的营收就有机会能够到41%到42%”。
- 谷歌GCP的82%增速“跟Gemini没有太大的关系”,主要是TPU直销,约100K张卡带来$2B收入。 直销毛利率约50%——单卡售价$20K、博通成本$10K——从Q3起往后6个季度可做$50B–$60B。Gemini二季度环比增速从超50%掉到20%–30%,两个瓶颈是合规导致买数据“比其他公司买数据慢半拍”,以及训练算力只有OpenAI和Anthropic的一半、约1GW;3.5 Pro不被认为能赶在9月或10月发布,Gemini 4大概率年底或明年初发布。
- 整个链条的关键变量是开源与前沿模型的差距:Kimi K3把美国投资人认知中的6个月差距缩到3–4个月,其冲击比年初DeepSeek V4和本周的V4 Flash/Pro更大。 若维持3–4个月,毛利份额将从模型公司流向CSP;若拉回6–9个月则可能逆转。前沿参数从去年底约1T增至5月、6月约3T,下一代可能到8T–9T;月底英伟达财报对情绪的影响“没有那么大”,Anthropic上市和各实验室ARR进展更重要。英伟达对存储的加价/价格安排,原文同时出现“按1:3甚至1:4比例降价”的表述,存在歧义;周默认为这会加大GPU产业链的CapEx压力。
Deep dive
1. July’s Semiconductor Crash: A Positioning Cascade “Worse Than the Internet Bubble”
- 周默’s view carries an explicit “possibly” qualifier: judging by the magnitude and speed of the pullback and participants’ experience, “even looking back at the internet bubble, we have never seen a pullback this violent.” The Magnificent Seven were all making new highs, while most semiconductor stocks fell by half or more in a single month.
- Why did it happen so fast? Two structural changes: with AI, “you can complete basic research on a company in 1 hour,” so investors process both bullish and bearish information faster; meanwhile, influencers on Twitter and Substack are “increasingly replacing the influence of investment-bank research,” compressing the path from rumor to discussion to price reaction.
- Positioning was the fuse. Several influential, influencer-style U.S. funds raised semiconductor exposure from a previous average of roughly 10%–15% to 30%–35% in May, making the pace of institutional repositioning “possibly the fastest month in history.” Hedge funds began adding shorts amid June’s “token maxing” debate, taking institutional net exposure lower, but retail semiconductor positioning was at an all-time high. Leverage was also elevated in parts of the market in July, so even a modest decline could trigger a fresh debate over fundamentals.
2. The ROI Narrative Reverses Again: Supply-Chain Inflation Puts CapEx Back Ahead of ARR
- The semiconductor rally through May rested on two pillars. First, AI ARR was growing faster than CapEx: Anthropic was up roughly 8x from the end of last year, OpenAI had more than doubled, and CapEx had not risen by anything close to the same magnitude, narrowing the gap in absolute dollars. Second, the market’s estimate of annual revenue from a 1GW data center rose from about $10B at the end of last year to $30B: the same amount of compute was now worth more.
- This quarter’s reversal is the mechanism flagged last quarter—higher memory and upstream prices—only “prices rose too much.” The cost of a 1GW data center is expected to rise from roughly $35B at the end of last year to $55B–$60B by the end of next year, regardless of whether it uses B300 or Vera Rubin, a roughly 50% increase. From Q2 onward, CapEx growth again outpaced ARR, primarily because of supply-chain inflation.
- At the micro level, Meta faces the sharpest scrutiny. After price pass-through, next year’s CapEx is expected to reach about $220B, equivalent to 70% of revenue and 2x profit. The debate has shifted from “spend all of its cash flow” to whether an investment of this scale can “build another Meta.”
3. The Full Downturn Chain: From Meta’s Compute-Rental Rumor to Kimi K3 and a Fund Liquidity Crisis
- The timeline is worth preserving. News on July 1 that Meta planned to rent out GPUs and compute triggered violent moves in both semiconductor stocks and Meta. Combined with Satya’s support for open-source models over the prior 2 months and his emphasis on their benefits to CSPs and ROI, the market began asking whether companies unable to absorb the spending might cut CapEx.
- Roughly 1 week later, Kimi K3 launched. 周默 believes its impact was “much greater” than that of DeepSeek V4 at the start of the year and the DeepSeek V4 Flash and Pro models released this week. Most U.S. investors had assumed the gap between frontier and open-source models was roughly 6 months; Kimi K3 cut it to 3–4 months, raising concern that open source could disrupt the current CapEx regime.
- In the second half of July, rumors circulated that memory could not command its usual prices in Asia, including reports that handset makers and BAT were rejecting memory quotes. Although North American CSPs had accepted the prices and had strong negotiating leverage, Twitter rapidly spread the narrative across the U.S. At the same time, a prominent fund was deleveraging and nearing a blowup; its high-gross book across the entire U.S. equity market may have touched hundreds of billions of dollars in semiconductor positions, producing the biggest leg down at the end of July.
4. Microsoft’s First Reversal: Copilot Accelerates Enterprise Penetration for the First Time in 2 Years
- The core line in 周默’s team’s pre-earnings report was: “This was the only quarter in the past year in which we saw Copilot growth accelerate and penetration accelerate in large enterprises.” The clearest evidence was a large U.S. financial customer moving directly from a trial involving several thousand seats to a 100K-seat purchase—“something we had never seen in the past 2 years.” Paid seats rose from 20M in Q1 to above 30M, with net additions doubling sequentially.
- The key product change was replacing the underlying model with a coding model, enabling longer reasoning chains, more checks before execution, and fewer errors on complex tasks. Copilot is beginning to build applications independently: PowerPoint has moved from almost unusable to producing usable outputs on command, while Excel’s error rate has also fallen visibly. Consumer and commercial Copilot are being combined as the product evolves toward a “Copilot super-app.”
- 周默 still draws a boundary. On frontier innovation and more native AI organizational capabilities, Copilot “still cannot match Anthropic and OpenAI.” But this quarter was precisely the window in which Copilot iterated fastest and had the best chance of closing the gap; whether that momentum continues over the next 2 quarters remains difficult to judge.
- The next metric to watch is enterprise paid-seat penetration. Activation and activity rates mattered more previously; the question now is whether more companies can move from 5%–10% penetration to 30%–50%. If so, Copilot enters its next phase. With consumption-based billing integrated through Copilot Studio, the price of a single paid seat at an AI-native company could exceed $30.
5. Azure Accelerates to 43%: OpenAI’s Three-Layer Contribution Plus GPU/CPU Price Increases
- OpenAI’s contribution to Azure improved on 3 fronts simultaneously in Q2. Training revenue from OpenAI’s largest supercluster previously was not counted in Azure; as OpenAI adds more clusters on Azure, the new training revenue can now be recognized. OpenAI’s own inference revenue also accelerated in Q2, with Codex feedback improving from May onward. Microsoft’s resale of the OpenAI API benefited at the same time.
- Supply and pricing both moved higher. The Wisconsin data center released 300–400MW of capacity, much of which can generate revenue. Q2 was the tightest quarter for compute: when Anthropic rented CoreWeave data centers, roughly 1GW could generate more than $30B of revenue. GPU prices for new customers rose about 10% from prior levels, while overall increases were in the high single digits to low double digits. Existing contracts generally rise only at renewal; CPU prices are being raised implicitly through tighter discounts, contract changes, and pressure to migrate to newer instances.
- Model sales provided another increment. Microsoft began selling Anthropic models in Q1 and more open-source models in Q2. Anthropic models accounted for roughly 10%–15% of Microsoft’s API sales this quarter.
6. Satya’s Three-Part ROI Framework: CSPs Hold Pricing Power for Open-Source Models
- On the earnings call, Satya compared 3 models: the ROI of training a model in-house, hosting someone else’s model and reselling it, and owning a model and selling it. The most notable point was an ROI of roughly 30% for partnering with open-source model companies, implying a payback period of about 3 years. This was his first earnings-call explanation of why Microsoft is pursuing open-source models.
- The core logic was direct: “The pricing of open-source models is fundamentally CSP pricing, not open-source-model pricing.” If Microsoft can diversify token demand and host more open-source models, its bargaining and pricing power increase.
- Could the RPO disclosure—commercial RPO up 84% to more than $600B, or 25% excluding OpenAI, with the incremental growth coming from customers outside the Frontier Labs—resolve concerns over customer concentration? 周默’s answer was blunt: “No, because most of the RPO still comes from OpenAI.” But the shift does address the gross-margin and ROI questions. Azure OpenAI’s gross margin has reached 85%, above OpenAI’s roughly 50% or slightly higher and above pure GPU rental at just over 40%. Hosting open-source models could keep part of model-company profit within the CSP layer.
- More structurally, Microsoft is “the only one among these CSPs and Meta to mention positive cash flow next year; the others have negative CapEx.” Positive cash flow gives Microsoft stronger negotiating power in financing markets and could allow it to fund more CapEx internally.
7. Meta’s Real Problem: Can $220B of CapEx “Build Another Meta”?
- First, 周默 rejects 曹卿云’s premise. Meta’s absolute CapEx is the lowest, but CSPs have substantial outsourced CapEx as well as network, CPU, and other spending. Meta’s CapEx “linked to cloud, GPUs, and models has the highest purity and the highest concentration” of AI exposure.
- The quarter’s ad revenue grew 27%, impressions rose 14%, and pricing increased 12%. Excluding a $2.4B litigation provision and $1.2B of severance costs, margin was still expanding. But the market is unwilling to look through those numbers because Meta must now answer a different question: can CapEx equal to 70% of revenue create another 70%–100% of Meta’s revenue? This quarter offered no clear evidence of AI revenue.
- The math implies a 6–7 percentage-point negative impact from CapEx on next year’s gross margin, with the effect continuing into the following year. Profit growth next year may be only in the low single digits. From here, every quarter will be judged on how Meta generates AI revenue and how its margins evolve.
8. Why Meta Cannot Sell Compute Immediately: The SLA and H-Card/B-Card Mismatch
- The “debunking” of compute rentals should not be read as a simple “no sale.” 周默 believes Zuck was explaining that selling its own tokens could eventually carry higher margins than selling compute, thereby justifying other CapEx investments. If Meta receives an attractive compute offer, it could still sell.
- The first obstacle is SLA capability. Cloud providers must promise “three nines” or “four nines” of uptime and compensate customers at the original price or even above it when something goes wrong. Meta has never operated this kind of service; its data-center hosting capability is “not just worse than the CSPs, but worse than the Neoclouds, and even worse than CoreWeave’s.” It therefore needs to bring in more cloud specialists, including some hires from AWS.
- The second obstacle is card mismatch. Most of Meta’s compute is used for training, with relatively little inference demand. Its recommendation algorithms are not especially compute-intensive, and internal token usage at Meta Superintelligence Labs is also far smaller than training demand. Meta wants to rent out older H cards, which are more cost-effective for inference, while keeping B cards for training. Buyers such as Anthropic, however, want access to both H cards and B200s; they will not accept “renting 1M H cards with not a single B card.”
- Negotiations could therefore cycle repeatedly, with a real risk that no deal is reached. 周默 expects Meta’s earliest AI compute-rental revenue in late this year or early next year, not this quarter. The market, however, wants to see the path to AI revenue this quarter.
9. Business Agent Is the Next Advantage+: Toward Fully Automated Ad Buying
- This quarter’s disclosures showed that some AI models lifted Facebook ad clicks by about 8% and conversion rates by roughly 15%. Advantage+’s fully automated ad buying now generates more than $75B in annualized revenue, with roughly 1/3 of ad revenue already flowing through AI automation systems.
- 周默 sees a larger shift in Business AI’s repositioning as Business Agent. It goes beyond Advantage+’s automated buying by ingesting advertiser data and performing data, retention, and content analysis, potentially even monitoring competitors’ campaigns. When creating ad assets, the advertiser can discuss optimization in a conversational interface.
- It is “not just the optimizer portion.” The system also handles pre-optimization analysis, decision-making, and data integration, with the goal of taking over most of the ad-buying workflow for SMBs rather than only the optimizer function.
- Agency feedback recalled their reaction to Advantage+ several years ago, and some believe Business Agent could become a product Meta mandates within 1 year. The rollout may follow the Advantage+ playbook: push first into SMBs, then into large customers. If successful, it would be advertisers’ first step toward fully automated ad buying.
10. Does Meta Need a Frontier Model? “Good Enough Is Enough,” but WhatsApp Holds Greater Leverage
- On Meta’s existing business model, the answer is: “Honestly, good enough is enough for Meta’s business model.” Content, ad buying, and basic search do not require sufficiently complex agent planning for a much stronger model to add substantial value to the current business.
- A stronger model could unlock a new growth vector. Meta owns extensive personal-behavior data, and WhatsApp and Facebook Messenger may be the best environments for a Personal Agent to emerge. Meta could charge a subscription fee, or even charge higher consumption-based fees at a Personal Agent price point.
- If users develop deeper interactions with Meta, information on behavior, e-commerce preferences, and travel preferences could be anonymized into cohort-level data. That could help Meta generate more sales leads and develop lead-generation or commerce-oriented advertising.
- The data ecosystem is becoming richer. More companies can supply reinforcement-learning data, including data at the pretraining level. That could help Tier 2, Tier 3, and open-source models close the gap with frontier labs. For Meta, the return on investing in frontier models is now better than before, but that conflicts with the pressure on CapEx and ROI, making it difficult for management to abandon higher spending.
11. AWS Back to 36.7%: Anthropic and the Delayed Recognition of Blackwell
- AWS growth accelerated to 36.7% from 28% last quarter, marking the fifth consecutive quarter of acceleration and the fastest rate in 18 quarters. Sequential revenue additions exceeded $4.6B, 80% above the previous historical record. “The main acceleration still comes from Anthropic”: about 70% of its API is sold through CSPs, nearly 90% of that runs through AWS’s backend, and AWS can take 40% in this model. Anthropic’s ARR at the end of June was up roughly 110%–120% from the end of March, directly lifting AWS as the base expanded.
- The second driver is Blackwell. AWS accelerated procurement from Q1, but shipments typically take about 3 months to become recognized revenue because the systems must be debugged and installed in data centers. Q1 procurement was therefore recognized mainly in Q2, while Q2 procurement will continue to affect Q3.
- 周默 forecast last quarter that Q2 growth could reach 36%–37%, and the result validated that call. Because Q2 procurement was larger than Q1 procurement, Blackwell’s contribution will be more visible in Q3. Its contribution to AWS revenue is almost the same as H200’s. 周默 expects AWS growth to reach 41%–42% in Q3 and Q4 before gradually declining next year.
12. AI Spills Into Traditional Cloud: Graviton Price Hikes and the Illusion Around AI Revenue Mix
- 曹卿云 summarized the mechanism as “every $1 spent on AI spills over into traditional cloud consumption,” because post-training, reinforcement learning, and agent tool calls rely heavily on CPUs. 周默 agreed and noted that average Graviton pricing rose roughly 2%–3% this quarter, reflecting AI-driven demand and pricing for CPUs.
- CPUs do not have the same hard-commodity characteristics as memory. Price increases should contribute gradually to cloud revenue, perhaps 2–3 percentage points of pricing or utilization change per quarter, rather than a one-time jump of more than 10 percentage points. CPU increases could be higher next quarter than in Q2.
- On the claim that AI represents 15% of AWS revenue, 周默 distinguishes chip revenue from AI revenue. Chip revenue includes Trainium, Graviton, and Nitro; AI revenue is more software-heavy and may actually be around 20%. The comparable figures are roughly 45% for Google and 30% for Azure.
- The ratio is driven mainly by the base of traditional cloud revenue and shortages of power and capacity. AWS’s 20% could eventually reach 40%, while Google’s 45% could reach 60%. Given AWS’s much larger base, adding 10 percentage points to growth is already extremely difficult.
- AWS’s margin near 39% mainly reflects supply tightness, not the permanent realization of Trainium economics. Trainium utilization rising from 40%–50% to 70%–80% is equivalent to having more capacity at the same cost. But training utilization is already high, so it cannot be expected to provide another major margin tailwind.
13. Gemini’s Two Bottlenecks: Data Compliance and Half the Compute of Its Rivals
- The apparent contradiction needs to be unpacked. GCP’s acceleration this quarter “does not have much to do with Gemini”; it was driven more by TPU and GPU sales. Gemini’s own sequential growth fell from potentially more than 50% in Q1 to 20%–30% in Q2, roughly equivalent to OpenAI and Anthropic’s monthly growth rates.
- The first issue is a squeeze in competitive positioning. Google was previously the only Tier 1.5 company; that label now fits xAI more closely after its acquisition of Cursor and the resulting access to large amounts of high-quality data. Google’s compliance requirements are strict, so it is “basically half a beat slower than other companies” in buying data. Since May, the data ecosystem has changed sharply: suppliers are no longer limited to Mercor and Surge AI, and the pool of high-quality data accessible to other labs has multiplied. Data has instead become Gemini’s bottleneck.
- The second issue is compute. Google used only half as much training compute as OpenAI and Anthropic in Q2, roughly 1GW. Less data and less compute left Gemini behind in the quarter.
- On coding, 周默 believes Gemini may not beat Kimi K3, GLM-5.3, or xAI 4.6. The value of 3.5 Pro is already limited, and he does not expect it to launch in September or October; Gemini 4 will most likely arrive at the end of this year or early next year. Model training has become highly industrialized, so once Google secures more data and increases compute, catching up with Gemini 4 may be a matter of time.
- On reports that Jeff Dean may leave and Hassabis may move into a secondary role, 周默 believes some researchers pursuing Nobel Prizes may leave to start companies without affecting the current work, because the effort is now led mainly by researchers focused on training frontier models.
- 周默 explicitly rejects the assumption that enterprise customers care more about Flash’s cost, latency, and stability. If the competition is only on Flash’s cost-performance ratio, it will be difficult to beat xAI or compete with open-source models. Gemini still needs frontier-model and coding capabilities to generate more revenue—ideally performance only 10%–20% slower at a 30%–50% better cost-performance ratio. Flash likely did not contribute to GCP’s acceleration this quarter.
14. GCP’s 82% Growth Comes From Direct TPU Sales: 50% Gross Margin and $50B–$60B Over 6 Quarters
- GCP growth reached 82% and backlog exceeded $500B, with the biggest contribution coming from direct TPU sales. Roughly 100K cards generated $2B of revenue. “You are becoming more like Nvidia, not more like a CSP.”
- Direct TPU sales are recognized upfront, but next year’s TPU shipments will be 3x this year’s and sellable units may grow by more than 3x. They should therefore continue contributing to GCP’s acceleration at least next year. By 2028, TPU growth may finally fall below the growth of GCP’s other businesses and drag down the overall rate.
- Direct sales carry roughly 50% gross margin: a $20K selling price per card against $10K of Broadcom cost leaves about $10K of gross profit, potentially above GCP overall and above the margin of its CPU and GPU cloud businesses. From Q3 through the following 6 quarters, direct TPU revenue is expected to reach $50B–$60B, roughly 10% of backlog. Direct TPU sales currently account for about 10% of GCP revenue, while TPU rental, Vertex AI, and Gemini are closer to Recurring Consumption.
- After GCP’s margin rose from 20% to 35%, the business will need to use Neocloud capacity next quarter to ease the supply shortage. Because Neoclouds mainly provide GPUs rather than TPUs, Google needs to add external GPU capacity. 曹卿云 noted that this rental is booked as OpEx rather than CapEx, potentially implying greater total investment and margin pressure. 周默 agreed on the impact but argued that if Google fails to fulfill locked-in demand, customers may move to other clouds.
- CoreWeave’s rental pricing is high, close to spot or on-demand pricing. Google can earn a marginal profit, but margins will be affected.
15. Google Advertising: Inventory Cadence Drove This Quarter; Conversational Discovery Ads Drive Next Year
- The divergence between Search growth slowing from 19% to 17% and YouTube growth rising from 11% to 13% “does not have much to do with the models.” A recommendation algorithm contributes 0.5 percentage points in a quarter at best and likely contributed less this time. The main issue is that Google did not release more ad inventory, while Meta released substantial soft-ad inventory in Facebook Reels resembling TikTok and Douyin. The dollar index also reduced dollar revenue for each company by roughly 1 percentage point on average.
- AI Mode monthly active users have exceeded 1B, but Google has not disclosed revenue per query on the AI interface. Traditional advertising’s low-hanging fruit has already been harvested: AI Overview can add ads outside traditional placements and optimize allocation query by query. More than 60% of searches contain no ad-related keywords and remain unmonetized; long queries are the opening where AI Max and AI Overview can better understand intent and convert it into advertising.
- The next step is native advertising in AI Mode. Conversational Discovery Ads work like a funnel: when purchase intent is weak, the system does not show shopping cards; as intent strengthens, it adds shopping cards and commerce ads; near the decision point, it provides fuller product descriptions and comparisons.
- This format is fundamentally different from existing advertising. Google must teach agencies and advertisers how to use the new AI Max promotion model, while users also need time to adapt. Large revenue contributions should not be expected within 2 or 3 quarters, but the format could add 3–5 percentage points to revenue next year.
16. Next Year’s Supply-Demand Math: $1.7T of CapEx Requires $600B–$800B of ARR, While Open-Source Revenue Is Impossible to Track Cleanly
- 周默’s framework treats supply as CapEx and demand as AI ARR. Ideally, ARR at the end of next year should reach 1/3 to 1/2 of next year’s CapEx—a level at which ROI is relatively easy to calculate. $1T of CapEx requires $300B–$500B of ARR; if the major labs plus open-source models can reach $300B by the end of next year, the case becomes highly persuasive.
- But supply-chain inflation pushes 2027 CapEx to $1.7T, requiring $600B–$800B of ARR. The major closed-source model companies may total only $200B–$300B by the end of next year, making open-source revenue and new use cases essential, including CapEx from quantitative finance.
- The problem is that open-source revenue is highly fragmented. It sits in the revenue of companies such as Neoclouds and Fireworks, as well as in enterprise on-premises deployments. No one in the market can aggregate open-source model revenue into a comparable benchmark against closed-source models.
- The total could reach $600B next year, but it will take several months, more analysts, and clearer datasets to measure. With the hurdle and tracking difficulty both higher, the associated concerns and strategic battles will intensify.
17. The Pricing-Power Pendulum and Nvidia’s Dilemma: Is the Open-Source Gap 3–4 Months or 6–9 Months?
- Without this major change in open source, model companies would retain bargaining power over the long term because they capture most of the revenue and gross profit. Gross margin is ultimately a reflection of pricing power. 周默 notes that model-company gross margins could rise from 50%–60% to 85%–90%, signaling substantial pricing power.
- As annual revenue from a 1GW data center rises from $10B to $20B–$30B, model companies, CSPs, and memory companies can all make money. Model companies, memory companies, and Nvidia would capture the largest increases in share.
- But if CSPs host more open-source models and route token demand to them, the share could shift toward CSPs because “pricing power for open-source models in the U.S. sits with the CSPs.” Model-company margins would face competitive pressure, and industry profit could move from model companies to open-source model companies and then to CSPs.
- The key swing factor is whether the gap between open-source and frontier models remains at 3–4 months or widens back to 6–9 months. The former means the current trend continues; the latter could reverse it. Frontier-model parameters rose from roughly 1T at the end of last year to about 3T in May and June, and the next generation could expand another 2x–3x to 8T–9T. Starting with the next generation, both open-source and frontier-model iteration are accelerating.
- Expectations for Nvidia’s earnings at the end of the month should be kept in check. Nvidia’s impact on overall AI sentiment “is not that large”; Anthropic’s IPO and ARR progress across the labs matter more. After memory and optical-module prices rose, Nvidia has become relatively inexpensive on valuation. But the source discussion also includes the wording that Nvidia’s storage pricing arrangement involves cutting prices at a “1:3 or even 1:4 ratio,” an ambiguous formulation. 周默 believes that if GPU and memory can be decoupled more like TPU, CapEx pressure would ease; whether Nvidia chooses that to make CapEx more sustainable or simply to capture more profit is the company’s dilemma.