Vol. 80: I Went on 大内密谈 to Talk AI with 相爷
Summary
- Three years after ChatGPT launched, AI has moved from L1 chat to L2 reasoning and is now transitioning from L2 to L3 Agents. What is slowing progress is no longer model capability alone, but permissions, CAPTCHAs, payments, and platform coordination. 庄明浩 calls L2 “thinking” in quotation marks, while L3 is “crossing from language into behavior”; but today’s internet was built for humans, and the history of autonomous driving—from usable demo to real-world deployment—suggests Agents may still need 5, 10, or even more years to mature.
- China still trails the US on frontier models, but the gap with leading open-source models has narrowed from 1 to 1.5 years to roughly 3 to 6 months. DeepSeek was the first to achieve L2 “in the truest sense” at the end of January 2025, publishing its training, data, and post-training approach. Its paper disclosed a final training cost of about $1.28M, challenging the US capital narrative of billion- and tens-of-billions-dollar spending while quickly turning reasoning into an industry standard.
- The most dangerous risk in AI is not technological stagnation, but a circular financing system in which OpenAI, cloud providers, and chipmakers rely on future revenue to support one another. 庄明浩 describes the risk as “ten power strips connected in a chain, but no electricity”; his example is OpenAI using five-year commitments to induce Oracle to build data centers, Oracle buying Nvidia chips, and Nvidia investing in OpenAI. OpenAI has currently made about $1.4T in external commitments, while he estimates its 2025 revenue at roughly $20B and its maximum annual revenue in 2030 at about $200B—the numbers “just don’t add up, however you calculate them.”
- Nvidia is the clearest pick-and-shovel winner of this cycle, but its profits are also driving the broader capital loop to expand. Its market cap briefly reached $5T, quarterly revenue exceeded $6B, gross margin topped 70%, data centers may account for more than 80% of sales, and it held over $100B in cash. Meanwhile, a 1 GW data center costs about $50B, the largest project under construction is roughly 2.7 GW, and AI has pulled the power grid, nuclear energy, cooling water, storage, and even smartphone memory prices into the trading narrative.
- The application layer has moved past “is it usable?” and competition is shifting toward controllability, workflow integration, and who gets commoditized first. 相征 used Photoshop’s integrated Gemini 3 Nano Banana Pro to generate images and attracted 400,000 to 500,000 views without anyone recognizing them as AI-made. Mathematics, Coding, and other fields with clear answers have been “flattened” by reinforcement learning; AI music is already good enough for BGM, but may be “devastating” for ordinary and mid-tier creators, while copyright enforcement is clearly lagging.
- 庄明浩 believes “there isn’t much of a bubble in AI technology; there is a bubble once AI is combined with finance,” and that debt—not stock prices themselves—determines the destructive potential. Compared with Cisco’s roughly 100x P/E during the dot-com bubble, leading AI companies currently trade at an average P/E of “perhaps” around 28x, and GPUs can be put to work as soon as they arrive. The counterpoint is that GPUs depreciate over only 3 to 4 years, debt in the circular financing system keeps growing, and the final 6 to 12 months of a bubble are usually the craziest—so knowing something is overvalued does not mean knowing when it will turn.
- The most effective personal response is not chasing daily industry news, but handing AI each task you repeat more than 5 times a week and accumulating experiences, emotions, and judgment that AI struggles to replicate. Choosing an AI-related major is reasonable, but chasing lagging degree titles is not. There is no single answer on whether to keep learning programming, though understanding underlying logic remains valuable. 庄明浩’s final takeaway was not a tool list, but “祝你复杂” (“may you stay complex”): the more general the model, the more valuable individual experience becomes.
Deep dive
1. 400,000 to 500,000 Views Without Detection: Image Generation Has Cleared the “Good Enough” Bar
相征 ran an almost blind test: using the latest Photoshop integration of Gemini 3 Nano Banana Pro, he generated a faded-marker rendition of 《浪客行》 and an image of Sakuragi Hanamichi playing alongside Michael Jordan. After posting them on social media, the images drew roughly 400,000 to 500,000 views, and nobody identified them as AI-generated.
More tellingly, a friend who had studied drawing and knew 相征’s actual skill level also had no doubts, asking only, “You can draw this too?” 相征 replied: “I can draw it, but not that fast.” AI did not change the ceiling of ability here; it changed delivery speed and the cost for an observer to detect the difference.
2. ChatGPT Is Still Accelerating Three Years In
The recording took place at the end of 2025, exactly 3 years after ChatGPT launched on November 30, 2022. 庄明浩 calls this “an impossible-to-overstate milestone”: many previous technology buzzwords were disproven within 2 or 3 years, while AI is still intensifying.
The discussion has expanded from model parameters to energy, power grids, capex, and the meaning of human work. 庄明浩 says that once people call it the “Fourth Industrial Revolution,” it naturally begins to touch “the ultimate questions of humanity.”
相征 asked whether ChatGPT was truly recognized across the industry as a dividing line. 庄明浩’s answer was unequivocal: “Super important. Impossible to overstate.” The shock was not that AI had appeared for the first time, but that ordinary people were seeing a machine continuously generate language they could understand for the first time.
3. Generality Took AI from Go Demonstrations to Everyone’s Desk
Deep Blue, AlphaGo, autonomous driving, and StarCraft AI were all astonishing, but most people are neither professional chess players nor race-car drivers or esports competitors. ChatGPT went directly into everyday tasks—writing emails, making images, spreadsheets, and PowerPoint decks—bringing the feeling of substitution close to every white-collar worker for the first time.
庄明浩 attributes this generation’s difference to being both “general” and “generative”: it does not merely recognize faces or play Go, but appears to know something about everything and can demonstrate that capability through the most familiar human interface of all—conversation.
相征’s own usage changed just as directly: AI has replaced search engines in most situations and entered his design, illustration, audio-editing, and plugin workflows. “Fucking impressive” is no longer a technology demo; it is an action that used to take time suddenly getting done much faster.
4. Natural Language Is Rewriting Search and Translation First
相征 was once tormented by 2 interaction problems in the new macOS: saving a file activated every external hard drive first, while switching between multiple windows required an extra click before he could do anything. He gave the entire conversational complaint to DeepSeek, followed its instructions, and said the problem was solved; the original conversation did not clearly establish whether both issues were resolved separately.
庄明浩 sees this as the leap large language models represent over keyword search: users no longer have to break a problem into machine-readable terms first. They can describe their circumstances, intent, and desired outcome directly, while the model handles comprehension, reasoning, and output.
Online translation has effectively been “taken away completely” by this AI wave. His blunt summary: the era moved on without giving translators any time to react. Language went from being one AI capability to becoming the foundation for rewriting the entry points of other software.
5. GPT’s Breakthrough Was Not Just the Model, but Chat Exposing It to Everyone
庄明浩 explains GPT letter by letter. G stands for Generative: the model generates text or images, almost like “spitting out characters.” P stands for Pre-training: feeding it huge amounts of human data in advance. T stands for Transformer, a technical architecture proposed by Google that did not suddenly appear in this cycle.
He uses a domestic helper to explain pre-training: before arriving at a new home, she already knows the basics of kitchens, cleaning tools, and the general flow of 2 or 3 hours of housework. She is not someone who “knows nothing.” Models likewise acquire foundational capabilities through training before encountering a specific user.
GPT-3 was already extremely strong within the industry. ChatGPT used Chat to package GPT-3.5, converting research capability into a product ordinary people could directly perceive. Its difference from Siri, voice search, and voice-dialogue bots was that the latter visibly matched preset keywords, while the former appeared to genuinely know what you were saying.
The “large” in LLM refers to parameters and knowledge coverage far beyond those of earlier small models; “language” carries much of humanity’s knowledge and thought. 庄明浩’s inference is that if a machine can understand human language through 0s and 1s, “intelligence seems to emerge logically.”
6. LUI Wants to Take Over GUI, but the Permission Wall Is Harder Than Conversation
GUI defined decades of personal-computer interaction through windows, icons, and a mouse. LUI—Language User Interface—lets users control systems through natural language. 庄明浩 cites Bill Gates saying he has seen only 2 absolute technological revolutions in his lifetime, the first being GUI; 庄明浩 summarizes the current one as LUI, although that was not an explicit formulation in Gates’s original words.
豆包’s phone assistant shows the ideal form of LUI: say, “Compare several food-delivery apps and order the cheapest fried rice,” and the assistant searches, compares prices, chooses, and places the order in the background. The user no longer has to click through each app.
相征 immediately raised the privacy issue, and 庄明浩 added permissions, data, and payment risks. Why would WeChat allow an underlying assistant to operate it on a user’s behalf? The platform only has to decide that the assistant triggered a security policy, and the entire Agent chain is cut off.
Apple’s sluggishness can be understood in the same way: full LUI requires control over accounts, content, and payments, while Apple insists on privacy and system boundaries. Technical capability cannot be separated from platform governance; the latter may be the slower layer.
7. Tokens Track Business Scale; a Five-Level Roadmap Tracks Industry Ambition
庄明浩 explains Token as the unit of data consumed while a model receives input, thinks, and generates output. Traditional storage is measured in KB, MB, and TB; AI products commonly use Token volume to describe access scale and business growth, and cloud providers and application companies regularly publish the metric.
OpenAI’s 5-level roadmap borrows from autonomous driving: L1 is a chatbot, L2 is reasoning, L3 is an Agent, L4 is a researcher, and L5 is a manager or organizer. The industry is not merely upgrading model version numbers; it is trying to define stages on the road to AGI or ASI.
L1 lets a machine answer questions. L2 lets people see a machine “thinking” in quotation marks. L3 requires it to actually make a PowerPoint, visit a webpage, log in, like a post, order food, or call a ride—what 庄明浩 summarizes as “crossing from language into behavior.”
L4 imagines reaching human-researcher level and then being able to train itself. L5 would organize multiple researchers. The most accurate description today is still the transition from L2 to L3; the final 2 levels remain largely future expectations.
8. L2 “Thinking” Means Being Told to Slow Down, Not Human Consciousness
Early chatbots answered complex questions almost instantly, but humans usually need to slow down when thinking through difficult problems. One core method in reasoning models is to require the model to work through its chain of thought step by step before giving the final answer.
相征 asked whether the “thinking” displayed on screen might simply be a performance. 庄明浩 admitted: “To some extent, it is performing for you.” It is not equivalent to human consciousness, but slow thinking improves accuracy on complex tasks.
OpenAI released the relevant reasoning model around September 2024. From the second half of 2024 through early 2025, every major player aimed to reproduce o1. DeepSeek then became the first company 庄明浩 described as having “truly and completely achieved L2 and fully open-sourced the path,” after which L2 became standard.
9. Once Agents Act, the Human-Centered Internet Becomes the Bottleneck
An Agent does more than generate answers: it opens websites, clicks buttons, fills out CAPTCHAs, and calls external tools. But the existing internet “was built for humans, not AI.” Anti-crawling systems, abnormal-traffic detection, and login checks treat Agents as attackers.
This means L3 is no longer an internal problem for model companies; it is a coordination problem involving payments, platforms, webpage design, and real-world rules. A language model can be extraordinarily smart and still stop at a CAPTCHA, permission confirmation, or blocked API, asking a human to take over.
庄明浩 uses autonomous driving to warn against underestimating deployment time. Around 2010, there were autonomous-driving demonstrations that could travel 100 or 200 kilometers; Tesla also had a road-testable demo around 2013. Yet by 2025, the industry still had not reached the idealized L4.
One school of thought therefore believes that Agents may need 5 years, 10 years, or even longer to adapt to the real world. In the past, piling on parameters and compute could improve models; ecosystem interactions cannot be solved along the same linear path.
10. Language, Multimodality, and Coding Form the Three Main Tables
庄明浩 divides the industry into 3 tables: natural language, multimodality, and Coding, which is like “half a table.” Multimodality covers images, video, music, 3D, and world models; the latter even aims to generate a space that can be navigated forward, backward, left, and right from a single sentence.
The debate behind multimodality is that although human civilization is organized through language, perception and learning often begin with vision. 相征 mentions 周梦清 and 韩阳’s 《废墟》 project, which uses photo scans and 3D reconstruction to recreate Soviet buildings, statues, and churches close to collapse. Even if current generation is imprecise, the original data can be reconstructed repeatedly as models improve.
Coding is seen by some as a subset of language, and it is also the language humans originally used to communicate with computers, so AI is consuming it rapidly. 庄明浩 estimates that more than half of new code at some leading technology companies is already generated by AI. It still writes bugs, but the trend is strong enough to pose a clear challenge to programmers.
11. Four US Front-Runners Compete as Microsoft and OpenAI Move from Binding to a Limited Separation
The core US model players are summarized as OpenAI; Anthropic, whose Claude excels at Coding; Google, which has been particularly strong recently; and Elon Musk’s xAI. Microsoft was late to launch its own model because it had simultaneously been a major OpenAI investor, cloud channel, and revenue-sharing partner.
庄明浩 says Microsoft did not launch its own model until around July 2025. By the end of November that year, the 2 companies had reconfirmed their relationship in something resembling a “peaceful separation”: they would continue working together, but OpenAI would no longer be required to use Microsoft exclusively, and Microsoft could develop its own models. He says they could still maintain a fairly close relationship in 2032 while OpenAI pursued its own path.
12. China’s “Six Little Dragons” Recede as Open Source Becomes the Real Mainline
After ChatGPT appeared, China produced the so-called “Six Little Dragons”: Kimi, MiniMax, 智谱, 阶跃, 百川, and 零一万物. As the market developed, the capital required to sit at the main table became too high, while businesses remained unprofitable and continuous funding became increasingly difficult.
By 2025, the original narrative had split. 零一万物 shifted more toward enterprise-model implementation and To B work; 百川 later moved into healthcare rather than continuing to chase the largest general-purpose models. The companies still clearly competing at the model table were mainly Kimi, MiniMax, and 智谱.
Media and developer attention shifted toward DeepSeek and 阿里千问. 庄明浩 says “perhaps” all of the top 15 names on current open-source rankings are Chinese companies. Some leading US internet companies, after testing them, have also acknowledged using Chinese open-source models.
US mainstream vendors had originally focused on closed-source models. The advance of Chinese models forced them to release lower-tier open-source versions within their product families. 庄明浩’s conclusion is specific: “China is relatively ahead in open-source models.”
13. DeepSeek Rewrote the Cost Narrative with a $1.28M Final Training Run
After o1 launched, every main-table player tried to reproduce L2. 庄明浩 considers DeepSeek the first model to “truly and completely achieve L2,” and says that at the end of January 2025 it opened its model, training methods, data-model construction, and post-training details to the world.
This was not merely a callable product; it gave others a path they could reproduce. L2 therefore stopped being an advantage held by a handful of companies and quickly became an industry standard.
The most disruptive number was the roughly $1.28M cost of the final training run disclosed in the paper. 庄明浩 stresses that this was only the final run among many, but it still collided dramatically with the US media narrative of $1B- and $10B-scale data-center investments.
DeepSeek did not just explode in China; it triggered concentrated discussion across overseas technology and financial media. 相征 initially wondered whether it had merely called someone else’s model, but after seeing the overseas reaction, he found that “Silicon Valley had gone completely crazy.” 庄明浩 says the gap between leading Chinese and US models narrowed from 1 to 1.5 years to roughly 3 to 6 months.
14. 幻方’s Quant DNA Explains DeepSeek’s Compute Efficiency
DeepSeek’s parent, 幻方, originally did quantitative investing—high-speed computer trading. Quant strategies must capture spreads that exist for only a few milliseconds, so the company bought large numbers of cards and built clusters and data centers early, creating an exceptionally high bar for compute efficiency.
庄明浩’s summary is almost legendary: “He bought a lot of cards early to do this; then he realized that putting all those cards together might also let him build a model.” The hardware and engineering capabilities accumulated through quant investing were converted into assets for large-model training.
He also emphasizes that 幻方 took no external investment and rarely communicated externally, promoted itself, or pushed commercialization. Its founder 梁文锋 therefore carries a kind of “great love” in quotation marks. 相征 wanted to invite the team onto the show, but heard that 梁 was under tight protection when returning to his hometown, and eventually gave up.
When DeepSeek took off, 冯骥 called it “something at the level of national destiny.” Some people in China also questioned whether it had merely called on overseas work. 相征 again looked at the overseas discussion and found that “Silicon Valley had gone completely crazy.” 庄明浩 says the gap between leading Chinese and US models had narrowed from 1 to 1.5 years to roughly 3 to 6 months.
15. Once Models Clear the Human Testing Line, “Who Is Stronger?” Gets Harder to Answer
庄明浩 uses 柯洁 to illustrate the evaluation crisis. When 鲁豫 asked whether a professional Go player worried that AI would learn his style, 柯洁 laughed and replied: “AI is so strong—why would it learn our data? Isn’t that just garbage?”
If an ordinary person’s Go strength is 50 or 60 points, a world champion’s is 80 or 90, and AI has reached “100 million points,” humans can perceive no difference between 100 million, 10,000, and 1,000. 庄明浩’s analogy: “A top student scores 100 because the test only has 100 points.”
Traditional scoring also breaks down. Once model companies know the question bank, they can optimize for the exam and score highly without necessarily feeling better to ordinary users. Yet ordinary users have different preferences, and subjective evaluations cannot form a single yardstick.
Mathematics and Coding have therefore become popular testing grounds: mathematics has clear answers, while Coding can be verified by whether it runs. 庄明浩’s judgment is that wherever a clear standard solution exists, AI has basically “flattened” the field; the remaining constraints are mostly compute and time.
16. Reinforcement Learning Creates Superhuman Specialists and Exposes Cracks in Generality
Pre-training feeds huge amounts of human data into a model at once. By 2025, the industry increasingly emphasized post-training and reinforcement learning: reward a correct math answer, punish a wrong one; reward code that runs, punish an error. After countless rounds of computation, capabilities can leap forward.
But the human data that is easy to obtain and digitize has largely been consumed. What remains is either difficult to acquire or complex. Continued reliance on rewards with standard answers can produce a “specialist” who has mastered enormous numbers of algorithms for programming contests but lacks transfer ability.
This conflicts with the original emphasis on “generality.” Someone else might score only 70 in a contest yet be able to learn A, do B, and adapt to a new environment. 庄明浩 believes the industry is producing more and more of the first type without necessarily gaining genuine generalization.
Data labeling has also moved from object recognition to complex judgment. The work requires lawyers, doctors, and experienced finance professionals to assess answer quality and align it with human common sense. The problem is that human judgment itself is highly individualized.
17. Scaling Law’s Marginal Returns Are Moving to the Center of the Debate
A core researcher who left OpenAI—context points to Ilya—argued that the current direction may be a dead end: continuing to pile on compute, energy, data, and time may no longer produce a leap in general capability.
庄明浩 explains the debate through marginal returns. If investment increases 10x but performance improves only 10%, and another 100x increase produces perhaps just 5% more improvement, scaling law is nearing failure. The next step may require returning to research and finding a new architecture or technical paradigm.
He does not accept the view wholesale, instead asking listeners to consider the speaker’s position. Ilya has left OpenAI, founded a new company, and needs to tell a new research story. The new companies founded by OpenAI figures Ilya and Mira both raised billions of dollars in angel funding without products, models, papers, or websites. “Everyone has their own position.”
18. White-Collar Work Will Not Vanish at Once; It Will First Be Broken into AI Function Blocks
Applications are already entering law, finance, healthcare, marketing, sales, HR, and accounting, then subdividing further into PowerPoint decks, copywriting, and emails. WeChat, Feishu, DingTalk, Office, browsers, input methods, meeting software, and design tools are all adding AI.
While designing a cover for the 《山海经》 project, 相征 needed a shape with an exact number of corners. Instead of searching for an asset and tracing it in Illustrator, he asked ChatGPT to generate code, copied it into Illustrator, and obtained the shape with exactly the right number of corners.
The example highlights controllability. AI must not merely look right; it must execute constraints inside a mature workflow—16 corners, a specified size, or a defined structure. The ability to accept constraints consistently has become the capability contested by image, text, and video startups alike.
庄明浩 believes the more realistic form today is not an all-powerful Agent suddenly taking over a company, but “the AI-ization of functions that already exist and have been defined.” The impact on jobs will first appear as tasks are fragmented, reorganized, and compressed.
19. Generative AI’s Probabilistic Core Makes Controllability the Product Battlefield
This year’s Coca-Cola Christmas commercial was used to illustrate the consistency problem. Take frame-by-frame images of the truck from different angles and the number of wheels changes: sometimes 2, sometimes 3, sometimes even 4. A single usable frame does not make a continuous narrative reliable.
相征’s experience is that even if 8 or 9 out of 10 generations fail, one success is enough. More subtly, he turns around and blames himself: “Those failures weren’t its problem; you didn’t describe it clearly enough.”
庄明浩 explains that generative AI is fundamentally predicting the most probable continuation given the preceding context. “Today the weather is really…” is likely to be followed by “good” or “bad”; given only “today’s weather,” the range of possibilities expands sharply.
The more specific the input and the clearer the constraints, the closer the output should theoretically come to the desired result. But probabilistic generation means it is never fully deterministic. Prompting skill, model capability, and workflow constraints jointly determine the outcome; there is no simple rule that always works.
20. Music Is Good Enough for BGM; Copyright and Mid-Tier Creators Face the First Pressure
相征 says the show once used music he named himself, generated by AI and lightly adjusted, without any listener noticing. For much BGM that does not carry the core narrative and only needs to match a scene, the technology is already “ready to use.”
The music at the end of each 《正经叭叭》 episode is also generated by the host with 豆包: discuss saving money today, travel tomorrow, and the lyrics and style can match the theme exactly—甚至 “100 songs in a week.” It may not produce classics, but it removes the cost barrier to custom music.
The copyright debate remains: can models be trained on copyrighted music, and will Hollywood screenwriters be turned into “means of production”? Both speakers are fairly pessimistic, believing litigation is lagging and the training has already happened. 庄明浩 mentions that Suno and Warner have reached a settlement-like agreement, while Spotify already contains a large amount of AI music.
相征 believes the very best creators remain irreplaceable, but ordinary and even above-average creators will face “something devastating.” Once pure operational skill is commoditized, experience, emotion, and aesthetic judgment become more visibly the remaining surplus value.
21. Visual Production Costs Have Been Broken by “Generate in Bulk, Then Pick One”
In “AI 山海经,” sharks wear Nike shoes and fighter jets emerge from crocodile mouths—things that could never exist in reality, but can spread rapidly in a unified absurdist style. The point is not realism; it is combining imaginative recombination with low-cost production.
相征 mentions a music video 《影视飓风》 made for a Yunnan band. Reusing similar prompts 2 or 3 years apart had moved the process from barely generating anything to laying out 50 or 60 candidates across the screen at once. “You can always pick one that’s okay,” and generating dozens no longer requires high cost.
In the segment “If Every Shot Were Fake,” the team initially wanted to generate even the first shot of the valley. Because the result was poor, Tim was filmed against a green screen, and the later transformations and environments were handed back to AI. For ordinary viewers, the detail no longer matters; for practitioners, the production process is worth studying step by step.
22. AI Technology Has Not Peaked, but AI Financial Pricing Already Has a Bubble
庄明浩 first defines “bubble” narrowly: an asset’s price far exceeding its actual value for a period of time. Applied to the current situation, his answer is: “At this stage, there are definitely some bubbles in the stock prices of AI-related companies.”
Pure technological development is still advancing rapidly and has not reached an obvious peak. But the financial system can instantly price in a narrative that might become the future. Railways, fiber optics, real estate, and Bitcoin all went through similar processes.
In the second and third quarters of 2025, people were still debating whether it was a bubble. By the fourth quarter, the question had become, “What kind of bubble is this?” Technology iterates on a monthly basis, while capital markets have already extrapolated revenue and orders all the way to 2030. That mismatch in time horizons is the source of risk.
23. Circular Financing Links OpenAI, Cloud Providers, and Chipmakers into Power Strips with No Electricity
庄明浩 says OpenAI used expectations of roughly $300B in cloud-computing costs and orders over the next 5 years to induce Oracle to take a huge contract away from Microsoft Cloud, Google Cloud, and AWS. Oracle then had to borrow, raise financing, buy chips, and build data centers.
He illustrates the chain with a hypothetical example: Nvidia is both a supplier and a potential OpenAI investor seeking to lock in demand. An investment of $100B could correspond to OpenAI committing to 10 GW of future data centers, with each GW using roughly $10B worth of chips. OpenAI might then attract AMD with equity and data-center demand, leaving the remaining demand to Broadcom. These were examples 庄明浩 used to explain the circular structure, not individually confirmed contracts.
Chips, cloud companies, data centers, and model companies thus all enter the same loop. 庄明浩’s metaphor is: “Ten power strips connected in a chain, but no electricity.” What is missing is sufficient cash flow from end users.
After Oracle received news of roughly $300B in orders, its stock rose about 40% overnight and its founder briefly became the world’s richest person. A month later, almost all of the gain had vanished as the market recalculated and found that construction and operating costs might not even cover the debt.
24. OpenAI’s $1.4T in Commitments Is the Hardest Account in the Chain to Reconcile
庄明浩 estimates that OpenAI’s annual revenue was about $20B in 2025, while it expected to earn perhaps $500B over the next 5 years and incur roughly $300B in cloud costs. More importantly, it has already made about $1.4T in external commitments involving cloud-provider orders, data-center construction, and related spending.
In February 2024, when OpenAI’s revenue was around $2B, Sam Altman reportedly said the company needed $7T. That figure seemed almost insane at the time, but OpenAI has continued expanding its external commitments.
Even if 2030 revenue is pushed to an extreme $200B, revenue is not profit. 庄明浩 says OpenAI is unlikely to be profitable before 2029 and may burn another $100B on its own business. “However you calculate this money, it will never add up.”
ChatGPT has roughly 800 million weekly active users, but the world’s population, paid-conversion rate, and monthly subscription fee all have limits. 相征 asked why it does not enter China. 庄明浩 answered with a joke: “You don’t attend Tsinghua—is it because you don’t like it?” If user revenue cannot fill the investment commitments, the gap may ultimately become debt.
25. Nvidia Is Both the Clearest Pick-and-Shovel Winner and the Capital Provider Keeping the Loop Turning
Nvidia’s market cap briefly reached the first $5T valuation in human history. Quarterly revenue exceeded $6B, gross margin topped 70%, data centers may account for more than 80% of revenue, and the company held more than $100B in cash.
Gaming GPUs actually have lower gross margins, implying that data-center GPU economics are even more extraordinary. 庄明浩 says that if he were in 黄仁勋’s position, leaving cash idle would make no sense; investing in customers and suppliers while maintaining the demand loop is a perfectly rational form of profit maximization.
Individual rationality does not guarantee system safety. Oracle needs to catch up in the cloud era, while AMD wants to take share from Nvidia’s roughly 90% data-center position; both have to accept aggressive orders. Every company has a reason to bet, but together they may create tail risk that nobody is responsible for.
Google is the notable exception. It owns the full stack—applications, models, cloud, data centers, and TPU—and also invests in Anthropic, letting it use Google Cloud and TPU. It can avoid paying Nvidia’s high margins and choose to “open its own table.”
26. Capital Markets Are Starting to Use “Quadrillion”; Storage and Phones Are Being Squeezed by AI
After thousand, million, billion, and trillion comes quadrillion, or 10 to the 15th power. 庄明浩 remarks that AI commitments and human GDP have forced financial discussions into orders of magnitude that were rarely used before, making ordinary corporate valuation language start to fail.
He uses Microsoft Flight Simulator’s roughly 2.5 PB of data as a comparison. In the past, PB described an entire virtual Earth; now similar magnitudes are appearing regularly in capital and infrastructure commitments.
AI data centers are even consuming storage supply. 庄明浩 says 2026 orders for Samsung, SK hynix, Micron, and other manufacturers are already sold out, with Micron no longer making consumer products. Even if storage accounts for only about 1% of data-center costs, that is enough to alter global supply and demand.
After the Redmi K90 launched in November 2025 and was criticized for its price, 雷军 explained that memory prices had risen too quickly. 相征 then understood 王自如’s warning to “buy now if you can”: AI is taking high-end memory, and phones may face both higher prices and lower specifications in the future.
27. A 1 GW Data Center Is a $50B Construction Project, Not Just a GPU Purchase
庄明浩’s rough formula is that a 1 GW data center costs about $50B: roughly 40% goes to GPUs, more than 10% to networking equipment, with the rest covering CPUs, motherboards, power supplies, fans, memory, and storage.
The largest project currently under construction is about 2.7 GW. One Meta data center under construction, according to 庄明浩, is roughly two-thirds the size of Manhattan. Shenzhen, by comparison, represents perhaps 10 to 15 GW.
A cluster of 100,000 GPUs, each costing tens of thousands of dollars, requires billions of dollars in hardware alone. It must also receive continuous power, transmission, and cooling 24 hours a day; the equipment cannot overheat, and expensive compute cannot remain idle for long.
Technology companies are therefore investing in small nuclear reactors and other self-built power sources. Behind one data center are land, power stations, grids, water, construction, and redundancy systems. Discussions about “chip shortages” often see only the first step.
28. The Base of US-China Competition Is the Grid, Land, Cooling, and Engineering Capacity
相征 says people living in China find it difficult to understand “power shortages,” because stable, cheap electricity is almost assumed. 庄明浩 notes that in some countries, major cities may lose power 3 days a week. AI is magnifying this infrastructure gap into a model-capability gap.
The US power grid is more fragmented and aging. Data-center developers must coordinate with local utilities and sometimes build their own power plants. Land approvals, environmental groups, local residents, and transmission infrastructure can all slow projects; announcing an investment number does not light up GPUs.
庄明浩 cites Microsoft CEO Satya Nadella for support: the problem is not only that you cannot buy enough cards, but that once you have them, “there is nowhere to put them and turn them on.” New data-center power demand in the US may exceed half of total new electricity demand; the corresponding share in China may be only one-tenth, reflecting different levels of preparedness.
This is why the US-China comparison is not a matter of arrogance. Competition now involves capital, energy, data, and national-scale engineering. Smaller countries can build projects in the hundreds of megawatts, and Saudi Arabia has money and energy, but very few players can sustain a GW-scale race.
29. The US Has Bet on Its Only Positive Lever; China Is Embedding AI in Manufacturing
庄明浩 summarizes the US path as continued escalation in capital markets, chips, data centers, and training capacity. With the US economy, inflation, and politics intertwined, AI looks like “the only positive lever.” If it fails too, the market loses a major growth narrative.
A large share of the gains in US stocks over the past few years came from a handful of AI technology giants. Household accounts, pension funds, and Social Security money are all tied to them. “Too big to fail” is not strong enough to describe the situation, because AI now connects government, markets, and household wealth at the same time.
China still has narratives around advanced manufacturing, electric vehicles, embodied intelligence, and robotics. 庄明浩 believes the US is one generation behind China in this wave of embodied intelligence, while China can reuse the components, controls, vision systems, and supply chains accumulated during the previous EV cycle.
黄仁勋 repeatedly emphasizes that China may win, but he also has a lobbying interest: he wants the US to loosen chip restrictions on China, expand sales, and preserve Nvidia’s monopoly. 庄明浩 warns listeners not to “just listen to what 老黄 says.”
30. The Chip Gap Is Real, but Clusters, Inference, and Overseas Training Offer Detours
Domestic GPUs from Huawei, 寒武纪, 摩尔线程, and others are catching up. 寒武纪’s stock briefly surpassed Moutai, while 摩尔线程’s IPO subscription multiple was extremely high. But 庄明浩 clearly acknowledges that they remain generations behind Nvidia.
A CPU is like an old professor capable of complex scheduling; a GPU is like a large group of interns who can each handle relatively simple tasks. AI training is well suited to assigning clearly separable computations to huge numbers of “interns,” so an individual core does not need to be as versatile as a CPU.
China can combine more relatively weak cards into clusters, then add cheap power, network interconnects, and algorithmic efficiency to reduce the single-card gap. As usage shifts from training toward inference, the demand for top-end compute will also decline. The 2 speakers describe this as “besieging Wei to rescue Zhao,” or taking a curved route to the same destination.
In practice, leading Chinese technology companies also operate data centers overseas. If the main obstacle is that advanced cards cannot enter China, model training can happen abroad and the results can then be used in products. This does not remove the bottleneck; it temporarily routes around it.
31. Mining Companies Become Neoclouds, Inheriting Both Extreme Cost Discipline and High Leverage
The core model of large cloud providers is virtualization: split one machine among many customers and use peak-and-trough demand to raise utilization. AI customers, by contrast, often need 1 card or 1,000 cards dedicated entirely to them, with the cluster guaranteed to work.
Former mining farms understand this business best. They had already built precise cost formulas covering power prices, cooling, compute intensity, machine depreciation, and efficiency down to the decimal point. After mining became less attractive, many companies shifted into Neoclouds, renting GPUs and building and operating data centers for others.
Even large cloud providers such as Microsoft sign contracts worth tens of billions of dollars with Neoclouds because demand for dedicated GPU clusters is so strong. But these companies rely on debt to buy cards, replace cards, and expand; some generate less revenue than their interest expenses.
Criticizing excessive leverage may not change their behavior. 庄明浩 jokes: “You tell a bunch of miners their leverage is too high? Come on.”
32. Nebius Shows That “Too Risky” May Not Stop People from Betting
Nebius is linked to Yandex, the Russian search engine, through its former ownership. After the Russia-Ukraine war, Yandex was sanctioned and its business became difficult to continue. Its people dispersed, and Yandex’s CEO reconvened more than 1,000 former core employees in Amsterdam to establish Nebius and enter Neocloud.
For this group, an AI data center was not merely a valuation trade; it was a chance to resettle the team and continue their careers. 庄明浩 says outsiders can easily criticize leverage, but a CEO trying to provide for “more than 1,000 brothers” has a reason to seize the opportunity even if it could become the spark for a bubble.
The 2 speakers also touched on Bitcoin’s 4-year cycle. 相征 mentioned Bitcoin falling from a high of roughly $120,000 to the $80,000-plus range, and 庄明浩 replied, “Don’t sell for now; wait.” The exchange captures a holder’s attitude toward cyclical drawdowns but did not disclose how much Bitcoin 相征 holds.
33. China’s Application Layer Gains a Different Diffusion Speed from Open Source and Supply Chains
Even if leading Chinese open-source models trail the US by roughly 3 to 6 months, their capabilities are already sufficient to support a large number of applications. Combined with more than a decade of mobile-internet experience, application-layer expansion may be more promising than simply chasing the strongest model.
AI toys are a typical diffusion path. Shenzhen’s supply chain can quickly produce a dog, cat, or dinosaur with cameras, voice, and multimodal recognition. 庄明浩 believes these products can answer more questions than Sony’s robot dog could in its era.
Autonomous driving and robotics originally focused on hardware structure, control, and movement; AI solves the problem of the “brain.” Once combined, both humanoid and non-humanoid robots can move up another level.
庄明浩 therefore sees China’s AI diffusion as more like a garden in bloom: if software does not work, build hardware; beyond models there are cars, robots, and manufacturing. Even if one area temporarily lags, China is not as dependent as the US on a single, more financialized narrative.
34. Apple and WeChat Are Using Ecosystem Control to Buy More Time
Among the “Magnificent Seven,” Apple has invested relatively little in AI infrastructure. The respectable explanation is privacy and an intact user experience. 庄明浩’s more direct speculation is that Tim Cook is nearing retirement and is not an aggressive, innovation-driven CEO, so he has little incentive to take a massive risk that could damage his legacy.
相征 bought a non-mainland iPhone specifically to use Apple Intelligence, yet after several years still had not truly experienced the promised AI capabilities. The everyday difference from upgrading from an iPhone 14 Pro was also minimal. The departure of Apple’s AI director suggested to him that “this job is impossible.”
Department heads must both build frontier AI and avoid violating Apple’s principles on privacy, permissions, and ecosystem control. If the boss will not actively take the risk, middle management cannot make the critical trade-offs. Apple can still use control over its devices to wait until the technology is more ready, but repeated delays will consume credibility.
WeChat faces a similar situation. Its clearly visible AI functions are currently limited to things like search and tagging 元宝 in comments. It handles enormous privacy and data concerns and holds a sufficiently strong monopoly position, so it has no need to move too aggressively during the most chaotic phase.
35. Even a “Good Bubble” Will Burst; the Difference Is Whether Damage Travels Through Equity or Debt
庄明浩 uses a story about Bezos living through the internet bubble to introduce the distinction between a “good bubble” and a tulip-style bad bubble. Another framework looks at whether financing is equity or debt and whether the technology raises productivity. AI initially sat in the safer quadrant of “productive technology plus equity,” but debt is increasing. The episode did not explicitly say that “good bubble” was Bezos’s original wording about current AI.
Compared with 2000, Cisco’s peak P/E was roughly 100x, while leading AI companies currently average “perhaps” around 28x. Around 97% of fiber-optic cable was not lit then, whereas GPUs are in short supply as soon as they are connected today. The Federal Reserve was hiking rates then, while it was cutting rates at the time of recording. All of these points were used to support the claim that “this time is different.”
The negative difference is that fiber could support the internet for decades, while a GPU may become obsolete after only 3 to 4 years when a new card is released. Data centers must earn back their costs during the depreciation period, while continued upgrades create another round of capex.
Bubbles are usually first supported by people who believe in linear gains and by those who understand the danger but do not want to miss the final profits. Then comes “the party must go on,” followed by the crash. In a Q3 survey this year of more than 3,000 core US fund managers, Bank of America found that 93% believed market prices were overvalued. Yet they could not easily sell because the final 6 to 12 months are often the craziest, and the turning point cannot be predicted accurately.
36. The Most Effective Personal Strategy Is to Use AI and Hold On to “祝你复杂”
庄明浩 tells ordinary people to spend less time chasing news about $50B or $500B commitments: “Whatever AI or not AI—just enjoy life.” A more practical rule than tracking every financing round is to take tasks repeated more than 5 times a week and ask AI, one by one, whether it can improve efficiency. If you do not know which tool to use, ask another AI directly.
The applications can be small: photograph food and ask about calories and health information, photograph a homework problem and ask AI to explain it to a child, or make images, memes, and short videos for fun. If the goal is simply to stay aware of the industry, listening to a few summaries a year and following a small number of channels such as 卡兹克, 赛博禅心, or The Way to AGI is enough.
It is reasonable for young people to choose an AI direction, but university majors are inherently slow to update. There is no single answer on whether to keep learning programming: people may no longer need to hand-write large amounts of code, but understanding underlying logic remains valuable. More reliable than chasing new degree titles are foundational abilities, genuine interests, and things that have generated long-term positive feedback.
庄明浩 ultimately brings the answer back to the individual. Even if skills are replaced, AI struggles to replicate a person’s experiences, emotions, subjective cognition, and judgment built over time. “If you’ve already found it, slowly make it bigger; if you haven’t, keep looking.” 相征’s “祝你复杂” (“may you stay complex”) thus became the episode’s most important conclusion.