Pioneers Insight Method Research Author
172: Revisiting 曹旭东 After Momenta’s IPO: We Just Want to Build AI Without End
Back to Episodes

172: Revisiting 曹旭东 After Momenta’s IPO: We Just Want to Build AI Without End

Summary

  • The endgame window for assisted driving may already be closed. 曹旭东 stands by his judgment from last October: competition will end in 2026, leaving 2-3 players in China and 3-4 globally; “the window of opportunity that could change the landscape … may be closed now, or may close by the end of this year.” Momenta and Huawei together now account for 90% of urban NOA and have strong pricing power; second-tier players have no bargaining power, only 1-2 customers, and not enough revenue to fund R&D, leaving them squeezed by first-mover and scale advantages.
  • Assisted driving is following a 10x-per-year Moore’s Law. By 2028, even an L2++ passenger car with driving safety “at least at human level” will turn the product from a nice-to-have into a must-have. The cadence is powered by training-compute investment: about $200M this year, $400M next year and $600M in 2028; gross margin rose from under 20% in 2023 to over 70% in 2025, while losses narrowed from nearly RMB1.1B to just over RMB300M; the plan is to break even in 2027 and turn profitable in 2028.
  • He is very much looking forward to FSD entering China. As Tesla’s entry accelerated the shakeout of weaker EV brands, FSD will pull the whole industry into competing harder on safety, quality and experience, but it will have no material impact on the endgame of 2-3 players in China and 3-4 globally; if it enters by year-end, Momenta’s R7 against V14 will be “evenly matched, each with its strengths.”
  • Robo’s 5-year ranking is Robo One first, Robo Truck second and Robo Taxi third. Truck means bulk logistics—hauling coal and ore over rough suburban roads for 30-50 km—with stronger demand and less competition; Robotaxi only provides an “AI driver” and earns a revenue share, with capex held by platform partners, so it does not compete with Didi or Amap. The fleet is projected to grow from 30-plus vehicles at year-end 2025 to hundreds in 2026, tens of thousands in 2027 and more than 100,000 in 2028.
  • Momenta will enter robotics in 2027, betting on the 2030 inflection point for home robots. The robot body needs to cost $10K-$20K, ideally $10K; edge compute might be solved in 2028 with autonomous driving paving the way; under the world-model paradigm, more than 80% of autonomous driving’s infra, talent and experience can be reused, while robotics R&D will require at least $10B and potentially several tens of billions, funded by assisted-driving gross profit rather than fundraising burn. A signal of the shift: Silicon Valley’s Generalist used pre-training on non-embodied data and post-training on embodied data to lift its success rate from 50% to 90%.
  • His data-driven conviction has held for 10 years, rooted in a simple dimensional estimate. To make autonomous driving 10x safer than humans means 1 major accident per 1B km; statistical confidence requires roughly 100 events, or 100B km of data, which no proprietary fleet could collect—hence the need for mass production and the resulting “1 flywheel, 2 legs.” The methodological foundation came from 孙剑: “top-tier work is often made, not thought up,” and “truly world-class people are smart people who put in the hard work.”
  • His aside on the large-model market carried unusual investment value. His first take in Silicon Valley in January was that Anthropic “might have been the cheapest AI company at that point”—an L2 coding agent is only a $10B market, but at L4 coding alone becomes a $1T market, while “a general-purpose L4 agent … is a $10T market.”

Deep dive

1. The Window May Already Be Closed: 2-3 Players in China, 3-4 Globally

  • 曹旭东 reiterated his judgment from last October and defined what “competition ending” means: “The window of opportunity that could change the landscape … may be closed now, or may close by the end of this year.” The scope is third-party suppliers, excluding automakers’ in-house efforts and Tesla.
  • The logic rests on only 2 points: “The first-mover advantage is especially pronounced, and so is the scale advantage. Put those 2 together and you get the conclusion above.”

2. 90% Share and the Pressure on the Second Tier

  • Current snapshot: Momenta and Huawei together account for 90% of urban NOA, and “the 2 of us have fairly strong pricing power; our prices are clearly higher than the second tier.”
  • The mechanism is a 10x-per-year assisted-driving Moore’s Law, funded by the leaders’ rapidly growing revenue and gross profit—“in practice, not through financing.” Second-tier players lack pricing power, can secure only 1-2 customers, face heavy development and delivery costs, and have “very, very limited” money left for core technology R&D, leaving them constrained by first-mover and scale advantages.

3. Concentration Will Continue to Rise: The 5-Year Logic of OEM Sourcing

  • 曹旭东 believes the high share is not only sustainable but “will actually strengthen.” When an OEM designates a supplier, the vehicle program must be mass-produced and receive OTA updates for 5 years. It first asks whether “this supplier can stay alive for 5 years,” then whether the technology gap with the first tier is “narrowing or widening” over those 5 years.
  • The deeper shift is a business-model migration: as one-off licenses become a subscription model, “it is no longer enough to rely only on the OEM’s brand.” Consumers need the supplier’s brand and product experience before they will pay.

4. Only 1 Slot Remains; Tesla Cannot Be a Supplier

  • Asked whether chipmakers vertically integrating solution stacks could break the market structure, 曹旭东 conceded it “would be a variable”—hence 2-3 players in China. “One is us, one is Huawei” are the 2 certain names; the third slot is undecided.
  • Tesla is explicitly excluded: “Tesla will not be a supplier … everyone cannot possibly buy Tesla; who would buy Tesla?” An automaker that builds its own cars and also wants to sell assisted driving is “actually a constraint” on scale effects.

5. From Fun Toy to Must-Have: Human-Level Safety by 2028

  • In response to the host’s challenge that assisted driving is still not a core car-buying factor, 曹旭东 said the timing was off: “That may have been true 2 years ago.” Two years ago it was “a fun toy”; now it has just crossed the threshold of being useful. At the current 10x annual pace, even an L2++ passenger car will have driving safety at least at human level by 2028, turning it from a nice-to-have into a must-have.
  • The host cited a Nielsen survey showing assisted driving still ranks behind price, exterior design and cabin space in purchase decisions. 曹旭东 accepted the point: “Getting it above price and exterior design is actually pretty hard …” so “making it 4th would already be impressive”—a few years ago it was near the bottom of the top 10.

6. A Top-Executive Project: Chairmen Now Discuss Experience Personally

  • The clearest customer-side signal is the rising seniority of whoever owns assisted-driving experience: from the assisted-driving lead to the head of intelligentization, then the head of R&D, and finally the company president or chairman. Assisted-driving experience has become a top-executive project—“he will talk to me directly about what the assisted-driving experience should be.”

7. R7 Versus R6: “Wow” in Construction Zones

  • Can experience create separation? 曹旭东 compared 2 generations of Momenta’s models in construction zones. R6 could handle them and was reasonably safe, but felt “a bit sluggish,” like a novice driver. R7 might catch only a glimpse of the scene and immediately begin avoiding or changing lanes; its timing was very early, the maneuver very smooth, and its interactions with other cars “reassuring and smooth.”
  • His own test-drive impression: if you were not paying attention, you might not even register that a construction zone was ahead. Only when the car began avoiding it would you wonder what happened, then look and realize: “It may have noticed it earlier than a human.”

8. Assisted Driving Is Like Batteries: Only Good and Better

  • Asked whether leading players would develop distinct styles like large-model companies, 曹旭东 offered the episode’s key industry characterization: “Assisted driving may be a lot like batteries: there is only good and better … it is not a matter of different tastes.” Style can be configured, but the good-to-better portion is 80%-90%; genuine differentiation may account for only 10%-20%.
  • The implication is concentration: industries defined by good versus better have very strong first-mover and scale effects, so the final player count will not be large—like batteries in relation to CATL. Industries where “different tastes” rule, such as apparel and car brands themselves, are the ones that can support many winners.

9. The Market Prediction from 4 or 5 Years Ago Was Wrong: Gross Margin Rose from Under 20% to Over 70%

  • The host revisited the old consensus that in the smart-EV era, car brands would consolidate like smartphone brands, while third-party suppliers would fight a price war and squeeze margins. Momenta’s gross margin instead rose from under 20% in 2023 to over 70% in 2025. 曹旭东 said the old view had always struck him as strange and inconsistent with the industry’s nature.
  • His counterexample was human nature: cars naturally have attributes of clothing and housing—“no one likes wearing a school uniform every day.” Car brands therefore will not converge to 2-3; it is suppliers that will consolidate.

10. Very Much Looking Forward to FSD in China: Evenly Matched by Year-End

  • FSD entering China “will have no material impact on the directional or endgame call, but may accelerate the process.” Just as Tesla’s China entry cleared out weaker EV brands, it will push the whole industry to compete harder on safety, quality and experience. 曹旭东 repeated twice that he was “very much looking forward to it.”
  • The host pressed him on a direct matchup: shouldn’t FSD theoretically be better? 曹旭东 acknowledged that the U.S. is currently testing V13 and that V14 “will certainly undergo a qualitative change in China,” but held to his conclusion: if it arrives by year-end, R7 against V14 will be “evenly matched, each with its strengths.”

11. In-House R&D Is an Internal Supplier; R&D Arms Race: GPU Investment 2→4→6 ($200M → $400M → $600M)

  • On the long-term pressure from OEM self-development, he framed it as a contest on the same field: “Treat it as an internal supplier; we are an external supplier.” The market will continue to consolidate. Keeping an internal supplier makes sense as a business strategy, but technology and mainstream adoption lie with third parties.
  • The R&D war chest: Momenta has spent “more than $1B but less than $2B” cumulatively on autonomous-driving R&D to reach its current share. R&D spending may continue doubling in coming years; the increment is not headcount but data centers for training—about $200M this year, $400M next year and $600M in 2028—to support 10x annual improvement.

12. We Are Not a Cost Item but a Profit Driver

  • Would price wars in the car market pass margin pressure on to suppliers? 曹旭东 reversed the narrative: assisted driving “is not given away for free; it is a paid option, and its take rate is very high.” Momenta helps customers sell more vehicles, at better prices and with higher gross margin—that is customer success.
  • As more supplier brands offer the service, would the uplift be competed away? He returned the question to the automaker’s core job: know who its users are, define products from the brand, derive technology from the product, and integrate it. Building the capability in-house versus using the best supplier “is not fundamentally different—just as building your own battery is not fundamentally different from using CATL.”

13. Vertical Integration Is Leverage; BYD Won with the Han

  • Asked about Tesla and BYD as counterexamples to vertical integration, 曹旭东 said the answer was “right and wrong.” BYD spent 10 years vertically integrating, then only gained real volume with the Han around 2020: “Vertical integration is a bit like using leverage to grow.” It caught the expansion of China’s new-energy market and made the right bet. But leverage was not the essence; the essence was that it built a product as good as the Han at that moment.

14. Brand = Trust = Consistent Over-Delivery; The IPO Was Not for Money

  • That led to a broader brand argument: China has many automakers with “products but no brand.” A brand is mindshare and trust, built by consistently over-delivering from 1 generation of product to the next. What gets over-delivered is “what you want but sometimes cannot articulate.” He singled out Huawei as having some brand effect, even though it is not a vehicle model.
  • For Momenta, the key strategic goal of this IPO was not fundraising but “branding and trust,” for consumers and investors alike. His self-assessment was blunt: awareness is average, reputation is decent. The image he wants to build matches the 10-year vision: safety and peace of mind.

15. Profitability Timeline: Break Even in 2027, Profitability in 2028

  • Losses narrowed from nearly RMB1.1B in 2023 to just over RMB300M in 2025. This year will still be a strategic loss, narrowing “a little more”; the plan is to break even in 2027 and turn profitable in 2028. It could choose to lose money more aggressively for longer, but as a listed company, “if losses keep expanding, it may harm investors’ interests.”
  • The pace assumes R&D spending follows the GPU curve above—“that would already be pretty good”—and he is confident Momenta can deliver 10x improvement every year. Narrowing losses is a balancing outcome under that first strategic objective.

16. The Robotics Ledger: At Least $10B in R&D, Funded by Gross Profit

  • Robotics is a “natural extension” of autonomous driving, but capital discipline is explicit: Momenta would rather fund R&D with gross profit generated by autonomous driving than “keep burning money or continually raise from the public markets.”
  • The scale estimate is unambiguous: robotics requires “at least $10B, potentially several tens of billions,” because the jewel of the application set may be the home robot, the most complex and commercially valuable setting. Was 2025 gross profit of more than RMB1.7B enough? “I believe sustained gross-profit growth will definitely support us in building the robot.”

17. 2 Conditions Changed His Mind: The 2030 Inflection Point and a Working World Model

  • When interviewed in 2024, he said Momenta would not yet do robotics. The first changed judgment is that home robots will hit an inflection point in 2030 and begin commercializing at scale. The second is the technical paradigm: world models can make the physical world’s massive data genuinely usable. In autonomous driving, middle training may learn driving common sense and post-training may align behavior; robotics follows a very similar R&D pattern, with “data infra and training infra highly reusable.” Momenta pre-researches 1 generation while mass-producing the next, and validated the approach last year.
  • Silicon Valley provided corroboration: conversations in the first half found that the best robotics companies were moving in this direction, with Generalist as the representative. It pre-trains on large volumes of non-embodied data and post-trains on embodied data, lifting success from 50% to 90%; the result “sent a major shock” through robotics and ignited a data industry, where 100,000 hours is entry-level and 10M hours is massive.

18. How to Project the 2030 Inflection Point: Edge Compute, the Physical Brain and a $10K Body

  • The forecast starts with hardware. Edge compute “should not be much of a bottleneck”; with autonomous driving paving the way, it can be solved—and solved well—in 2028. Architecturally, “the physical brain must sit on the edge”: cloud latency and instability are unacceptable. The language brain can be in the cloud or on-device, though a good experience also needs a small edge model. Privacy is therefore manageable: physical-brain data naturally stays off the cloud.
  • The other leg is the robot body. Reliability, precision, cost and supply chain should all be “pretty good” by 2030. The mass-market threshold is a body cost of $10K-$20K, ideally $10K—corresponding to the “around RMB100K” price he later quoted to his aunt.

19. Why Wait Until 2027: 3 Pots, 2 Lids

  • The reason for not entering in 2025 or 2026 is the order in which capabilities spill over. 2028 is the key year for autonomous driving: Momenta must first get the data flywheel upgraded into an L4 flywheel and a Robo flywheel and turning smoothly. Once R&D and organizational capabilities have spilled over, robotics becomes “much more of a natural progression.” Force it before then and it becomes “3 pots, 2 lids—however you cover them, something is left uncovered,” which would be painful.

20. Same Age, Different Path from the Company That May Be Unitree: Start with the Brain, Design but Do Not Build the Body

  • Momenta was founded in 2016, the same year as a company that may be Unitree and that also plans to go public this year. That company likely approaches robotics from the body; Momenta starts from the use cases it wants to build, from the brain, and then designs the body. It will design the body itself but not manufacture it: China’s supply chain and ecosystem are complete enough that it does not need to build hardware in-house.
  • Asked whether it would follow the platform model—Google and Nvidia building an Android for robotics—or pursue a full-stack hardware-software model, he declined to choose a camp. The test is whether the work adds value: if someone already does it well, Momenta need not; if it matters to success and no one is doing it, Momenta will do it. The answer remains open: “we’ll know when the time comes.”

21. Versus Embodied-AI Newcomers: Reuse 80%+, Add a Cash Cow

  • Compared with embodied-AI newcomers founded in 2024 and 2025 and raising huge sums in the short term, Momenta’s hurdle is that the task needs both a physical brain and a language brain, implying the $10B-scale R&D budget described above. Its advantages are that training infra, data infra, accumulated talent and world-model experience can be reused “80% or more,” plus a cash-cow business; autonomous driving, he expects, will see explosive growth in the next few years.
  • Asked about relative weaknesses, he gave a deliberately vague answer: “There may simply be more factors to consider.” Resource allocation among the 3 lines is discussed very little internally; the company works like a startup, validating “1 minimal closed loop at a time” rather than splitting resources top down.

22. Robo Ranking: Robo One First, Truck Second, Taxi Third

  • He ranks the businesses by whether they will be good businesses over the next 5 years: Robo One—same-city delivery using a 4.2-meter Jinbei van plus a small three-wheeler—first, Robo Truck second, Robo Taxi third. Truck does not mean trunk-line logistics—“trunk-line logistics is the last piece of the autonomous-driving puzzle,” extremely difficult and, if done at all, last. This truck is bulk-commodity logistics—coal and ore, 30-50 km short hauls and 100-200 km on longer runs—on remote, even barren roads with potholes and animals crossing; “in some places, wolves even wander along the road.”
  • The reason for choosing these segments is purely commercial: demand is stronger, while competition should be lighter than in Robotaxi.

23. Robotaxi Only Provides the AI Driver: Partners Carry Capex

  • When comparing his business with Pony.ai and WeRide, whose 2025 gross margins were about 15% and 30%, respectively, he described the difference as a strategic choice: Momenta is not building a platform to compete with Didi, Amap or T3. Its role is an “AI driver,” and it earns the AI driver’s revenue share. The fleet and capex are held by third-party platform partners. The logic is still adding value: “I do it when I can create incremental value; when I can’t, I don’t”—which also avoids unnecessary arms races.
  • The sudden entry of JD.com, Meituan and Hellobike into Robotaxi is related to “Waymo making very good progress,” along with Waymo securing a very large financing round and a strong valuation.

24. 1 World Model Supports Every Application

  • The foundation that makes all 3 curves possible is the world model. Momenta has validated that “1 world model can support every autonomous-driving application”: mass production, Robo One, Robo Truck and Robo Taxi all work. If the foundation model improves 10x, every application could improve 10x as well.
  • In the language-model context, he said, this sounds obvious: 1 model can do math and coding. The new validation is in Robo One and Robo Truck; Robotaxi and passenger cars are “exactly the same” under Momenta’s technical route, while the truck scenario “still has differences.”

25. After Clearing All 3 Hurdles, What Is Momenta?

  • Once all 3—mass production, Robo and robotics—are working, he does not have a carefully prepared definition: “Momenta is just AI … everything points toward physical AI; at bottom, it is an AI company.” How large it will be is “hard to say.”
  • He would rather talk about the original impulse. When he founded Momenta in 2016, the team was not the most glamorous—he was 30, colleagues were 24-29, and 200-300 companies were in the field. The driver was not becoming number 1, but that after years in perceptual intelligence he wanted to do cognitive intelligence; autonomous driving was a good entry point with huge customer and social value. The 3 curves are like a video game: after clearing the first level, it is fun to enter the second.

26. His Aunt’s Eyes Lit Up: The Moment for Home-Robot Demand

  • The choice of the next level was partly designed—like schooling, primary school before middle school, with each step building for the next—and partly triggered by a moment that made the demand real. He told the full story of his aunt: when he explained autonomous driving in 2023, she said “excellent”; when he got to robots, “my aunt’s eyes began to light up.” Could it take a child to kindergarten, buy groceries and clean the house? When he returned in 2025, her first question was: “Have you built that robot yet?”
  • He estimated a price of around RMB100K. His aunt said it was “not expensive … if it can do all those things, RMB100K is good value.” What shook him was that she remembered it: people often forget cutting-edge technology after a conversation. That is the huge value home robots could bring to older people and children amid aging and falling birth rates. Momenta cannot treat robotics as a supplier job because there is no industry consensus on what a home robot is, no off-the-shelf product or form factor; it has to define the value proposition end to end.

27. The Starting Point: Statistical Physics, Railway Cameras and a Circuitous Route to MSRA

  • His interest in AI began in 2006, his 2nd year of university. Trained in mechanics, he loved statistical physics, which led to statistics and then statistical learning: “learning knowledge from data, even distilling wisdom from it, felt miraculous.” He never learned mainly through classes: “I start from the problem, learn whatever I need, and if the implementation does not work, I go back and look again, iterating until the problem is solved.”
  • His first internship was not at a marquee AI lab but at a state-owned research institute doing image processing for railways: high-speed cameras filmed train undercarriages to identify loose components and stone damage. Microsoft Research Asia rejected him for lacking a formal AI background and track record; after detours through Sony Research and Yahoo Research, he joined MSRA in 2010. “Frankly, I was never that fond of studying … I was not that interested in publishing papers; I preferred solving practical problems.” A new method unlike prior work earned him 孙剑’s job offer, and he did not pursue a PhD.

28. 孙剑’s Legacy: Top-Tier Work Is Made, Not Thought Up

  • Before joining 孙剑’s group, he saw himself as a logic-and-theory thinker and imagined good work as Einstein discovering relativity. 孙剑 placed enormous weight on experiments, especially phenomena “very far from what you expected,” changing his view: “top-tier work is often made, not thought up.” There are too many smart people in the world; everyone knows the top-level principles and frameworks, so the unknown that only you can see is often hidden in anomalous data.
  • That became his test for people: a smart person can make any argument logically self-consistent, but that does not mean it matches the raw data; sometimes smart people do not like looking closely at the data. “Truly world-class people are smart people who put in the hard work.” The standard matters even more in the large-model era: iteration cycles are longer, wrong assumptions cost more and create longer stalls, so “never take things for granted.”

29. The SenseTime Years and the Anthropic Aside: It Comes Down to Value Creation

  • In 2015, 汤老师 brought him into SenseTime, where he led R&D for every deployed product line. “At the beginning of the month you brag about how good the technology is; by the end of the month you have to deliver.” He recruited interns from Tsinghua to build and ship, often working until 3 or 4 a.m., sometimes sleeping at the office—the most physically intense period of his life. The all-in instinct is innate: he focused so hard on his senior thesis that he developed dry eye. “Your nature can make you exceptional in some areas; non-native strengths can be learned and can reach first-rate, but first-rate is sometimes the ceiling.”
  • Comparing the capital outcomes of 2 AI waves—the twists and turns of the “4 CV dragons” versus Zhipu’s market cap breaking RMB1T after more than 6 months as a listed company—he reduced it to 1 phrase: “value creation.” In January, during a Silicon Valley trip, he already knew Anthropic’s AI business was surging. His first call was that Anthropic “might have been the cheapest AI company at that point”; the PS math made it clear. Turning an L2 coding agent into L4 takes the market from $10B to $1T for coding alone; “a general-purpose L4 agent is not a $1T market—it is a $10T market.”

30. The Origin of the Data-Driven Conviction: Disappointment at Google and a 100B-Kilometer Calculation

  • In 2016 he chose autonomous driving over robotics because the cognitive-intelligence complexity of general-purpose robots was too difficult at the time. He also considered joining Google, but left disappointed: in discussions, people did not think from first principles about what was essential to making it work; ask a question and they jumped into specifics of an algorithm module. “If you only see problems and solve problems, you may never reach the moon.” His own conclusion was simple: make the architecture data-driven and secure massive data.
  • 2 things fed that conviction. One was a long interest in biological intelligence and On Intelligence—he read the Chinese edition titled The Future of Artificial Intelligence, while 孙剑 read the English edition; after recommending them to each other, they discovered it was the same book. The other was product experience: applications that genuinely worked at human level always used massive data. His scale yardstick was 1M for a demo, 10M for a good demo, 100M for a passing product and 1B for a good product.
  • The anchor was a dimensional analysis in the style of statistical physics: humans have roughly 1 major accident per 100M km, and a new technology must be 10x safer—1 per 1B km—to be accepted. Testing 1B km of safety is like testing a coin’s probability; 1 toss is useless, and statistical significance takes about 100 events. That means 100B km of data, impossible to collect with a proprietary fleet. The next question was where to get it: mass production. That is how “1 flywheel, 2 legs” emerged, and the idea has barely changed in 10 years.

31. The Era of “Scammers”: Build It First, Then Talk

  • The time lag in consensus was concrete. In 2016, data-driven perception was nowhere near common sense. In 2020, when Momenta used deep learning for planning, fundraising colleagues returned from external meetings with the feedback: “Any company saying it can use deep learning for planning and decision-making is definitely a scam.” Momenta’s response was quiet: “Our culture is that we only talk about something after we have built it.”
  • He is colder toward outside skepticism: explanations are useless, especially to external investors; only results earn belief. Internally, the company selected for people willing to believe. He also rejects the idea that this requires faith over a 3- or 5-year cycle: feedback loops are designed to be short—1 week to 1 month, or 3 months at the long end. If positive feedback takes 3 or 5 years, “your R&D system is designed incorrectly.”

32. The Technical Rhythm of Moore’s Law: 2023 on the Eve of End-to-End, 2026 World Model

  • The technical cadence is 1 step per year: deep-learning planning in 2023, end-to-end in 2024, reinforcement learning in 2025 and world model in 2026. The names change, but underneath is the same: the model absorbs more data and produces better results, with performance improvements that are predictable. When he coined assisted-driving Moore’s Law in 2024, it meant 10x over 2 years; it is now accelerating to 10x per year.
  • The full-driverless leg is validated by speed, not installed fleet. From outside, the fleet had only 30-plus Robo vehicles by year-end 2025. Internally, the goal was not this year’s scale but the pace of improvement: hundreds in 2026, tens of thousands in 2027 and more than 100,000 in 2028, driven by 10x annual technical gains.

33. 2 Troughs: Overpromising and the Bloody First Delivery

  • The first trough ran from late 2018 into 2019. The company looked more like a loose research institute and had badly overpromised investors on product and commercialization progress—so badly that outsiders could see it without internal access. The correction combined a new principle and personnel changes: make the product work first, make technical innovation serve product value, set benchmarks, decide who rises and falls and who stays and leaves. The company was turbulent for 6 months; today he thinks he could solve it in 1 month. The difficulty was that every issue looked like a thousand loose ends, with no way to separate priorities or distinguish problems of the work from problems of people. That judgment “emerged” in him much as it does in a large model once the data volume is sufficient.
  • The second trough ran from 2021 through early 2022: the pandemic hit just as the first mass-production project arrived. The car, actuators, sensors and chips were all new, and every new component had problems. The R&D and delivery system was primitive; locating a single issue could take weeks. “It was like wartime, with weapons not advanced enough: a suicide squad charging with broad swords. It was truly bloody.” Today, the same question can be diagnosed in 1 second or 1 minute and resolved automatically soon after.

34. The Guiding Principle: Too Many Chimneys Mean Being Torn Apart

  • The delivery system’s core doctrine, distilled from the postmortem, was “reuse the mainline, develop the mainline”: solve problems by reusing what already exists in the mainline; if that fails, innovate within the mainline architecture, rather than opening a side branch. A Huawei veteran’s warning stayed with him: “徐总, whatever you do, don’t build so many chimneys. If you have that many chimneys, you’ll be torn apart by 5 horses.”
  • The 3 accompanying rules are “mainline razor”—Occam’s razor, because smart people like to pile on assumptions and the answer should use the simplest method to solve the key problem; “mainline synergy”—a method that works alone but cannot fit the same architecture is a low-priority direction; and “mainline accumulation”—capabilities must accumulate over time and align organizationally.
  • The numbers show the change: the first customer required more than 400 people for more than 1 year. Momenta now has 100 mass-produced vehicle models, while adapting the full urban-NOA feature set to a new brand and model takes about 10 people and 3 months. Its VVP (Vehicle Verification Platform) evolved from rule-based validation to large-model validation; it launched a “GPT bounty” at the end of 2022, and today an agent can crawl the data and identify the root cause when something breaks.
  • Can competitors catch up? Feature and configuration coverage can be closed with time and greater volume—actually ship 100 vehicle models and you gain the corresponding accumulation. But fast, high-quality delivery requires a data-driven delivery system and end-to-end automation. The real barrier is whether a company truly believes in data-driven development: when it hits difficulty, it must change the architecture and system creatively to make the approach work, rather than retreat to another method.

35. 10 Years of Survival: 2 Cultural Principles and a Gestation Period for Decisions

  • 2 cultural practices explain how a relatively unglamorous team survived 10 years: customer value first, and low-cost, short-cycle validation. A directional thesis is not validated by making it sound self-consistent; you must dig down several layers and actively find data points that can test it. He learned the term zoomality from investors: the ability to zoom in and out quickly across layers. “4 layers up and 5 layers down—zoomality of 9—and that person is very capable.” It keeps the company from fatal mistakes and lets it detect and correct errors quickly.
  • Asked whether he had decisions he regretted, he said it was hard for that to happen: major decisions were never made on the spot. Robotics was discussed as early as 2020 and 2021; when the time came to act, the decision was no longer whether to do it but the entry point, pace and route. External claims that he is forceful in decisions belonged to the immature 2019 period. Now he talks through an idea every day; by execution, a short-cycle idea may have been 1 year old, a long one 3-5 years old. The cost is slowness, so “you need to think through major decisions earlier.”
  • His view of tools is consistent: he will talk through a critical issue with multiple large models, but their performance on key questions remains far below that of the company’s strongest executives. They are useful when he wants to think until 3 or 4 a.m. but feels awkward calling people. “The CEO’s most important job is not efficiency but cognition and judgment.” Bad judgment usually comes from being too far from the front line and lacking first-hand information, or having information that is incomplete, biased or wrong.

36. A 100-Year Mission: Build AI Without End

  • The origin of the 100-year vision, “Better AI Better Life,” was a calculation of time. He once estimated autonomous driving would take 10 years to mature technically and 20 years for a commercial ecosystem. By 2036 he would be around 50 and the team in their 40s, so “we need a 100-year mission, one that is enough for us to work on for 100 years—do to last, to build an enduring legacy.” Perceptual intelligence will not be the whole of intelligence; entering cognitive intelligence and then robotics, a more advanced and general form of cognitive intelligence, are the next stages on the same road.
  • His closing self-portrait is hedgehog-like. He decides whether to do something based on 2 questions: is the incremental value large enough, and does he genuinely like it? Then he asks whether there is an insurmountable fatal obstacle. “When I am not deliberately trying to do better than others, I may simply get better and better with time.” “The fox knows many things; the hedgehog knows 1 big thing”—he gradually found the lesson useful. His longtime Luffy avatar is another footnote: he likes “One Piece, One Team,” and the IPO project code name was OP.