Pioneers Insight Method Research Author
118. Li Xiang’s Second 3-Hour Interview: CEO Models, VLA, Wisdom
Back to Episodes

118. Li Xiang’s Second 3-Hour Interview: CEO Models, VLA, Wisdom

Summary

  • Li Xiang draws AI’s commercial dividing line at whether it can become a production tool: not giving advice, but replacing professional work and generating paid-for results. He divides products into information tools, assistive tools, and production tools; the first two do not change employees’ most important eight hours or KPIs, while the third must have action and “unite knowing with doing.” So far, only Cursor and OpenAI Deep Research appear to be crossing the line, because programmers, business analysts, and strategy staff are willing to pay out of pocket.

  • A general-purpose Agent that handles everything will not appear in the next five years; what is more likely is an Agent OS on which professional Agents can grow. Drivers, doctors, lawyers, and programmers need entirely different data, CoT, tools, and alignment responsibilities. A real product must outperform the corresponding profession and bear consequences involving income, property, and even personal safety. Li Xiang is therefore betting only on a “driver foundation model”; internally, customer service, sales, and R&D teams will train their own Agents on a unified Agent OS.

  • DeepSeek changed Li Auto’s capex and R&D path: open-source language capability moved its VLA plan forward by roughly nine months and probably saved several hundred million yuan. Li Auto did not stop building foundation models; it instead raised its training-card purchases to roughly 3x the original plan. Office work and Li Auto Classmate will use a roughly 300B model, autonomous driving’s VL backbone will use 32B, and Li Auto will train its own models for 3D vision, traffic data, and joint VL data. In return, Li Auto quickly decided to open-source its self-developed automotive operating system, which Li Xiang defined as “purely thanking DeepSeek,” not as company strategy.

  • Li Auto’s VLA roadmap is already specific on model size, the training loop, and the compute ceiling; the key proof point is an L3-capable product, not a concept launch. The 32B cloud VL backbone is distilled into a 3.2B, 8-expert MoE model; after action is added, it reaches roughly 4B. Post-training retains only 2–3 CoT steps and predicts the next 4–8 seconds, then uses human takeovers and world-model-generated data for reinforcement learning. Li Xiang believes current compute can support roughly L3, with the capability appearing as early as Q3 of the interview year and no later than Q4, although regulation remains a variable. L4 may require a 32B model running directly on the vehicle.

  • World models cut the economics of autonomous-driving validation from RMB170K–180K per 10,000 km to roughly RMB4,000, and will become the training ground, examination system, and operating system for L4. They can precisely recreate corner cases involving the ego vehicle, traffic participants, and unusual roads, rather than relying on real vehicles to reproduce them by chance. Reinforcement targets can also be quantified through G-force comfort, traffic compliance, collision avoidance, and reaching point B from point A. “Precisely because the industry has run into problems,” Li Xiang calls the current stage “the darkness before dawn.”

  • Li Auto is redefining itself as a “globally leading artificial-intelligence terminal company,” but says its valuation must be earned through L4 and organizational efficiency. Li Xiang sets four conditions for an AGI terminal: 360-degree perception, cognitive decision-making, action, and reflection and feedback. The company may expand from SUVs into family sedans, MMPVs, and other terminals, but expansion remains constrained by scale. The test is not the narrative; it is whether Li Auto can be first to build L4 and generate $100B in revenue with fewer than 100,000 employees, perhaps even 50,000. If it cannot, “we should be valued like an automaker.”

  • Behind the technology roadmap is an anti-human-nature operating method: capability must pass through research, development, expression, and business, while business must pass through analysis, goals, strategy, and review. Li Xiang sees DeepSeek as having “used humanity’s best practices in the simplest possible way,” and admits that organizations most often skip research and jump straight to changing strategy. He is designing teams of 3–7 people as MoE-like “more powerful brains and more powerful hearts.” His longer-term view is that AI will surpass human capability, but intelligence is not wisdom—“wisdom is our relationship with all things.” If humans are to keep leading civilization, AI must free up time and energy, not make the workday longer.

Deep dive

1. Humans Are Small Models; AI Should Carry Information Complexity

  • Li Xiang rejects the idea of simply converting the human context window into tokens: humans are physiologically bad at processing complex information, and the brain consumes energy, so the essence of methodology and tools is that “humans should reduce entropy, not increase it.”

  • His analogy is that reading and learning knowledge resemble pretraining, converting knowledge into business resembles inference, and inference then calls tools; but humans are more like professionally trained small models than giant foundation models covering all knowledge.

  • AI can process data that grows from 15T to roughly 30T, with the latest Llama 4 cited as an example. But entering a professional domain still requires higher-quality data and more specialized CoT. The most rational relationship is not a replacement narrative, but for “each side to do what it does best.”

2. Chinese Models Closed the Gap with the US in 130 Days

  • Looking back over the 130 days between the two interviews, Li Xiang says the biggest progress was not in himself but in China: DeepSeek and Qwen brought foundation models, reasoning, and multimodal capability “basically close to the US, or basically onto the same level.”

  • He particularly emphasizes the training and inference efficiency of Chinese teams and the confidence created by deep engineering improvements. Breakthroughs by Manus and Genspark in Agents also convinced him that China can continue making better AI products.

  • But language models are only one part of the world. For a car or robot to operate across the digital and physical worlds, it also needs vision, action, an operating system, tools, and alignment. That is why Li Auto continues to build VLA.

  • Zhang Xiaojun asked whether Li Xiang himself had become smarter. His answer was “not by that much”: AI had advanced rapidly, but his own working hours and those of people around him had not fallen, nor had their output materially improved. That paradox became the starting point for the interview.

3. Chatbots Are Creating a New Kind of Entropy

  • Li Xiang’s criticism of current Chatbots is that they must output the next token and a determinate result, while users usually begin with online search. The sources retrieved by RAG may already be distorted, so the model can “reason very seriously” while both the process and conclusion are fundamentally wrong.

  • When DeepSeek first launched, people turned its reasoning processes into images and posted them on Xiaohongshu. Several months later, nobody did so anymore because the content had become “all the same.” In Li Xiang’s view, that shows the reasoning process itself does not automatically create lasting value.

  • The more practical constraint comes from organizations: even if AI gives advice on stocks, careers, or work, employees still prioritize the company’s KPIs and OKRs. The advice spins in their heads without entering the workflow or changing the output of their eight-hour workday.

4. Willingness to Pay Defines the Three AI Tool Categories

  • Li Xiang divides AI products into three tiers: information tools provide reference, assistive tools improve an existing experience, and production tools genuinely replace professional labor. In-car voice, navigation, music, and assisted driving are all tier two because the human still cannot leave.

  • His commercialization test is direct: “You often aren’t willing to pay for an information tool, while you see an assistive tool as something a product should come with.” Only a production tool makes users pay separately because it changes the work result.

  • The two early production tools validated by Li Auto employees paying out of pocket are Cursor and OpenAI Deep Research. The former entered programmers’ coding workflows; the latter entered the research workflows of business-analysis and strategy teams. “The moment artificial intelligence becomes a production tool is the true moment AI takes off.”

5. Action and Tool Calls Matter More Than Longer Answers

  • A production tool must “have action; it cannot only know, it must act.” No matter how intelligent O2, O3, or DeepSeek R1 is, strategy simulation alone does not complete the job. Controlling a computer, software, vehicle, or robot is what creates real value.

  • Li Xiang explains why models and tools are complementary through digging a hole: “You’re ten times smarter than me. You use a spoon to dig one hole, and I use a shovel to dig one hole. No matter how dumb I am, I’m still more efficient.” Tools bring greater certainty with less energy and fewer tokens; a stronger brain can design stronger tools.

  • He credits Manus with taking a meaningful product step toward becoming a production tool. Using a virtual machine and tools, it analyzed Tesla by going directly into SEC filings, investor-relations sites, and leading analyst reports rather than merely indexing sources through RAG. The challenge is that its general-purpose coverage is too broad, which also explains why startups with less specialized division of labor sometimes assemble products first.

6. Three Variables Should Change Only Around Scale

  • For a company of more than 30,000 people, Li Xiang would not immediately change the organization because of a single technology update. He puts scale at the center, with user demand, technology and products, and organizational capability as three dynamic variables around it. The organization must follow only when user and technology changes genuinely connect.

  • This diagnosis also suppresses hot-topic reorganizations: “Adjust things today because this came out, adjust again tomorrow because something else came out” only creates confusion. Over the past 130 days, the technology variable was driven mainly by Chinese AI, and the international environment also changed, but Li Auto’s organization did not instantly become more intelligent.

7. DeepSeek V3 Shows a Four-Step Capability-Building Method

  • Li Xiang sees DeepSeek V3 as a minimalist sample of humanity’s best practices. The 671B MoE model combines multiple experts, but the first step in building expert capability is not development—it is research.

  • The complete path is “research—development—capability expression—business value.” Research resolves understanding and direction; development turns that understanding into usable capability; the product then makes the capability visible; finally, it must enter real business combat.

  • Li Auto follows the same path in end-to-end driving, VLM, VLA, chips, and operating systems. He cites papers being referenced by researchers including Fei-Fei Li as evidence that research is not decoration but a prerequisite for later development efficiency.

  • Capability expression is not marketing packaging. Li Auto shows inside the car how end-to-end driving selects trajectories, what attention focuses on, and how VLM understands ETC gates, traffic lights, and complex roads so users and the organization can see how the model works.

8. DeepSeek R1 Compresses the Business Loop into a Reasoning Chain

  • Li Xiang reads another best practice in R1’s chain of thought: first analyze the indexed information, then determine the objective from the direction supplied by the user, output a strategy and simulate its execution, and finally reflect on the gap between the result and the objective.

  • Mapped onto business, that means analyzing users and the market, setting goals, forming and executing a strategy, and reviewing the outcome. If the gap comes from strategy, continue analyzing and adjusting; if it comes from capability, return to research and development rather than forcing the business forward.

  • He retains one key limitation: R1 provides a workable strategy but does not control a machine or computer to execute the action. It is therefore not complete unity of knowing and doing, but an excellent “brain system.”

9. Organizations Love Skipping Research and Changing Only Strategy

  • In capability building, people most easily jump straight into development without research, capability expression, or real-market combat. When business hits a problem, they most easily change strategy directly, without reviewing, analyzing users, or jointly resetting the objective.

  • Li Xiang admits that strictly following best practices “runs against human nature”; doing whatever one feels like doing satisfies human nature. Exceptional individuals and organizations must therefore continually fight laziness, shortcuts, and personal arbitrariness rather than assuming that writing down a methodology means they possess the capability.

10. Liang Wenfeng’s Edge: Discipline and High-Probability Research

  • Li Xiang spoke with Liang Wenfeng only once, in September 2024, several days before OpenAI released o1. His first impression was extreme self-discipline: Liang could hold to principles he believed in and fight humanity’s tendency toward laziness and shortcuts.

  • His second judgment was that Liang studies best practices and methodologies globally, then systematizes his own high-probability path to success: “Research first, analyze first, then act, and the probability of success is very high.”

  • On young people doing research, Li Xiang says mature researchers often already have their own frameworks, which can become obstacles. He also relayed Liang’s point that Chinese tutoring materials contain complete solution processes and feedback, making them an excellent reinforcement-training system. Li Auto later mapped the idea onto traffic rules, comfort, and driver takeovers.

  • Li Xiang says Liang estimated at the time that DeepSeek was roughly one year behind OpenAI, but only three or four months passed from the eve of o1 to R1’s release. He still acknowledges that OpenAI is strong across research, development, products, and distribution. Its Ghibli-style image generation and “more than 400M visits in one week” illustrate that combined capability; how long it will remain ahead is “very hard to say.”

11. DeepSeek Chose to Train Capability Rather Than Absorb All Traffic

  • Faced with the surge in traffic during the Spring Festival, Li Xiang speculates that DeepSeek did not fully absorb DAU because it had limited cards: using all of them for inference would leave no capacity to continue training and improving capability. Some user queries are valuable, but excessive traffic is not necessarily equally useful.

  • He does not believe he should have built DeepSeek himself: “I can only be the best version of myself.” Liang extended from AI at Zhejiang University through quantitative models and engineering into foundation models; Li Auto’s natural extension is products, cars, and terminals in the physical world.

  • DeepSeek’s direct benefit to Li Auto from open-sourcing was to bring forward language capability originally scheduled for September of that year by roughly nine months and probably save several hundred million yuan. Li Xiang calls it “standing on the shoulders of giants,” not abandoning Li Auto’s own model capabilities.

12. Li Auto Accelerated Its Own VLA with an Open-Source Language Layer

  • During the Spring Festival, Li Xiang first discussed the issue with Xie Yan: the language model Li Auto developed by September might not be stronger than DeepSeek V3 plus R1. Since DeepSeek’s openness was “basically closer to Linux than Android—far more open,” Li Auto should use it directly as the language component of VLA.

  • He initially worried that model lead Chen Wei would resist. Instead, Chen was more determined than management, arguing that DeepSeek should be the foundation for accelerating VLA and end-to-end multimodality. Research, compiler, and chip teams simultaneously optimized training and inference efficiency.

  • The organization then formed three different requirements around three businesses: internal office work needs language plus Agent OS; Li Auto Classmate in the car and on the phone needs real-time voice and vision in end-to-end VL; autonomous driving and factory robots need VLA. Different teams therefore need to adapt models for the three businesses.

  • Li Auto still has to train its own models because automotive applications require 3D vision, high-definition 2D vision, traffic and household data, and joint VL data. OpenAI and DeepSeek do not have Li Auto’s vehicle data, scenarios, or business objectives. Open-source language capability is one part of VLA, not the complete answer.

13. Foundation-Model Investment and Product Promotion Moved in Opposite Directions

  • Li Auto raised its training-card purchases for that year to roughly 3x the original plan. The model used by Li Auto Classmate is expected to be around 300B, or 300 billion parameters; autonomous driving’s VL component will use a 32B version, while 300B models will also be used in real work. The exact size will vary by demand.

  • For the Li Auto Classmate app, Li Xiang instead required the company to postpone paid user acquisition and marketing, focusing first on capability and understanding demand. The in-car version combines information and assistive-tool attributes, while the phone version was still mainly an information tool at the time. Computers and the web are future extensions. Li Auto would first finish these versions for the more than 1.2M users who need connectivity.

14. VLA Is a Continuous Evolution from Insects and Mammals to Human Drivers

  • Li Xiang calls the current stage of intelligent driving “the darkness before dawn”: the concentration of industry problems is precisely what signals the arrival of the next generation of capability. Li Auto’s work on range extenders, 5C charging, and operating systems also began with problems the industry could not solve.

  • Rule-based algorithms plus machine-learning perception resemble insect intelligence: the model has only several million parameters, depends on high-definition maps and extensive restrictions, and ultimately approaches “rail transport,” completing tasks only within predetermined rules.

  • End-to-end driving resembles mammalian action learning. It imitates human driving, sees 3D imagery, senses its own speed, and outputs a trajectory. Generalization improves substantially, but it does not truly understand the physical world; in unfamiliar complex scenarios, it often can only stop.

  • VLA comes closer to a human driver: it reads the 3D world, high-definition 2D detail, and navigation software at the same time, uses language and CoT to understand semantics, then executes and communicates through action. Li Auto therefore does not call it an abstract foundation model, but a “driver foundation model.”

15. VLA Pretraining Starts with a 32B VL Backbone

  • The first step is to train a 32B VL backbone in the cloud. Visual data is split between the 3D physical world and high-definition 2D imagery. Compared with common open-source VLMs, the recognition distance for the latter improves by roughly 3–5x, solving the problem of insufficient long-range detail.

  • The language component is not a simple extension of general web data, but traffic- and driving-related data. More difficult is joint VL data—for example, combining navigation maps, the driver’s judgment of the map, and the vehicle’s recorded actions into training data.

  • The 32B backbone is then distilled into a 3.2B edge model with 8 experts in an MoE. Running the full 3.2B model directly would not deliver the token output rate and frame rate required for real-time driving on dual Orin X or Thor-U.

16. Post-Training Turns a Knowledge Model into a Roughly 4B Driving Model

  • Li Xiang compares pretraining with learning about the world and traffic, and post-training with going to driving school. Action is added at this stage, using imitation learning to connect vision, understanding, and action end to end; the model expands from 3.2B to nearly 4B.

  • Safety latency means its CoT can retain only 2–3 steps; it cannot copy long-chain reasoning. Traffic and robotics scenarios cannot wait for the model to conduct lengthy internal debates.

  • After outputting action, the model also uses diffusion to predict the environment and trajectory over the next 4–8 seconds. Li Auto treats this prediction as part of the driver’s own capability, rather than creating a separate “world model.”

17. Reinforcement Learning First Aligns with Society, Then Tries to Surpass Humans

  • The first reinforcement layer is RLHF: driver takeovers, non-takeovers, habits, and safety judgments teach the model to follow traffic rules while incorporating normal driving behavior on Chinese roads, avoiding the behavior of a novice who obstructs other vehicles.

  • The second layer is pure RL based on world-model-generated data. It no longer receives step-by-step human feedback, but only needs to get from point A to point B, constrained by three outcomes: G-force for comfort, traffic-rule violations, and collisions.

  • Li Xiang’s progression is clear: pretraining is learning knowledge; post-training is driving school; RLHF is entering real society and aligning with humans; pure RL uses world-model-generated data to make the vehicle drive better than humans.

18. The Driver Agent Turns VLA into a Product Users Actually Hire

  • Humans do not need to learn machine instructions: “Speak to it the way you would speak to a normal driver.” Short instructions are handled directly by the vehicle-side VLA; complex long instructions first go to the cloud-based 32B VL model for interpretation, then are sent to the edge model for execution.

  • The north-star metric is not a single takeover rate, but three judgments about the human driver: Is professional capability strong enough? Does professionalism guarantee safety and comfort? Do memory and communication build trust? The end state is “I say the first half, and it knows the second half.”

  • Li Xiang compares super-alignment with professionalism. A racing driver may be highly capable but unsuitable for carrying passengers every day. A professional driver must drive well, follow rules, care for passengers’ experience, and understand the employer through long-term memory.

  • Pricing should reference the labor being replaced. If a human driver costs RMB10,000 per month, users might pay RMB2,000–3,000 to use an AI driver. Insurance, automatic charging, or a certain amount of electricity could also be bundled in the future, provided total cost is genuinely lower.

19. Traffic Is the First Testing Ground for Professional VLA

  • The traffic world is “complex but determinate”: cars can only travel along roads, rules and objectives are clear, and control mainly involves two or three degrees of freedom—forward/backward, left/right, and slight rotation. A humanoid robot starts with more than 40 degrees of freedom, making the problem entirely different.

  • Cars are also naturally suited to imitation and reinforcement learning. Sensors record the world a person sees, the vehicle records the person’s actions, and takeovers are misalignment feedback. Comfort, compliance, collisions, and reaching B from A can all be quantified.

  • Zhang Xiaojun summarized this as “building a driver.” Li Xiang agreed, emphasizing that L2 and L2+ remain assistive tools. Only taking on the driver’s job makes a product a production tool. Whether it is better than a professional driver is a far stricter standard than whether it can demonstrate a few features.

20. The General Agent of the Next Five Years Is More Precisely an Agent OS

  • Li Xiang gives a clear judgment: “There will be no general Agent within five years; there will be an Agent OS.” Drivers, doctors, lawyers, and programmers depend on entirely different vision, language, action, professional data, CoT, and liability systems. One generalized model cannot handle everything directly.

  • Agent OS provides tools, virtual machines, and the foundations for developing professional Agents. Professionals then add their own data, corpus, workflows, and chains of thought to create Agents that can take over high-frequency work.

  • The professional Agent Li Auto is best positioned to build externally is the driver Agent; education should be handled by specialist teams such as Yuanfudao and TAL Education. Internally, Li Auto needs Agents for customer service, programming, sales calls, and validation experiments, and should allow one professional to manage multiple Agents.

  • The intelligent-business, foundation-model, and operating-system teams should jointly build the internal Agent OS, but the customer-service Agent must be developed by customer service, and the sales Agent by sales. “There will not be one team that builds everything for everyone.” The platform can only make it easier for professional teams to do their own work.

21. End-to-End Was Not Abandoned; It Became VLA’s Action Layer

  • Zhang Xiaojun asked whether end-to-end had already been replaced after one year in production. Li Xiang’s answer was continuous evolution: the execution capability trained by end-to-end is precisely VLA’s A layer; Li Auto is simply adding stronger 3D and high-definition 2D vision plus language.

  • Jumping directly to VLA is “not possible,” because “you can’t take the tenth dumpling without taking the first nine.” If rule-based algorithms have not been done well, there is no understanding of how to do end-to-end; without taking end-to-end to its limit, there is no understanding of how to train VLA.

  • He calls the belief in one-step transformation “practicing the Sunflower Manual.” DeepSeek’s efficiency likewise came from long-term optimization of clusters, pipelines, and infrastructure, not from suddenly discovering a shortcut. Li Auto began self-developed intelligent driving in 2021 and researching end-to-end in 2023; the accumulated foundation cannot be deleted.

22. Compilers, Chips, and Operating Systems Determine Whether Models Can Run in Cars

  • Orin did not originally support the language models Li Auto needed, so Li Auto’s compiler team rewrote the lower layers to run VLM at INT4. Li Xiang compares this with DeepSeek’s use of FP8 to optimize training: both use engineering capability to break through hardware constraints.

  • Dual Orin X and Thor-U can run VLA at the same model scale as other platforms, but this still depends on boards, chips, operating systems, and real-time computing capability. Li Xiang’s conclusion is that small scale allows companies to bypass fundamentals, but at large scale “fundamentals and capability can never be bypassed.”

23. VLA Upgrades “Stop” into Understanding, Communication, and Rerouting

  • Complex road construction is the classic corner case. A rule-based algorithm might crash; end-to-end might stop and hesitate; VLA can understand the complex scene and handle a feasible path. Even if it cannot do so initially, Li Xiang says generated data can train it to handle the scenario within 3 days.

  • A bus lane that remains ambiguous over a long stretch illustrates the communication gap. After being temporarily pushed back into the middle lane, end-to-end may enter the bus lane again; a driver Agent can understand the persistent constraint: “Stay in the middle lane from here on until the next navigation node.”

  • When it misses a turn, end-to-end can easily lose the objective. VLA can first roam through a residential compound or open space, then reconnect with a replanned route. Li Xiang has no so-called aha moment: animals suddenly learning an action is surprising, while “one person making something good” seems normal.

24. The World Model Serves as Examination Hall, Data Factory, and Operating System

  • Li Auto defines a traffic world model as a complete physical world constructed through reconstruction and generation, containing roads, fixed objects, and all traffic participants. VLA is like a driver: it can drive in the real world and enter the simulated world for an examination.

  • It has three stages: first an examination system for the model, then a generator of training data, and finally the operating system for L4 autonomous vehicles. Traditional IT software cannot manage large numbers of vehicles without drivers; the world model must handle real-time operations.

  • This differs from some robotics papers that call predicting the next few seconds a world model. Li Auto puts the 4–8-second diffusion prediction inside the driver’s capability and calls the runnable, interactive traffic environment itself the world, deliberately aligning AI product concepts with human understanding.

25. Simulation Cuts Validation Cost per 10,000 km to Roughly RMB4,000

  • In the past, having human drivers cover 10,000 km, including the vehicle and various costs, ran to roughly RMB170K–180K. After the world model, the cost falls to more than RMB4,000, with compute as the main remaining expense, while the efficiency of solving problems is actually higher.

  • Real vehicles struggle to simultaneously recreate the same ego vehicle, multiple participants, speeds, positions, and unusual roads, so validation is often only approximate. A world model can “reproduce the exact same real scene 100%,” allowing precise testing of the repaired model.

  • A black box therefore does not mean an untestable box. The model can repeatedly take exams in a simulated city—going from A to B, maintaining comfort, obeying rules, and avoiding collisions. Failed scenarios can then enter the generated-data and training loop.

26. Super-Alignment Constrains Strong Capability into Professional Driving

  • After reaching 10M km, Li Auto began building a super-alignment team of more than 100 people. The typical problem is not inability to drive, but behavior such as repeatedly cutting into traffic during congestion—efficient, yet unsettling for passengers—because the model learned the wrong driving behavior.

  • Li Xiang separates capability from values. Whether a collision occurs is mainly a capability issue; whether the vehicle cuts in, follows human customs, and respects the passenger experience is alignment. “The greater the capability, the greater the responsibility,” just as highly capable people need stronger professional constraints.

  • The three dimensions used to recruit exceptional employees are also transferred to Agents: professional capability, overall professionalism, and the ability to understand others and build trust. The model handles the first, super-alignment handles the second, and the Agent’s memory and interaction handle the third.

27. Compute Determines the Capability Ceiling for L3 and L4

  • Li Xiang believes VLA can solve fully autonomous driving, but it may not be the most efficient ultimate architecture. It is still based on the Transformer and currently has the strongest capability, but requires substantial compute. A more efficient architecture will “most likely” appear in the future.

  • Dual Orin X or a single Thor-U has roughly 64GB of memory. It might theoretically fit a roughly 30B model, but cannot achieve the frame rate required for driving. The edge device can therefore currently run only the 3.2B version, or roughly 4B after action is added, putting a ceiling on capability.

  • His projection is that if a 32B model can eventually run directly on the edge, L4 may become possible. Cloud VLA could then also expand to 320B. The current generation of compute is “basically at the level of L3.”

  • His timing judgment retains conditions: a product with genuine L3 capability could appear as early as Q3 of the interview year and no later than Q4, but regulation and subsequent problems could still affect commercial delivery.

28. Rule-Based Algorithms Remain Efficient Tools in the Model Era

  • A VLM can handle two or three ETC entrances, but becomes confused by more than ten because its positional judgment is weak. After the team added data continuously for 3–4 months without solving the issue, Li Xiang ultimately required a rule-based algorithm for up to roughly 15 entrances.

  • The program was completed in less than one week, after which ETC performance stabilized. Li Xiang uses the example to push back against the idea that models solve everything: deterministic rules should be retained when they complete a task with less compute, fewer tokens, and higher accuracy.

  • His analogy is that humans memorize multiplication tables. They are essentially a rule-based algorithm, yet save significant brainpower. Brains and tools are never contradictory; in real deployment, capability, determinism, and energy consumption must all be optimized together.

29. Strategic Reasoning Works Backward from Scale to Users, Products, and Organization

  • Li Xiang uses the scale target for roughly three years around 2027 as the starting point for the CEO foundation model’s long-chain reasoning. Scale implies time; from there he works through user demand, technology, product form, and organizational capability, and everything must eventually become action. “Having an idea and being able to execute it remains the biggest gap.”

  • Li Auto generated roughly RMB145B in revenue last year and was expected to generate more than RMB100B in the interview year. If it continues toward RMB300B, RMB500B, and beyond, relying only on family SUVs could hit the category ceiling.

  • Annual sales moving from roughly 500,000 toward more than 1M would bring the user base close to that of BMW, Mercedes-Benz, and Audi. Li Auto would need to reach people it had not reached before and change how it communicates. Globally, the family positioning could remain, but it could not mean only SUVs.

  • The product conclusion is to enter family sedans and a richer range of MMPVs, not sports sedans, while continuing to control SKUs and avoiding a disorderly exchange of model count for scale. This is expansion derived from demand and scale, not from seeing a new technology.

30. Li Auto Chooses to Become a Terminal Company in the AGI Era

  • Asked what Li Auto’s identity would be in 2030, Li Xiang answered: “a globally leading artificial-intelligence terminal company.” Compared with the vague label of an AI company, adding “terminal” means choosing an integration of software, hardware, and services rather than only a model platform.

  • An AGI terminal must meet four conditions: perceive the physical world in 360 degrees, make cognitive decisions, execute action, and reflect and provide feedback on results. Cars are evolving from smart terminals into AI terminals and could become the highest-revenue terminals of the AI era, reaching $100B in revenue.

  • Li Auto may not only make cars in the future, but it will not make ordinary terminals from the previous era. Any new product must satisfy the four conditions and cover the main scenarios in users’ lives or work.

31. Organizational Capability Learns in Stages at RMB10B, RMB100B, and Larger Scales

  • During the Li Auto ONE phase, the company learned Toyota’s working methods, GM’s R&D processes, and Google’s OKRs. It completed the development and delivery of its first vehicle, sold more than 200,000 units cumulatively, and generated more than $10B in revenue.

  • Moving from RMB10B to RMB100B, the company focused on Huawei’s IPD, finance, processes, and three-pillar human-resources system, alongside L-series platform development and a larger sales organization, rapidly reaching more than RMB100B in revenue.

  • Once moving beyond RMB300B, Apple became understandable again. Li Xiang admits that when the company was smaller he could not understand Apple; only after reaching RMB100B in scale did he gradually see how Apple expanded from the Mac to the iPod, iPhone, and services ecosystem, and why large-scale organizational capability requires long-term construction.

32. Software, Hardware, and Services for AGI Terminals All Need to Be Rebuilt

  • The software side needs three layers: a model that understands the physical and digital worlds, a high-performance real-time operating system, and deterministic tools called by Agents. Traditional Android-style background queuing and AUTOSAR-style chained structures may not suit real-time AI terminals with NPU and MCU parallel execution.

  • On the hardware side, the first layer is the body and drive-by-wire system, followed by a combination of centralized and distributed computing rather than faith in a single central brain. This can reduce wiring and improve transmission efficiency.

  • Li Xiang compares the NPU with the heart. At the same 10Hz, if one terminal can run only 3B while a competitor can run 30B, that is a 10x performance gap. The cloud can stack cards, but the edge is constrained by power, space, and real-time performance, making independent chip and compiler capabilities more important.

  • On the services side, two problems must be solved: how to operate large numbers of Agents running in the physical world, and how one person can manage multiple Agents at once. The world model, Agent OS, and connectivity mechanisms are not auxiliary software; they are the AGI service itself.

33. Factories Should Become Robots Rather Than Simply Add Humanoid Robots

  • Li Xiang’s vision for manufacturing AI is to “use AGI to produce the AGI terminals”: the factory as a whole should become a kind of “Cosmic Emperor” robot, simplifying the entire process through perception, decision-making, and automatic feedback.

  • He opposes simply placing humanoid robots into existing human workstations. Labor costs are not as large a share of factory costs as outsiders imagine, so simple replacement may not be economical; employment will still be necessary over the long term. The larger opportunity is to improve the efficiency of the entire production system.

  • Wearable terminals are not mature. Glasses may offer 360-degree perception, but waveguides or centralized displays, batteries, independent computing, and communications have not reached the standard for long-term use. Research can move ahead, but product timing cannot be determined by imagination.

  • Household robots have also not converged on a form. They might be humanoids using human spatulas, or a unified perception-and-brain system combined with redesigned smart cookware. Li Auto will not announce a path in advance; it will wait for research and industry capability to determine which route works.

34. Whether Expansion Is Justified Depends Only on Scale and Capability Returns

  • Li Xiang believes small companies should converge as much as possible, while large companies must expand. If Google had only search, Microsoft had no Office or cloud, or Apple made only the Mac, each could have become a different company.

  • Li Auto already has more than RMB100B in revenue and is moving toward RMB200B. In Li Xiang’s view, this is the right point to build operating systems, chips, and models. When Apple launched the iPod in 2001, it had only several billion dollars in revenue but already owned a computer, operating system, and software ecosystem.

  • Developing its own operating system has cost roughly RMB1B over 4 years. Li Xiang estimates cumulative savings of RMB5B–6B, while also improving support for domestic chips and development efficiency. “It can strengthen capability and lower costs,” so this is not an overextended empire but an economy-of-scale investment.

35. AI Valuation Must Be Earned Through L4 and Productivity

  • Li Xiang does not want the market to award Li Auto an AI premium before its capability is visible; “it actually makes me nervous.” The first test is whether it can be first to achieve L4, allowing users to commute without holding the wheel while sitting at a desk working or eating.

  • The second test is organizational productivity. An organization approaching 1M people currently generates more than RMB700B in revenue; if Li Auto can generate $100B in revenue with fewer than 100,000 people, perhaps even 50,000, that would prove its AI strategy is creating real value.

  • If it fails both tests, “you should value us like an automaker.” This is the interview’s clearest valuation discipline: technology, terminal identity, and grand narratives must ultimately land in sales, revenue, labor efficiency, and chargeable services.

36. Li Auto Will Disappear if Demand, Technology, or Organization Breaks

  • Li Xiang gives three possible causes of death: failing to grasp user demand, failing to master the best product technology, or suffering a major organizational-capability failure. The three must be diagnosed together; examining any one alone gives a distorted picture.

  • He admits the company thinks far ahead but says it remains “grounded” in the present. In the programming era, companies competed on features; in the AI era, they compete on capability and on how capability becomes business. Skipping capability to seize the result is still “practicing the Sunflower Manual.”

  • Strategy will always face opposing views in real work; some people even continue validating projects privately after they are halted. If the result proves them right, Li Xiang will agree with them again. He does not equate alignment with an absence of debate.

37. Debate with Energy Creates Organizational Intelligence

  • Li Xiang distinguishes debate from internal friction not by volume, but by whether the connection still has energy. When energy is present, arguments, discussions, and quarrels form “a more complete brain”; once the connection disappears, the same actions become internal friction.

  • When Autohome was at its most difficult, he, Qin Zhi, and Fan Zheng formed a three-person support structure: unified externally, they completed one another’s judgments through internal debate, then became a powerful heart of shared resources and will once a decision was made.

  • Li Auto uses the same mechanism. The group expanded from Li Xiang, Shen Yanan, Ma Donghui, and Li Tie to include people such as Xie Yan and Zou Liangjun. The point is not for the CEO to act on personal whim, but for multiple judgments to be stronger than one brain and for nobody to collapse alone during execution.

  • In his experience, 3–7 people is the most stable structure. Two are too few; more than seven makes close connection difficult. This is not a quantitative study but a long-held intuition, and it resembles MoE: multiple small experts form the support of brainpower and heart, then connect to a larger organization.

38. Caring About Users and Those Around You Enables Energetic Connection

  • Li Xiang reduces connection to “two forms of caring”: caring together about customers and users, which creates shared values, and caring about the people beside you who carry responsibility together—“put people first, then do the work,” rather than speaking abstractly about being objective about the task and not the person.

  • He actively brings core leaders together, then trains collaboration through repeated real events. Only by jointly experiencing scarce resources, judgment disputes, and responsibility for results can a 3–7-person structure become genuine support rather than an organization chart.

  • This is also his biggest recent personal growth: extracting the partnership model from which he benefited and helping more teams form “a more powerful brain, a more powerful heart, and stronger energy.”

39. AI Organizations Can Coexist with Manufacturing Rather Than Eliminate Manufacturing Management

  • Li Xiang believes high-dimensional management can accommodate low-dimensional processes. Digitalization can contain factories, sales, and IPD; an AI organization can likewise let different businesses use different methods. Factories, operating systems, complete vehicles, and intelligent driving do not need to share one organizational template.

  • Li Auto’s true end-to-end team has roughly 200 people, and its model team more than 100. Competitors’ rule-based algorithm teams may have 2,000–6,000 people. Additional people support the model, research, and data periphery, but the small core teams show the productivity difference of the new architecture.

  • Roughly 60%–70% of the intelligent-driving and model teams come from campus recruitment, with little reliance on so-called industry stars. Li Xiang connects this practice to Liang Wenfeng’s influence: research needs young people to break through existing frameworks, while scenarios, data, sustained funding, and whether the company truly believes in AI matter more than titles in attracting talent.

  • He clarifies that he does not spend 80%–90% of his time looking only at AI. Roughly 60% of his working day goes to organization and people, including interviews and training; the rest is split approximately evenly between the automotive business and AI. As CEO, he believes “AI is not science; AI is engineering,” and helps teams clarify architecture through structural questions.

40. Ten Years of Memory Rewritten as Capability Growth Rather Than a History of Suffering

  • Li Xiang’s clearest image of happiness is the first launch of Li Auto ONE in 2018, its formal display and price announcement at the 2019 Shanghai Auto Show, when it drew the largest crowds, and the L9 in 2022, a product later imitated by at least 5 companies.

  • The peak of L9 was followed by a wave of online smear campaigns, rumors that the company would fail, and a quarterly loss of nearly RMB2B. But he chose to remember that the crisis exposed capability gaps and drove revenue to nearly 3x growth in 2023, reaching roughly RMB120B.

  • His memory strategy is to “retain as much as possible only the valuable, beautiful fragments,” and to forget bad things quickly once they have been resolved. By the clock, entrepreneurship certainly contains more hardship, but “there is no need to be miserable about it.” Hardship and sweetness are two sides of the same coin.

41. Growth Means Adding Capability, Not Changing Yourself

  • Li Xiang says roughly 90% of who he is today resembles his high-school self: when a problem appears, he solves it, especially user problems others do not want to solve, and he finds partners to compensate for his weaknesses. What changed are the problems, users, and organizational scale.

  • He sees nearly being forced out of a company in 2008 as the starting point of wisdom because it forced him to seriously address his relationship with himself. Autohome’s setback from underestimating capital taught him to add FA, law firms, equity, voting rights, and cash governance when building Li Auto. That was not a change in product instinct, but the addition of capital capability.

  • For himself, he advocates accepting all of his strengths and weaknesses, because a weakness is often the other side of a strength. Someone skilled at high-dimensional decisions may be poor at detailed operations; a lazy person may become an excellent product manager, while a strong operator may not excel at cross-dimensional judgment.

  • “Replacing change with growth” prevents people from becoming worse versions of others or worse versions of themselves. Growth itself creates energy and allows a person to gain missing capabilities through partnership structures.

42. People Are for Being Used, Not Changed

  • After becoming a father, Li Xiang first confirmed that every child is entirely different: “People are for being used; people are not for being changed.” The role of an organization is to combine different strengths, personalities, and weaknesses into a complete brain rather than force everyone into one template.

  • His second realization was that “I need them” matters more than “they need me.” His children, partner, and executives who share responsibility make him better. Expressing that need boldly makes others feel needed and turns relationships from control into mutual support.

  • The third role of children is helping parents see themselves. Children do not conceal themselves and cannot be chosen, forcing parents to understand different kinds of people and learn how to handle unavoidable relationships.

  • This understanding leads to more proactive behavior: focus on what children need and what they are good at, and communicate before problems worsen. Others can give you energy, but “the initiative over your energy” remains in your own hands.

43. A Three-Person Family Support Structure Creates New Energy in Intimate Relationships

  • Li Xiang and his wife previously had limited mutual support. Around the time of the interview, his 14-year-old elder daughter became a third pillar: she had developed a relatively complete worldview and could discuss life plans, preferences, family problems, and travel, creating a three-person structure of mutual support.

  • He believes intimacy begins with need rather than independence. If you have no need for the people around you, it becomes easy to ignore, evade, compete with, or fight them internally. Families and companies alike should allow members to provide one another with brainpower, heart, and energy.

  • Intimacy should not be generalized. Immediate family, a small number of longtime friends, and people at work who share responsibility belong to the core circle. “The only people who can hurt us are those in intimate relationships,” because only these relationships carry both real value and real energy.

44. Work Creates Social Value; Family Creates the Value of Affection

  • Li Xiang also sees long-term colleagues as part of intimate relationships because work carries personal social value: better products and services benefit more users, allowing work and companies to create social value.

  • Family corresponds to the value of affection, whose endpoint is happiness. Happiness is defined as “high-quality companionship.” Work and family are not mutually sacrificing: better work supports life, better life creates energy, and that energy improves work results.

  • His long-term drive remains “to control my own destiny and challenge the limits of growth,” from his personal website and Autohome to cars and AI terminals. He sees this as aligned with AI’s direction of continuously increasing capability, compressing knowledge, and expanding action.

45. Wisdom Comes from Relationships with All Things, Not from Stacking More Intelligence

  • Li Xiang defines wisdom as “our relationship with all things.” If someone has never lived in a forest, they understand wood only as chopsticks, paper, or a table. Without spending sustained time with children, they cannot truly understand intimate relationships.

  • Building relationships with all things first requires time. If humans spend all their energy on complex information, repetitive labor, and involution, they lose the opportunity to experience the world, understand others, and gain wisdom. That is the deeper meaning of AI as a production tool.

  • He sets a concrete goal for Li Auto’s sales system: appointment calls must be handed to an Agent. If this is still not achieved by the end of the interview year, intelligent business and sales-and-service work will be “unqualified.” Handing over this work to Agents could save 20%–30% of time and reduce substantial energy consumption.

  • Humans’ limited compute should handle relationships, wisdom, and entropy reduction; AI should handle complex data, knowledge compression, and automatic action. The two are partners. Machines already handle welding, painting, and stamping, but that does not mean machines have replaced human leadership.

46. The Current Transformer Has No Consciousness, but a Stronger Architecture Remains Unsolved

  • Li Xiang clearly believes that the current token and next-token architecture has no autonomous consciousness and cannot self-evolve. To fundamentally change its output, the model still needs to be retrained by humans; query and prompt can produce only relatively limited changes.

  • He therefore judges current large models to be “pretty safe” for humans. Risks to information, property, and personal safety can still be constrained through training and alignment. If an architecture beyond the Transformer emerges with autonomous evolution, he can only answer whether humans will still control it with “I don’t know.”

  • AI solves intelligence, not wisdom; intelligent people may be not wise at all. Li Xiang expects AI capability to surpass humans within 5–10 years. The human challenge then will not be continuing to compete with machines on intelligence, but developing wisdom in dealing with oneself, others, and all things.

47. Whether Humans Continue to Lead Civilization Depends on Whether Wisdom Can Be Trained

  • Li Xiang sees trade conflicts and geopolitical disputes as evidence that wisdom is growing slowly: corresponding examples for what happens today can also be found in 1930. He therefore asks whether wisdom can become education and training like capability, rather than something acquired randomly through fate.

  • Children first come to know themselves through their parents, adolescents through classmates and teachers, and adults through society. This process is currently highly random. He imagines turning “knowing oneself, understanding relationships, and becoming a better self” into an educational module that can be developed through guided dialogue.

  • Faced with the prediction that AI may cause humans to surrender cognitive sovereignty, he admits this is the biggest unsolved problem since humanity appeared. If AI solves the problem, it may rule humans; if humans solve it, humans can continue leading civilization.

  • He rejects filtering out so-called bad human nature for AI: “Without the bad, there is no good.” Strengths and weaknesses, culture and personality are all traits of life, and all human nature is worth preserving. What must improve is not turning humans into pure models, but understanding and using these traits.