Pioneers Insight Method Research Author
92. Academician Zhang Yaqin on Consciousness, Longevity, and Robots
Back to Episodes

92. Academician Zhang Yaqin on Consciousness, Longevity, and Robots

Summary

  • Zhang Yaqin’s AGI roadmap is not a single leap, but a staged rollout: information intelligence in roughly 5 years, physical intelligence in roughly 10, and biological intelligence in roughly 20. The information layer will first reach human-level performance in text, images, video, reasoning, and multimodality, and exceed humans on some intellectual tasks; the physical layer must clear regulatory, environmental, and operational hurdles; the biological layer will move from medical repair to extending memory, intelligence, and lifespan. “Reading ten thousand books” can produce a smart brain, but learning to swim still requires getting in the water.
  • Foundation models will increasingly resemble capital-intensive “ecosystem operating systems,” while schools and companies will differentiate themselves through vertical models, Agents, and edge deployment. Zhang Yaqin argues that foundation models depend on enormous amounts of compute, data, and real-world scenarios, and are primarily engineering innovations that scale “from 1 to 100, even from 100 to N”—work better suited to companies; AIR has chosen intelligent transportation, IoT, and AI+Life Science, while putting the development of future CTOs and system architects ahead of research and industrial impact.
  • Robotaxi has moved from technological conviction to validated capability, with the commercial model still up for grabs; companies that have already been running for 8-10 years may have strong prospects over the next 5 years. Zhang Yaqin sees Waymo and Baidu Apollo/萝卜快跑 as the leading proof points in the US and China, and says 萝卜快跑 is “at least 10x safer” than human driving; but he notes that FSD might still require a takeover every 45 minutes, and while moving from L2+ to L4 is technically possible, it has not yet been achieved.
  • Large models have given embodied intelligence a reusable “brain” for the first time, but they have not eliminated the long cycle of physical deployment. Synthetic data eases the shortage of samples, common sense improves corner-case generalization, and end-to-end frameworks reduce dependence on fragmented modules; RSR then attempts to close the loop from Real to Simulation and back to Real. Yet a home robot still has to handle countless details involving doors, stairs, force, safety, and personal habits, so mainstream general-purpose humanoids may still be 8-10 years away; “adding a brain” does not mean the body immediately knows how to act.
  • The robotics industry will not produce a single dominant player, but will divide into three layers: operating services, vehicle platforms, and chips and components. Tesla can integrate software, hardware, and Robotaxi services end to end, while automakers such as XPeng can build L4 vehicles without operating a platform; Waymo- or Didi-style companies can focus on services. Zhang Yaqin expects many players to survive, with competition shifting from feasibility to who moves faster, who moves slower, and how the business model is executed.
  • Biological intelligence will begin with medicine and longevity, and only later become an enhancement layer for healthy people. Zhang Yaqin is more bullish on non-invasive interfaces and believes biological intelligence can be achieved within 20 years, likely beginning with interventions for blindness, deafness, neural damage, and Alzheimer’s before extending toward near-infinite memory and higher intelligence; he believes silicon and carbon-based systems will merge, but does not believe a silicon system will generate consciousness on its own: “Will we become immortal? I don’t know.”
  • The most radical social revaluation may come in 30 years rather than 10: driving, household devices, working time, lifespan, and education will all be rewritten. His baseline picture is that autonomous driving will be widely adopted within 10 years and home robots will be as common as refrigerators; 30 years from now, people may drive the way we ride horse-drawn carriages today, 100 may become the norm, some people may live to 120-150, and the workweek may shrink to a single day. Humans may become a carbon-silicon “new species,” while education shifts from memorization to “how to ask questions.”

Deep dive

1. AI Did Not Appear Overnight; Disparate Fields Converged Over the Past Decade

  • Zhang Yaqin recalls that when he returned to China at the end of 1998 to join Microsoft Research China, artificial intelligence was still in a “dormant” phase: Kai-Fu Lee worked on speech recognition and synthesis, Zhang worked on video compression, Tong Xin on multimedia, 沈向阳 on computer vision and graphics, and Professor Huang on NLP—but “not one of us was called an artificial-intelligence researcher.”

  • In his account, the capabilities now grouped under AI emerged from the convergence of decades of work in vision, video, speech, natural language, statistical learning, and other fields. AI truly moved from exploration to “a useful science” over the past 10 years, rather than at the instant any single model was released.

  • That history also explains his view of research strategy: China’s early research capabilities lagged the US by a wide margin, and some fields had no obvious leading figures, so the approach was often “find the right people and let them do what they want”; after 25-30 years of development, every field now has world-class talent, and the harder question is “what should we choose to work on?”

2. The Prerequisites for a World-Class Research Institution Are Talent Density and Long-Term Goals, Not KPIs

  • When Microsoft Research China was founded, the central question was whether a world-class computer research institute could be built in China, where research still lagged the US by a wide margin. Zhang Yaqin says the team, Microsoft, and Bill Gates all had confidence, but during the first 2-3 years, “80% of our time was spent looking for people”—first assembling top young researchers from China and abroad.

  • Asked how “first-class” should be quantified, Zhang Yaqin relayed the answer from Turing Award winner Raj Reddy: “When you stop asking that question, when you stop talking about it, you’ve become first-class.” The goal can be explicit, but actually reaching it may not be confirmed by any single metrics system.

  • He applies the same test to ChatGPT: within half an hour of using it for the first time, his instinct was, “This ChatGPT has passed the Turing test.” He also cited a remark by another Turing Award winner—possibly Jim Gray, as he remembers it—“When you pass, you know,” arguing that certain phase changes first appear as an overall experience rather than as a score crossing a threshold.

  • The three priorities at the time were attracting first-rate talent, doing work with major impact on Microsoft, industry, and academia, and training a generation of researchers; talent development initially looked more like an outcome. At Tsinghua AIR, the order was explicitly reset to “talent first, research second, industrial impact third.”

3. Foundation Models Are Capital-Intensive Platforms; AIR Is Betting on Talent and Industrial Verticals

  • Zhang Yaqin compares foundation models to a horizontal “ecosystem operating system”: they require huge amounts of compute, data, and concrete scenarios, which schools struggle to provide independently. Universities can build them with companies, but this work is largely about going “from 1 to 100, even from 100 to N”—part algorithmic innovation, but with a vast amount of engineering innovation.

  • AIR chose three industrial directions: intelligent transportation, IoT, and AI+Life Science. In later terminology, these roughly map to embodied intelligence, biological intelligence, edge intelligence, and Agents. The institute does not serve a single company; it aims to influence entire industries through corporate partnerships and new-company incubation.

  • Whereas Microsoft Research centered on scientific research and corporate technology strategy, AIR’s first priority is to develop “top-tier architects and CTOs.” Zhang Yaqin believes China is not short of technical talent; its shortfall is leaders with international perspective and the ability to architect large systems. His assessment of AIR’s PhD students is blunt: “Absolutely world-class,” and no weaker than those at MIT or Stanford.

  • When 张小珺 asked whether the opportunity to build the brain for information intelligence had already been taken by large companies, Zhang Yaqin disagreed: the horizontal foundation layer is more concentrated, but vertical information applications such as image generation and programming still offer opportunities. The fundamentals of business have not been rewritten by AI either: “You still need to do what you should do.” The first thing to change is productivity.

4. Information Intelligence Will Near Human-Level Performance Within 5 Years, but High Intelligence Does Not Equal the Ability to Act

  • Zhang Yaqin’s 10-year framework is that next-generation AI will combine information intelligence, physical intelligence, and biological intelligence. The three are related, but correspond respectively to the information world, physical infrastructure, and living organisms; sharing AI technology does not make them the same problem.

  • His timeline is that within 5 years, information intelligence will pass the Turing test and reach human-level performance in natural content such as text, images, and video. ChatGPT has already convinced him that the text component has “basically arrived”; what remains is stronger reasoning, multimodality, and Sora-style video generation.

  • Information intelligence is like a person with an exceptionally smart “head”: it can write, draw, program, solve mathematical equations, and even outperform mathematicians or physicists on corresponding intellectual tasks. That is what he means by reaching human-level “IQ”; it does not mean every real-world skill will arrive at the same time.

  • Zhang Yaqin uses swimming to illustrate the boundary: “No matter how many books you read, you still won’t know how to swim.” Information intelligence is like “reading ten thousand books,” while physical intelligence requires “getting in the water” and “traveling ten thousand miles”; even if a model understands the knowledge behind a skill, it still has to acquire the skills in a real environment.

5. Physical Intelligence Will First Land in Closed Tasks; Home Robots May Be the Hardest

  • Zhang Yaqin estimates that high-level physical intelligence may still be about 10 years away. By then, robots may outperform humans at most tasks, but “whether they will be better than humans at everything is certainly not guaranteed.” The first major deployment will not be the general-purpose humanoid, but autonomous driving, because driving is a relatively controllable closed problem, while humanoid robotics is an open problem.

  • Autonomous driving can be viewed as “a robot that drives”: it does not need to write poetry, sing, or understand biology; it only needs to master driving. A general-purpose robot, by contrast, must understand common sense, human behavior, and constantly changing environments, with far less clearly defined boundaries than road driving.

  • He divides robots into three categories: household, industrial, and social. Industrial robots in mines, factories, or dangerous environments have fixed objectives and may not need a humanoid form; social robots include police, security guards, delivery workers, and drivers; household robots must care for the elderly, do housework, adapt to personal habits, and meet safety requirements, making the fragmented home environment “possibly the hardest.”

  • When 张小珺 pressed him on whether household or social robots were harder, Zhang Yaqin refused to offer a facile ranking: “It depends. Neither is easy.” His real distinction is between specialized machines, which already exist in large numbers, and general-purpose humanoids that can be bought and put to work at home, which will require patience.

6. The Value of General-Purpose Robots Lies in a Shared Brain; Humanoid Form Is One Shell Adapted to Human Society

  • Zhang Yaqin wants the backends of household, industrial, and social robots to be 70%-80% identical: a massive general-purpose multimodal model in the cloud, a vertical model for driving or a specific task in the middle, and a smaller edge model running in the vehicle or robot. The architecture resembles the layering of an operating system, a super app/Agent, and end-user applications.

  • Today, foundation models and autonomous driving still use two separate architectures. He expects them to converge: models will learn from the physical world, form understanding and decisions through large models, and then execute through different bodies. The front-end form can vary—“a machine can be anything”—whether a person, dog, cat, or another specialized structure.

  • The first advantage of the humanoid form is interaction. Elderly people are more likely to chat naturally, confide in, and treat a humanlike robot as a housekeeper or companion; second, stairs, buttons, doors, and other basic infrastructure were designed for people, so a humanoid body can use the existing environment directly.

  • He even suggests that within 10 years, there “may be more robots than people,” with each person potentially having a copy or “double.” That double could be smarter than its owner and handle tasks on the owner’s behalf, but under his formulation it must “belong to you” and “listen to you,” rather than possess independent authority.

7. Waymo and 萝卜快跑 Have Proved Robotaxi Feasible; FSD Has Yet to Clear the L4 Threshold

  • Zhang Yaqin recently rode in a Waymo in San Francisco and found the vehicle more “smooth” than a human driver. More important, local residents were already willing to ride in it and no longer viewed it as a novelty. That shift in public acceptance is a key reason he believes autonomous driving has moved beyond the validation phase.

  • On the controversy surrounding 萝卜快跑 in Wuhan, he distinguishes technical performance from social acceptance: passengers rate it highly, while the wider public still treats it as something new. His strong view is that 萝卜快跑 is “at least 10x safer than human driving,” but he acknowledges that acceptance will still require time in operation.

  • Asked by 张小珺 about the L2+ versus direct-to-L4 paths, Zhang Yaqin said Tesla’s approach “can rise to” L4, while emphasizing that “it cannot do so today.” Existing FSD, he noted, may still require a takeover every 45 minutes, whereas a Robotaxi must basically require no takeovers—and eventually may not even have a steering wheel.

  • In his view, Waymo and Baidu Apollo/萝卜快跑 are furthest ahead, having respectively demonstrated in the US and China that “the technology is feasible.” Two years ago he believed it would eventually happen but could not say when; now he is “very clear,” and the remaining questions are who runs faster, who runs slower, and how the business model will be executed.

8. Autonomous Driving’s Business Model Will Split into Operations, Vehicles, and Core Components

  • Zhang Yaqin breaks the industry into three categories: Didi- or taxi-style operating-service providers, companies that manufacture the vehicles, and companies supplying chips and core in-vehicle components. The three layers can operate independently or integrate vertically; technical feasibility does not mean value will necessarily accrue to one company.

  • Tesla can choose to handle the vehicle, autonomous-driving system, and Robotaxi service itself; XPeng can develop L4 technology and sell vehicles without operating a Robotaxi platform. Zhang Yaqin expects many automakers to continue doing what they do today—selling cars—and does not expect a single company to dominate. All of these positions can coexist.

  • His view of new entrants is notably more restrained: starting another generalized autonomous-driving company from scratch is no longer realistic, because the existing players have already been running for 8-10 years. 萝卜快跑, Pony.ai, WeRide, and Horizon Robotics are beginning to show validation through listings, products, or services, and the next 5 years may be when the sector enters a stronger growth phase.

9. Large Models Ignited Embodied Intelligence by Loosening Three Locks at Once: Data, Generalization, and Architecture

  • Traditional autonomous driving faced three fundamental locks. The first was insufficient data: testing could not cover every dangerous scenario; the second was the endless emergence of corner cases and weak model generalization; the third was the fragmented architecture in which maps, vision, language, lidar, perception, planning, and decision-making each operated independently and were then awkwardly stitched together with rules—“a complete mess.”

  • The first change brought by large models is the ability to generate more training scenarios from real data, sharply increasing simulation speed; the second is stronger generalization through common sense and simulators, making the system more likely to handle situations it has never seen; the third is end-to-end modeling, placing complex modules inside a unified input-output framework and using rules mainly as a fallback.

  • Zhang Yaqin twice preserved the boundary: these challenges “cannot be described as completely solved”; they have only been rapidly and substantially mitigated. Large models have not erased the safety problem, but they have accelerated development paths that were previously difficult to scale.

  • At the architectural level, the industry is moving from different algorithms for different inputs toward unified Transformer and token-based processing. Musk’s end-to-end model uses BEV and has given the industry a glimpse of light; other teams are adopting the same direction. Asked why embodied intelligence is so hot this year, his answer was that everyone has finally “seen the light.”

10. RSR Brings Simulation Training Back into the Real World, While Large Models Turn Natural Language into Action Sequences

  • Robotics has even less data than autonomous driving, and traditional reinforcement learning often produces simulation policies that “don’t work” when transferred to a physical machine. Zhang Yaqin’s team proposed RSR, or Real to Sim to Back to Real: analyze real scenes to create simulations and digital environments, expand them in Simulation with generative AI, Stable Diffusion, or planning tools, and then return to the real world for validation.

  • This loop can be called a world model. Its purpose is not to replace reality with simulation, but to make strategies learned in digital space usable more quickly and compensate for the shortage of real-world data. It connects the part of the pipeline most prone to distortion: the gap between model training and physical deployment.

  • Large models also “add a brain” to robots. When a user says, “Take these dirty clothes down, wash them, take them to the dry cleaner, and bring them back for me,” the model can first understand the intent and decompose the task, then hand the actions to the machine. The hardest layers of natural-language understanding and common sense now have a general-purpose foundation.

  • The details still determine whether deployment works: a robot may know how to drive but not how to open a door, may enter a home and fail to open the microwave, or may not know how much force to use. In the past, every step required an explicit rule; now the hope is that the robot can reason from what it sees, know to handle hot objects carefully, and remember that the user does not eat lamb. Those capabilities still have to be trained on specific bodies and in specific environments.

11. Humanoid Robots Are Still Where Autonomous Driving Was 10 Years Ago; Mainstream Consumer Products May Take 8-10 Years

  • The difficulty of physical intelligence lies in more than the model. Autonomous vehicles must deal with road access, regulation, and interactions with human-driven cars; delivery robots must handle stairs and doors. Every real deployment point requires a physical-world model and new operating skills, making the cycle inherently longer than for information products such as phones or computers.

  • Zhang Yaqin compares today’s humanoid robots with autonomous driving 10 years ago: simple industrial robots are already widespread, but the hardest general-purpose humanoids may still take another 8-10 years. 周古月’s Discover Robotics already has products, but its customers are still primarily research and educational institutions and small companies, not mainstream households.

  • Automakers moving into robotics do share some capabilities: both must perceive their surroundings, understand scenarios, and act through Vision-Language-Action models. But robots must learn far more tasks; autonomous driving involves a narrower task set, yet it cannot tolerate a mistake within several hundred milliseconds. Taking too long to make breakfast is fine; reacting a beat too slowly while driving may not be.

12. Biological Intelligence Will Start with Treatment, Then Extend Human Memory and Intelligence

  • Zhang Yaqin puts biological intelligence on a 20-year timeline and says he is “more bullish on the non-invasive route”: connecting the brain and machines through more sensitive sensors and new types of brain-computer interfaces without requiring every application to follow a Neuralink-style implantation path. He believes he will see such systems become real within his lifetime.

  • The first stage will not be enhancement for healthy people, but treatment: stimulating relevant neurons to restore sight for the blind and hearing for the deaf, repairing central nervous-system connections in people with disabilities, and attempting to address age-related forgetfulness, Alzheimer’s disease, ADHD, and autism in children.

  • Only afterward would capability expansion begin: storage and memory could approach infinity, and people could become more intelligent. Because this would no longer be an external tool but something closely connected to the living organism, ethics, identity, and how humans coexist with new technologies would become unavoidable constraints.

  • On lifespan, Zhang Yaqin believes organ replacement, better brain health, drugs, and other technologies will extend carbon-based life, while cancer and age-related physical and mental diseases may gradually become treatable. But he refuses to equate longevity with immortality: “I don’t know whether we’ll become immortal, but our lifespans will definitely be longer.”

13. Carbon-Silicon Fusion May Create a New Species, but Consciousness, Society, and Education Will Not Change at the Same Speed

  • When 张小珺 asked whether becoming sufficiently intelligent should automatically produce consciousness, Zhang Yaqin explicitly disagreed. Humanity still does not know how its own consciousness arises, and more compute and data have no inherent connection to consciousness; echoing the point Richard Feynman made, his boundary is that we cannot create something we do not yet understand.

  • On Vitalik Buterin’s view that combining biology and silicon may be one way for humans to participate in super intelligence and avoid being dominated by independent computers, Zhang Yaqin accepts the direction of fusion but insists that “the conscious part will still be our carbon-based selves.” The new form would remain an extension of humanity, not a silicon soul taking control of people.

  • He calls this extension “a new species”: 30,000 years ago, humanity extended itself through fire and stone tools; today, through phones, computers, and the internet; 100 years from now, these technologies may look to us the way stone tools do today. The Industrial Revolution extended physical strength, the information society expanded information capabilities, and AI is rapidly expanding intelligence, making evolution nonlinear and exponential.

  • The 10-year picture is relatively concrete: large numbers of vehicles will become autonomous, household robots will be as common as refrigerators and televisions, and AI hospital, medical, and education applications will become real, while people themselves will not fundamentally change. The 30-year picture is more radical: people may drive the way we ride horse-drawn carriages today, perhaps requiring a special license; 100 may become the norm, and some people may live to 120-150.

  • Longer lifespans and higher productivity will rewrite work and demographic structures. Zhang Yaqin extrapolates from the seven-day, six-day, five-day, and four-day workweeks in parts of Europe to a future in which people may work only one day a week, rather than a small minority working seven days while everyone else is unemployed; falling birth rates, population aging, and healthier old age will emerge together.

  • Education will also have to shift from “student-teacher” to “student-human professor-AI professor.” An AI assistant may know more than a professor, but the core of learning will no longer be memorization and drilling; it will be asking good questions, critical learning, and debate: “Ask it anything and it may answer randomly; ask an exceptionally good question and it will give you a very deep answer.”

  • His reading list centers on longevity and memory. Auto Life discusses AI, Life Science, longevity, and lifestyle; In Search of Memory explains genetic memory, short-term memory, and long-term memory. Zhang Yaqin emphasizes that current AI is still “not very good” at abstracting and grouping knowledge or building hierarchical memory, making the human learning mechanism an important gap for the next stage.