Pioneers Insight Method Research Author
World Models and Physical-World AI with BAAI Dean Wang Zhongyuan
Back to Episodes

World Models and Physical-World AI with BAAI Dean Wang Zhongyuan

Summary

  • Wang Zhongyuan’s clearest conviction is that artificial intelligence will formally move from the digital world into the physical world. BAAI is exploring that direction through its “Physical World” model series. Its roadmap runs “language → multimodality → embodied intelligence and AI for Science → world models,” on the grounds that the large-language-model track is relatively mature, “should be left to companies,” and “large-model development is nowhere near its end.”
  • There are 3 ways out of the text-data bottleneck: synthetic data, post-training/inference scaling, and multimodality—with multimodal data offering especially substantial volume. Citing Ilya’s view that “there is only one internet dataset in the world,” he says audio, images, video, 3D and neural signals could amount to 100x, 1,000x or even 10,000x the volume of text data, and have not yet been effectively used to train large models. That is the core reason BAAI is betting on native multimodality.
  • Most current multimodal models take a shortcut: they use a language model as the core and map other modalities onto it, with the result that a PhD becomes a high-schooler. A typical symptom is that they cannot tell whether 3.1 or 3.2 is larger; by contrast, the human brain becomes more capable as it encounters more modalities. His technical conviction is that “the simpler and more unified the architecture, the more likely it is to have staying power”—native multimodality should use one unified architecture for every modality from day one.
  • Embodied intelligence is hot, but he publicly pours cold water on it: it is a 5-to-10-year cycle, possibly longer, and embodied intelligence is not the same as humanoid robots. “People should absolutely not expect robots to be everywhere within the next 3 years”; the hype has built too quickly and “could very likely fall into a trough of the bubble within the next 1-2 years.” His expected path to deployment is a multimodal/world-model foundation model plus reinforcement learning, illustrated by a 2-year-old girl learning to unwrap candy from short videos.
  • His defense of scaling law takes a 70-80-year view: BERT’s 100M parameters were 1M times smaller than the human brain, GPT-3 was 1,000x smaller, and GPT-4’s 1.8T was 100x smaller. “When our models one day reach the number of parameters in the human brain, then we can debate whether scaling law is effective or not”—he remains firmly committed to the neural-network path, while saying any replacement for Transformer would be a long-cycle event.
  • He only partly agrees with Yao Shunyu’s view that AI has entered halftime, with evaluation becoming more important than training: the thesis may hold for large language models. Domestic models did enter an application boom after catching up with GPT-4, but “the arrival of the entire physical-AGI era still requires several major technological breakthroughs.” On Meta’s world model V-JEPA 2 topping the rankings, he says BAAI’s hard requirement for its embodied team is “deployment on real machines,” not leaderboard performance.
  • The world-model definition has not even converged, and the Turing laureates do not share a single path: Yann LeCun rejects autoregression, Bengio rejects reinforcement learning and advocates scientist AI with “intelligence but no self,” while Sutton remains convinced by RL and the age of experience. “Some are climbing the southern slope, others the northern slope.” BAAI has chosen a simple, unified architecture that can scale up to train a world-model foundation model, while acknowledging that “whether it can ultimately deliver the expected results remains to be verified.”
  • BAAI’s organizational model is itself an investment signal: a post-1985 dean, a 29-year-old head of Emu3, a culture of flagship work rather than seniority, and genuine open source. Its BGE embedding models have been downloaded hundreds of millions of times and once topped Hugging Face’s monthly rankings; hundreds of institutions across more than 30 countries and regions use its data and frameworks. Under Zhang Hongjiang’s philosophy that “success need not belong to me,” BAAI sees itself as the soil in which towering trees grow, not the fruit itself.

Deep dive

1. BAAI’s DNA: Giving Tens of Millions in Resources to Young Researchers Who Couldn’t Make Associate Professor

  • Host Wei Shijie opened by revisiting her interview 3 years ago with founding board chair Zhang Hongjiang. After GPT-3 emerged in 2020, BAAI followed OpenAI’s example, cut every side project and went all-in on large models, scaling from 80 GPUs to a 10,000-GPU cluster. Liu Zhiyuan had not even qualified for associate professorship at the time, yet BAAI put him in charge of a project with tens of millions in resources—“completely impossible in the traditional research system.”
  • The Wudao project produced figures including Tang Jie, Yang Zhilin, Liu Zhiyuan and Huang Minlie, many of the eventual central figures in China’s large-model startup wave. BAAI was consequently dubbed “the Whampoa Military Academy of Chinese AI.” 3 years later, the institution has appointed its first post-1985 dean, Wang Zhongyuan.

2. Six Months of Interviews, Ending with “Choosing the Youngest One”

  • Wang Zhongyuan first witnessed BAAI’s formation in 2018 as a representative of Meituan, a governing institution. He was introduced to the selection process in July or August 2023: “For a research institution, the usual choice would be an academician or a big name. I was only 37 or 38 at the time—would I really have a chance?” He says, “I was excited, but I couldn’t believe it.”
  • The evaluation covered the full stack: research achievements, the ability to bring systems into the real world—“not just publish papers,” but make research results accessible to ordinary people—and management ability: “setting strategy, building teams and delivering results requires a complete system.”

3. Why AI Is a Young Person’s Business: Past Success Can Become Innovation’s Constraint

  • Citing Hinton’s line that “many of my innovations were actually done by my students,” as well as remarks by Sora’s technical lead at a BAAI conference, he says today’s AI breakthroughs “have overturned all previous paradigms.” The deeper someone is embedded in traditional research conventions, the harder it is to break free; “young people are unconstrained, and young people have no failures.”
  • The mechanism is explicit: no seniority-based promotion, no emphasis on titles, and a flagship-work culture. The head of the native multimodal world model Emu3 is only 29 and “can receive support on the order of tens of millions.” Selection focuses more on whether someone has “an ideal for technology,” echoing Zhang Hongjiang’s philosophy that “success need not belong to me.”

4. The Little Girl’s Question: The Moment He Left Industry for BAAI

  • A few days after GPT-4 launched in March 2023, Wang taught an AI class to a third-grade elementary-school class. The room erupted at the idea that they might not need to do homework in the future. Then a girl stood up and asked: “Uncle, if artificial intelligence can do everything, what will we do in the future?” “Such a simple question went straight to my heart.” He realized that AGI, which he had assumed the next generation would encounter, might arrive in his own generation.
  • His macro view is that large models could drive “an industrial-scale revolution, or at the very least an industry-wide revolution”—10-20 years at the short end and 30-40 years at the long end. At a company, however, “it is hard to put 100% of your energy all-in on AI”; as he wrestled with that, BAAI extended an offer.
  • His answer to the girl—one he admits is still not fully accurate—is to understand and use the technology, build a worldview, philosophy of life and value system, and develop the ability to learn independently and interact with others. Universities will no longer be primarily about testing knowledge; they will matter more for building learning and interpersonal skills.

5. Fundamentals and Details: From Microsoft’s Ethics Standard to Facebook’s Bad Case

  • The only line he remembers from Microsoft onboarding was: “Hold yourself to the highest ethical standards.” Wang Xing, or “Brother Xing,” offered the methodology in a company-wide letter titled “Hone the Fundamentals”: “If you stretch business out over a sufficiently long cycle, many failures ultimately come from fundamentals that were not solid enough.”
  • Applied to large models, why do models trained by different companies perform differently? “At the end of the day, it’s the fundamentals”: data quality, chip utilization, handling every network failure, and “not letting any anomalous fluctuation in the loss curve go without investigating the cause.” Evaluation must likewise leave no detail unexamined.
  • His best example is Facebook’s entity linking. When someone searches for “apple,” is it the company or the fruit? The team initially focused only on the metric and ignored the details, producing poor accuracy. He went deep into every bad case: “Once we look closely at every detail, the solution emerges naturally.”

6. Leaving Microsoft: A Shangri-La and the Fear of Being Left Behind

  • In 2016, BAT was already established while Meituan and Toutiao were rising and “all kinds of skyscrapers were appearing from nowhere.” Research institutes felt like “a safe harbor, somewhat like a Shangri-La, a utopia, an ivory tower.” He sensed that the world was changing violently and worried about “being left behind by the times,” so he chose a more commercial, product-oriented company.
  • Facebook’s philosophy combined looseness and pressure: 4 months of parental leave, working from home on Wednesdays and fully subsidized meals, but weekly scrum, frequent OKR reviews, incentives and a performance-exit mechanism that “made you constantly prove yourself.” “Everyone around you was among the best researchers and engineers in the world. If you didn’t advance, you fell behind.” The host’s summary, which he endorsed: “Loose on process, strict on results.”
  • He brought the model to BAAI: no clock-in or clock-out, sharply higher year-end bonuses for S- and A-level employees—“even higher than at internet companies”—while retaining the performance-exit mechanism.

7. Speed and Patience: The Llama 4 Lesson and the Positive Feedback Loop of Low-Hanging Fruit

  • That morning, the team had been discussing why the Llama series had recently been less impressive. According to a source among his former Facebook colleagues—a claim he stressed “may not be correct”—one group was replaced every few months, leaving too little accumulation and institutional memory. “My first reaction was that this fits Meta’s move fast and break things culture very well.” That culture can produce unique breakthroughs in some areas, but is not universally applicable.
  • BAAI’s answer is to move quickly on what is certain—“if you don’t move fast, someone else will do it first”—while giving 2-3 years of patience and room for trial and error to long-term exploration. He fully agreed with the host’s description of the mechanism: prove yourself through short-term results, earn more resources and then “fight a bigger battle.” Without that loop, he asks, how can investors and resource owners gain confidence when someone proposes a major innovation?

8. Updating the Mind: Satya Nadella, the 18th Floor and the Growth Mindset

  • 2 ideas from Satya Nadella’s Hit Refresh had a lasting impact: empathy and growth mindset. “A strong person constantly updates their cognition through the world, the people they meet and the knowledge they acquire.” The opposite, a fixed mindset, is simply “this person is too stubborn.”
  • “Old Wang’s” elevator analogy from Meituan is worth preserving: “When the elevator you’re riding reaches the 18th floor, did you have the ability to reach the 18th floor yourself, or did your ability to bounce around inside the elevator reach the 18th floor?” In a review, one must identify not only what was right, but also reflect clearly on what was wrong.
  • His career summary: Microsoft taught him how to do research, Facebook taught him how to deploy quickly, and Meituan taught him management. “There are common traits behind different excellent companies. All of them require you to keep updating your cognition.”

9. Training the Mind: Peak of Ignorance, Valley of Despair, Slope of Enlightenment

  • The fourth element in Meituan’s management framework—“set strategy, build teams, deliver results, train the mind”—is the last because “even if you give it everything, you can still fail, and your mind needs to be strong enough.” A line from Old Wang remains with him: “If we succeed, we achieve it together; if we fail, we help one another become better.”
  • The Meituan path works like this: people who succeed may mistakenly attribute the combined result of timing, location and human factors entirely to their own ability, placing themselves at the peak of ignorance. Managers must “push you into the valley of despair,” then help you climb the slope of enlightenment—learning what can be achieved through ability, what requires moving with the broader trend, and what should perhaps be abandoned.
  • Asked whether he had been through the valley of despair, he replied verbatim: “Of course I have, but I don’t want to share it.” His way out was a trip to Norway’s Lofoten Islands: “You realize how large the world is. Why confine myself to my own cocoon? There can be many paths through life; there doesn’t have to be only one.”

10. The Kuaishou Breakthrough: Transformer Was Headed Toward Unification

  • At Meituan, he grew from a one-person operation into the leader of a 400-500-person search and NLP organization. At Kuaishou, after taking over the MMU team, his first major push was to upgrade the entire visual backbone to Transformer. At the time, Transformer remained controversial in industry because of its heavy resource consumption.
  • The key insight came as NLP—BERT and GPT—vision and audio all moved toward Transformer: “I began to realize that deep learning in AI might be heading toward a grand unification.” Previously, each field had its own encoding methods, data formats and network structures.
  • His analogy for the host: “It’s as if our country has 56 ethnic groups, but they all share a common language, Mandarin. Once we can all use Mandarin, different modalities can communicate, potentially fuse and produce unexpected effects.”

11. Cross-Modal, Multimodal, Full-Modality and Native Multimodality: Deflating the Terminology

  • His distinction is that “this multimodality is not that multimodality.” Most image- and video-understanding models are composites: a large language model at the core, with other modalities mapped into language—CLIP-style. Text-to-image uses diffusion; text-to-video uses diffusion transformer. The underlying technical approaches differ, leaving non-specialists unsure which kind of multimodal model is being discussed.
  • Full-modality is closer to grand unification: every modality can serve as both input and output. Like the human brain, it can “imagine a scene with its eyes closed, almost like a video,” making it more human-like. Native multimodality means “designing a unified architecture from the start,” rather than bolting modalities onto an existing language model.

12. Large Models Are Nowhere Near the End: 3 Exits from Data Exhaustion

  • BAAI’s strategic debate was whether an institution should continue working on large models now that the technology has been industrialized. The conclusion was that “the mature and converged technology path is essentially limited to large language models, which should be left to companies.” Ilya said at NeurIPS that “there is only one internet dataset in the world,” suggesting the pre-training phase may already be over.
  • There are 3 exits. First is synthetic data: once machine intelligence reaches or exceeds human intelligence, it could create higher-quality data than humans do and feed it back into training—important, but “not yet fully solved.” Second is post-training and reasoning models: slow thinking revealed “a new scaling-law curve,” leading to o3, o4, DeepSeek R1 and the R2 everyone is waiting for. Third is multimodality.
  • The strongest quantitative case in the episode is the scale of multimodal data: audio, images, video, 3D and neural signals could amount to 100x, 1,000x or even 10,000x the volume of text data, and have not yet been effectively used to train large models.

13. Can Multimodality Increase Intelligence? An NLP Researcher Corrects Himself

  • He acknowledges the debate across academia and industry, as well as his own “natural pride as an NLP person”: “Language is the complete system unique to humans,” and many researchers believe “language is the boundary.”
  • His position shifted: “Whether it increases intelligence depends heavily on how you define intelligence.” Many animals have no language system but still possess intelligence—foraging, communicating, climbing and distinguishing things by smell. Some even have capabilities humans lack.
  • The more practical point is that in real-world deployment, “language alone is far from enough.” Presentations, flowcharts, design drawings, MRI and X-ray images, and teachers’ lesson plans and notes are all multimodal. “Whether or not it increases intelligence, multimodality is a direction AI must break through.” Human learning also precedes language: “We cannot speak at birth, but we already begin learning through vision. Systematic language learning starts in kindergarten and elementary school.”

14. The Cost of the Shortcut: A “PhD” Becomes a “High-Schooler” After Multimodal Inputs

  • The strange behavior of today’s language-model-centered multimodal systems is that a language model trained to a doctoral level appears to regress to university or even high-school intelligence after other modalities are added. “The problem we often see is that it can’t tell whether 3.1 or 3.2 is larger.”
  • His analogy is pointed: “It’s as if you earned a PhD in a closed environment, then someone suddenly told you what the world is really like. The shock is so great that you become somewhat mentally unstable, lose intelligence and muddle through basic problems.” The human brain, by contrast, becomes smarter as it encounters more knowledge and does not suddenly forget how to speak. That is the fundamental neural-network question native multimodality is meant to address.

15. What Is a World Model? Predicting the Next Action, Not Describing a Static Image

  • Existing multimodal models are weak at spatial and temporal perception and mostly describe static scenes: “There are 2 people communicating here, both wearing black.” But people rarely communicate with the real world that way. Our first instinct is to predict the next scene and the next action.
  • His example is straightforward. A water bottle sits near the edge; one accidental touch could knock it over. If the cap is off, the situation is more dangerous because the water will spill everywhere. A human therefore caps it quickly and moves it closer to the middle. We do not think, “There is water at the edge in a transparent plastic bottle with no label.” That is not how we reason.
  • Yann LeCun proposed world models at the 2023 BAAI conference, but “the industry still has no clear definition.” With even the term unsettled, technical practice naturally diverges. BAAI’s choice is to train a world-model foundation model with a “very simple, easily extensible structure that can scale up.” “Whether it can ultimately deliver the expected results remains to be verified”; additional modules may be needed.

16. 3 Turing Laureates, 3 Different Paths

  • In conversations with Bengio and Yann LeCun this February, he observed that LeCun rejects autoregression and has challenged the large-language-model route in many settings as incorrect and incapable of reaching AGI. Bengio rejects reinforcement learning as lacking generalization, and as a pessimist advocates scientist AI with “intelligence but no self,” built with relatively unified controls to ensure safety.
  • Sutton, the father of reinforcement learning, took the optimistic view in a brightly colored shirt: “From birth, humans constantly interact with the world, receive rewards and then grow.” AI, he argues, must enter the age of experience.
  • Asked whether BAAI is in Sutton’s camp, Wang rejected the framing: “We can’t divide the camps so simply.” “Some are climbing the southern slope, others the northern slope. Perhaps they can all reach the summit eventually, but the scenery along the way will be different.”

17. Why Embodied Intelligence Is So Hot This Year—and Why to Fear a Bubble

  • He admits, “I don’t know why embodied intelligence is so hot,” but identifies 3 drivers. Humanoid robots can finally walk, and walk stably, making the idea that carbon-based and silicon-based life might coexist seem plausible. The role of reinforcement learning in language models has made researchers believe robot RL can also break through. Large models are helping robots “see the world,” with policy support adding another tailwind.
  • The cold-water message is unambiguous: “Embodied intelligence is a 5-to-10-year cycle, possibly longer”; “embodied intelligence is not the same as humanoid robots”; and “people should absolutely not expect robots to be everywhere within the next 3 years—that is a completely unrealistic expectation.” The hype has risen too fast and “could very likely fall into a trough of the bubble within the next 1-2 years.”
  • On BAAI’s position, he says: “We are not an institution that follows the trend. We didn’t work on embodied intelligence because it became hot. Quite the opposite: we worked on embodied intelligence—and to some extent, we may have made it hot, though I’m not sure.” The strategic logic is that AI should benefit humanity: it should not only write poetry, but also do things. To do things, it needs a body—whether a robotic arm, a wheeled single arm, a wheeled dual arm or a humanoid.

18. The Embodied-Intelligence Stack: Foundation Model Plus RL, and a 2-Year-Old Learning to Unwrap Candy

  • His route—explicitly “just one school of thought”—is modeled on the success of “foundation model + RL post-training” for large language models. Embodied intelligence will likely use a multimodal model or world model as its foundation model, then collect real-world data or learn through experience, continually unlock capabilities and retain the skills it learns.
  • The real-world example came during the Lunar New Year. A 2-year-old girl, without instruction from an adult, learned to unwrap candy and thread blueberries onto a toothpick. “We were all shocked.” She had watched large numbers of livestreamers unwrapping candy on her phone; her brain learned the skill, and she then practiced it in the real world. She could not tear the wrapper at first, but eventually realized that the serrated section was the point to pull and learned how to do it.
  • The historical limitation of RL-only approaches is also clear: RL-based autonomous driving and robots picking up cups have “remained demonstrations” and lack the generalized decision-making and action capabilities of humans.

19. The Physical World Series and RoboBrain: The Real Test Is Deployment, Not Rankings

  • At this year’s BAAI conference, the institution released the Physical World series. “World” refers to breaking through the boundary between virtual and physical worlds, continuing the naming lineage of Wudao—a pun on Wudaokou. The lineup includes the native multimodal world model Emu3, the neuroscience multimodal general foundation model Brain-Mind, the cross-embodiment large- and small-brain collaboration framework RoboOS and embodied brain RoboBrain, all iterated to 2.0, as well as the all-atomic microscopic life model OpenComplex.
  • The “most powerful brain” claim is based on evaluations of RoboBrain’s spatial understanding and task planning against industry vision-language models and other embodied-brain models. “It has indeed demonstrated that it exceeds those models on these capabilities,” and it has been open-sourced. Spatial tasks include “bring me the reddest apple,” “the largest apple” or “the apple on top”; when an obstacle blocks the water, the robot must go around it or move it. These are easy for humans and extremely difficult for current robots.
  • On Meta’s world model V-JEPA 2 topping Hugging Face’s rankings, he was restrained: “The technical routes are not quite the same, and the significance of rankings in this matter is not even that great.” BAAI’s “very important requirement” for its embodied team is deployment on real machines, validating systems in real-world scenarios rather than chasing rankings. “These headlines will keep appearing. We will continue moving forward at our own pace and along our own technical route.”

20. The Ledger of Genuine Open Source: Hundreds of Millions of BGE Downloads, but No Way to Track the Users

  • BGE’s general-purpose embedding model topped Hugging Face’s monthly download rankings in October last year and was the most-downloaded AI model on the platform from 2023 through the end of last year. It has been downloaded “hundreds of millions of times.” Many corporate contacts have privately told BAAI, “Your BGE model is really good; we use it internally,” including familiar large internet companies and prominent startups in China and abroad—“though of course they won’t say so publicly.”
  • Its open-source data has been downloaded millions of times. Registration records show users across more than 30 countries and regions and hundreds of institutions. Genuine open source means more than releasing weights: BAAI releases “the code, data, model itself and evaluation methods.” “There has always been a deep belief that success need not belong to me.” The seedlings it grows do not have to remain in BAAI’s soil; handing them to the market and to companies is also a meaningful contribution.
  • He does not hide the resource constraint: “As a nonprofit institution, I often feel deeply the bottleneck in the resources required to grow, and many times we have no choice but to make difficult decisions.” But if the technical route is right, an individual BAAI team may achieve better results with fewer resources than a large company.

21. Partial Agreement with “Halftime,” and a Long View on Scaling Law

  • Yao Shunyu argues that AI has entered halftime, shifting from “training greater than evaluation” to “evaluation greater than training.” Wang says, “I partly agree—perhaps this is true for large language models.” He predicted last June that domestic models could catch GPT-4 by year-end and that Agents would become a major direction; DeepSeek confirmed the view, though another domestic model could have caught GPT-4 even without DeepSeek. Once models become usable, evaluation and user insight become critical. Research institutes, however, must continue pushing the frontier: “The arrival of the entire physical-AGI era still requires several major technological breakthroughs.”
  • His most quotable scaling-law data chain is this: BERT had 100M parameters in 2018, 1M times fewer than the human brain; GPT-3 had 175B, 1,000x fewer; GPT-4 had 1.8T, 100x fewer. “When our models one day reach the number of parameters in the human brain, then we can debate whether scaling law is effective or not.” Compute is advancing under Moore’s law, and multimodal data will expand the pool. “I still believe in scaling law, and I believe in the neural-network technical route.”
  • The architecture question also needs a long-cycle lens. Transformer “has been proven to be a highly general architecture capable of scaling data up.” Efficiency issues are real, but replacing it is closely tied to chip acceleration and operator optimization and will require a long period of continued proof and validation.

22. Seeds Beyond Large Models, and a Pragmatic Willingness to “Wrestle”

  • The difference between how the human brain and large models learn is why BAAI keeps planting seeds. The brain can perform extraordinarily complex comprehension, reasoning and interaction at roughly a dozen watts, and recognize cats and dogs after seeing a few images. Large models must read the entire internet and train on enormous clusters. BAAI is therefore allocating small amounts of resources with ample time and room to explore brain-inspired intelligence, digital-twin hearts, proteins and biomolecular modeling. Brain-Mind came from “a very small team” putting EEG, functional MRI, two-photon and other signals into one architecture and producing “unexpected results.”
  • He is blunt about the gap with the world’s leading institutions: “Whether in resources invested or talent density, we are still far from doing enough.” China’s advantages are its “vast market, huge user base and complete innovation and entrepreneurship system.” “I have very, very strong confidence in the long term. No force can prevent our development in artificial intelligence.”
  • The closing posture is the episode’s credo: “Reaching the summit of Mount Everest is our pursuit; climbing step by step is what we are doing.” As the saying goes, “without taking small steps, you cannot travel a thousand miles.” And his clearest conviction remains: “Artificial intelligence will formally move from the digital world into the physical world.”