80. A 3-Hour Interview with Academician 梅涛: Language, Video and Embodied AI Will Ultimately Converge—Why World Models?
80. A 3-Hour Interview with Academician 梅涛: Language, Video and Embodied AI Will Ultimately Converge—Why World Models?
Summary
- 梅涛’s model worldview can be compressed into one sentence: “Images are the entry point, video is the process, and the endgame may be a world model.” Images are 2D signals that encode spatial relationships and physical laws. Customers spend 50%-60% of their time on the platform creating a single image, so the priority is to make the image model the world’s best, then move up the stack from “image → video → all modalities → world model.” His team published the world’s first text-to-video paper in 2017; apart from one professor from the University of Science and Technology of China, every author is now at 智象未来 and none has left.
- 智象 is betting on architectural innovation rather than a compute arms race, with its proprietary UiT (Unified Transformer) removing the encoder and decoder so every modality can interoperate through native tokens. DiT’s lossy compression into latent space creates mismatched levels of abstraction, leading to hallucinations and making pixel-level control difficult. Startups have only 2 options: innovate on model architecture, then “use one-tenth the cost and 10x the efficiency to approach or surpass the big labs.” Its open-source 8B HiDream-I1 has been downloaded more than 2 million times, ranking first globally among open-weight models on Artificial Analysis; its closed-source models are now pushing up the rankings.
- Multimodal AI still has no mature scaling law and remains somewhere between GPT-2 and GPT-3—precisely the opening for startups. There is still no unified definition of a visual token: one image contains roughly 100,000 tokens and a 15-second video roughly 1 million. His early experiments showed that a visual “vocabulary” needs more than 1.2 million tokens to produce good results, versus only 5,000-6,000 common English words. Training remains at the 1,000-GPU and near-10,000-GPU level, with the architecture not yet converged; “multimodal AI is not in an arms race today.” Machines may not even be consuming data as fast as humans are producing video.
- Commercialization is arriving before intelligence emergence in multimodal AI: “Intelligence hasn’t emerged yet, but expressiveness has.” Global AIGC revenue was approximately $20B in 2024, with roughly 60% coming from images and video. 智象’s platform consumes about 3T tokens per month; revenue was RMB100M last year and “definitely several hundred million yuan” this year, with To C and overseas revenue each accounting for roughly one-third. AI short dramas will incur nearly RMB10B in production costs this year alone. Short-video advertising from more than 1 million TikTok merchants is a long-duration opportunity, monetized by selling the “shovel” and taking a commission on GMV.
- A world model is not one model but 3 unconverged paths: LeCun’s JEPA predicts the next state, 李飞飞 reconstructs 3D environments, and his own camp is driven by video data. The test is to ask where the most data is. Today’s so-called world and embodied models are only tens of billions or 100-200B parameters, while video models have already reached 100-200B, and 智象 processes about 100 PB of data a year. Real-robot data may amount to only 1 million hours, “and possibly not even that.” The sensible move is to build a business around embodied-data augmentation: 智象 is working with 诺亦腾 to turn teleoperation data into more first-person data while preserving millimeter-level or comparable accuracy.
- His response to Seedance 2.0 and the big labs is differentiation plus downside protection: “Don’t sit in the tire tracks of the big labs… I have only one magazine of bullets, and once it’s gone, it’s gone.” Sora burned $5.5B, while Seedance’s investment is certainly at the RMB10B level. Competing head-on on compute is “neither rational nor worthwhile.” The strategy is to make its closed-source image model a global top-3 product and its open-source model the global No. 1, while taking leadership in verticals such as short dramas and short-video marketing. The big labs are even more anxious: once they fall out of the first tier, the impact compounds exponentially.
- In China, he leans more toward Sam Altman than Dario: “Starting a company in China is much harder than starting one in Silicon Valley,” and investors need you to keep closing orders along the way. Chinese capital “only invests in consensus”; once consensus forms, there is little room for non-consensus bets. The LLM compute gap is “poverty limiting your imagination”—Chinese big labs may train on only tens of thousands of GPUs, versus hundreds of thousands or even 1 million for their rivals. But Chinese companies have performed well in video multimodality, with advantages in short-video data and supply chains. The ticket to stay in the game is “whether you can build a good self-developed model, and whether it ranks on the leaderboard.” Only about 3 multimodal startups may remain.
- The embodied-AI sector has “a little bit of a bubble,” and today’s video models still cannot train robots. “Some companies go a month without anything happening, yet their valuations triple or rise several times.” Seedance merely “looks physically plausible”: real-world data may show a person grasping an object with 3 fingers, while generated video may always use 5; its millimeter-level point-cloud precision also fails to line up. 梅涛’s end-state definition of a world model is to learn from the world, reconstruct it, reshape it and edit it—not simply create a world.
Deep dive
1. The Setup: Images Are the Entry Point, Video the Process, and World Models the Endgame
- 梅涛 opened by showing all his cards: his team published the world’s first text-to-video paper in 2017. “Every author in the paper, except that professor from the University of Science and Technology of China, is at our company, and none of them has ever left.” He sees that stability as the first reason to compare the team with Anthropic.
- The show’s governing thesis was clear: “We have always believed that images are the entry point, video is the process, and the endgame may be a world model. It still comes down to who has the most data and where.”
- The host set the scene: Google unveiled Gemini Omni at I/O on May 20, OpenAI quietly cut back Sora to focus more on coding, Silicon Valley was simplifying while China’s multimodal battlefield remained unsettled. 梅涛 is the entrepreneur taking the fight down to the underlying architecture.
2. Building “Anthropic for Video”: A Stable Team and a Firm To B Commitment
- His positioning for the company is direct: “We want to be Anthropic in the video space—a global, high-tech, platform company.” The 2 things worth emulating are the team’s zero attrition and its “firm commitment to enterprise services.”
- His industry read is that Anthropic “wasn’t as prominent as OpenAI at the beginning,” but enterprise services gradually took off, with “repurchase rates and average contract values getting better and better.” His conclusion: “The ultimate destination for AI is to provide better services to individuals inside enterprises.”
3. The “Blade” in Video Is the Image
- The host’s question was precise: Anthropic chose coding as its blade; what is the blade in video? 梅涛’s answer: the image. “An image is a 2D signal, completely different from language… It is a visual capture of a moment in the physical world.” It encodes the spatial and logical relationships among people as well as physical laws; extend it along the time axis and you get video.
- The operating data supports the sequencing: customers spend “50%-60% of their time creating a single image” on the platform. The underlying architectures of images and video are also broadly similar, both rooted in Diffusion Transformer. “If you can’t make the image model good, you won’t be able to make the video model better later.” He cited Google’s image model and Veo 3 as examples.
4. A Full-Time Academician-Entrepreneur Is a Rare Profile
- 梅涛 says he has spent more than 20 years in industry—from Microsoft to JD.com and then 智象—and “has never taken a second paycheck.” In China, “there aren’t many teams with an academician title willing to enter the arena personally, go all in, and think about the company 7×24.”
5. From Elementary School to USTC to Microsoft Research Asia: The Seed of a Scientist
- His starting point was a mathematics competition in third grade, where he won his school’s only prize. The principal publicly said, “This student will become a scientist one day.” “I had no clear idea what a scientist was, but I thought I could work toward it.” The University of Science and Technology of China’s culture of “1 academician for every 1,000 students” reinforced that seed.
- He interned at Microsoft Research Asia during the SARS outbreak in 2003. 颜水成 and 余凯 sat in the neighboring cubicles, with 沈抖 next door. Of Microsoft’s 1,000 researchers at the time, 5 had won the Turing Award and more than 300 were Fellows. “That atmosphere made me feel that doing scientific research was an incredibly happy thing.” The alumni network was equally striking: 周博文, 寒武纪 co-founder 陈天石, and 余嘉辉, whose $100M transfer fee was “more expensive than Cristiano Ronaldo,” were all USTC senior schoolmates.
6. “Plum Blossoms Bloom After a Bitter Winter”: The Host Questions His Lack of Hardship
- 卫诗婕 challenged him: “I looked through your résumé and couldn’t find much evidence of a bitter winter.” His answer was methodology, not a hardship story: “When I do something, I put in 120% effort and make sure of its success from every angle.” At his most extreme, he visited investors and customers in 5 cities over 4 days, crossing China from north to south. Running half-marathons and lifting weights was about keeping himself “energized and in high spirits.”
- He also rejected the template life: “A person needs to put in 99% effort to become part of the 1%, but many people are satisfied at 50%-60%.” Breaking through oneself against human nature means that “even negative feedback is positive feedback to me.”
7. The Most Important Negative Feedback of His Career: A Slow Promotion
- At Microsoft, he became frustrated when his promotion lagged his peers. A mentor’s advice changed his trajectory: “Sometimes being slower isn’t necessarily bad. If you’re a little slower today, you may move very quickly once you clear this hurdle, because everything is accumulating.” 6 months later, his progress became “very smooth” and accelerated.
- The general lesson he took away: “Many people overlook their long-term progress and long-term value while magnifying their short-term value. A long game begins with repeated breakthroughs—or a refusal to stop breaking through—the boundaries.”
8. USTC as “Another Huanggang High School”: First Impressions of Hefei
- He only realized after arriving that his classmates were all working through Gimmyadovich and preparing for the TOEFL, and that “I needed to plan my life early and work harder than I did in high school to speak with these people on the same level.” When he filled out his application, he thought USTC was in Beijing; he learned it was in Hefei only after the exam.
- One vivid arrival scene stayed with him: his parents asked for directions at the West Campus gate, and the person they stopped happened to be the campus district chief, 黄继虎. “He said, ‘The fucking West Campus gate of USTC is right here.’” The campus culture was plain and austere. The sports field was usually empty; students were in classrooms, libraries and labs, unlike the more active students he later encountered at Beihang in Beijing.
9. 张宏江’s First Assignment: 200 Papers and Feet on the Desk
- On his first day at Microsoft, his mentor 张宏江 assigned him multi-camera sports-video analysis: read 200 papers and present a summary in a month. During the first presentation, 张宏江 had his “feet on the desk,” flipped from the last page backward without looking at the first, finished the 30-minute presentation in 5 minutes, and asked him to explain only 1 or 2 pages. “That’s when I realized my mentor was a founding figure in the field and examined problems in such detail.”
- The result was that “nothing came of it” for more than 6 months. CMU’s Takeo Kanade had used around 60 cameras to build the EyeVision system for the Rose Bowl Super Bowl, while China had no such system. “Even the cleverest woman can’t cook without rice.” The failure stayed with him for a long time.
10. The 2019 Reckoning: JD.com’s Automated Sports-Directing System
- In 2019, working with JD.com, Migu and others, he built an automated directing system using 16 cameras at a soccer stadium to mimic a human director. It was “possibly more advanced than the EyeVision system Takeo Kanade’s team built at the time.” The work won the ACM Transactions Best Paper of the Year award in 2020. “It was my way of answering for not completing that first internship project Microsoft gave me.”
- Looking back, the assignment was fundamentally about using AI to analyze and understand video—“the earliest starting point for the later development of multimodal, all-modal and world models”—and it set his direction in computer vision. 智象’s current HiClips product, which takes an interview and a prompt and automatically edits, generates and switches languages, is the modern version of the “virtual camera positions plus automated directing” concept.
11. Why Obscure Fields Matter: Without Video Compression, There Would Be No LLMs Today
- On the shift from a few hundred attendees at an annual CV/multimedia conference to “everything is generative AI,” he offered a defense worth preserving: “If people hadn’t worked on video compression back then, today’s large models couldn’t have succeeded. Multimodal systems process hundreds of billions of images and tens of millions of hours of video. Without good compression algorithms, you can’t store or load it.” Many fields laid the groundwork for today’s AI. “Other fields will come up too after some time. That’s how things develop.”
12. From R-to-D to “Model Is Product”
- The Microsoft conversion chain was laborious: turning a paper into technology required multiplying the effort by 10—1 researcher plus 10 engineers—and turning it into a standalone product required another 100 engineers. “Without more than half a year, it was impossible.” Today is different: “萨提亚 said the other day, ‘model is product.’ You keep innovating inside the model; the model itself is the product and needs no further translation.”
- That is why researchers have become more valuable. Research and development are now tightly coupled, and improvements come from a researcher’s domain knowledge, judgment and ability to iterate quickly.
13. Why He Is Committed to To B: A Mirror Image of 闫俊杰’s Choice
- The host pointed out the contrast: MiniMax’s 闫俊杰 disliked doing To B at SenseTime and chose To C when he started his company. 梅涛 went the other way: “To B enterprises are capable of lasting for generations.” It comes down to team DNA. His team spent years at To B companies and knows how to manage key accounts and channels and control costs. “The curve is a little slower, but the ceiling is high and it creates real ecosystem barriers.”
14. To B in the Generative Era Is Not SaaS-Era To B
- In response to the criticism that To B businesses become heavy, project-based and unglamorous, he pointed to product leverage: “Look at Sweep. Its valuation has risen dramatically, into the tens of billions of dollars, while the team hasn’t grown.” 智象 serves merchants on TikTok and China Mobile’s offline stores. It has made “a lot of To B attempts, but the company hasn’t added people; we’re still only 200-300 people,” and it “insists on not doing customized R&D.”
- There are 2 models: serve platform customers and use platforms to reach SMBs and consumers; then maintain long-term relationships with central government ministries and direct customers, using those direct relationships to build channels.
15. Pro C as a To B Multiplier
- The B/C boundary is blurring. “The ordinary C user on the mobile internet was a free user. Today’s AI users are all pro C: they pay, subscribe to tokens and consume tools to create work.” To C is “direct output of model capability.” Consumers are the first to sense whether a model is strong, and they also provide branding and feedback. Hence the strategy is To B-led, with tools for prosumers alongside it.
- There are 3 business lines: commercial agents for short-video marketing, creator tools for short-drama teams, and social-media ad design. The first 2 are core To B businesses; the third is not a particularly large big bet.
16. 3 Keywords for Starting a Company in China: Globalization, Going Abroad, and Software-Hardware Integration
- “If I were only building a domestic company, I might not have left.” The least-discussed of the 3 keywords is software-hardware integration. “If you don’t use China’s supply-chain advantage when starting a company in China, you’re competing with Silicon Valley using your weakness.” Silicon Valley’s strengths are capital, compute and talent density. Pure software is “not impossible, but extremely challenging.”
- The practical example is HiFans, an advertising display similar to a 3D holographic screen, which sold more than 10,000 units in Q1. The hard part is not the display but understanding offline businesses: mom-and-pop stores, bubble-tea shops, hair salons, hotels and internet cafés. The company built templates covering more than 80% of offline scenarios, allowing owners to create ads themselves and returning design capability to them.
17. AI Is Basketball, Not Soccer: The 4-Quarter Framework
- His stage framework: “AI is a basketball game with 4 quarters. The first is the competition among large models, which will converge quickly. The second is agents—general and vertical. The third may be devices, traffic entrances, and software-hardware integration. The final quarter may be another stage; I haven’t figured out what it is yet.” The industry is nearing the end of the first quarter and entering the second.
- The pace will be faster than the internet. Large models have a minor iteration every 3 months and a major iteration every 6 months. “The agent layer will definitely iterate much faster than the large-model layer,” while the device layer will be faster still. Software-hardware integration is his early positioning for the third quarter.
18. Sam and Dario: 2 Forms of Idealism, Neither Wrong
- On OpenAI’s split, he refused to take sides: “Dario wasn’t the company’s No. 1, so he didn’t know how the business was being run. Sam was thinking more strategically and sustainably, making the company walk on 2 legs. Dario saw a point of technical breakthrough from inside the research team and believed a single point could break through. It was right for him to leave then. Looking back, neither side was right or wrong.”
- On the popular claim that Dario is the real god and Sam a false one, his distinction is that Dario likes studying sociology and history and represents personal idealism, while Sam came from investing and is more pragmatic. “Over the long term, both of them may be right.”
19. Playing Sam in China: Financing Conditions Determine the Narrative
- His self-positioning is explicit: “For now, I’m closer to Sam Altman. If I were in Silicon Valley today, I’d definitely prefer Dario’s approach—following a conviction. But we’re in China, where financing isn’t as large as in the US. Chinese investors need you to keep placing orders along the way and proving your ability to commercialize. If I say it will take a long time to erupt, I may die before that day arrives.” His blunt conclusion: “To be honest, starting a company in China is much harder than starting one in Silicon Valley.”
- But he cautioned against freezing the lead Anthropic is building toward 1T in revenue: “In AI, don’t look at a single moment. Look at the span before and after, left and right. Being No. 1 today doesn’t mean being No. 1 tomorrow.” He agreed with the host’s summary: benchmark the business model against Anthropic, learn from Sam’s founder profile, “build a narrative long enough to attract more resources, and run a long race.”
20. Cognition Is the Ceiling: “You Can Only Make Money Within the Range of Your Understanding”
- Asked why Anthropic focused on coding, he attributed it to the technical cognition of the person at the top: “The thing I worry about most is cognition. Once something falls outside your range of understanding, you basically don’t look at it.”
- His hypothesis is that Sam needs a grand narrative that includes AGI, language, images, video, 3D and coding. Dario may believe coding has rigorous logic and a closed loop, that strong coding is necessary for a strong language model, and that more powerful agent capabilities can then be built on top.
21. 2017’s TGAN-C: A “Brave” Beginning
- He jokes that the world’s first text-to-video paper was poorly named: T for temporal, GAN for generative adversarial network and C for caption. By today’s standards, the result was primitive: around 80 frames at 48×48 pixels, producing a 3-4-second GIF—“a chef making soup, but the soup was so muddy you couldn’t tell what kind it was.” The model had only tens of millions of parameters.
- But the venue itself was a statement. The paper was submitted to ACM Multimedia’s “Brave New Ideas” session, where even weak or nonexistent results might be accepted. “Whether you question it or dismiss it, we took a step in that direction—from zero to one.”
22. 4 Stages of Video Generation: Seed, Sprout, Development and Explosion
- His periodization is as follows: the seed stage was the GAN era in 2017; the sprouting stage began with Stable Diffusion in 2022, when “image generation showed huge potential,” followed by the company’s founding in 2023; the development stage arrived with Sora in 2024, showing that video could be made stable enough for early adopters and professionals; the breakout stage began with Seedance in 2026, which “crossed the second gap,” bringing stable video creation from professionals to ordinary and semi-professional consumers.
- The prehistory goes back to roughly 2011-2015, when the field was called cross-modal rather than multimodal and focused on video-to-language description generation. Reversing the direction in 2017 was the real challenge: “Going from a high-dimensional signal to a low-dimensional one is many-to-one and easy. Going from low-dimensional to high-dimensional has no standard answer and no evaluation.”
23. From GAN to U-Net Diffusion to DiT: The Technical Essence of Each Leap
- The GANs of the seed stage used a CNN backbone in a duel between 2 networks, one generating and one judging real versus fake. Diffusion changed the physics: images are gradually noised into white noise, then a neural network simulates each denoising step. The backbone, however, remained a convolutional network.
- DiT’s leap was replacing convolution with Transformer in the backbone. “Its scale-up capability is stronger, its context processing is better, and it simulates the visual world better.” That is why OpenAI described Sora as “a simulator of the world.” Seedance is also based on DiT, but with “more refined data, better scale and engineering taken to the extreme.”
24. UiT’s Original Design: Removing Encoders and Decoders for Native “Childhood Sweethearts” Integration
- DiT’s problem is that text, images and video each pass through a separate encoder and are compressed into latent space before cross-attention. “Compression is lossy, and reconstruction during decoding creates hallucinations, making the system uncontrollable.” UiT, or Unified Transformer, has “neither an encoder nor a decoder. Every modality enters as tokens and interoperates directly in discrete space. The ceiling is higher, pixel-level and token-level control is better, and it can support many downstream tasks.” The team spent more than half a year training it.
- He gave the intuition a distinctly Chinese metaphor: childhood sweethearts. “When 2 people grow up together without reservation, they develop more trust and默契 as adults. We want signals to enter a native all-modal Transformer architecture at the native level, without any compression or encoding.” He was careful about the claim: this is an architectural innovation—GPT, DiT and UiT sit at the same level—but “we did not overturn the Transformer backbone.”
25. Why Bet on UiT: Latent-Space Loss and an Obsession with Pixel-Level Control
- The technical intuition is that different encoders operate at different abstraction levels: “Text is more abstract, images less so, and they don’t line up.” You can roughly fix a person’s eyes or nose, but “it is hard to fix them at the pixel level because the model learned in latent space.” Pixel-level methods had already shown promise at academic scale, but “you can’t scale them up at a university, so you can’t verify whether the idea is real.” The team took the idea into large-scale video generation.
- He was candid about the risk: “We thought this path was somewhat dangerous. DiT had already been validated, and stacking data and compute on top of DiT would produce predictable results. That’s how the giants operate—resource accumulation. As a startup, we had to take another path. We needed sufficient compute and data, but the architecture had to innovate, using one-tenth the people and resources to approach or even surpass the results.” The fulcrum is architectural innovation paired with efficiency and lower cost.
26. Open Source Means Being “Taken Apart, Crushed Up and Laid Bare in Front of Everyone”
- Asked whether big labs could copy the architecture quickly, he did not flinch: “Architectural or algorithmic innovation can give you an advantage of only about 6 months. But we’re not worried at all, because we can keep innovating. Look—we open-sourced it.” Competing on open-source leaderboards requires one-to-one model comparisons, and the community can quickly tell whether the architecture is genuine or just a pile of data and compute.
- Open source delivers 3 returns: branding that attracts talent—“the biggest value of a brand is that it brings in talent”; community feedback to improve the next model; and an 8B model that can be distilled again for edge deployment, driving token consumption. The strategic meaning is clear: “Even if they have 10x or 100x our resources, we will use open source, efficiency and cost to hold them off until we grow from a startup into a big company.”
27. Artificial Analysis as a Challenge Ladder and the First-Tier Philosophy
- Getting onto the Artificial Analysis leaderboard resembles a challenge ladder. After submitting a model API, the model competes one by one against anonymous models: beat No. 10 to become No. 9, continuing until it can go no further. The anonymity and head-to-head format make it fair. But he offered a correction: “No. 1 on a leaderboard doesn’t necessarily mean No. 1 in product capability, and No. 1 in product doesn’t necessarily mean No. 1 in commercialization.” Being No. 1 in open source forces closed-source models to improve “by a whole level,” or there is no reason for closed-source models to exist.
- His competitive posture is marathon-like: Chinese companies first need to remain in the first tier. “Whether you’re occasionally first, sometimes second, sometimes fifth—it doesn’t matter.” Then surpass rivals through product strength, and pull further ahead through commercialization.
28. The China-US LLM Gap Is Real: “Poverty Limits Your Imagination”
- On the view that Chinese LLMs remain at the edge of the second tier, he offered an unusually blunt admission: “Objectively speaking, domestic models still lag overseas models in large language models, especially closed-source models.” The reasons are a shorter history and compute. xAI, for example, has 200,000 GPUs; almost no Chinese company has more than 100,000, and Chinese big labs may train on only tens of thousands, versus hundreds of thousands or even 1 million at OpenAI and Anthropic.
- He preserved the painful metaphor: “The difference in compute is a bit like poverty limiting your imagination. Perhaps you could have gone to high school and taken the college entrance exam, but now you can only make it through middle school because you don’t have the money.”
29. Multimodal Is Different: The Arms Race Has Not Started
- Chinese companies perform well in multimodal AI: Kuaishou’s Kling and Seedance are both strong. “The compute requirement here isn’t as high, and Chinese companies have more advantages in data—the world’s best video platforms are basically in China.” More importantly, the architecture has not converged. “That means everyone can innovate on architecture within a certain compute constraint. Once the architecture is fixed and everyone stacks compute on the same path, the big labs will win.”
- The scale is still manageable: video models can be built with tens of thousands or thousands of GPUs; they do not require 1 million GPUs. “Multimodal AI is not in an arms race today. I think we’re still in the GPT-2 to GPT-3 phase, not GPT-4 or GPT-5. There are still many opportunities during the climb.”
30. The Short-Video Data Debate: The Basic Training Unit Is 5-15 Seconds
- He technically clarified a Peking University scholar’s claim that short video is useful for training video models. Virtually every model company uses large quantities of 15- and 30-second online videos, and long videos must be cut down during training. “The basic training unit is 5, 10 or 15 seconds, extending to 30 seconds or 1 minute. Beyond that, the compute runs out because you need to support a longer context.”
- The current state is clear-eyed. Since Seedance, models have stronger narrative sense, “know how to storyboard and switch within a shot,” look physically plausible and have higher first-pass success rates. But “even today’s best video models can’t meet every customer requirement. You still have to draw cards.” In 2023, images required 4 draws; later the team produced 2 images at once because “the draw probability improved.”
31. Why They Missed Transformer in 2017: “Different Fields Are Like Different Worlds”
- He admitted the miss: “We noticed it. At the time, we thought it was a different field—that was something from natural-language processing, and the fields hadn’t been connected yet.” The pattern had repeated at Microsoft: machine learning produces a breakthrough model, the first wave comes in NLP, then images, then video. “Text always runs ahead of images because text is a 1D signal, so it’s easy to prove at small scale whether it works.”
32. The GPT-3-Era Debate Over Large and Small Models: Small Models Reached Production
- He was at JD.com when the industry debated scaling laws after GPT-3. Alongside 陶大成, he proposed “large models plus small models.” Training a large model in the lab to top a leaderboard is fine, but in an industrial-inspection plant in Changzhou, accuracy must exceed 90%. “You find that the large model doesn’t work. It doesn’t have enough generality. A small model trained on a large model actually has higher fault tolerance and precision.” The model that reached deployment was the small one; the large model was simply the expert model hidden behind it.
33. Why Multimodal Has No Scaling Law: Visual Tokens Still Cannot Be Defined
- The core argument is straightforward. Each character in language “has accumulated a specific meaning over thousands of years. It is highly distilled knowledge and already structured.” English has 20,000 words and a countable set of combinations. But “visual signals still have no defined concept of what a visual token is.” His rough estimate is 100,000 tokens for an image and 1 million for a 15-second video.
- His strongest empirical evidence came from work around 2010 on visual vocabulary—image dictionaries formed by clustering patches, for which he won 2 Best Paper awards. He found that “the vocabulary for pure images had to reach a scale of more than 1.2 million words to produce good results.” Compare that with 5,000-6,000 common English words, then multiply again by 100 for video. “If you discretize tokens entirely the way language does, you need vastly more data than language, and 10x-100x the compute. It can’t be done.”
- The reverse comparison is telling: LLM researchers say general-purpose data may run out in 2027 or 2028. “In multimodal, we may not even be consuming data as fast as it is generated. The consumption rate of machine training may lag the new visual data humans create every day.”
34. Pixels Have Little Statistical Meaning; Patches Are a Makeshift Token
- He broke down from first principles why a pixel-level vocabulary does not work. RGB has 3 channels, each ranging from 0-255, so “256 cubed or even to the fourth power—calculate the order of magnitude of that vocabulary.” Pixel-to-pixel relationships are also weak and sometimes random. Tokens therefore use patches: 3×3, 5×5 or 7×7 for images, with a temporal dimension added as a cube for video. “Even that is not a fully accepted definition.”
- On a DeepMind scientist’s claim that multimodal is 2 years behind language, he preserved the hedge: “A scaling law will definitely appear, but whether it will be in 2 years I can’t judge right now. Many things in AI that seem true today will later prove to be wrong.”
35. Video Models Do Not Produce Intelligence; They Replicate—Until All Modalities Arrive
- His self-imposed limit is important: “I don’t think pure vision is a form of intelligence. Vision models and DiT are not responsible for producing intelligence. What we do is replication—reproducing the world we see in full, and combining it into new things. But does it contain intelligence and world knowledge hidden inside it like an LLM? I think that’s very difficult.”
- The turning point is an all-modal model. “Once all modalities emerge, language knowledge will enter the system. Visual signals plus language descriptions of physical laws could produce new intelligence. I don’t dare judge when that will happen, but I don’t think it will take particularly long.”
36. Self-Supervision vs. Strong Supervision: CV’s Fate Is Pair Training
- He agreed with 谢赛宁’s view that LLMs use self-supervision while CV uses strong supervision, and explained why. GPT performs next-token prediction and “has a correct answer.” Video generation requires pair training: “An image must be paired with the most accurate possible text description, and a 15-second video with a very, very long description.” That requires more human labor and potentially more compute, especially for embodied systems, which must record joint x, y and z positions, 6-7 degrees of freedom and force-control data.
- But “as much human input as intelligence” is only a phase. “Human labels are needed for the zero-to-one breakthrough. Multimodal actually does 2 things: generation and understanding. Once understanding and generation are unified, the model can label and generate itself, and you won’t need so many annotations. After training an initial model, it can label continuously and learn from its own labels.”
37. A World Model Is Not One Model: A Map of 3 Paths
- He first corrected the common misconception: “Many people think a world model is one model and that the entire world will have only one. That’s wrong.” The 3 paths are 杨立昆’s JEPA, which predicts world states and understands interaction between people and the world, with embodied intelligence as the underlying goal; 李飞飞’s 3D reconstruction approach, using Gaussian splatting and other 3D data, with relatively little human-world interaction; and the video-generation camp, which believes real video can teach more physical laws and more general intelligence—a data-driven belief.
- It is now widely accepted that an LLM alone is unlikely to become a world model. Without interaction, “it is more like a brain.” How the 3 paths converge remains unsettled. “From the end-state perspective they will definitely converge, but we haven’t seen that architecture yet.”
38. Video Models Must Be the Foundation of World Models: Scale and Data Make the Case
- The sharpest argument is model size. Today’s so-called world and embodied models are only tens of billions or 100-200B parameters. “If you want an embodied model with reasonably general capabilities, how could it be so small?” Video models have already reached 100-200B, and possibly 200-300B, trained on general first-person and third-person video. “Only after seeing that much video can you generalize more when you have sufficient generality.”
- The data comparison is even more lopsided. “A company like 智象 has 100 PB of data a year. It’s enormous.” Real-robot data may total only 1 million hours, “and possibly not even that.” Simulation data will not be much larger. “Without a solid video model as a foundation, it is difficult, in my view, to build an embodied model with good generality.”
39. Reading the Bitter Lesson Correctly: Not No Humans, but an Elegant System
- He corrected the host’s reading of Sutton: “He wasn’t saying there should be no human intervention. He was saying the algorithm should be simple and elegant enough to scale through data and compute. Human participation means designing the algorithm well and feeding it good initial data—like teaching a baby. At the start, you have to give the baby clean, labeled data so it can learn on the right path.”
- DiT and UiT are his examples of elegant architectures. “The more complicated something is, the harder it is to scale up. If humans load in too many tricks, each layer introduces decay and creates more error.”
40. Multimodal vs. All-Modal: Only Any-to-Any Is Native
- His definition is exact: “Multimodal is 1 plus 1—adding an extra modality makes it multimodal. All-modal means training all modalities together, natively, without any encoding.” The test of nativeness is any-to-any: signals enter and exit in their raw form, with any input point and output point selectable. One model can solve many tasks and absorb every application.
- Gemini Omni’s release “validated our judgment,” but he also questioned its substance, as discussed in Section 44: integration is not the same as native unification.
41. Differentiating from ByteDance: No Head-On Fight at the Foundation Layer
- The targets are specific: make the closed-source image model a global top-3 product and the open-source model the global No. 1, standing alongside Google’s Imagen and GPT as 1.92-meter competitors; make video a leader in short-video marketing—the people, products and places of commerce—and short-drama generation.
- Asked whether he feared Kling or ByteDance more, he dismissed the premise: “If you won’t do something because a big lab is competing, you shouldn’t start a company. Every era and every field has big labs. And today’s big labs were small labs once.” The approach is built around 3 curves: the company’s development curve, the AI industry’s curve and his own cognition curve. “Our cognition has to be steeper than the industry’s. We are more efficient, lower-cost and quicker to turn.”
42. Time and Resources Are the Biggest Variables in Fighting Big Labs
- He agreed with and amplified the host’s observation. Training has a cycle, and big labs can make decisions away from the front line because of architecture, hierarchy and competing interests. They may follow the wrong path until the model is released and performs poorly, at which point senior management notices. “Time cost includes the cost of doing the right thing and the cost of doing the wrong thing. Public-market investors make money from cognition, and so do we. Technical judgment has to be agile, and action has to be agile.”
43. Why He Started in 2023: 2 Restraints and 5 Years as an Apprentice
- His motivation was unusually direct: “I had been a technologist at a company now worth $1T and an executive at a company worth $100B. The next stop was to found a $100B company.” Many people urged him to leave in 2015-2016, but he judged the timing wrong. Xiaoice’s conversations “fell apart after roughly 7-8 rounds,” the AI Four Tigers were narrowly focused on facial recognition, and there was a gap between being an MSR researcher and commercializing a company.
- He spent 5 years at JD.com learning company strategy, management, product development and how to meet customers. By 2023, both conditions and mindset were ready. But he offered an honest hindsight: “If we had come out 6 months earlier, things might have been different. But there is no going back. We can only keep moving forward.”
44. Understanding and Generation Still Use 2 Architectures; Omni Is Still an Integration
- The technical map is clear. The understanding branch—GPT-4V and later GPT-4o—has converged with the LLM’s autoregressive architecture across text-to-text, image-to-text and video-to-text. The generation branch uses Diffusion Transformer. Why can’t they be combined? “The understanding architecture processes discrete tokens and cannot turn discrete into continuous. Autoregressively generating images and video cannot match diffusion on quality and detail.” Sora’s second significance is exactly this: “Sora is Sora and GPT is GPT. Generation and understanding still haven’t been unified, including today.”
- His judgment on Gemini Omni was blunt: “Omni is actually an integration. Clearly, Veo 3 has been integrated into the larger Omni model hub. Putting video and images into a single model is still difficult.” UiT is intended to solve the problem, but it has not been fully solved; for now the focus remains on using UiT for generation.
45. What to Watch at Google I/O: The Agent Philosophy and Proprietary Test Sets
- The next battleground for foundation-model companies has 2 parts: the intelligence frontier on public leaderboards and agent capability. “Is it a smarter architecture, or merely a fixed workflow built with industry know-how? People inside the field need to understand the reason. Only when you see the essence can you know which direction your own capabilities should take.”
- How can you tell whether a rival has made a foundational innovation when there is no technical report? “Every company has its own specific test sets. A public exam like the gaokao cannot measure a model’s boundaries. If its boundary capability hasn’t broken through and it has merely been nominated on a leaderboard, that doesn’t mean it has made a foundational breakthrough.”
46. His Talent Philosophy: The Humiliation of “Why Wasn’t It Me?” and the T-Shaped Path
- The people he recruits share a technical sense of mission. “When I see a great new paper that someone else made, my first reaction is: Why wasn’t it me? It’s humiliating, really. I ask whether the atmosphere was wrong, my ability insufficient, the organization poorly structured or the person lacked freedom.” What technical people care most about is not title, level or cash, but “whether they can establish their standing in the field, whether they have work they can show off. It’s the same with artists.”
- 2 pieces of advice shaped his own development. First, take the vertical stroke of the T as deep as possible—“until you can’t go any deeper”—then build horizontal capabilities in organization and strategy. Second, look to the future in China: the country is developing quickly and has strong momentum, and Chinese people are more likely to succeed in a Chinese environment.
47. Admitted Failures: Shelf E-Commerce and the Print Shop
- Asked for failures to make the success credible, he offered 2. First, he initially tried shelf-based e-commerce, “still thinking in the JD.com framework.” Friends invested several million yuan out of personal loyalty, but “there was no token consumption, so we eventually exited.” It took more than a year of trial and error to shift to TikTok-style content commerce, where demand for images and video was huge. Second, he tried to use an investor’s network covering roughly 60% of offline stores nationwide to upgrade print shops, then discovered that print shops were a sunset industry and slow to update their understanding.
- The meta-lesson was that “looking for a nail with a hammer is a necessary stage. Without going through it, you can’t break through your own cognition. Entrepreneurship has no standard answer. You navigate through fog, while many people create noise—and those people often haven’t done it themselves.” The company’s narrative has not changed since day one: models and applications as 2 engines.
48. The Cold of 2023 and the Heat of “Finding China’s Sora”: A Lesson in Consensus Culture
- Financing in 2023 was possible but “not especially smooth.” Top-tier institutions had all bet on LLMs. “You talked to investors about video and they talked about images; you talked about models and they talked about commercialization. You were never on the same channel.” Only after Sora arrived in 2024 did the second wave of investors return.
- His verdict on market sentiment: “Chinese investment only funds consensus. Everyone piles in together; being in a crowd feels safer. If it flips, everyone flips together.” His response is to think in terms of both floor and ceiling. “If big labs absorb all model capability, how does our company survive? A strategy needs an upper and a lower bound. If you believe your model will always stay in the first tier, you’re hallucinating. That’s why big labs are actually more anxious than startups.”
- On the claim that the first 5 pages of the fundraising deck were written in 戴若黎’s office, 梅涛 did not accept it wholesale, but confirmed 2 discussions there. He learned how to frame an entrepreneurial narrative: “Entrepreneurship is one part narrative, but I think it should be called grounding—you need to tell a narrative that reaches the sky and stands on the ground, while building the foundation underneath.”
49. Why He Abandoned Embodied AI for AIGC: The Accounting Lesson of JD.com’s Robotic Arms
- He had 2 projects in reserve when he started the company. The robot project was rejected after lessons from 2 robotic arms at JD Logistics, with 3 requirements: it could not disrupt the existing workflow; the robot had to run stably 7×24; and ROI had to be calculable within 12 months. “If you ask them to calculate 18 months, they won’t even talk to you.” The economics also had to be based on large-scale deployment: one robotic arm replacing one person could never work because someone still had to watch the robot. “Only when one person watches 10 robots can you make the numbers work.”
- He remains skeptical of today’s general embodied-AI narrative: “Which company’s robot can work 7×24 without interruption in a factory? Figure AI is only sorting; it doesn’t take a package out and put it into a specific bin. It’s still far from a real business scenario—but it is already remarkable.” The trigger for shifting to AIGC was a Midjourney image of a space opera winning an award. “There was a humanistic feeling in it,” alongside Midjourney’s 11-person team generating $200M in annual revenue and Canva’s valuation surpassing $40B. The appeal was to build a general-model foundation while closing commercial orders along the way.
50. Midjourney’s Ceiling and the Shock of Seedance 2.0
- His diagnosis of Midjourney is that its data capability and aesthetic services target “the artists and designers at the top of the pyramid,” but it lacks a large-model advantage. The middle tier—advertising, posters and real-person photos—has been taken by GPT and other large-model companies, forcing Midjourney to become increasingly specialized.
- His reaction to Seedance 2.0, which appeared during the Lunar New Year, came in 2 layers. First, surprise: he had assumed US peers would be ahead. Then less surprise: ByteDance has high talent density and enormous data scale, with overwhelming advantages in data quality and quantity—“when you apply enough force, strange results appear.” The impact on his own company was real but manageable. Sora burned $5.5B; Seedance’s investment is “certainly at the RMB10B level.” A friend warned him not to sit in the big lab’s tire tracks. “As a startup, I have only one magazine of bullets. Once it’s gone, it’s gone. Sitting in its tracks makes you cannon fodder. We have to stick to original innovation in model algorithms, compete at lower cost and win time.”
51. 2 Commercial Engines: A RMB10B Short-Drama Market and Cross-Border E-Commerce Commissions
- The scale is substantial: AI short dramas may spend “close to RMB10B a year on production alone,” excluding distribution revenue, with heavy token consumption. 智象 is selling the “shovel.” Its product 真赞—“every frame is excellent”—targets professional teams making premium short dramas, while large institutions with their own workflows use MaaS and buy APIs.
- Advertising and marketing are the longer-term market. TikTok has more than 1 million merchants. “An ordinary merchant needs several thousand short-video ads a month. A top customer needs hundreds of thousands a year, even millions or tens of millions.” The appeal is dual monetization: “On the one hand, you can sell the shovel; on the other, you can sell the result and take a commission on it. We have already made that model work.”
52. The Multimodal Opportunity: Expressiveness Emerges Before Intelligence
- He does not expect one company to dominate multimodal AI. “There will definitely be several companies. Some will be strong in film and television, some in people-product-place retail, and some in advertising and animation. It is unlikely that one company will be good at everything.” His maturity score is candid: LLMs are at GPT-5 or GPT-5.5, while multimodal is “at most a 3 or 3.5.” Intelligence has not emerged, but that does not stop commercialization from being close.
- The data supports the thesis. Global AIGC revenue was approximately $20B in 2024, “with around 60% coming from multimodal.” 智象’s platform consumes about 3T tokens a month. Image and video are token-intensive, and users are willing to pay—unlike LLMs, where 1 million tokens may cost only $1. The current state is “usable and functional; there is still a little distance to truly good.”
53. Harness and the 3-Layer Token Value Chain
- He defines a harness as “the OS underneath an agent—a form of intelligent orchestration that combines APIs, skills, basic model capabilities and industry know-how.” Compared with last year’s model orchestration, the upgrade is safety guardrails and cost control: “After OpenClaw came out, without this layer, the bill exploded.” Skills will become shareable and sellable products.
- The value chain has 3 tradable layers: underlying compute—GPUs, bare metal and AIDC; the middle layer of tokens, which is “getting thinner and becoming a commodity, like water and electricity today”; and the customer value created by agents. “Selling tokens is the big labs’ business—Alibaba Cloud, Google, OpenAI and Anthropic are selling water and electricity. Startups need to make the customer-value layer large enough to earn commissions and co-create value.”
54. The Compute Reality: Training at Thousands of GPUs, Inference as the Future Compute Hog
- Multimodal training is “close to 10,000 GPUs, but still at the 1,000-GPU level,” which remains a game startups can afford. But the cost structure will reverse: training is somewhat cheaper, while inference will require enormous capacity. Generating a 5-minute video currently “may cost RMB100-200 in inference.” The good news is that by the time multimodal reaches that stage, better inference chips may exist.
- Supply-side anxiety is real. “Global GPU capacity has been insufficient for the past few months. OpenAI has said it will buy as many GPUs as it can and needs 100x more.” If China’s compute reserves remain one-tenth of America’s, how can the country keep up? Domestic compute still lags in cost-performance and performance. “The next 2 years will be difficult. How to get through them intelligently is also a question.”
55. Tickets to Stay Afloat, a Death Wave and a Thinner Application Layer
- His definition of the ticket to stay in the game is unambiguous: “Every company building a large model has a ticket. The essence of the ticket is whether you can build a good self-developed model and whether it ranks on the leaderboard. Otherwise, you’re truly nobody.” A death wave may hit multimodal startups in 2024-2025. Companies without visibility or model capability will lose investor support; those that can commercialize only at the several-million-yuan level, rather than several hundred million, may also die. “The survivors will have basic model capability and commercialization”—including 智象, roughly 3 companies in total.
- He accepts the consensus that as models improve, the middle layer gets thinner. “After Claude Code came out, the value of US SaaS-style labor companies fell sharply, and SaaS companies are trembling. Should we only do the application layer, or cut into the model layer too? For me, we definitely have to cut into the model layer and build a commercialization moat in applications.” 智象’s self-assessment is that To C represents about one-third and overseas revenue about one-third. “Last year was already RMB100M; this year will definitely be several hundred million yuan.”
56. Second-Market Divergence: The Logic Behind Comparing MiniMax and Zhipu
- He offered 2 explanations for the divergence among the first listed foundation-model companies. One is a temporary difference in foundation-model capability—“the rankings may look different again in a few months.” The other is the comparison framework: “Zhipu has a good coding agent, so it is benchmarked against Anthropic. The second market sometimes operates through a powerful comparison logic.” Chinese companies should be valued in line with US companies after exchange-rate conversion, but the gap has not closed, leaving room for further upside.
- The second market no longer distinguishes To B from To C. The real dividing line is the business model: “Do you want to do private deployment of the model? That’s not a particularly good business. We never do private deployment or customized development.” Overseas, the company pushes SaaS tools; in China, it sells technology plus services—delivering creative assets and taking a commission on the GMV those assets generate.
- Taking state-owned capital was a necessity in 2023: “Dollar funds didn’t dare invest in Chinese AI companies. It wasn’t that we didn’t want the money.” He also mentioned 深创投’s investment in 智象 and its positive view of the multimodal space and the To B/To C combination.
57. The Embodied-Data Business: A Closed Loop with 诺亦腾
- The business exists because embodied AI lacks high-quality data, with demand estimated at “tens of millions to hundreds of millions of hours.” 诺亦腾 supplies real-world teleoperation data captured from humans at millimeter or sub-millimeter accuracy across 6-7 degrees of freedom. 智象 augments it: “He gives me one set of data and we turn it into 100 or 10,000 sets—changing the background and skin tone, turning gripper and glove data into real first-person human operating data, while keeping the accuracy at the same level.”
- Pricing must be tied to a closed loop: “We also need to verify that it helps the VLA or World Action Model products on the market. Otherwise, how could I price it?” He deliberately avoids building an end-to-end embodied model. “The data isn’t mine, and I don’t have the copyright to commercialize it. I only need a lightweight model to calculate the delta. I want to be the hidden master.” He accepted 戴若黎’s advice: “With data and model parameters at your scale, using a broadsword to swat a small mosquito is overkill.” But he left open the possibility of a pivot: once the data becomes large enough, the know-how could help train his own World Action Model. His view on embodied-AI hot money is cool: “Some companies go a month without anything happening, yet valuations triple or rise several times. There is a little bit of a bubble.”
58. 3-Finger vs. 5-Finger Grasping: Why Seedance Cannot Train Robots
- The host returned to an earlier question: can physically plausible video generation be used to train embodied systems? The answer remains no. “There are no physical laws inside it. It only looks physically plausible.” The first synthetic-data lesson was highly specific: finger precision was inadequate and object interaction was off. “Our synthetic videos always showed 5 fingers joined together to grasp, but in the collected data people grasped with 3 fingers. The other 2 fingers weren’t touching anything. Which one is wrong or right? There is no answer.”
- The threshold for embodied use is quantitative: verify millimeter accuracy on 3D point clouds across x, y, z and t, and verify that the degrees of freedom meet customer requirements. “Seedance cannot do this yet. Precision matters.” He also offered an industry black box: companies truly training large models do not show their cards. Ask about parameter counts and many will not say; they may have only tens of billions or a few billion parameters.
59. A Preview of the Future Architecture: One Backbone, 3 Video Models and a Switch
- He unusually disclosed part of the roadmap: “Underneath is a new architecture that I can’t discuss yet, derived from UiT. On top are several video models: one for consumer video generation, one for interactive generation—somewhat like 李飞飞’s approach, allowing users to adjust angle, direction and camera movement within a video—and one for high-precision, highly controllable generation for embodied AI. It works like a switch that can lean left, center or right.”
- The product targets are 3-dimensional: quality at the level of the first tier, including pixel-level control, physicality and logical reasoning; duration progressing from 5-15 seconds to minutes and then finished films—1-minute and 3-minute generation is already possible; and real-time performance. “Sora takes more than 10 minutes to generate 15 seconds. We also need several minutes, and users may not wait. Even a real-time preview would let them interact.”
60. An AGI-Scale World Model May Be 5-10 Years Away: Compute Must Rise 10x-100x
- He offered a clear order of magnitude for a unified world model. Training embodied x, y, z and t, 6 degrees of freedom, touch, temperature, video, images and language in one model could require 10x-100x today’s compute. “No one has that much data, no one has that architecture, and we don’t have that much compute. That model may come in 5 years or 10 years.”
- His underlying belief is a division of labor between 2 fields. LLMs embody accumulated human knowledge; with long context and extended reasoning, “they are beginning to feel a little intelligent,” and their ceiling has not been reached. Multimodal systems are not primarily for the brain or intelligence; they are for humans to interact with the world and simulate its physical states. They are difficult to train together because one is self-supervised and the other strongly supervised. The talent pools also remain separate: OpenAI’s Sora and ChatGPT teams are independent, as is Google’s Veo 3 team. “The upper layer has domain knowledge, while the middle-layer infrastructure is shared.”
61. Hefei’s Twin Stars and a Self-Driven Organizational Culture
- The host called 智象’s headquarters the neighbor of USTC and dubbed it one of Hefei’s “twin stars” alongside iFlytek, which is also a shareholder. 梅涛 said the 2 companies have substantial business synergies. Hefei Industry Investment has invested in ChangXin Memory Technologies and BOE, and hopes to replicate that model with 智象未来. “Hefei and we are running toward each other.”
- His senior schoolmate 刘庆峰 offered one piece of advice: “Scale the company quickly. Timing is critical in AI. Miss the timing and you can’t stay at the table.”
- He also admitted to a management mistake. The company used to make annual plans, then realized “this was complete nonsense.” AI startups cannot plan a year out, so it shifted to quarterly planning, with product and technical teams proposing their own OKRs. His founder philosophy is that a company is a barrel and cannot have an extremely short stave; individuals can maximize their strengths. “As the No. 1, you are lonely. My only anxieties are whether my cognition is sufficient and whether the team is complete. I interview people even on weekends.” His first step with candidates is to discourage them: “Entrepreneurship is a 9-in-10 failure proposition. I’m not a babysitter. You have to think it through yourself.”
62. Why Zero-to-One Innovation Is Hard in China: A Fundamental Question About Education
- He did not avoid the issue: “At present, zero-to-one innovation in China is somewhat difficult. Our education system likes teaching definite answers. Students get used to following the teacher’s instructions and become anxious when nothing happens after a period of work.” He contrasted that with the US PhD system as he observed it: a paper is revised more than a dozen times, every technical detail is discussed, and students graduate with a framework for their own research. Domestic professors supervise too many students, while students remain too far from industry; research methods are harder to internalize.
- His response was to bring Microsoft’s intern system into the company. Large numbers of PhD students read arXiv every day, publish 20-30 papers a year at top conferences and are sent to conferences. “Without that feel for the field, you may miss this generation and fall behind cognitively by several months.” UiT emerged from that process: pixel-level methods had appeared in academia and worked reasonably well, but could not be scaled at a university, so the team brought the idea into a massive system and built its own foundational innovation. Hiring is based on peer recognition: recruit the Top One in the computer science department, “not the person with the highest GPA, but the person the entire class agrees is the Top One.” The average R&D age is about 30.
63. Naming, Globalization and Localized Aesthetics
- HiDream and 智象 have deliberate origins. HiDream was selected after research with native speakers for its 2 syllables and open vowels. “智” means artificial intelligence; “象” means both imagination and modality—“one gives birth to 2, 2 to 3, and 3 to all things.” He also drew an uncommon distinction between globalization and going abroad: globalization means serving customers around the world through platforms such as VivaGo and Hike Leap, while going abroad means helping Chinese companies globalize while the customers remain in China.
- Localization of aesthetics belongs at the product layer, not the model layer. “The model layer will only be generalized. If the training data contains data from that country, the model has the capability. Differentiation happens at the product layer: Brazilian users see a different homepage from US users, with different effects, templates and UX. You can’t unify every value system and aesthetic.”
64. Faith in AI and the Commercialization Imperative
- He does not like discussing faith, but admits to having it: “Superintelligence will definitely be achieved. AGI is difficult to define, but I think superintelligence is already being realized.” His personal posture is to start from zero: “I climbed one mountain in academia. Now I’m coming down and climbing another. I’m willing to reset my identity. I don’t need additional prestige or wealth.”
- The business rule he learned at JD.com runs alongside that belief: “A company must create value.” No matter how difficult things become, technology must be pushed through and commercialization cannot be abandoned. There must be a closed loop: “If one day nobody invests in you, you have to generate your own cash. In China, it is difficult to make people believe through sentiment or faith in AI alone.”
65. The Ultimate Curiosity: “Machines Will Definitely Understand the World Better Than People”
- The force pulling him through decades of CV research is a concrete scene: driving at night in heavy rain or fog, relying on assisted driving. “Its eyes are better than yours. You can see only 5 or 10 meters; it can see 30. So I believe machines will see farther and understand more than humans. I want machines to help humanity understand the world, then help us create a better one.”
- Returning to the question of 李飞飞 and 杨立昆, he said: “I hope to learn from the world, reconstruct it and then surpass the world as it exists—not simply create a world, but mold it, build it, transform it, and make it possible to modify and edit according to our imagination.” Asked whether video was not the most fundamental medium because even blind people can understand the world, he replied: “That’s a corner case. There’s no point discussing it now. We still have to look at where the most data is. For now, the largest data volume is definitely in video.”
66. Epilogue: Hard Things About Hard Things, Murakami and Her
- The 2 books he mentioned were opposites in tone. The Hard Thing About Hard Things is an “operating manual”: “Even in Silicon Valley, with such good conditions, starting a company is extremely difficult.” 1Q84 is his mirror for AI: Murakami can spend 2 pages describing a person’s psychological changes in minute detail. “That is something AI cannot do today and is still hard to surpass.”
- The final detail was unexpectedly on point. During roughly 3 years as a campus projectionist at USTC, he screened more than 10 films a week and 6 showings, but “never watched an entire film, not even Titanic.” “I can only do one thing at a time, with 120% commitment.” The 2014 film Her showed him an early version of the future. “What happens in Her is exactly what is happening with today’s large language models—except they haven’t yet resonated with multimodal systems and built a bridge. AI may eventually evolve fast enough to catch up with science fiction writers’ imaginations.”