姜哲源 & 宁慕楠 on Embodied AI, Seedance 2.0 and the Robot-Heavy Gala
姜哲源 & 宁慕楠 on Embodied AI, Seedance 2.0 and the Robot-Heavy Gala
Summary
- The rumored price of entry to this year’s CCTV Spring Festival Gala robot sponsorship battle: Zhiyuan bid RMB60M, while Unitree reportedly bid more—RMB100M, according to an unverified rumor—and after Zhiyuan withdrew, it staged its own “Robot Wonder Night.” 宁慕楠 said Wonder Night offered longer exposure at lower cost; the show ran for about an hour and featured more than 200 robots, most serving as “audience members.” Zhiyuan said the brand exposure recovered several million yuan in costs. The host also noted that RMB100M could fund several joint-factory projects; the final number of companies appearing on the gala is inconsistently given as 4 or 5 in the transcript.
- 宁慕楠 sees 2026 as a critical year in the embodied-intelligence race, with the EV industry as the template. His view: “BYD has already appeared; whoever becomes the ‘Nio-Xpeng-Li Auto’ of embodied intelligence has to fight for this one shot in 2026.” 姜哲源 said the industry completed its “zero-to-one transition from demos to mass production” in 2025, and companies that have yet to reach production will “most likely pivot or lose their chance to get a seat at the table.” Leading companies should have at least RMB1B on their balance sheets, while mid-tier and smaller players should still have at least RMB100M-200M. The industry “won’t die that quickly, but it won’t have an easy time either.”
- The key reality from the front line of the gala: no company actually put its “brain” to work onstage. The host raised the point and 姜哲源 agreed that no company has truly brought the so-called brain onto the stage; most performances still rely on motion control. All dialogue was prerecorded and the robots were operated by humans because AI onstage cannot guarantee absolute safety. Song-and-dance acts can switch to backup footage if something goes wrong, while skits depend on actors to keep the timing and carry more pressure. Asked about contingencies, 姜 answered: “To be honest, we didn’t make one.”
- Seedance 2.0’s value for embodied-intelligence training is the episode’s biggest strategic fault line. 宁慕楠 believes a video-generation model with high physical consistency could serve as “some kind of world model” and become a “low-cost, continuously operating training-data factory”; he likened Seedance to a “DeepSeek moment.” 姜哲源 said Seedance renders convincingly but gets “the physical process very, very wrong,” meaning robot training on such data could introduce systematic errors. On the idea that multimodal generated data will greatly accelerate robot training, he said, “I definitely don’t think it will,” while acknowledging that the industry has no consensus.
- 宁慕楠 says video models have a different commercial structure from language models: they are difficult to open-source, and distillation struggles to preserve quality, potentially strengthening big-tech dominance. Training on private data makes open-source deployment expensive, while high-fidelity video still requires native large models and large servers. He therefore expects the technical monopoly of major internet companies to deepen “at least in 2026.” Diffusion models’ fixed output-frame length also gives short-video data a structural advantage.
- Embodied AI is easier to explain to investors than AI because its ROI can be calculated, but the industry also has a “fundamentally To LP” dynamic. 宁慕楠 agreed that some orders and financial results may exist to support fundraising and may involve related-party transactions, a shared feature of the AI and robotics eras. A robot costing about RMB400K could replace 2 or 3 workers earning roughly RMB70K a year, potentially reaching payback in 2 years. The To Lab model is not durable because school budgets and 2-3-year project cycles constrain repeat purchases; the incremental growth in To Lab orders “comes overwhelmingly from China.”
- 姜哲源’s differentiation strategy is a battle for scenarios, not a price war, with bets on K-12 and bionic humanoids. He believes “second place has no meaning; only becoming number one, with roughly 80% market share, may produce attractive compounding.” Unitree has established itself in motion control and Deep Robotics in inspection. Songyan’s RMB9,998 Xiaobumi is “the first high-performance bipedal humanoid robot at the RMB10K price point,” with the price cut designed to open a new customer segment rather than take share through industry-wide price cutting. Its bionic humanoid targets tour guiding and visitor navigation; 姜 believes earlier wheeled guide robots underperformed because they lacked attention and topicality, and says Songyan is globally ahead in the broader bionic-humanoid technology direction.
- Mass production has become the hard metric separating winners from the rest. Unitree’s own 2025 disclosure put its humanoid-robot production at more than 5,500 units; Songyan produced more than 1,000, reached monthly capacity of 500 units at the end of last year, and expects to exceed 1,000 units per month this year. The difficult parts are body and process design, plus the supply-chain triangle of cycle time, cost and quality; inventory of core components is used to absorb delivery volatility. 姜哲源 nevertheless acknowledged that Unitree “is still the world’s number-one robot company in motion control,” and that its gala-rehearsal performance was “far beyond expectations.”
Deep dive
1. The RMB100M Ticket: Zhiyuan Leaves the CCTV Gala and Stages Its Own Wonder Night
- Host 卫诗婕 disclosed the industry’s nickname for the bidding battle—the “RMB100M ticket”: Zhiyuan bid RMB60M, while Unitree reportedly bid more, rumored to be RMB100M, although the host said she had not verified the figure with Unitree. Zhiyuan withdrew and declared, “Going on the CCTV Spring Festival Gala is not as good as putting on my own gala.”
- 宁慕楠 said the gala offered a short, uncontrollable exposure window at a very high cost, with no guarantee that the outcome would match expectations. The host added that RMB100M could instead fund several joint-factory projects or support a core technical team of several dozen people.
- Wonder Night’s cost structure was straightforward: roughly an hour-long program, supposedly featuring more than 200 robots, although “most were so-called audience members,” with perhaps only one-tenth actually appearing onstage. Rehearsals relied more on robot performance companies than on Zhiyuan’s own engineers. Zhiyuan said brand exposure had already recovered several million yuan in costs and brought the event to breakeven, and that it hoped to turn it into its own annual “robot gala.”
- 宁慕楠’s assessment was blunt: the event had limited reach and appeared mainly in press releases, and was indeed “not as attractive as a gala with real people.” The host compared it with the striking effect of dozens of robots performing yangko and handkerchief routines at last year’s CCTV gala. Ning said that if robot-made programming eventually develops its own ratings, a robot gala “could become a format worth building out in the future.”
2. From Motion-Control Showmanship to “Brain + Cerebellum”: VLA Is This Year’s Draw
- Ning’s framework is that multiple companies are now pitching a “brain plus cerebellum” architecture. The brain is VLA—Vision-Language-Action—giving robots GPT-, Gemini- or DeepSeek-like capabilities in understanding and dialogue. The cerebellum consists of specialized small models for 3D detection, robotic-arm control and similar tasks. VLA “counts as one of this year’s highlights.”
- The two practical endpoints are clear. In factories, simple interaction between robots and workers could “greatly reduce danger or improve efficiency.” In the home, the desired machine is not a quiet robot that folds clothes, but “one that can communicate with you and provide emotional value.”
- There is an ironic twist: if the CCTV gala showcases interaction capabilities, viewers may look back at Wonder Night and realize Zhiyuan had already demonstrated them. In that sense, “the so-called RMB100M gala may still have had its value.”
3. From a Two-Horse Race to a Four-Way Contest: 2026 Is About the “Nio-Xpeng-Li Auto” Moment
- The host said companies had been competing since May or June. The contest began as a two-horse race between Unitree and Zhiyuan before becoming a four-way battle. She listed the final companies appearing on the gala as 银河通用, Unitree, 魔法原子, 松延动力 and 追觅.
- Ning’s timing logic is that the robotics sector has spent money for a year and investors now want results, even if that means spending more for visibility. His more important analogy was: “BYD has already appeared; whoever becomes the ‘Nio-Xpeng-Li Auto’ of embodied intelligence has to fight for this one shot in 2026,” because the most critical period from EV breakout to commercialization lasted “only about 2 years.”
4. The Three-Leader Road Map: Unitree on Hardware, 银河通用 on Simulation, Zhiyuan on Joint Factories
- 银河通用, represented by 王鹤, is taking the simulation-data route. Real-world collection requires a body costing at least RMB200K, an operator to program and record the task, and then data labeling—a heavy process. The company uses an improved Isaac Sim physics-simulation model to generate realistic videos of robotic arms grasping objects, folding quilts and performing similar tasks in virtual environments, reducing reliance on real-world collection.
- Zhiyuan’s approach is joint factories through government and enterprise partnerships. Equipment enters the factory, operators wear data-collection devices, and contracts stipulate that 80%-90% of the data will be collected back. The factory receives an additional data input beyond its normal output. Zhiyuan is “single-mindedly focused on industrial deployment” and recently also worked on a smart port.
- Unitree focuses on the most visually impressive hardware and is globally ahead in body design. The host said its ecosystem is another major moat: the technology behind last year’s Year of the Snake gala handkerchief routine came from cooperation with Nvidia and top laboratories at CMU, “an ecosystem that is very difficult for other domestic embodied-intelligence companies to catch up with.” Ning said the host’s assessment of the value of collecting real-world data was “absolutely correct.”
5. Newcomers 松延动力 and 魔法原子: The Technology Is There, but the Commercial Label Is Not
- Ning’s view of 松延动力: it has links to Tsinghua and Peking University and appears to have support from the Beijing municipal government. Most of its engineers are graduates of Tsinghua or Peking University born in the 1990s, and its 小顽童 robot once placed second in a robot marathon. The technology is “nothing to worry about,” but the company “has not yet emerged with a particularly distinctive commercialization path.” Like the first 3 companies, it needs to bind itself to application scenarios and promote them.
- The host suggested that 魔法原子 may have the lowest valuation among the companies, citing a valuation in the tens of billions of yuan and relatively limited product exposure. Ning did not directly verify that assessment, emphasizing instead that such companies need to tell a credible commercial story and build the right ecosystem. On its relationship with 追觅, he said “the market seems to have already settled on that view,” but offered no definitive confirmation.
- Ning’s underlying judgment is that China’s engineer supply means “the gaps in technical shortcomings between companies will not be especially large.” The industry is highly homogeneous. Whoever develops a differentiated product, brand or technical strength can create a dividing line; “once it has its own distinctive characteristics,” interested investors or users will find it on their own.
6. The Pragmatic Logic Behind Zhiyuan’s Withdrawal: Every Day in a Factory Is a Spring Festival Gala Test
- Ning’s reversal of perspective is that the kind of high-pressure test represented by the gala exists “at every moment” in a factory. VLA models—whether Gemini, GPT or Claude—inevitably hallucinate. A hallucination in dialogue may merely send someone to the wrong paper; in a factory, it could cause property damage or even loss of life, creating enormous reputational damage. Ning compared the risk with the public-relations fallout from accidents or fires involving new-energy vehicles.
- His conclusion: “A company has to understand who its product and service are for.” Unitree wants to showcase the most advanced motion control, so appearing at the gala was the right choice. Zhiyuan wants to enter factories and concentrate its money on industrial partnerships, which Ning also considers “the correct choice.”
7. Retreat and Consolidation: Investors Will Run the EV Playbook Again
- On the claim that the first robotics-sector retreat already occurred in 2025, Ning qualified it: saying the market has fully retreated may be premature. Entering now as a small company “is not a wise choice,” while discussing an exit mechanism for mid-sized and large brands is “still premature”—although the window is “not as long as people imagine.” He cited the actual exits of Neta, WM Motor, HiPhi and other EV brands.
- Investors with EV experience will follow the same path to identify the projects “most likely to replicate the pattern.” Such projects need a business model and target users; most important, the target user must know, “If I want to achieve my goal, I have to buy your product.”
- Entrepreneurs continue to pile in for 2 reasons. The theoretical market is close to unlimited: 1 robot per household or 10 per factory would each represent an enormous market. And China has strong engineering talent and supply chains, making it relatively easy to build a specific product and then use it to raise capital. That is Ning’s description of “the reality of 2025”; whether it remains true in 2026 and 2027 is unknown.
8. “Fundamentally To LP”: The AI-Era Bubble and Embodied AI’s Calculable ROI
- The host put forward the industry line that there is no real To B or To C business model and that the sector is “fundamentally To LP,” with many orders and reported revenues tied to financing projects and related-party transactions. Ning said, “I believe that’s exactly right,” calling it a shared feature of the AI and robotics eras. Musk has said that without the current AI frenzy, the US “might be on the verge of bankruptcy”; OpenAI, Oracle and Nvidia finance one another and inflate one another’s bubbles. “That’s the environment, so there’s no need to be overly harsh.”
- Ning believes China’s embodied-intelligence bubble is somewhat healthier than America’s AI bubble because the ROI is calculable. Assuming annual worker pay of RMB70K and a robot price of RMB400K, payback would take 6 or 7 years. “But if that robot can replace 2 or 3 positions, it may pay for itself in 2 years.” The economics are visible and tangible, with clearer commercial scenarios than AI.
- AI’s payment model remains unresolved. Big companies make their apps free to lock in users; Ning cited Tencent recruiting 姚振宇 with RMB200M and said it would be difficult to recoup that investment through Yuanbao alone. Video models cannot be open-sourced, while pricing them too high drives users away and pricing them too low makes it difficult to recover the cost of data and servers.
9. Of the Four Business Models, To Lab Has a Ceiling and Performance Rentals Look Like DJI
- Ning is explicitly bearish on the durability of To Lab. School budgets are not unlimited: a RMB1M project may spend nearly all its money after buying 1 or 2 robots, and the project may not renew within its 2- or 3-year cycle. After Trump took office and froze substantial education funding, US projects became even harder to apply for; “the overwhelming majority of incremental To Lab orders come from China.” China itself cannot escape the cycle problem, and even outside top-tier 985 universities, it is difficult to secure RMB several-million projects every year.
- Within To B, Ning believes rental and performance models “could be a fairly good model,” using DJI as a reference. The company sells drone swarms overseas and provides emotional value to Arab billionaires by drawing their portraits.
10. Ning’s 3 Predictions for 2026: Human-View Data, Cross-Scenario ROI and Autonomous RL Evolution
- Data collection will become more standardized. Operators will wear cameras on their heads and wristbands or wrist-mounted cameras, capturing “action logic” rather than motor parameters. Motor parameters become difficult to reuse when the body changes, while human-view data can be mapped by a model and generalized across robots, allowing the same data to be used after later iterations. Zhiyuan has already promoted this approach through its joint factories.
- Cross-scenario capability is the key to calculating ROI. In the past, a palletizing robot could only palletize and a transport robot could only transport, making factory returns impossible to calculate. With VLA, a single robot that can “replace workers across an entire production line like a real person” could finally generate meaningful purchase demand.
- The AlphaGo-style question is whether robots can improve through reinforcement learning. Robots began entering factories in 2025; can they become stronger through repeated learning, the way AlphaGo improved through play? “A human who does one job for 20 years will certainly become an expert in that field. Can a robot also become an expert by doing one job for 20 years?” If robots can search for better solutions autonomously after deployment, their rate of improvement could accelerate.
11. Seedance 2.0, the Bull Case: Physically Consistent Video Generation Is a Kind of World Model
- Ning’s central proposition is that a model capable of generating videos consistent with real-world physics “can be considered some kind of world model.” He sees Seedance 2.0’s breakthrough in physical consistency: even large movements show no obvious body penetration or geometry errors. His examples included the small belly beneath Trump’s white shirt bouncing up and down as he rode a horse, and highly dynamic fight scenes.
- The attraction for embodied AI is obvious. Replace an Ultraman battle with a robotic arm grasping objects or folding clothes; as long as physical consistency holds, the model becomes a “low-cost, continuously operating training-data factory.” It eliminates the cost of real-world collection and labeling and can keep generating data without leaving the office.
- Ning described it as “a bit like the DeepSeek moment before Chinese New Year in 2025,” when a tool for a specific user group became “a toy that can do everything.” He said Seedance supports native 2K resolution, that its physical consistency can be directly compared with Veo 3, and that at the current stage it has “already surpassed overseas technology.”
12. Video Models Are Hard to Open-Source and Distill: 2026 Could Deepen Big-Tech Monopoly
- The key difference from DeepSeek is that video-model data and compute costs are extremely high, while training relies on private data. “Open-sourcing is inherently a money-losing proposition.” Video models also cannot preserve performance through distillation as easily as language models; post-distillation quality may be too poor for users, leaving native large-parameter models that must run on large servers.
- The implication is that major internet companies will still lead world-model and video-model upgrades “at least in 2026,” further “deepening the technical monopoly advantage of the major companies.” Providers will build closed environments and charge users for access. Sora once claimed it could replace 90% of work in the CG and advertising industries but has not fully done so. Ning remains “fairly optimistic” that Seedance can replace much of the low- and mid-budget CG and advertising-generation work.
- Private-domain data explains the domestic landscape. Language-model training relies heavily on public data, while multimodal models depend more on each company’s private video data. ByteDance’s Seedance and Kuaishou’s Kling lead partly because their parent platforms are short-video or video-social companies.
13. Why Short-Video Companies Are Winning: A Technical Lesson in Fixed Frame Length
- Ning’s analogy: a language model predicts the next token and can continue generating indefinitely. A Transformer-plus-Diffusion video model is more like, “Here are 32 sheets of paper; fill all 32 sheets, and the video is done.” Output length is fixed, which happens to favor short video. Training on long videos creates temporal discontinuities and attribute mismatches.
- Longer generation works by chaining clips together: use the final frame of a 30-second clip as the starting point for the next round to build 60 seconds. The cost is uncontrolled drift in small details—“by the end, the person may not even look the same, or their personality and movements may have shifted substantially.”
14. The Hottest vs. Most Important Direction in 2026: Video World Models and Embodied Scaling Up
- The hottest direction is a video-model-based world model. Within the constraints of physical laws, it offers “what you see is what you get,” enabling short dramas, CG and game concepts, while generating first-person, third-person and full-factory data for industrial use. A robot’s “first moment on the job would already be that of a veteran expert with several hundred days of experience.” Factory commissioning could fall from 6-8 months to “6-8 weeks, or even 6-8 days.”
- The most important question is whether embodied AI can achieve Scaling Up. Today’s VLA “is fundamentally living off the old foundation of large language models,” fine-tuned from a capable base model with extensive manual labeling. If a large volume of first-person or robotic-arm data could train “the robot’s brain model” from scratch—like intelligence emerging in a child after tens of thousands or even millions of hours of video—then moving from Factory A to Factory B would take only “half a day” to become familiar with the environment. “That would be a new era for embodied intelligence.”
15. 姜哲源 on the Gala Front Line: The Full Product Family, “Grandma’s Favorite” and a Hidden Title
- 松延动力 did not send a single machine but its full product family. Two 小顽童 N2 units performed running, front flips and cartwheels. The 1.4-meter E1 walked anthropomorphically, performed magic and extended its neck higher than 王天放 in a running joke. The RMB9,998 Xiaobumi received a relatively large role, alongside the bionic-humanoid product line.
- The program was a skit titled “Grandma’s Favorite,” made with 蔡明 and 喜人奇妙夜 champion 王天放. Barring surprises, it was scheduled as the third program at around 8:30 p.m.: a grandson who has not returned home for a long time is repeatedly roasted by robot 蔡老师, competes for attention with several small robots, and ends by dancing with them.
- The bionic 蔡明 was the show’s biggest gag: a 1:1 replica with very high degrees of freedom and biomimetic facial expressions. Celebrities walking through the backstage area would greet it with “Hello, Ms. Cai,” receive no response, and only then realize that it was a robotic version of 蔡老师.
- The naming puzzle reflects 4 different public titles. 银河通用 called itself the “embodied large-model robot”; Unitree used “Spring Festival Gala robot partner”; 魔法原子 used “strategic partner”; and 松延动力 used “Spring Festival Gala humanoid-robot partner.” 松延动力 was originally set to be the “exclusive bionic-robot partner,” but that wording would have spoiled the program, so “humanoid” was used instead of “bionic.”
16. No Real Brain onstage: Prerecorded Dialogue, Human Remote Control and No Contingency Plan
- The host first argued that “no company currently intends to truly bring the so-called brain onto the stage.” 姜哲源 agreed and added that “everyone is still mainly focused on motion control.” These AI systems could not be allowed to run in real time because he “couldn’t guarantee absolute safety.” Dialogue was therefore prerecorded and the robots were operated by humans. When pressed on why, he called it “a reality, not something I can explain the reason for.” Fully autonomous running is not expected until the marathon demonstration in April this year.
- The risk structure creates the pressure. Song-and-dance acts run to tight timing and can switch to backup footage if something goes wrong. Skits depend on the actors to maintain the rhythm and are harder: “That is the point putting the most pressure on us.” Asked about a contingency plan, 姜 answered, “To be honest, we didn’t make one.” The practical safeguard was to stay up late before every appearance and conduct a full inspection of each machine.
- 姜 described distinct roles for the 4 companies. Unitree handled martial-arts showmanship, 松延动力 took the dialogue-heavy skit, and 魔法原子 was reportedly assigned a song-and-dance routine. 银河通用 worked with 沈腾 on a micro-short drama showing some brain and upper-limb manipulation capabilities, although 姜 acknowledged that such abilities receive less public exposure and attention than dancing and bipedal movement.
17. The Commercial Logic of Bionic Humanoids: Tour Guides, Spectacle Value and the Similarity Debate
- The answer to the killer scenario is tourism and visitor guidance. A shortage of tour guides “is not about wages”: the job requires memorizing tens of thousands of Chinese characters and can involve fines for mistakes, which puts people off. A bionic robot “can remember the entire script very easily.” The current form is a wheeled base with a human head, already capable of elevator control, autonomous planning and autonomous recharging.
- The host asked why the robot had to look human and whether 大白 or a robot dog could do the same job. 姜 argued that purely wheeled guide robots from OrionStar and 小笨 had not performed particularly well because they “do not inherently have attention or topicality.” People visit scenic sites and museums to see spectacles; a guide with a human face creates “a natural attraction.”
- Asked whether excessive human likeness could be frightening, 姜 acknowledged that when he first saw Hanson Robotics’ Sophia, he felt that even reaching out to touch it might be “a little disrespectful.” But he believes the shock is “more positive than negative.” He distinguished this from household robots designed to do chores, where adding a bionic face “may not be 100% necessary.” Film and television represent another use case; 姜 said Disney makes the world’s best bionic faces.
18. Linkage Drive vs. Cable Drive: The Confidence Behind “World-Leading”
- 姜 described the cable-drive system used by Sophia as “possibly a previous-generation technology.” Motors pull wires to move the facial skin, but repeated use can cause irreversible plastic deformation—the cable may gradually stretch, making bionic control less precise. Disney’s cartoon bionic characters, such as Chief Bogo from Zootopia, require relatively few degrees of freedom. On that basis, he said 松延动力 is globally ahead in the overall bionic-humanoid technology direction, potentially by a generation over most competitors.
- 松延动力’s proprietary approach combines linkage-driven facial skin with a bionic humanoid face. Linkages cannot bend and route freely like cables, making the design substantially harder. The facial skin creates a second trade-off: if it is too thick, small motors cannot pull it; if it is too soft, it collapses onto the underlying skeleton. 松延动力 uses platinum-cured silicone and has shifted from imported material to domestic material.
19. No One Had Bought the Gala Set Pieces in More Than 30 Years; Unitree Was “Far Beyond Expectations”
- During the first rehearsal in CCTV’s Studio 1, 蔡明 and 王天放 “performed extremely well,” while 松延动力’s robots, in plain terms, had not yet learned their positions. At 2 a.m. that night, 姜 took the skit’s stage set pieces back to the company and reproduced the entire gala stage at 1:1 scale for further rehearsals. He later heard that in more than 30 years of Spring Festival Gala history, nobody had bought and taken away the main-stage set pieces; 松延动力 “may have been the first to do it in the true sense.” 蔡明, 王天放 and the director’s team also came repeatedly to rehearse, and the directors called 松延动力 “the most hardworking” of the 4 companies.
- On competition, 姜 said outsiders assume the companies are locked in an intense battle, but in reality they are “on different tracks.” “We don’t pay much attention to what others are doing.” Helping the actors and production team deliver a strong program was “the most important—and the only important—thing.”
- His praise for Unitree was particularly strong: “far beyond expectations.” Its front flips, vaulting over a pommel horse, wall-running and drunken boxing were all impressive, and “Unitree is still the world’s number-one robot company in motion control.” 姜 bought a Unitree robot dog while he was a student, added 王兴兴 on WeChat and visited Hangzhou several times to seek his advice.
20. Xiaobumi at RMB9,998: A New Market, Not an Arms Race
- Xiaobumi stands more than 90 centimeters tall and is “the first high-performance bipedal humanoid robot at the RMB10K price point”—姜 said that before it, only robot dogs had reached that price level. It targets K-12 programming education, storytelling and companionship, with the goal of becoming “the first humanoid robot in the true sense to sell more than 10,000 units.”
- 姜’s pricing philosophy was the interview’s most complete piece of reasoning. A price cut can either take share in an existing scenario—“which has no meaning; someone can always cut prices further, and eventually nobody has any margin”—or open an entirely new customer segment. At RMB9,998, the product “opens up the entire C-end customer base at once, especially families with children and relatively high net worth.” The aim is massive incremental demand rather than cannibalizing existing demand: “The price cut was not for an arms race.”
- The difference from Unitree is scenario-based. Humanoid-robot applications remain fragmented, without a single massive killer use case, and the overlap between the 2 companies is limited. Unitree serves research, education and performance markets at higher average ticket sizes; 松延动力 targets K-12 and companionship, choosing scenarios and price bands with less direct competition. 姜 also stressed that its products are distinct from pure-AI companionship products and do not target exactly the same users.
21. The Value of Mass Production: Dual-Layer Design and the Supply-Chain Triangle
- The numbers come first. Unitree’s own 2025 disclosure “should be” that it produced more than 5,500 humanoid robots. 松延动力 produced more than 1,000, reached monthly capacity of 500 units at the end of last year and expects to exceed 1,000 units per month this year. 姜 acknowledged that Unitree has “a competitive advantage from body design through the entire production process, as well as barriers built through years of accumulated experience.”
- What makes mass production difficult? 姜 asked whether the problem could simply be solved by hiring more workers. The first challenge is design: beyond the body drawings, process engineers must design tooling, inspection stations and the entire production line. The second is the supply-chain triangle of cycle time, cost and quality. The biggest pain point is delivery time, addressed by building inventories of core components when cash allows, absorbing volatility and shipping directly from finished-goods inventory.
- The supply-chain ecosystem changed materially in 2025. Most robotics suppliers overlap with consumer electronics and automotive supply chains; joints are the one major category specific to robotics. Suppliers now offer different service models, including custom development, retail distribution and KA key-account support. 姜 sees the sector becoming more mature, which is “definitely a good thing” for 松延动力.
22. RL Has No Scaling Law, and a World Model Cannot Cook Tomato-and-Egg Stir-Fry: The Seedance Bear Case
- Responding to 王兴兴’s view that a Scaling Law for robot RL has yet to emerge, 姜—who trained in Deep RL—went further: “Deep RL by itself does not have what is called a Scaling Law.” RL data comes from a Replay Buffer and simulation Rollout rather than external information; “talking about RL plus Scaling Law is a little strange.” He believes Scaling Law may become a meaningful concept when large volumes of human-action data are used to train full-body manipulation RL Agents, as well as upper-body VLA systems. 西湖大学 and Unitree are doing well in these areas, while 松延动力 has preliminary results expected after the holiday. The training paradigm shifted from manual tuning throughout 2024 and the first half of 2025 to human-action-data-driven training in Q2-Q4 2025—and “still now.”
- His fundamental challenge to the world-model camp is that a world model equals a renderer plus a physics engine. Consider “cooking tomato-and-egg stir-fry”: can the model learn the process by which egg liquid turns solid at different temperatures? “If it can’t learn that, the information it provides is biased.”
- After the host relayed Ning’s Seedance view, 姜 said directly, “I don’t particularly agree.” Seedance renders convincingly, but “the physical process is very, very inaccurate.” Even Isaac Gym, designed specifically for the physical world, is not highly precise, and the Sim-to-Real Gap remains large. The day before, he had seen a Seedance 2.0-generated toy train on Xiaohongshu that flew and bounced around randomly and frequently disappeared—behavior inconsistent with physical laws.
- His conclusion was unequivocal. Will multimodal generated data greatly accelerate robot training? “I definitely don’t think so. Of course, if I’m wrong, I’ll pay for my mistake—the payment is the opportunity cost of the opportunities I missed.” Both sides acknowledged that this remains a route-level dispute without industry consensus.
23. Consolidation Is Nearing a Verdict and the Scenario War Has Begun: Second Place Has No Meaning
- 姜’s verdict on 2025 is that, beyond technology, the industry’s biggest change was “the zero-to-one transition from demos to mass production.” Companies that have not achieved scaled production and sales will “most likely consider a pivot or lose their chance to get a seat at the table.” The constraint is the combination of funding and time: companies arriving too late simply do not have enough time to build a deliverable product.
- Capital-market conditions support the host’s reporting. 姜 believes leading embodied-intelligence companies “should all have at least RMB1B” on their balance sheets. Mid-tier and smaller companies also have several hundred million yuan in cash, or at least RMB100M-200M. “This industry won’t die that quickly, but to be honest, it won’t have an easy time either.” The remaining companies will either disappear, pivot or differentiate in a niche, unless a “new technological black swan” changes the field again.
- The price war has not started; the “scenario war” has. “Second place has no meaning. Only by becoming number one, with roughly 80% market share, may a company achieve attractive compounding.” 姜 sees Unitree and Deep Robotics as scenario leaders, with Deep Robotics established in inspection. 松延动力 is targeting the K-12 education ecosystem. He still describes the company as “on the edge of the table,” but expects its position to stabilize somewhat after the CCTV gala.
- 松延动力 spent 2025 fixing organizational weaknesses. In 2024 or early 2025, it was “all technical people,” without even human-resources or sales functions. The company spent the year building an organization and making itself look like a real company rather than a pure laboratory. 姜 believes it has now reached the point where it needs to submit its first small report card, with mass production and commercialization as the tests.
Verification Notes
- The host listed 5 companies appearing on the gala in the first part of the transcript, while the host and 姜哲源 repeatedly discussed 4 companies later. This edition does not use that discrepancy to determine whether 追觅 belonged to the same group of gala participants.