157: Embodied AI Quarterly 26Q1 with 陈哲
Summary
Unitree’s core assets are not its Spring Festival Gala exposure, but production consistency and de facto status as the research-platform standard. Twenty-one Unitree humanoids simultaneously executed leaps, flips, and giant loops, reflecting the concentrated deployment in 2025 of reinforcement learning, imitation learning, and sim-to-real, on top of its experience designing, producing, and mass-producing motors at million-unit scale. Humanoid revenue rose from less than 2% in 2023 to 27% in 2024 and 51% in 9M25, with humanoid gross margin around 63%; 陈哲 cautions that the high margin primarily shows customers are still concentrated in the price-insensitive research market, while the fact that G1 still has no true challenger roughly 20 months after launch suggests that “a hardware company’s advantage may last 12, even 24, months.”
This quarter’s biggest shift in 陈哲’s thinking is that he is no longer convinced bipedalism is simply an expensive, unnecessary form factor. A humanoid occupies only about 40×60 cm of standing footprint, yet can work from floor level to roughly 2 m; Atlas can reach about 2.3 m at maximum. Boston Dynamics’ wheeled Stretch weighs roughly 800 kg to 1 ton to move 20-kg boxes, while a humanoid carrying the same load might weigh only 70-80 kg. A wheeled design with four-wheel steering, lift, and dynamic pick-and-place would also require a large number of motors and a heavy chassis; “complexity and cost are not necessarily lower than a humanoid’s.”
Q1’s tennis, kung fu, and whole-body-control demos show the capability frontier moving outward, but they do not yet prove general-purpose autonomy. Galbot’s humanoid had to handle tennis balls traveling at up to roughly 100 km/h, with real-time perception, decision-making, and whole-body closed-loop control; it likely relied on external cameras, motion-capture data, and non-edge compute, yet still led 陈哲 to focus on “whether it can be done is the critical first step; optimizing how comes second.” By contrast, Unitree’s Spring Festival Gala performance was pre-choreographed, and Figure’s videos were carefully trained and edited; the genuinely new trend is that locomotion and manipulation are beginning to be controlled by a unified model rather than two isolated systems.
Dexterous hands are emerging as the next major frontier after quadruped locomotion and control, VLA, and humanoid locomotion—and may produce a G1-like standard research platform. Sharper demonstrated autonomous windmill assembly at CES; its design includes 22 DoF, tactile input, and hierarchical System 0/1/2 control. A hand priced around $50K is expensive, but already fits the pricing logic of subsidized research. The counterexample is Optimus Gen 3’s high-DoF tendon-driven hand: motors in the forearms would control 40-plus tendons, making consistency, creep, maintenance, and 10,000-unit production difficult; the sharpest objection is, “Once you’ve used motors, it’s not muscle—what first principle does that come from?”
NVIDIA’s WAM is not a simple relabeling of VLA; it pushes robot policy from static mapping toward visual prediction of the future. DreamDojo is like a video-based simulator, while DreamZero generates policies directly from the environment and task; both seek to use video generation to establish a causal chain between actions, environmental changes, and outcomes, rather than behavior-cloning mappings between images, text, and teleoperation trajectories. DreamZero currently runs at about 7 Hz, and physical consistency remains constrained by the video foundation model; tactile sensing is also inherently absent. World models and language models are therefore more likely to complement each other than one to replace the other.
The data bottleneck has only found a possible way through; it is far from solved. EgoScale broadens diversity with more than 20,000 hours of first-person human manipulation data; 陈哲 describes a data pyramid that descends from high-quality, expensive real-robot teleoperation to UMI/Dex UMI, first-person video, and internet video, with simulation either in the middle or used for augmentation. Sunday expanded the two-finger UMI into a three-finger capture device matched to the target embodiment. Chinese companies are chasing “1 million hours,” but usable data still needs cleaning, labeling, and extensive training experiments; he has no conclusion on whether “1 million hours is enough.”
The dividing line in the US-China embodied-AI race is not model quality alone, but how hardware, foundation models, and compute are re-coupled. 陈哲 believes China already leads in robot bodies, dexterous hands, supply chain, and mass production, while the US retains an edge in top talent, compute, and data, with Pi as an example; if research becomes increasingly tied to complex embodiments, China’s advantage could widen. The risk is that world models depend heavily on video foundation models: many robot world models currently use Alibaba’s Wan 2.1/2.2, and even NVIDIA may not need to train from scratch, but the next-generation recipe will require massive GPU experimentation, so an early-stage company could run a bill of hundreds of thousands of yuan simply by copying several PB of data.
The IPO wave will strengthen embodied AI’s national-strategy narrative while exposing the huge gap between valuation and commercialization. 陈哲 estimates that China already has more than 20 embodied or humanoid companies valued above RMB10B or $1.5B, versus only 4 or 5 companies at that level during the most euphoric phase for foundation models in 2023-2024; even Unitree, the sector leader, generates only a little over $200M in revenue. He sees Unitree as the benchmark for “real user value, real revenue, and highly efficient operations,” but expects the industry to go through absorption and shakeout. Asked whether the humanoid is the optimal form for a general-purpose robot, his final answer remains: “I don’t have an answer today.”
Deep dive
1. Q1’s five key advances now span the full stack, from isolated moves to a complete technology stack
陈哲 emphasizes that this Top 5 is not an academic ranking, but a set of “strong subjective picks” made by a longtime robotics investor weighing technical, product, and commercial potential. It happens to cover embodiment and motion control, end-effector manipulation, foundation models, systems engineering, and mass-production design.
The first item is more than 20 Unitree humanoids performing kung fu at the Spring Festival Gala, representing the highest level of Chinese robot hardware and motion control. The second is Sharper’s autonomous windmill assembly at CES, which gave the industry a view of the SOTA in high-DOF dexterous hands.
On the model side, the protagonists are NVIDIA’s DreamZero and DreamDojo. They build on ByteDance’s GR-2 at the end of 2024, which used internet video, and attempt to understand environments and generate robot actions through video generation rather than text generation.
The other two are Galbot’s real-time whole-body closed-loop tennis demo and the new-generation production electric Atlas. The former expands the range of what seems possible; the latter uses extreme modularity and 360-degree rotation to redefine how a humanoid can surpass the constraints of human anatomy.
2. Unitree’s Spring Festival Gala threshold was not a single difficult move, but consistency across 21 machines
Unitree brought to the stage work that matured rapidly only in 2025: first using motion capture and imitation learning to map human martial arts onto a robot, then using reinforcement learning in simulation to turn rough trajectories into a more stable, robust policy, and finally transferring it to the physical robot.
曼琪 noticed that details such as the shuffling footwork, leaps, and continuous flips looked “more human.” 陈哲, however, focused on consistency: 21 production-grade machines performed complex actions simultaneously, each facing different environmental disturbances. The test was hardware quality and software stability, not tuning one prototype to the limit.
Compared with early Boston Dynamics parkour, the underlying paradigm has shifted from expert hand-tuned MPC and other classical controls to simulation, reinforcement learning, and end-to-end policies. Older videos were often carefully edited from large amounts of failed footage, and a single machine’s continuous success rate cannot be directly compared with 21 machines executing in sync.
The lidar mounted on the robot’s head also sends another signal: recent dance and parkour demos are beginning to add visual feedback, mapping, and localization. They may no longer be fully open-loop playback of joint trajectories, but they remain far from complex autonomous decision-making.
3. A successful stage performance did not bridge the gap between manipulation and autonomous decision-making
陈哲 draws a clear boundary around the Spring Festival Gala excitement: all the dancing and kung fu were still preprogrammed fixed sequences. Under a sufficiently large unexpected disturbance, the robot still cannot understand the situation, replan, and resume the task autonomously as a human would.
The demonstration focused mainly on whole-body movement and lower-body control. It did not prove upper-body manipulation, complex contact, or long-horizon task understanding. Many of the hardest problems in embodied research lie in manipulation and in closing the loop around object state and contact changes.
曼琪 compared the industry’s reactions across two Spring Festival Galas. Last year, many practitioners thought the handkerchief twirl relied on a mechanical retrieval device and had limited technical content; this year, recognition rose markedly. The change came not only from flashier moves, but also from the industry’s ability to recognize the reinforcement learning behind them and the reliability of multi-robot execution.
4. The prospectus shows humanoids are now Unitree’s main line, but 63% gross margin still reflects research-market economics
Unitree’s humanoid revenue share rose from less than 2% in 2023 to about 27% in 2024 and about 51% in 9M25. The product transition from H1 to G1 has already made humanoids the company’s core future growth engine.
Under the latest figures, humanoid gross margin is about 63%. 曼琪 views that as an unusually high margin for an integrated hardware-software product. 陈哲’s correction is that research equipment often carries 70%-80% gross margin, so 63% is not implausible: the research market is small, volumes are low, orders are fragmented, and customers are far less price-sensitive than industrial mass-market buyers.
This market is not fully elastic. A $50K hand or a cheaper research robot will not automatically make the customer base 2x or 3x larger. Customer numbers and demand are relatively clear; pricing is driven mainly by relative product competitiveness and the scarcity of substitutes.
Both agree that early robotics is highly supply-driven: without reliable mass-produced products, latent demand cannot turn into revenue. Once reliable supply appears, research, performance, and industrial-pilot demand are released. But 陈哲 estimates that the current humanoid research market itself may be only about RMB1B.
5. G1’s 1.3-meter form factor defined the research robot; later entrants cannot catch up by copying specs
Unitree’s first-generation H1 was about 1.8 m tall, essentially “standing a large quadruped dog upright.” 王兴兴 once said development took “three and a half people”—3 engineers plus half of himself—because he did not yet believe in humanoids or treat them as a core strategy.
G1 was designed from the outset for the education and research market. Its height fell to about 1.3 m, sharply reducing weight and easing the demands on motor power density, batteries, and motion control. For research, the questions that can be studied at 1.3 m and 1.8 m differ little, so downsizing did not sacrifice the core use case.
Starting at about RMB99K, G1 combined scenario, size, and cost unusually well from day one. Even if a later entrant produces one or two prototypes with higher specifications, it is difficult to find enough differentiation in a small research market that investors do not particularly like.
The harder barrier is mass-production history. Before humanoids, Unitree had already sold tens of thousands of quadrupeds, each with at least 12 motors—equivalent to years of designing, producing, and validating motors at million-unit scale. 陈哲’s comparison is that a foundation-model lead may last only 3-6 months, while “a hardware company’s advantage may last 12, even 24, months.”
6. Wang Xingxing’s restraint kept Unitree alive until the wave arrived; next it may defend its standard through follower-mode AI
When 陈哲 met 王兴兴 in Hangzhou in 2019, the company had around 10 people, sold 10-20 robot dogs a year, and generated about RMB10M in revenue, yet was already profitable. “Why is Unitree a profitable company? Because it had to be profitable; if it wasn’t, I wouldn’t have seen it in 2019.”
That is also why he puts Unitree alongside DJI. The founders did not enter after first spotting a huge market; they pursued a direction they had long been passionate about and believed in. The program recalls that when 王兴兴 was raising money in 2017, investors asked, “What can this actually do?” At the time, he genuinely could not say.
Unitree’s R&D expense in 9M25 was more than RMB90M, potentially far below many peers’ spending. Yet its prospectus plans to raise RMB4B, with roughly RMB2B earmarked for the “brain.” 陈哲 believes its VLA and world-model work remains largely follower-mode, and that the core management team lacks an AI partner who can independently lead the function.
His view is not that Unitree must become the strongest model company. It can follow for the long term: as long as the latest global research continues to use G1, both closed- and open-source models will feed back into the hardware ecosystem. Humanoid shipments exceeded 5,500 in 2025, with a 2026 target of 10K-20K; performance and rental demand after the Spring Festival Gala could make the actual market larger, leaving capacity investment as the main constraint.
7. Galbot’s tennis demo mattered because it first showed a full-size humanoid entering a high-speed closed loop
Tennis balls can travel at up to roughly 100 km/h. The robot has to judge the ball’s trajectory, move its body, swing the racket, and make contact in an extremely short window. Even a dedicated wheeled practice robot would find this difficult, let alone a humanoid with more degrees of freedom and a longer control chain.
During the Chinese New Year holiday, Galbot rented a large tennis court, set up motion-capture equipment, collected data, and repeatedly trained with reinforcement learning. 陈哲 believes its academic novelty may not be the strongest; it is more a piece of complex systems engineering. But it clearly demonstrates execution.
The system most likely relies on external cameras for high-frame-rate perception and does not use edge compute exclusively, so it cannot be transferred directly into a product that works anywhere. It is still autonomous execution on a real robot, not CG. Andrew Cote questioned on X whether the video had been generated by AI, which itself shows how far the capability exceeded prior expectations.
The computer-science lesson 陈哲 retains is: “As long as something can be done, humans will find ways to optimize it.” Proving that real-time closed-loop control works first, then optimizing dependencies such as external vision and compute, is a more rational sequence than debating commercialization at the outset.
8. Locomotion and manipulation are converging from two systems into a unified whole-body model
In 2024-2025, the industry generally treated lower-body locomotion and upper-body manipulation as separate problems. As robot stability and data scale improve, Figure, AgiBot, and NVIDIA Sonic are beginning to let a single model coordinate movement, posture, and manipulation.
Figure’s January and March demonstrations showed highly natural whole-body motion, while Galbot’s tennis demo necessarily coordinated the legs, torso, and arms. The videos were specially trained, choreographed, and edited, but the results were still autonomous execution on real robots; careful production alone is not grounds to dismiss them as fake.
陈哲 believes this unified paradigm has only just begun to emerge and could compound rapidly over the next 12 months. He says he cannot even imagine what humanoid performances might appear at the 2027 Spring Festival Gala. Once the paradigm is validated, more companies can iterate around a common framework instead of building the upper- and lower-body control stacks from scratch.
9. The cost disadvantage of bipedalism is being recalculated; wheeled bases are not as simple as they look
陈哲 used to think warehouses and factories needed two hands, not two legs. This quarter’s shift came from a supply-side judgment: once a capability can be delivered reliably, use cases that were previously invisible can emerge quickly. Form-factor value cannot be defined solely by what robots cannot do today.
In a structured environment, a humanoid needs only about 40×60 cm of standing footprint, yet its whole-body degrees of freedom let it pick objects from the floor and reach roughly 2 m upward. It can dynamically adjust its legs, waist, and center of mass to handle 10-20 kg boxes, with particularly strong space efficiency in narrow aisles.
Boston Dynamics’ Stretch uses a large AGV base and one arm to move roughly 20-25 kg boxes, with a total weight of about 800 kg to 1 ton. A full-size humanoid might weigh only 70-80 kg. A wheeled design with four independently steered wheels may need 8 motors just for drive and steering, before adding a lift mechanism, heavy base, large battery, and higher-power actuators. Its complexity and cost are not necessarily lower.
10. New Atlas’s “superhuman” architecture puts mass production, maintenance, and motion freedom into one design
The new electric Atlas does not custom-design a complex actuator for every joint. It uses roughly two types of standardized rotary motors as its main actuators, trading performance redundancy for structural simplicity—a modular shift similar to collaborative arms versus traditional industrial robots.
Atlas is not constrained by the human skeleton: its head and torso can rotate 360 degrees. A person turning from north to south needs 3 or 4 steps; Atlas can simply rotate at the waist. Its left and right legs can be interchangeable, and the hands should follow the same modular logic. The robot does not need to preserve the human concepts of front, left, and right.
The design lowers production assembly, spare-parts, and field-maintenance costs at the same time. 陈哲 connects it to the US shortage of skilled technicians: when assembly and maintenance resources are expensive, modularity, replaceability, and performance redundancy are needed to make robots easier to build, rather than pursuing biomimetic form at all costs.
11. The US humanoid field is highly concentrated, while Optimus Gen 3’s timeline remains hostage to hardware
The most closely watched US companies remain Tesla’s Optimus and Figure, which have the highest funding and valuations. Other active players include Boston Dynamics, Norway’s 1X, and Apptronik’s Apollo in Texas. Pi, Sunday, and Generalist are more focused on models and data than on building full-stack humanoid platforms.
Optimus Gen 3 is reportedly design-frozen, but supply-chain information cited on the program suggests the launch plan has slipped from 2026 Q1 or March to April, and may move again to late June. Production originally planned for October could also slip into 2027. 陈哲 summarizes Musk’s pattern: “Elon is always right, but his timing is always wrong.”
Even after the target was scaled back from more aggressive historical guidance to at least 10,000 units, it remains extremely challenging as April approaches. The delay is not simply a software problem: mass production, reliability, and maintainability of a high-DOF dexterous hand may be Gen 3’s key hardware bottleneck.
12. Optimus’s tendon-driven hand pushes biomimicry to the limit—and magnifies reliability problems
Optimus uses tendon drive: many motors sit in the forearm, with tendons transmitting force to the fingers. Sharper’s direct-drive design instead places motors near the finger joints. These are not minor structural differences, but two different philosophies—biomimicry first versus engineering modularity first.
A high-DOF tendon-driven hand may route more than 40 tendons through the wrist, palm, and other confined spaces. Any slack, creep, wear, or assembly error can break consistency. Once something fails, rewiring, recalibration, and repair are close to a complex surgical procedure.
陈哲 relays a friend’s objection: “Once you’ve used motors, it’s not muscle—what first principle does that come from?” Human muscle and tissue can recover, and have high energy and torque density. Wear in tendons, motors, and gears is irreversible; reproducing muscle structure with non-muscle materials does not reproduce its performance.
He does not therefore declare Musk wrong. Pure vision and end-to-end autonomous driving once faced strong internal and industry opposition, but were eventually realized through engineering. The testable questions are whether Optimus changes its hand architecture and whether the current route can support reliable 10K-scale production in 2026.
13. Figure’s presentation style warrants caution, but its talent and technical results can no longer be ignored
曼琪 summarizes the complex view many Chinese founders hold of Figure: it keeps producing good results, yet appears “very over-the-top.” 陈哲 jokes that the company deserves an Oscar for best visual effects because its videos are often carefully framed, trained, and edited.
The style is connected to founder Brett’s background. He previously founded the eVTOL company Archer, which, if memory serves, went public in 2021 and saw him leave not long afterward; he had also sold startups before. The market therefore questions whether he is committed to long-term engineering or is better at chasing hot themes, raising money, and exiting quickly.
When Figure was founded in 2023, it did not understand the full robotics industry, and a number of early members from Boston Dynamics and elsewhere have since left. But 陈哲 acknowledges that over the past 1-2 years the company has attracted many excellent software and hardware engineers, and that its new body and whole-body-control results have “real substance.”
On the model side, Helix uses a layered structure connecting low-frequency planning, mid-frequency actions, and high-frequency control. Internal feedback and public videos both suggest leading capability. Given how few US companies besides Optimus are simultaneously building a full-size complex body and its models, Figure represents the top tier of the local market.
14. The US lacks not humanoid ambition, but the industrial density to support multiple hardware startups
陈哲 believes the weakness of US supply chains and manufacturing explains why China can “cobble together a humanoid prototype for a few million yuan,” while the US struggles to support multiple startups and must concentrate resources on a few leaders. Figure’s large funding rounds reflect both the company’s needs and policy expectations around manufacturing reshoring.
Google worked with Apptronik on Apollo in its early days, but the feedback 陈哲 heard was that hardware precision, reliability, and consistency were inadequate. Researchers spent much of their time “making the machine usable” rather than studying models. The shift to Boston Dynamics was a choice of a more stable embodiment for basic research.
Boston Dynamics’ controlling shareholder, Hyundai Motor, can provide automotive-grade production, testing, and supply-chain resources; some assembly and initial testing could also move into Hyundai factories. Using allies such as Japan and South Korea to supplement US manufacturing is one route, but 陈哲 estimates that producing the same product in the US system could cost 2x-3x as much.
曼琪 cites Meta’s smart glasses against the claim that internet companies have no hardware DNA. 陈哲 separates sales from commercial quality: Meta Ray-Ban sells for about $299-$399, while cumulative losses and subsidies at Reality Labs have reached more than $10B and potentially several tens of billions. Selling well does not prove that a sustainable complex-hardware capability has been built.
15. Embodied foundation models are becoming physical agents with memory, tools, and reflex layers
In addition to launching Pi-0.6, Pi’s Q1 work continued on long-sequence memory, cross-embodiment, dynamic-environment adaptation, and online reinforcement learning on real robots. One approach is similar to OpenClaw: use text to record the current state over the long term, then continually reflect and update.
This led 曼琪 to argue that an embodied model is no longer just a VLA. It is closer to an agent in the physical world, requiring task orchestration, skills, tools, memory, and state management around the model. The systems engineering of virtual agents is entering robotics with much stricter real-time and safety requirements.
Sharper’s System 2/1/0 framework makes the idea concrete. System 2 is low-frequency, high-dimensional, and language-oriented macro planning; System 1 takes in images, robot state, and task information, then outputs relatively coarse arm and joint trajectories; the new System 0 uses tactile input and higher-level intent for the highest-frequency fingertip closed loop.
System 0 does not obtain tactile information before contact. It intervenes after the hand has touched the object but the position or force is still inaccurate. 陈哲 compares it with the brain, cerebellum, and peripheral nerves: a person can perform many actions with their eyes covered, but block fingertip sensation and even picking up a familiar object steadily becomes difficult.
16. Dexterous hands may become the next standard research platform, but the winner must first restrain its commercialization impulse
陈哲 maps the migration of the research frontier: 4-5 years ago it was quadruped locomotion and control; 1-2 years ago it was VLA, two-finger grippers, and UMI; humanoid locomotion was then partially solved. Since 2025, almost every researcher he has spoken with has treated dexterous hands as the next frontier.
World models are equally fashionable, but are more likely to be led by big tech because video-generation models consume far more compute than text models. Dexterous hands, tactile fusion, and end-effectors are areas foundation-model vendors cannot do on startups’ behalf, leaving startups a relatively independent technical arena.
Overseas users have commonly relied over the past year on 心动纪元’s roughly 12-DoF hand. Sharper offers a roughly 22-DoF hand excluding the wrist, close to the human hand, for about $50K, and is placing it in labs through subsidies. Its autonomous windmill assembly at CES gave researchers a direct view of what high-DOF dexterous hands can do.
陈哲 believes the market will compete to build the “G1 of dexterous hands”: reliable, sufficiently high-DoF, affordable, and supported by a complete sensor suite and development environment. Sharper’s founding team emphasizes AI and general-purpose robots as the end goal, but he warns that chasing full-system autonomy too early and refusing open research collaboration could cost it the chance to become the global default platform.
17. Mini Cheetah and Unitree’s history show platform opportunities belong to long-term builders, not earliest fundraisers
MIT’s Mini Cheetah open-source work introduced quasi-direct-drive QDD motors and a complete control stack, allowing many non-specialist teams to build quadrupeds quickly. It helped spawn Xiaomi’s “铁蛋,” XPeng’s PengXing, and a large number of startups.
One new company raised about $100M in its first round at a valuation of roughly $500M, while Unitree attracted little attention. The company that ultimately went furthest was 王兴兴’s, by continuing to sell research equipment and refine its motors and supply chain. 陈哲 sees this “staying grounded and persevering” as the real precedent for a dexterous-hand platform.
XPeng once tried to consumerize a robot horse, prompting 曼琪 to joke, “The memories of the robot horse’s death are coming back to attack me.” 陈哲 believes the failure was not that quadrupeds had no value, but that the company sought a consumer closed loop too early while the technology still needed R&D.
18. China already leads in complex hardware; the US still controls top-tier model resources
Asked whether China and the US are starting embodied AI from the same line, 陈哲 gives an answer more granular than a slogan: Chinese companies already lead in complex hardware such as robot bodies and dexterous hands, while the US retains a clear advantage in top talent, compute, and data on the “brain” side, with Pi as an example.
Unlike large language models, the value of an embodied model cannot be separated from the body. If future research becomes increasingly tied to high-DOF hands, full-body humanoids, and large-scale real-robot data, hardware supply, manufacturing speed, and research-platform density will feed back into algorithms. China’s advantage could continue to widen.
This means the competition is not a static division in which the US handles software and China handles manufacturing. 陈哲 believes China could move from catching up to original research and leadership, but if world models remain dependent on expensive video foundation models, compute and training recipes could again widen the advantage of US big tech.
19. World models and VLA represent two kinds of intelligence; the relationship is complementary rather than substitutive
陈哲’s simplest definition of a world model is: based on the current observation, predict what will happen next. Autonomous-driving simulation, Sora’s positioning of itself as a world model, and researchers’ work on explicit physical rules are all different implementations of this broad concept.
VLA evolved from multimodal language models: first give the base model an understanding of text and images, then add an action head and attach joint trajectories from a specific scenario to the image-and-text state. It is therefore good at semantics and instructions, but retains the character of “behavior cloning with a description.”
A world model uses video generation as its backbone and tries to learn how the environment will change after an action. It is closer to visual intelligence, time, and physical interaction; VLA is closer to language reasoning, task descriptions, and planning. 陈哲’s view is that human intelligence needs both language and vision, so the two routes will not simply exclude each other.
王兴兴 also expressed a preference for the theoretical ceiling of world models in a GTC presentation. 陈哲 uses a sharp analogy: “Even if you’re illiterate and don’t understand any written language, you can still survive quite well in this world.” Reading text without vision, by contrast, makes it difficult to truly understand many natural phenomena.
20. DreamDojo “fills in the world”; DreamZero turns that imagined world into action
DreamDojo can be understood as a video-based simulator. Given the current frame and conditions, it predicts and renders what the world will look like next, creating an interactive future for policy training.
DreamZero starts from the task and current environment, then outputs the robot’s policy and actions through video generation. It is not simply “observe one image and output one joint sequence”; internally, it constructs a causal process linking action, environmental change, and outcome.
NVIDIA calls this framework the World Action Model, or WAM. The robot first “imagines” how a given action will change the world, then selects a trajectory consistent with the task and physical laws, moving beyond the VLA control logic centered on text and action cloning.
ByteDance’s GR-2 at the end of 2024 was an important precursor: it introduced internet-scale video into pretraining and directly generated action outputs. By Q1 2026, NVIDIA and others had organized video generation, simulation, and policy generation into a clearer technical route.
21. World models’ upside is time and causality; VLA’s weakness shows up under out-of-distribution feedback
曼琪 summarizes the difference by saying that “world models have a sense of time.” VLA’s image and text tokens mainly describe a static state, with action sequences learned through a mapping; a world model explicitly predicts the next moment and the more distant future, making it easier to represent whether an action has succeeded.
She cites RoboChallenge’s Table 30 QR-code task. The robot does not merely raise the scanner to the right position; it must see the change on the screen and determine that “the scan has succeeded.” A VLA can bolt on a detector, but the model itself may not understand the causal chain linking action, feedback, and task completion.
Behavior-cloning-style mapping also explains fragile generalization: training uses a blue cup and testing uses a red one; training places the cup on the left and testing moves it to the right. The policy may fail in both cases because the relevant sample distribution is missing. A world model trained on broader video may theoretically have a higher ceiling, but 陈哲 does not present “theoretically” as something already achieved.
22. 7 Hz is not the hardest problem; the video backbone and tactile gap set WAM’s ceiling
DreamZero currently runs at about 7 Hz on a robot, which is slow. 陈哲 sees speed as an optimization problem: “The core of computer science is that we first find a way.” Early GPT-3 and GPT-3.5 generation was also slow; later improvements raised speed by 100x and even 1,000x.
The more structural constraint is that embodied models depend heavily on progress outside robotics: better VLMs produce better VLAs, and better video-generation models produce better WAMs. A robotics company may not be able to independently fix the foundation model’s spatial inconsistency, long-sequence instability, or physical errors.
Even if a future video model could “imagine” the next 30 seconds with its eyes closed, visual data would still lack touch by nature. Whether a cup slips, a soft object deforms, or grip force is sufficient cannot be closed-loop controlled reliably from generated images alone. That is where world models and high-DOF tactile hands must converge.
陈哲 remains confident that video backbones will eventually follow physical laws more closely, but leaves the timing and path uncertain. He does not claim that world models have replaced VLA; only that they open a higher and more expensive research frontier.
23. EgoScale expands the data base but does not erase the domain gap between embodiments
NVIDIA’s EgoScale uses more than 20,000 hours of first-person human manipulation data, combined with finger-motion records from Manus data gloves, and maps the policy onto a high-DOF dexterous hand. The data can support both VLA and world-model pretraining.
陈哲 describes a data pyramid whose top is real-robot teleoperation: the most precise and effective source, but scarce and expensive, with every joint and motor state recorded. The next layer is UMI and Dex UMI-style embodiment-matched capture, followed by more open first-person human video, with massive internet video at the bottom.
First-person data is closer to manipulation than ordinary third-person YouTube footage, but the human head, wrist, upper limbs, and fingers have far more degrees of freedom than a robot. Many actions cannot transfer directly. High-DOF hardware can narrow the gap, but cannot make human and robot embodiments identical.
Simulation is not a single independent tier in the pyramid. Fully virtual generated data can sit between first-person video and UMI, while augmentation of real-robot data is also a simulation technique. The tradeoff is always among quantity, diversity, physical accuracy, and correspondence to the target embodiment.
24. Sunday’s three-finger UMI improves data quality, but “1 million hours” remains a collection target
Sunday’s founding team worked on UMI at Stanford, and Generalist later improved the approach. Sunday then expanded the traditional two-finger embodiment-matched gripper into a three-finger device and added tactile and feedback signals, keeping the capture device matched to the robot end-effector and reducing domain-transfer loss.
The extra finger is not decorative. Rotating, supporting, and repositioning objects that a two-finger gripper cannot handle may become straightforward with three fingers. The team went through many iterations before converging on the current form; only then was low-cost, large-scale replication possible.
陈哲 admires a design that “did not exist in the world, but seems reasonable in hindsight.” China’s advantage is that once a route is clear, follow-on development and engineering move quickly. In the past 6 months, many similar UMI, three-finger, and first-person capture solutions have appeared.
More solutions do not mean the data bottleneck has disappeared. Many Chinese companies have set a target of 1 million hours of real data this year, but the data must still be cleaned, labeled, stripped of invalid segments, and checked for coverage. Even after obtaining 1 million valid hours, 陈哲 still has no conclusion on whether that is enough to train a general-purpose manipulation model.
25. Video-generation models optimize for commercial visuals, not necessarily the physics robots need
Google Genie 3, Seedance, and similar models initially optimize for high fidelity and visual appeal, emphasizing artistic style and watchability without necessarily strengthening object permanence, contact mechanics, or long-horizon physical consistency. A consumer “good video” is therefore not automatically the “good world” a robot needs.
The leading video models today include Google Veo, ByteDance’s Seedance, Grok Imagine, and Kuaishou’s Kling. Their parent companies have enormous content pools, compute, and experimentation resources. Owning YouTube, TikTok, or Kuaishou does not mean having frictionless access to all the data, but the correlation between resources and results is clear.
Many robot world models currently use Alibaba’s open-source Wan 2.1 or Wan 2.2 as a base. 陈哲 says “if not all, then most,” and NVIDIA has no need to train one from scratch for early research. Several researchers believe Wan was not optimized for robot video generation; it is simply the best available open-source option for the current stage, with the most debugging and adoption.
Leading vendors’ decision to stop or reduce open-sourcing is partly driven by the enormous cost of video models. But 陈哲 believes that as the technology matures, laggards will still have an incentive to use open source to erode the leader’s premium, just as DeepSeek and Kimi did in language models. Open-source supply may not dry up permanently.
26. World-model startups are “big-tech friendly”; the real expense is repeated experiments on data and training recipes
At GTC, 陈哲 encountered the Robo AI team, which announced roughly $450M in funding and chose world models as its startup direction. Even with an open-source base, it needs to collect large amounts of egocentric video and conduct continued training rather than simply plug in the model.
Chinese teams have also claimed to be exploring the world-model route, including 机加科技 mentioned on the program. 陈哲 does not have enough information to judge its results. Benchmarks, model boundaries, and product metrics are all still vague, so the financing narrative is easier to establish before comparable performance.
The difficulty is concentrated in data and compute. The head of a Chinese multimodal team once told 陈哲 that copying several PB of data once could generate a bill of hundreds of thousands of yuan—“before doing anything”—with costs far beyond handling a few TB of text tokens.
The larger hidden barrier is the recipe. Even with tens or hundreds of thousands of hours of cleaned data, training π0.6 or a high-quality WAM requires many attempts. Startups need more than the budget for one training run; they need a sustained GPU budget to support both the number and scale of experiments.
27. Edge compute is spilling over from automotive stacks; NVIDIA’s dominance will fade as cost constraints bite
For real-time VLA or world-model inference on a full-size humanoid, the default is not a low-end Jetson but NVIDIA’s automotive Orin, or even the higher-end Thor. Fitting an autonomous-driving model onto one Orin already takes substantial effort; embodied models will not be less complex.
陈哲 believes the industry has not yet entered the compute-saving phase. It will first seek maximum compute within roughly 100-200 W of power consumption, and only then pursue extreme compression. NVIDIA therefore remains an early infrastructure provider, while Chinese automotive-chip companies such as Horizon Robotics are also beginning to support humanoid projects.
Mass-produced vacuum robots, commercial robots, and drones are highly sensitive to cost, power, and supply. They commonly use D-Robotics, Allwinner, or Rockchip solutions, with Jetson holding almost no share. As cloud GPUs, automotive chips, and robot compute move down-market, NVIDIA’s margin economics and dominance will weaken.
Tesla plans to use the same chip in self-driving cars and Optimus, validating technical continuity from automotive to robotics. XPeng, Huawei, Li Auto, and Xiaomi also have mass-production capabilities in automotive chips and could migrate their existing compute stacks into embodied AI. Huawei’s variable is constrained capacity: cloud and smartphones take priority, so robotics may rank lower.
28. General-purpose robots and chips may both consolidate heavily; feature robots remain the more realistic startup entry point
Drawing on historical experience, 陈哲 believes complex chip markets usually end up with 2 major suppliers holding shares close to 80/20. The many players in robot edge chips today therefore imply a fierce shakeout over the next few years; not every automotive-chip company will retain its position.
Humanoid hardware may also converge sharply. If one general-purpose architecture is good enough, its scale will be enormous and winners will not remain as dispersed as feature-robot companies are today. Cars appear to have many brands, but by country and corporate group the market is already concentrated; AI, autonomous driving, and software-hardware compounding will only raise the bar.
In practice, 汪滔 started with drones and Roborock with robot vacuums, while the market has also produced single-task robots for lawns, warehouses, and pools. Feature robots can build successful companies, but they do not automatically evolve into general-purpose robots: the organizational culture for one task differs from that required for a multi-task, software-heavy platform.
陈哲 does not find it surprising if Sharper or DJI moves into general-purpose robotics. They may already have “50% of the recipe” in opto-mechanical-electrical engineering, precision manufacturing, and mass production; the remaining work is to add models and systems. Embodied-AI companies must fill the other gaps, which may be harder but is not impossible.
29. IPOs and competitions will keep driving the heat, but whether humanoids are the optimal form should remain open
陈哲 sees robotics as a national-level technology priority for China’s next decade, and Unitree’s IPO could open a new phase for hardware companies with global competitiveness. Unitree is not a concept company but a benchmark with real users, real revenue, and highly efficient operations; successful exits would bring more capital and talent into the sector.
The risk is a mismatch between valuation and deployment. He estimates that China already has more than 20 humanoid or embodied-AI companies valued above RMB10B or $1.5B, versus only 4 or 5 companies at that level during the peak of the foundation-model frenzy. Even leading Unitree generates only a little over $200M in revenue. If regulators restrict listings by undifferentiated companies, primary-market financing will be affected, but the larger shock could come from technical progress falling short of euphoric expectations.
Over the next 1-2 quarters, three things are clear to watch: which team can use world models to produce results meaningfully beyond π0’s current VLA route; what new manipulation capabilities emerge from high-DOF hands with tactile sensing; and whether the Beijing E-Town humanoid competition can, like F1, use high-density competition to generate technology that can be transferred into mass-produced products.
The final question remains open. 陈哲 once leaned toward multiple robot forms coexisting, but this quarter was shaken again by the spatial, payload, and dynamic-balance advantages of bipedalism. He sees the constraints imposed by mechanics, motors, and energy, while also believing breakthroughs can cross a threshold nonlinearly. “I don’t have an answer today” is the judgment most worth carrying into the next quarterly review.