Pioneers Insight Method Research Author
Did Robots Take a Wrong Turn? 苏度's 韩铮 on 3D Data for Embodied AI
Back to Episodes

Did Robots Take a Wrong Turn? 苏度's 韩铮 on 3D Data for Embodied AI

Summary

  • Core contrarian call: hierarchical architecture will return to the mainstream in 2H this year. 苏度科技 CEO 韩铮 believes Silicon Valley’s end-to-end VLA route, represented by Physical Intelligence, is running into a data shortage: “Starting in 2H this year, this kind of hierarchical architecture may return to the mainstream—and may ultimately be the way to do it.” High-level reasoning and low-level manipulation need to be decoupled, while low-level manipulation cannot rely solely on human teleoperation data; Sim2Real is unavoidable.
  • Data cold-start is the key obstacle for the end-to-end route. Tesla accumulated FSD data after selling millions of cars more than a decade ago, but a robot “is unlikely to sell 5 million units to offices before it has any functionality,” and nobody wants to wear a motion-capture suit at home to teach a robot how to cook tomato-and-egg stir-fry. Cold start requires millions, potentially tens of millions, of open-environment data points, so “Sim2Real is unavoidable.”
  • The proof point is 苏度 R1’s zero-shot demo. After 60 minutes of continuous filming and more than 100 objects outside the training set, it completed 240 grasp-and-place attempts at nearly 98% success per attempt, reaching 100% after closed-loop adjustment; at ICRA Vienna and CVPR, potentially close to 1,000 attendees then tested it at random with badges, phones and toys. Compared with a company that rented an Airbnb a month in advance to collect data and rehearse a coffee demo, this was a public test that gave the team “almost no time to prepare”—something no one had previously pulled off.
  • Manipulation is 2-3 orders of magnitude harder than locomotion—an important dividing line for valuation. Sim2Real has already validated locomotion for Unitree and Boston Dynamics, but object manipulation requires structured 3D data, sub-millimeter geometric understanding and predictive reasoning about the environment. Figure’s recent production-line sorting demo involved only two nearly identical object shapes and remained on its own test line; “object generalization is actually limited.” 泓君 heard that the dishwasher demo took a long time to produce a single usable take.
  • The contrarian view on the competitive field: the biggest threat is DeepMind plus an electrified Atlas from Boston Dynamics, not PI or Skild. Skild’s valuation jumped from $4.5B to $14-15B in 7 months, but its 2 founders “used to be firmly against simulation.” DeepMind has MuJoCo, Genie 3 and “almost unlimited GPUs,” and has moved Boston Dynamics’ core hardware talent to London alongside the model team. “Its bet on AI for the physical world may be the most determined of any major tech company.”
  • For hardware mass production, watch Unitree: it shipped more than 5,000 units in 2025, versus roughly 150 for Tesla and Figure. Unitree is only beginning to tackle object manipulation. 苏度’s business model is closer to iPhone+iOS than Android: sell the robot, foundational model and SDK as an integrated stack, then let 1,000 developers produce 100 deployment scenarios and 5-10 super-applications deployed at 10,000-plus units each. “The Sim2Real route is right; that is a byproduct.”

Deep dive

1. 苏度’s positioning: full-stack, with the bet placed on “manipulation,” the hardest layer

  • 韩铮’s positioning is full-stack: the brain and the body. The team has spent nearly 10 years using machine learning to drive robots, but only decided to build its own hardware when it began semi-commercial operations in 2H 2022, because “if you do not handle the body yourself, it is difficult to demonstrate what the brain can do.” The Silicon Valley comps are Optimus, Boston Dynamics, Figure and 1X, which combine hardware and models, while their model capabilities lag PI; Skild focuses on models.
  • 泓君 opened with the industry’s core test for this episode: can a demo tell you whether a robotics company is good? Most companies adapt in advance to a specific environment; 苏度 is willing to demo in a completely new environment it has never trained on.

2. Dexterous hands are not the cost-effective choice: 90% of manipulation tasks do not need them

  • 韩铮’s breakdown: teams building manipulation models, including PI, first push the capabilities of a two-finger gripper to the limit. “Of the object-manipulation tasks robots need to perform, perhaps 90% do not necessarily require a dexterous hand”; only the remaining 5%-10% need multiple fingers. The humanoid trend was driven by the public imagination sparked by Optimus in 2022, not by technical inevitability.
  • For now, the strategy is to decouple manipulation tasks from lower-body dexterity: bipedal robots remain unstable and fall frequently. Over time, the two may converge again; the general-purpose robot configuration “may be wheel-legged,” moving efficiently on flat ground and switching to legs on difficult terrain.
  • In-hand Rubik’s Cube rotation and walnut cracking do require multiple fingers, but threading a needle may not. Stability, reliability and cost—an early Shadow Hand cost several hundred thousand dollars per unit—remain unresolved. “It is an option, but not the most cost-effective one.”

3. From ImageNet to ShapeNet: the lineage of 3D datasets and the brutal gap in scale

  • 泓君 introduced co-founder 苏昊 as one of ImageNet’s core authors. At Stanford’s Leo Guibas lab, he later brought the “large dataset plus structured descriptions” concept into 3D, starting ShapeNet around 2013-2014; PointNet was also one of his core contributions. The scale gap is stark: 韩铮 estimates ShapeNet has several hundred thousand 3D models, versus 14M image-text pairs for ImageNet. Open, usable 3D models remain extremely scarce online.
  • Annotation depth falls off at each layer: ShapeNet has basic geometry and surface textures; PartNet, which adds part-level labels, has only tens of thousands of objects; PartNet-Mobility, which adds dynamics constraints—whether a bottle cap turns left or right and whether it can be removed—has fewer than 3,000 objects. By 韩铮’s recollection, it remains the largest open dataset with both part annotations and dynamics constraints. That sharply constrains synthetic-data training and video generation with physical consistency.

4. Why not simply learn from video: 10 centimeters is enough to drive, but manipulation requires sub-millimeter precision

  • Responding to the idea of training on vast volumes of open-world video, 韩铮 said that 2D images and video contain 3D information, “but not completely and not very accurately.” Even one of Sora’s training pipelines renders 2D data backward from 3D models in Unreal or Unity.
  • Precision is a hard constraint. Around 10 centimeters is sufficient for autonomous-driving navigation, “but a 10-centimeter error is absolutely unacceptable for robot manipulation.” Inserting a cable into a port on a tape recorder requires sub-millimeter accuracy, which only 3D-model-based datasets can provide.

5. The pioneers’ foundation: simulator lineage and Boeing-style systems engineering

  • The team’s industry coordinates are unusual. “Embodied AI” should date to a research field with evaluation standards jointly proposed by 苏昊 and several North American professors in 2019. The major commercial simulators—Isaac, MuJoCo and even some newer startups—have “countless connections” to the team; members also contributed early work to parts of Isaac’s reinforcement-learning engine.
  • 韩铮’s description of the core asset is worth remembering: Sim2Real is “a bit like Boeing building an airplane.” It is difficult to say whether the engine, avionics or materials matter most; the accumulated asset is the team that can assemble the complex system. The leaders of the core groups have worked together for 7-8 years.
  • The historical lesson came after Andy Rubin joined Google and led the Google X project: between 2012 and 2016, Google acquired Boston Dynamics and a number of other robotics companies, but “the pieces were not yet all in place.” In 2022, 韩铮 and 苏昊 decided to start a company after DALL·E appeared. It convinced them that lifting 2D into 3D might be possible: “Rebuilding the objects and environments of the entire world in a simulator … was, in theory, becoming feasible.”

6. Simulator philosophy: sacrifice visual appeal for millions of parallel environments and physical consistency

  • The design logic of next-generation robot simulators differs from traditional finite-element simulation. Instead of using a large machine to approximate reality ever more closely, the goal is to run hundreds or thousands, even millions, of parallel environments on a single GPU. Resolution is “very, very low,” but physical consistency is pushed as close as possible to reality. The design necessarily makes major trade-offs between realism and efficiency.
  • The ultimate ambition, in the team’s words, is to put into a simulator “the millions of years of evolution through which humans and all living things came to understand and interact with the environment and manipulate objects,” then let robots undergo the full evolutionary process again. The name Sapien is a reference to Sapiens and its account of evolution from apes to humans.

7. Data: structure is the moat; there is no single right answer on quality versus scale

  • The dividing line is structured versus unstructured data. Allen Institute’s ObjectVerse mined every 3D file on GitHub and amassed more than 10M records, but much of it consists of game armor and weapons, which “may not be very useful for robots doing object manipulation.” 苏度 is building its own structured dataset, but cannot formally disclose its scale for now.
  • 韩铮 is candid about the trade-off: “We do not have a standard answer either; we are always balancing quality and scale.” A single beverage bottle does not need to reproduce its materials and label text with 100% fidelity, but structured coverage must reach sufficient scale. The team’s self-assessment is that its structured understanding “may be the best among all the teams,” while dynamics constraints remain “the most deficient data in the entire industry.”

8. 苏度 R1: the first Sim2Real demo to close the open-world gap

  • The demo released in April had already worked internally in November or December last year: 60 minutes of continuous operation, more than 100 objects that had never appeared in the training data, and 240 grasp-and-place attempts. “The success rate for grasping and placing that object each time was nearly 98%”; after a failure, closed-loop control quickly adjusted the strategy and could reach 100%. 韩铮’s characterization: this may be the first attempt in robotics to use Sim2Real to achieve zero-shot performance and close the Sim2Real gap.
  • 泓君 pressed the obvious research-community objections: were there controls, and had the objects appeared in the training set? 韩铮’s answer was unequivocal: “No. Actually, absolutely not.”

9. Public random testing versus renting an Airbnb a month in advance

  • The response to skepticism was public random testing. At the ICRA conference in Vienna in early June, potentially nearly 1,000 attendees used badges, phones, earbuds and small toys to test the robot at random; the team repeated the exercise at CVPR. “No one had ever been able to do a demonstration like this at similar academic conferences,” in part because it leaves “almost no time to prepare.”
  • 韩铮 described the industry’s contrasting playbook plainly: before an academic conference, one company would rent an Airbnb nearby, put its robot there to collect large amounts of data, and potentially move in a month early. It then repeatedly demonstrated the coffee task. Change the coffee machine or the environment and its success rate “would definitely fall quite a bit.”
  • The manipulation community considers open-world zero-shot success on short skills more important than long-horizon tasks. “If your short skills are stable and reliable enough, and generalize well enough, you can combine them into all kinds of long-horizon manipulations.” That is the logic behind the ManiSkill framework; in 2023, the team also worked on CoTPC, which specifically handled chaining short skills around the same time chain-of-thought emerged.

10. White box or black box: hierarchical pretraining, end-to-end deployment at the edge

  • The fundamental difference from PI’s narrow VLA—an end-to-end black box that imitates human actions—is that 苏度’s model must have a strong understanding of and ability to predict objects, environments and their future changes in the physical world, while remaining interpretable. When a bottle cap will not turn, the model should know whether to apply more force or whether the cap’s rotation direction is designed differently in that country. The architecture is hierarchically pretrained, but remains end-to-end when deployed at the edge.
  • 韩铮 described the pathology of imitation learning bluntly: it does not need to know an object’s precise location or geometry in the environment. “The reason it does this … is that in that environment, someone once teleoperated the robot to perform that action. It is simply imitating mechanically.”
  • His most vivid analogy is a child opening a bottle for the first time: turning left, then right, trying randomly while watching the parents. Some motor capabilities are “encoded in the genes inherited by humans.” For robots, “we actually need to reproduce the entire evolutionary process once again in a simulator.”

11. The cold-start deadlock and the contrarian call for hierarchical architecture to return in 2H

  • The fatal missing premise in the end-to-end data-stacking route is scale. Tesla sold millions of cars more than a decade ago without FSD; “millions of drivers drove every day on Highway 101,” and those drivers continued helping label data. Robots cannot do the same: “It is hard to imagine buying a robot and needing to wear a motion-capture suit just to cook tomato-and-egg stir-fry.” Chinese companies may already have hundreds or even more than 1,000 units deployed, but that is nowhere near 5M. Cold start therefore requires millions to tens of millions of data points, and “Sim2Real is unavoidable.”
  • A bit of industry lore explains the shift. Core people at PI and Generalist came out of Google Robotics, where there was once an implicit consensus around a hierarchical structure: stable low-level manipulation plus high-level task understanding. When startups formed in 2023 and 2024, however, they chose the playbooks they knew best and stopped talking about hierarchy. 韩铮’s contrarian view is that this structure “may return to the mainstream starting in 2H this year,” and could ultimately be the commercial solution—his personal view only.
  • 泓君’s follow-up elicited an even sharper judgment. Silicon Valley companies are better at high-level reasoning, but “hundreds of thousands of hours—I think that may already be the limit for today’s manipulation—of real human manipulation data” are nowhere near enough to deliver stable, reliable operation across open environments and open-ended objects. Low-level manipulation requires tight software-hardware integration, but “whether only Chinese companies can do it, I do not think that is necessarily the case.”

12. Tight hardware-software coupling, and a manipulation problem 2-3 orders of magnitude harder

  • Closing the Sim2Real gap requires hardware and software to be designed together. A simulator is not万能; hardware selection must favor configurations that simulate well, and simulation must account for the fact that every machine performs differently every time it powers on. Open-source workflows that import a robot in URDF format are “extremely simplified”: they ignore motor noise, response curves and hysteresis. The low-level manipulation model is therefore tightly bound to the robot body, making it difficult to buy the software and install it on another platform.
  • Simulation-based reinforcement learning has been thoroughly validated for locomotion on Unitree and Boston Dynamics, “but if a robot is to manipulate objects, the difficulty is 2-3 orders of magnitude higher.” Posture control only needs to focus on the robot itself; manipulation requires structured data, environmental understanding and prediction—the most important branch of the World Model. 泓君 connected the dots to a question she had previously asked an Nvidia scientist: “Has control of dexterous hands still not been solved?” This conversation, she said, seemed to provide the answer.

13. The business model is not Android; it is iPhone+iOS

  • The hardware looks like a car, with a battery, motors and compute, but the business model looks like a smartphone: sell the robot, foundational model and SDK, then let developers build applications on top. “Someone will build Uber, someone DoorDash … someone Meituan, someone Didi.” 泓君’s correction was accepted: even Android is too weak a comparison, because robot hardware and software together form the manipulation system and software cannot simply be sold to fit someone else’s body. It is closer to iPhone+iOS.
  • The developer funnel has specific targets: “Of the 1,000 developers we serve, 100 will be able to find application scenarios for small-batch deployment. Of those 100, perhaps 5 or 10 will build a super-app that may need to deploy 10,000 units or more.” 苏度 is not betting on any single vertical. It is building the foundational model for low-level skills, APIs and composition tools; the general logic for upper-layer applications is called Adapt internally.

14. The field: Skild, PI, Figure, Optimus and Unitree

  • Skild raised $1.4B at a $14-15B post-money valuation, up from $4.5B just 7 months earlier. 韩铮’s disclosure: founders Deepak and Abhinav “used to be firmly against simulation,” but when they first raised money, robot simulation was the main fundraising pitch. The actual business “may be somewhat sensitive.” PI’s latest valuation, according to him, is also around $12-14B and it is raising again. The US first tier also includes Generalist and the Chinese-founded Sunday and Dyna.
  • Figure’s comeback needs a discount. Its production-line sorting demo “was indeed somewhat better than before,” but it handled only boxes and soft packages, with almost identical colors and weights, and remained on its own test line. Teleoperation-based VLA is “extremely sensitive” to changes in lighting and environment. 泓君 heard that the dishwasher demo took a long time to produce a single usable take.
  • 苏度 worked with Tesla’s automation team before 2025. At the time, Optimus did not meet production-line requirements and should only have been performing a fixed-position battery-pick task within its own team. Even so, “Optimus will definitely be the biggest driver of the US robotics industry in the future.” Tesla has strong hardware-manufacturing capabilities; in Musk’s style, it may wait for a relatively good industry solution to emerge and then fill in the model layer, which could still be timely. Unitree, meanwhile, has “delivered the market’s best answer on hardware mass production”: more than 5,000 units shipped in 2025 versus roughly 150 for Tesla and Figure, with real users receiving the machines. Its object manipulation work, however, is only getting started.

15. The biggest threat is DeepMind plus an electrified Atlas; in 10 years, the goal is to be remembered for iOS, not for choosing the right route

  • “Over the next 2-3 years, in terms of manipulation and overall capability, our biggest competitor may be DeepMind plus Boston Dynamics.” 韩铮’s chain of evidence: he believes Demis has made “the most determined bet on AI for the physical world” of any major tech executive; he cited the recruitment of Boston Dynamics’ core hardware talent to DeepMind. 泓君 added that DeepMind acquired the MuJoCo team several years ago, has continued improving the simulator and has Genie 3, while 韩铮 said it has “almost unlimited GPUs.” Hyundai is believed to be Boston Dynamics’ controlling shareholder—if 韩铮 remembers correctly, SoftBank retains 10%-20%—and Google may “repeat the Android playbook.” Amazon is one of the biggest spenders, with Kiva, Agility and Covariant, but provides the application setting only to eventually acquire a company or build the capability itself.
  • Why do Chinese companies repeatedly retreat into vertical markets? 韩铮’s diagnosis, and his standard for 苏度, is that like the AI Four Little Dragons, teams may not have a clear idea of how to build a foundational model and may simply want to experiment as they go; commercial pressure then forces compromises. 苏度’s confidence comes from judgments that can be tested within 12-24 months: “The phrase ‘maybe’ almost never appears in our team.” He acknowledges that the economics of vertical applications do not work today: deploying one robot to replace a worker may actually require the maintenance cost of 3-5 PhDs.
  • Asked what he wants to be remembered for a decade from now, 韩铮 chose the same answer: “The Sim2Real route is right; that is a byproduct.” He wants this period to look in retrospect like the iPhone+iOS moment, when “new Ubers, Didis, Meituans and new DoorDashes will emerge.” The hope is that at least part of the hardware, underlying system and software used by those applications will come from 苏度.