112: Talking Embodied Intelligence with 千寻’s 高阳: How Someone Who Looks Like a Robot Builds Robots That Look Human
Summary
高阳 defines embodied intelligence as an industry that has moved from open-ended research into engineering convergence, rather than a distant concept still waiting for a flash of inspiration. In 2023, he felt the room for truly groundbreaking research was narrowing, so he chose entrepreneurship as a way to engineer mature paradigms into products. To an outside world that still sees a jumble of competing approaches, his view is: “The ingredients are all there; it just takes time to build it.”
The industry’s key capability inflection point is L2, not bipedal locomotion or flashy single-task demonstrations. 高阳 defines L2 as a robot that can move around a single setting such as an office and complete roughly 5-8 tasks; L3 means covering 70%—80% of the tasks humans can perform in that setting. At the time, he expected L2 by “next June,” with L3 requiring another 1.5-2 years; the serviceable market could expand 10x or even several dozen times from L1 to L2.
The core technical path is end-to-end VLA paired with fast and slow systems, while the priority form factor is a wheeled base with 2 arms. The slow system understands vision, language, and intent; the fast system processes tactile and low-level visual feedback at 50—100 Hz. Bipedal robots will eventually need to be added, but 高阳 argues that “the main bottleneck is what your hands can do”; wheeled dual-arm robots may already cover roughly 80% of the scenarios robots can solve.
Data scale will be the key constraint on embodied-model scaling. 高阳’s team has observed a log-linear relationship between data volume and the optimality gap. Roughly speaking, “collect 10x the data, and performance gains another 9.” Following its roadmap to an embodied “3.5” would require roughly 10B useful internet videos, 100M teleoperation trajectories, and tens of millions of RL trajectories, over an estimated 4-5 years.
Folding laundry has become the industry’s shared demo because it tests deformable-object state understanding and generalization at the same time, not because housework is the first commercial use case. A shirt’s wrinkles cannot be reduced to a handful of fixed states: the robot must first understand “how it got messed up,” then decide how to unfold and fold it. The next hurdle is to change the fold on command or hide an envelope under the left sleeve—multi-task combinations that mark the transition from a single-task “IQ test” to L2.
China’s immediate edge is hardware manufacturing, repair, and the speed of experimental iteration, but the biggest value still lies in integrating the AI brain with the body. Robots in Tsinghua’s high-intensity experiments may be shipped back to Hangzhou every week, while 千寻’s in-house engineers typically resolve most issues within a day or 2. Even a US team that bought 100 robots could find that “the pace of repairs is hard to keep up with the pace of experiments”; but building only the brain leaves the system without the body-specific “muscle memory” it needs.
Commercialization should be delivered progressively with capability levels, rather than demanding that a prototype robot immediately function as general-purpose labor. 高阳’s response is that “you cannot demand GPT-1 have commercial capability while it is still GPT-1”: today’s L1 can enter a limited number of high-volume industrial stations such as automotive final assembly, and some customers are willing to pay RMB100,000-plus for a robot. The company’s real priority should be to reach “GPT-4.0”-level capability while using limited commercialization to improve resilience.
Deep dive
1. The More the Paradigm Converges, the Clearer the Startup Window
高阳 returned to China from Berkeley in August 2020 and joined Tsinghua. The moment he truly decided to start a company came in 2023, when he began to feel that “there wasn’t that much left to do in this research,” as paradigm shifts gradually turn open questions into engineering problems.
His reference point was GPT-4, which already handled everyday tasks very well. Outside areas such as AI safety, it had become difficult to find fundamentally new research space around ordinary tasks. A shrinking research frontier does not mean the technology has lost value; it means society can finally begin to enjoy its dividends.
王与桐 asked why scientists felt the field had converged when outsiders could still see so many points of disagreement. 高阳’s answer was that the academic topic space may have narrowed from roughly 500 questions to roughly 100. It is not that only 1 route remains; rather, many paths have not been fully disproven but are viewed as unlikely to lead far, making truly influential new work harder to produce.
王与桐 described 韩峰涛’s article on embodied intelligence as “only so-so in writing,” while arguing that its core judgment was exactly right: hand-coded programs constrained traditional robotics, whereas large models provided the first technical foundation for general-purpose tasks. 高阳 said 韩峰涛’s summary was more detailed than his own, and that it was exceptionally rare for an entrepreneur from a traditional industry to be this open-minded and willing to believe in the thesis.
2. Figure 02 Turned the Fast-Slow System into an Engineering Gap
The recent US example 高阳 admires most is Figure 02. 2 humanoid robots face a bag of jumbled groceries, recognize that milk belongs in the refrigerator and fruit in the basket, then converse, divide the work, and apply common-sense knowledge.
Figure has explicitly discussed its use of fast and slow systems, while the robot’s industrial design, compliance, and fluidity of motion also impressed him. China discusses fast and slow systems widely, but 高阳 has yet to hear of a company that has truly made the architecture work and integrated it into VLA; the gap is largely the sustained engineering needed to optimize response and motion.
China’s countervailing advantage is hardware manufacturing, repair, and experimental iteration. Because of high-dynamic experiments, Tsinghua’s robots are shipped back to Hangzhou almost every week, while Unitree may take roughly 1.5 weeks to repair and return them. US teams can only ask Chinese suppliers to ship spare parts and repair the robots themselves. Even if Physical Intelligence bought 100 robot sets, its repair cycle could still struggle to keep pace with experimentation.
3. The Industry Has Cleared L1 but Has Not Truly Reached L2
千寻 defines L0 as an industrial robot with little intelligence; L1 as intelligently completing a single task from a fixed position, such as driving screws on a factory line; and L2 as moving around a single setting such as an office and completing roughly 5-8 tasks, including making coffee and clearing a desk.
The real leap is L3: covering 70%—80% of what humans can do within a single physical environment. L4 means completing all human tasks in 1 setting, analogous to Waymo’s ability to drive anywhere in San Francisco; L5 removes the constraint of offices, homes, convenience stores, or factories altogether.
高阳’s judgment was that 千寻 and the industry’s best teams had already cleared L1 and were approaching L2. At the time, he expected L2 to arrive as early as “next June,” with L3 requiring another 1.5-2 years. This was a projection from the technical chain, not a validated industry timetable.
4. Wheeled Dual Arms Take the First 80%; Bipeds Handle the Final 20%
Embodied intelligence does not have to mean humanoid robots. L1 can be handled by a single robotic arm; by L2, most tasks may require 2 arms working together and a mobile platform, with the simplest combination being “2 arms plus a wheeled base.”
高阳’s bottleneck diagnosis is blunt: a robot with mobility but no manipulation still needs humans to load and unload it at both ends, creating little value. Wheeled bases are already mature yet have not entered offices at scale, which shows that locomotion is not the missing piece—the missing piece is what the hands can do.
Asked whether a wheeled platform merely postpones the eventual need for bipedal robots, 高阳 agreed that bipeds would need to be added later, but argued that wheeled dual-arm systems could account for the largest share of shipments for a considerable period and cover roughly 80% of use cases. The remaining 20%—stairs, sports fields, and the outdoors—would require bipedal locomotion.
王与桐 noted that children seem to learn to walk later than they learn hand manipulation. 高阳 still considers bipedal technology “relatively simple”: his lab has already achieved walking, swallow-like balance, and single-leg kicking, while industrial-grade stability still needs engineering work. But there is “no fundamental blocker.”
5. Humanoid Form Follows the Dimensional Constraints of the Human World
A robot does not necessarily need a humanoid head, but its cameras need to be mounted high enough to observe the scene. Tables are typically around 75 centimeters high, and indoor spaces are designed around human dimensions. Mimicking the human form at least ensures that existing environments remain physically workable.
王与桐 imagined a 1.2-meter-tall robot with arms extending 2 meters like trekking poles. 高阳 said that was feasible too, but unnecessary for most settings. Humanoid form is not mere conformity; it follows from the fact that “the world is designed for humans.”
A four-legged robot with 2 arms—a “centaur”—could also become a product category. It would be far more stable than a biped, but would occupy more space. The final configuration will depend on the setting, not on a requirement that all general intelligence replicate the human appearance.
6. Generalization and Specialization Both Serve Lower Costs
Assembly-line specialization and general-purpose robots are not contradictory. A dedicated solution for every task carries a fixed cost, while a general-purpose robot can reuse the same hardware and AI and add capabilities for different tasks, lowering the overall cost.
高阳 does not believe general-purpose robots will replace every dedicated machine: “We would not use one to produce plastic cups,” because a mold remains the fastest solution. The more likely future is coexistence between general-purpose robots and specialized robotic arms.
Installing headlights and seats on automotive final-assembly lines still requires human labor, making these processes candidates that traditional robotic arms struggle to handle. Humans may be cheaper for now, so the first targets should be high-volume stations. The economics only work if a robot can be used for years and cover the cost of several years of 1 worker.
7. End-to-End VLA Is the End State; Hierarchy Is Today’s Shortcut
高阳 treats end-to-end and VLA as equivalent paths: vision and language go in, actions come out directly, instead of first converting information into an intermediate representation and then generating a specific operation. In theory, this can cover manipulation problems broadly; in practice, the boundaries are set more by sensor capability.
Hierarchical systems are easier to engineer today, but he believes the end state “will definitely be end-to-end.” The lesson from more than 10 years of autonomous-driving development is that manually designed hierarchies are unreliable, and the industry is moving toward end-to-end systems.
This view does not simply follow the current fashion. Around 2016, 高阳 and 许华哲 worked on end-to-end autonomous driving, potentially using roughly 100x the data of Nvidia’s contemporaneous work. The paper’s technology is obsolete, but the underlying philosophy carried into their robotics research.
8. Folding Laundry Tests State Understanding, Not Grip Strength
千寻 demonstrated folding multiple garments in succession because deformable objects do not have the stable, easily describable states of a cup. The robot must understand each random wrinkle pattern before deciding where to unfold and fold; the difficulty is “not force,” but recognizing exactly “how the garment is messed up.”
高阳 compared this with the difficulty a 2- or 3-year-old has folding clothes. 王与桐 therefore called it an embodied-intelligence “IQ test,” and suggested it may represent a cognitive challenge closer to that of a 4- or 5-year-old. The task combines difficult manipulation with generalization across colors, garment types, and initial states.
Hugging Face, Physical Intelligence’s π0, and 李飞飞’s team have all demonstrated or published work on folding clothes, handling knives and forks, and similar tasks. 高阳 stressed that the industry’s focus on laundry does not mean housework is expected to commercialize first; it is simply one of the hardest technical benchmarks among single-task problems.
9. The Slow System Decides What to Do; the Fast System Keeps the Grip
千寻’s next step is to have robots handle ad hoc requests: use a different folding method, hide an envelope inside the garment halfway through, or place it precisely under the left sleeve. These combinations are what begin to approach L2.
The fast system receives tactile and low-level visual feedback at roughly 50—100 Hz, handling short-cycle corrections such as “did it grasp the object?” and “if the grip is unstable, regrasp immediately.”
The slow system searches for the envelope, determines left from right, understands language, and forms an intent. 高阳’s division is simple: the fast system makes actions reliable, while the vision- and language-driven slow system determines task structure.
Fine manipulation such as inserting a pin may not be more demanding of the “brain.” If the task logic is simple, the bottleneck may simply be tactile-sensor precision. Embodied capability cannot reduce every difficulty to model scale.
10. Generalization Is the Threshold from a Handful of Tasks to a Whole Setting
The leap from L2’s limited task set to L3’s 70%—80% coverage cannot come from collecting every task one by one. 高阳 believes the central challenge for robots today is generalization: “I see a new object—can I know how to handle it?”
A standard test is whether a robot trained to grasp A, B, C, and D can grasp an unfamiliar object just bought from a store. Folding clothes also requires adjusting the strategy across colors, garment types, and wrinkle patterns; generalization is already embedded in the task.
11. Internet Videos Provide Common Sense and Motion Priors, but Only Roughly 1% Is Usable
千寻’s data pipeline starts with internet images, text, and human-operation videos, adds real-robot teleoperation data for SFT, and then uses reinforcement learning to improve success rates. The goal is to “use all the data we can make use of.”
Video-platform content varies widely in quality, and only about 1% is closely related to human manipulation; first-person video is the best source. The team trains models to predict future object and hand trajectories so they can learn how objects “should be operated.”
王与桐’s challenge was important: watching many videos does not necessarily teach a person how to perform the task. 高阳 acknowledged that internet videos can only give a model a rough sense of which action is correct. Vision-language models have largely learned object recognition already; true action precision still requires real-robot training.
12. SFT Teaches the Form; RL Lets the Model Learn by Doing
Imitation learning is like watching an installation video and then assembling the part once yourself, turning an action from “roughly like this” into something precise. But if a human always holds the robot’s hand, it will still fail in roughly 5%—10% of cases.
高阳 uses cooking to explain RL: even after watching his mother cook for 10 years, a person can still misjudge the salt or the heat the first time they cook alone. “To do this well, you have to learn by doing it yourself.”
Reinforcement learning often shows a threshold effect. A robot may be unable to fold clothes or insert a USB drive for a long time; once it succeeds by chance, subsequent successes become increasingly easy. “Suddenly it gets it once, and then it keeps getting it in the future.”
The more valuable effect is cross-task generalization. Experience inserting one object may transfer to other insertion tasks, while turning screws, door handles, and light bulbs share motion structures. How quickly the first success appears depends on whether imitation learning and SFT were good enough beforehand.
13. The Data Roadmap Has Not Converged; Team Endowments Shape Beliefs
Simulation, teleoperation, and video learning remain genuine points of disagreement in embodied intelligence. 高阳 believes the choice reflects both an intellectual view and accumulated team capability: a company that excels at simulators will naturally place more faith in simulation. 王与桐 also cited Tesla’s extensive use of teleoperation.
千寻 places different data types at different stages, with the largest volume still coming from internet images, text, and video. 高阳 invoked the large-model playbook: “The most important step for every large model is pretraining.” Base-model quality is an important foundation for downstream capability.
14. The Data Scaling Law Sets the Order of Magnitude—and Reveals the Data Ceiling
In a data range of roughly 100,000 to several hundred thousand samples, 高阳’s team observed that the log of data volume has a linear relationship with the model’s optimality gap. Put roughly, a 10x increase in data might move performance from 99.9% to 99.99%.
Following its technical roadmap to an embodied “3.5,” 高阳 estimates it would need roughly 10B useful internet videos, 100M teleoperation trajectories, and tens of millions of RL trajectories.
If only about 1% of data is usable, 10B useful videos would require screening roughly 100x that amount of raw material. His estimate is that the total stock of usable internet video itself is only around 10B items, meaning the team would have to “learn from almost all of it,” a process expected to take 4-5 years.
Simulation data cannot be plugged into the same law by counting samples. A simulator can run indefinitely, but its diversity is limited; what matters is how many real-world tasks it covers, including transparent glass cups, clothing, and chairs that are soft on top and hard underneath. Teleoperation still follows the scaling law if its diversity is sufficient, but costs more.
15. Brain and Body Must Build “Muscle Memory” Together
Building only the brain can support methodological research, but makes it difficult to solve the real problem. 高阳 believes humans themselves do not have strong cross-embodiment capability; a model never trained for a specific body lacks the “muscle memory” needed for a tennis swing or rapid manipulation.
Building only the body has the opposite problem: most of the value sits on the brain side. Over the past 10-20 years, body capability has not changed fundamentally; the reason the market suddenly expanded was the emergence of a different AI brain. “The greatest value is on the brain side” is the broad consensus 高阳 sees.
千寻 therefore chose brain plus body from day 1. Combining them lets a model develop muscle memory for the characteristics of a specific body, rather than forcing 1 model to adapt rapidly to any robot.
16. The Supply Chain Will Eventually Resemble Autos, but Vertical Integration Is Still Necessary
高阳 expects the embodied-intelligence industry eventually to look like autos: the OEM defines the spec and integrates the system, while most components are sourced externally or co-developed. Tactile sensors, dexterous hands, and chips are among the difficult components in the body.
千寻 wants to work through open collaboration, but many components “still are not particularly good,” so the company has to build them itself for now. As the supply chain matures, more parts can be developed jointly with partners; only finer specialization will allow the final product to improve.
17. Commercialization Should Be Delivered Step by Step from L1 to L2
王与桐 relayed 朱啸虎’s challenge: the industry has broad consensus but no clear customer—who will spend RMB100,000-plus on a robot to do these jobs? 高阳 agreed that GPT-1 cannot be expected to have mature commercial capability while it is still GPT-1: “You cannot demand commercial capability from GPT-1 at the GPT-1 stage.”
His priority is to push the technology to “GPT-4.0” while deploying it selectively at each capability level to improve the company’s resilience. L1 can already work at some industrial stations, and some customers are willing to pay RMB100,000-plus; the problem is simply that there are not many viable use cases.
Once the industry reaches L2, the number of commercializable tasks could increase 10x or even several dozen times. When the roadmap truly converges will also depend on results: once 1 company reaches L2, or even L3, the rest of the field will naturally converge around that path.
18. AI Is the Shortest Stave, but Execution Is the Risk for Scientist-Founders
高阳 describes the data scaling law as the theoretical foundation for an embodied ChatGPT moment, not the moment itself. It is closer to OpenAI proposing scaling laws and producing GPT-4 2 or 3 years later. Robot data is harder to obtain and the industrial chain is longer, so the wait may be longer.
In his “barrel,” the hardware is also incomplete, but AI remains the shortest stave: “If we can fill in the AI,” it is equivalent to completing the barrel. A ChatGPT moment for robotics also requires the body, components, and engineering system to mature together.
A scientist’s background is not an exemption from execution risk. 高阳 openly agrees that scientist-founded startups may not work: understanding the technology is not the same as knowing how to divide engineering work, manage a team, or pace applications; all of that requires experience. The investor view relayed by 王与桐 was that 高阳 and 韩峰涛’s complementary fit and reasoning process appeared pragmatic.
Asked about his decision to retain his Tsinghua faculty position, 高阳 said he was simply taking the same work from research through engineering deployment, not splitting his attention between 2 unrelated pursuits. He also rejected the idea of trading long-term sleep deprivation for output: “Methodology matters more than putting in 1 or 2 extra hours every day,” because a single fatigue-induced decision error could force the people working with him to spend 2x as long.