Pioneers Insight Method Research Author
Vol.232 Industry Watch 44 | Embodied Intelligence + AI4S: What New Species Could Emerge?
Back to Episodes

Vol.232 Industry Watch 44 | Embodied Intelligence + AI4S: What New Species Could Emerge?

Summary

  • The embodied-intelligence industry has yet to converge: WRC showed “hundreds of companies running hundreds of parallel experiments.” 连文昭 sees broad agreement on the industry’s ceiling and future expectations, but the routes—point-to-point preprogramming, VLA, and generative world models—remain diverse. The biggest change from last year is not capability (“the capabilities on display are actually about the same”—still folding clothes and boxes), but the industry’s shift from showcase demos toward deployment and industrial value.
  • VLA has been singled out as “a shortcut around understanding,” while world models cannot stop at pixel-to-pixel prediction. Directly regressing actions from observations leads to overfitting and lowers the bar for generalization: “If a robot can still grasp an object after it moves 0.5 centimeters to the left, that counts as generalization; I have never considered that generalization.” 原洛’s answer is to center the model on objects, incorporate force as a modality, and close the loop, because force depends on the robot’s own impact on the environment and is tightly coupled with action.
  • The episode’s most counterintuitive call is a mismatch between hardware and software progress: hardware is advancing faster than it feels, while software demos are moving faster than paradigm innovation. The repeatability and absolute positioning accuracy of the latest robot bodies “can already approach the level of legacy industrial robots,” a prospect that was unimaginable 2 years ago. Algorithms may look dazzling in demos, but development still follows the large-language-model playbook of next-token prediction; over a longer time horizon, progress has been slower than expected. Robotics combines logic, geometry, and physics, making it harder than large language models.
  • AI for Science’s bottleneck is shifting toward wet-lab experiments, with robots providing a physical API across the generation-to-validation Gap. 李丰 cited this week’s significant Phase 3 results for Moderna’s personalized cancer vaccine and the use of AI in screening. Once discovery efficiency improves, the system needs an experimental data loop that is sufficiently large, fast, and standardized. 连文昭 says more than 80% of common experimental workflows can already be replicated; laboratory-instrument utilization may historically have been only 30%-40%, and robots can be embedded into existing labs through lightweight, even zero-cost retrofits.
  • The end state is a GPU-like “physical substrate.” Robot loops can generate standardized, trusted data and become infrastructure for scientific intelligence—“After GPUs, I can do things I never dared imagine before”—unlocking experiments that previously could not be designed. Cost savings and efficiency gains are easiest to underwrite commercially, and higher instrument utilization and throughput are broadly accepted; the value of “AI proposing better hypotheses” is not yet widely accepted, but could become the largest source of value.
  • China’s overseas robot demand is real, but it has to pass the “power-on test.” WRC saw a clear rise in overseas visitors from Europe, the Middle East, Japan, South Korea, and Southeast Asia. In 连文昭’s small sample, many came with concrete plans to act as integrators, distributors, or resellers, suggesting China is “gradually becoming the de facto global center of robotics.” But the scale-up filter is whether a robot “runs for a day and still wants to power on the next day”; overseas engineers are scarce, geopolitics is uncontrollable, and market expectations “are indeed on the high side.”
  • 原洛’s self-assessment: hardware sets the floor, models raise the ceiling; today, “the floor is stable, and the ceiling is deployable and deliverable.” In cooperation with 华大智造 and the Chinese Academy of Sciences, cell culture and cytotoxicity testing are already running at customer sites. A concrete example of the capability boundary is tearing plastic wrap—a real customer need involving material that is “highly transparent, highly reflective, thin, and tears immediately”; the task is technically achievable, but “really requires pushing through with gritted teeth.”

Deep dive

1. The return moment: WRC, Unitree’s listing, and a 2-year review of robotics

  • 李丰 opened by resetting the clock: when robotics “first began attracting attention” more than 2 years ago, he had invited 连文昭 on the show. Now the Beijing World Robot Conference (WRC) is taking place alongside robot sports competitions, and Unitree Robotics also just listed this week. Its post-listing performance has been volatile, but regardless, the event is a milestone for robotics and a useful moment to reassess the industry.
  • 连文昭’s current setup: 原洛科技 was founded in 2023 and builds both humanoid robot bodies and brain models. It is rolling out “OPN—a physics-centered, physics-native model,” using broad AI for Science laboratory scenarios as its proving ground and test. He is also a professor at Shanghai Jiao Tong University’s School of Artificial Intelligence, where he researches safe robot motion in unknown environments and generalization across scenarios.

2. Three U.S. chapters: from Vicarious’s brain-inspired networks to Figure’s data stack

  • After completing his PhD, 连文昭 joined Silicon Valley AGI company Vicarious—an early AGI business reportedly regarded as a peer of OpenAI and DeepMind. He worked on brain-inspired neural networks and deployed robots with the U.S. Postal Service and other logistics customers. He later joined Intrinsic, Google X’s early robotics project, where he worked on an Android-like operating system for robots across different hardware platforms. Intrinsic was folded back into Google earlier this year.
  • He joined Figure in 2023, building some of its earliest data-collection pipelines and imitation-learning frameworks. Figure subsequently became a benchmark company in the U.S. robotics industry.

3. WRC read: from demo theater to deployment; methodology still unconverged

  • 连文昭’s central observation was that “hundreds of companies are running hundreds of parallel experiments.” The industry has no consensus and the methods are proliferating. There is agreement that the ceiling—or future expectation—is very high, but everyone is still working out how to deliver against it.
  • The right way to read the expectations against last year is that the capabilities on display “are actually about the same,” while this year’s demonstrations are much more scenario-specific and focused on industrial value. Last year, and especially the year before, most exhibitors showed one small capability and stopped there. Bipedal dancing remains popular—“emotional value is a real value”—but 连文昭 is watching upper-body manipulation move from capability demos toward reliability.

4. More overseas visitors: real partnership demand, but the power-on test comes first

  • Non-Chinese visitors were visibly more numerous, including people from Europe, Japan, South Korea, the Middle East, and Southeast Asia. In the small sample of visitors 连文昭 hosted and spoke with, a high proportion wanted to bring robots to local markets as promoters, integrators, distributors, or resellers. Qatar’s Al Jazeera also visited 原洛’s booth to film and exchange views.
  • His feasibility assessment is that “the vast majority can be done, but the question is whether you are willing to pay enough.” Overseas engineering capacity is lower than China’s, coordination may require more resources, and geopolitics cannot be controlled.
  • The scale-up filter is reliability: after sending a robot to a customer, “it runs for a day—does it still want to power on the next day? If it willingly powers on by itself, the customer’s willingness to pay will definitely be strong.” Industry expectations “are indeed on the high side,” but a larger expectation also creates a larger demand pool to screen.

5. From intuitive physics to world models: VLA as a shortcut around understanding

  • 李丰 recalled that 连文昭 was already discussing “intuitive physics” 2 years ago. World models and physics models have since become hot startup themes, and “we have invested in several of these startups…whose valuations have risen very quickly over the past 6 months.”
  • 连文昭’s criticism of VLA is that, after collecting data, it regresses directly from observation to action. “It is essentially a shortcut: it bypasses understanding and goes straight to generating actions.” The result is prone to overfitting and generalizes less well than expected. He objects to a standard under which a robot’s ability to grasp an object moved 0.5 centimeters to the left is called generalization: “I have never considered that generalization. That is simply a capability it should already have.”
  • He is also skeptical of pixel-to-pixel world models. Predicting and reconstructing every pixel is “too cumbersome and somewhat counterintuitive,” forcing the model to spend a large share of its capacity restoring pixels rather than modeling the states that actually matter.

6. 原洛’s approach: an object-centric, force-modality, closed-loop world-dynamics model

  • Force is the critical modality because changing the camera angle changes visual information without materially involving the robot itself, whereas force “depends entirely on the robot’s own impact on the environment.” The greater the force applied, the greater the feedback, making force tightly bound and coupled to action.
  • The model predicts both future images and future forces, then compares the predicted force with the force actually measured after the robot applies an action and feeds the error back into the loop. It cannot be feed-forward only. As the data set grows, estimation error should decline and the model’s understanding of object-interaction dynamics should improve.

7. Work backward from the end state: zero-/few-shot generalization and measurable safety boundaries

  • The end state is zero-shot or few-shot generalization: put a robot in a new environment and, after a single demonstration—or even without one—it can move, understand, and complete the task. It must also be “absolutely safe, know its own boundaries, and handle anomalies.” That requires higher data efficiency and a way to assess the uncertainty bounds of its own outputs.
  • The field is already trying to lower data-collection costs through egocentric, first-person data, which 原洛 has collected in the lab since 2023; increase information density by centering on objects and implicitly deriving 3D from 2D or multi-view 2D information—“3D is a more fundamental understanding of the environment”; and collect force and tactile data, including through tactile gloves and other sensors. 原洛 is also exploring vision-tactile fusion and has built its own data-collection pipelines.

8. Misaligned progress: hardware faster than it feels, software demos faster than paradigm innovation

  • Hardware “does not look as if it is developing that quickly, but engineering progress has been rapid.” The repeatability and absolute positioning accuracy of new-generation robots “can already approach the level of legacy industrial robots,” something that was unimaginable 2 years ago. Government investment and broad experimentation have matured upstream machining, electronics, and sensors, while protocols and standards are taking shape. Hardware iteration may already be approaching software iteration speed.
  • Software feels as if it is improving rapidly when judged by demos and visual polish, but over a longer horizon the gains are not as fast as expected. The development paradigm has not been radically overturned and still follows the large-language-model approach. Robots face an extremely complex input space—vision, force, touch, joints, and states that cannot be observed visually, such as whether water is full or empty and whether an appliance is on or off—while the output is the robot’s own control system. That is fundamentally different from a large language model.

9. Defining embodied intelligence: beyond logic lie geometry and physics

  • 李丰 offered an outsider’s framework, noting his nontechnical background: language expresses logical relationships, and the same meaning can be phrased in multiple ways. Robots face hard physical and environmental constraints: a hand cannot pass through a table, and a bottle cannot be squeezed until it breaks.
  • 连文昭 fully agreed, adding that many local optima in a large language model can still be good answers. Robotics is not just a logic problem; it also involves geometry—kinematics, including how many degrees a joint turns and how many centimeters the end effector moves—and physics—dynamics, including how much force to apply, the level of friction, and how far or how many degrees an object will move. It is therefore “a harder problem than a large language model.” His informal definition is “how to change the state of a target object under many constraints”; in many cases, there is an exact answer, and either the robot achieves it or it does not.
  • If that framing is correct, robots can borrow next-token prediction from LLMs and use correlations as an approximation for causality. But text itself has no time dimension, whereas the physical world does, so robots also need causal constraints and an explicit time dimension.

10. Why go all-in on AI for Science: in the impossible triangle, sacrifice speed and choose the hardest test

  • The scenario-selection logic starts from an impossible triangle of generality, reliability, and speed. Speed only needs to clear a threshold; generality can be narrowed to solving dozens, hundreds, or more tasks. But the scenario cannot cap the technology’s ceiling or pull development toward a short-sighted route. That is why 原洛 chose long sequences to stress logic, high precision to stress geometry and physics, and highly flexible general-laboratory tasks. “If we can do these well, we will certainly be able to solve many of the problems” in future home scenarios, “甚至是1万件事情.”
  • Reproducibility is the industry’s first major pain point. One person may take a sample from a 4°C refrigerator and leave it in a 25°C environment for an hour before mixing at 25°C; another may pour it directly. Someone might even sneeze during the experiment. These details make experiments hard to reproduce, so “10 years, RMB1B, and a 10% success rate” remain painful economics. Robots can execute every step faithfully and at fine granularity while recording every condition and parameter.
  • Expensive instruments are tied to people’s 9-to-5 schedules, with utilization potentially only 30%-40%, stretching the R&D cycle. A robot lab assistant can connect and schedule the instruments, delivering both higher quality and higher throughput; hazardous experiments also no longer need to be performed by people.

11. Moderna signal and the wet-lab bottleneck: robots as the physical API for the “generation-to-validation Gap”

  • 李丰 cited this week’s signal: Moderna announced significant Phase 3 results for a personalized cancer vaccine, including a drug combination, and news reports said AI was used in the screening process. The takeaway is that once AI changes the paradigm of scientific discovery, “the biggest bottleneck moves to wet-lab experiments,” which need to send data back into the prediction loop in sufficient volume, at sufficient speed, and in sufficiently standardized form.
  • 连文昭 said 原洛 can already handle liquid, solid, powder, and granular samples, as well as common instruments and workflows including centrifuges and PCR. “More than 80% can already be replicated,” at least enough to form small local loops. He compared it with L2-assisted driving after entering a highway: assistance, not replacement, for scientists. An experiment can be scheduled for 6 pm before leaving work, eliminating a 2 am return to change the liquid; by 9 am the next morning, the result is ready to read.
  • The robotics industry often talks about the “think-to-real gap.” In this setting, the problem is the “generation-to-validation Gap”: digital AI can propose hypotheses and experimental protocols across a huge search space, but validating each one through human labor is impossible. 原洛 supplies a schedulable physical API to close that gap.

12. Hardware sets the floor, models raise the ceiling; plastic wrap marks the capability boundary

  • “Hardware improvements protect the floor.” Early in the company’s development, both externally sourced and internally developed robot arms shook badly and lacked precision; hardware now meets the requirement. “More of the upside comes from software models”—the question is how much it costs to teach and train a model to perform a new task. Today, “the floor is relatively stable, and the ceiling has reached a deployable, deliverable state,” while the company continues to reduce the number of samples required for new tasks.
  • High flexibility remains the harder problem: estimating and controlling tiny states such as whether something is tightened, fully inserted, or perfectly aligned. These tasks require substantial data collection and cleaning, as well as changes to the training architecture.
  • 原洛 is working closely with 华大智造 and the Chinese Academy of Sciences; cell culture and cytotoxicity testing are already running at customer sites. At WRC, 连文昭 did not see any company particularly close to 原洛’s direction. The venue was large, but like a shopping mall, much of the content on display looked similar.
  • Tearing plastic wrap is a real customer request and a concrete boundary case. “It can be done if we force it, but it really requires pushing through with gritted teeth”: the film is highly transparent, highly reflective, thin, and tears immediately, making both perception and force control difficult. The strategy is to expand organically and incrementally from the current capability set.

13. End-state vision: the “physical substrate” for scientific intelligence, as GPUs are for compute

  • “A human is the best API and can schedule every interface.” An intelligent robot can likewise be embedded into an existing lab through a lightweight, even zero-cost retrofit while remaining compatible with heterogeneous old and new instruments. The larger value is the standardized, trusted data generated by the closed loop: “After GPUs, I can do things I never dared imagine before.” That could unlock experimental topologies that were previously impossible to design and drive a reinforcing dry-lab/wet-lab loop.
  • Commercial acceptance is tiered. Cost reduction and efficiency gains are easiest to underwrite; higher instrument utilization and experimental throughput are broadly accepted. The value of better experimental quality and “AI proposing better hypotheses” is not yet widely accepted, but may ultimately be the largest source of value. The AI for Science boom “is already showing early signs.” Traditional wet-lab researchers and AI-native researchers “are not so divided; they are mixed together,” with the latter potentially more sensitive to whether data can accumulate and cycle. 李丰 also compared the effort with earlier experiments in lab automation and dry-lab/wet-lab loops at 晶泰 and 剂泰, companies he had backed.
  • Over the next year, the “Origin Point Plan” aims to embed the physical substrate into more partner organizations and build shared spaces. Within permitted limits, physical skills and data will be shared to support AI for Science R&D, while the company expands into agriculture, fine chemicals, manufacturing, and other industries. The closing self-assessment was “long-term optimistic, short-term pragmatic,” with an open invitation to R&D partners willing to pursue the real-world value of robotics over the long haul.