Pioneers Insight Method Research Author
Six Little Dragons' First Listing: 群核's 黄晓煌 on 15 Years of Hard Tech
Back to Episodes

Six Little Dragons' First Listing: 群核's 黄晓煌 on 15 Years of Hard Tech

Summary

  • 群核’s hardest asset is an accidental goldmine of physical-world data: InteriorNet, an open-source dataset released in 2018, was discovered by Silicon Valley giants after the pandemic shut down offline data collection and has since made the company a major data provider to embodied-AI “brain” companies and multimodal-model players. 黄晓煌 admits, “I had an instinct that it would be useful, but at the time I didn’t know what it was useful for”; inspired by ImageNet, the team open-sourced it anyway. The clearest sign of commercial validation is whether customers keep buying more data, and repeat purchases are “pretty decent.”
  • The core strategic logic was choosing a lane under China’s compute constraints: large language models showed that more data means more intelligence, but with compute restricted in China, 群核 moved into physical AI, where data is less constrained, data scarcity is greater, and compute requirements are less extreme. New models must also improve the legacy business—using large models to infer the physical parameters of objects in images and replace manual work. “The old business is our cash cow,” and that is the line against a second-startup-style transformation.
  • The company deliberately avoids red oceans: “I’ve been hit by giants myself—I’ve been running from them ever since. Why would you go looking for a market packed with giants? That’s suicide.” The major tech companies’ physical-AI efforts are still “small teams exploring,” unlike the company-wide push behind large language models. 群核 is entering through 3D data, combining 3D and video models as a complement to video foundation models; 黄晓煌 says he has not seriously studied Meta’s V-JEPA 2, while 群核’s stated edge is spatial consistency.
  • AI is rebuilding the business model: after seeing OpenAI charge by the token in 2023, 群核’s first move was to revise its annual and monthly SaaS plans, because under the old model “every time someone used it, I lost money,” while compute-heavy features such as video generation were blocked by the CFO. SaaS value is also being measured in reverse—from “users multiplied by the time they spend each day” to “the less time I occupy, the more value I create.”
  • The company’s closest brush with disaster came in the first half of 2021, when “everything was exceeding expectations,” hiring surged toward 3,000 people, and the combination of property, the pandemic, and capital markets suddenly reversed course by year-end. “If I had been even more reckless then, I probably would have been finished.” Since then, the hard rule has been break-even at the floor and never exhaust the ammunition: even while highly bullish on embodied AI, “I still won’t burn money recklessly,” and when investors push for faster growth, his answer is, “We’ll grow steadily. I’m not burning.”
  • Organizationally, 群核 separates a “process army” from an “innovation army,” with the latter accounting for roughly one-tenth of the company; the test is not business maturity but whether “you can’t even formulate the KPI, in which case you need an innovation team.” His view of shareholders is equally blunt: “A shareholder just wants your money, not your life.” For shareholders who do not share the company’s direction, “don’t drag it out—find a way, whatever the cost, to buy them out and get them out as fast as possible; any price today is still the smallest price compared with the future.”
  • The Six Little Dragons effect has materially changed the talent pipeline: resumes from C9 universities are up about 9x this year, overseas-returnee resumes are 20x last year’s level, and after the property downturn many people “barely dared to ask for the salary you were offering,” leaving projects indefinitely delayed. 黄晓煌 has made recruiting his top priority. His interview process is to hand candidates a paper and an open-ended problem; some work from 9 a.m. until 1 or 2 a.m., and “I stay with them.”
  • The operating philosophy has stayed consistent: “take a hammer and look for nails,” while commercial timing follows academia—public-facing hot topics are “2 or 3 years, even 3 or 4 years behind,” with robotics hot in academia in 2022–23 but pushed by industry only in 2024–25. “If scholars haven’t solved it, don’t go messing around with it yourself.” 群核’s biggest success case is when technology works across industries and then catches 1 or 2 industries in a breakout phase; for now, robotics is the sector he likes best.

Deep dive

1. The Six Little Dragons Debate and 15 Years of Technical DNA

  • The host opened by asking about the claim that 群核 was “a little dragon that got shoved into the lineup.” 黄晓煌 did not dodge the role of luck: “Hangzhou actually has a great many excellent technology companies—maybe you could count 60 of them. Being selected as one of the Six Little Dragons definitely involved some luck.”
  • But he stressed that the company had been technology-led from day 1. Before GPUs became fashionable, it recognized that GPU clusters could accelerate physical rendering—“taking a hammer and looking for nails, searching everywhere for applications.” Even today, he says, the technology from 10 years ago would still have been highly advanced; people simply “couldn’t really understand it” at the time.

2. InteriorNet: An Accidental Data Goldmine

  • The project began with pure curiosity. A small research team was “just doing it for fun, with no real business pressure,” and, inspired by 李飞飞’s ImageNet, open-sourced InteriorNet in 2018: “I had an instinct that it would be useful, but at the time I didn’t know what it was useful for. We didn’t know how to train on it ourselves, so we thought: why not open-source it and let others try?”
  • The pandemic became the unlikely turning point. Once offline collection of training data for robots and AI devices was blocked, researchers at Silicon Valley giants began searching everywhere for solutions and found 群核. In the early open-source phase, “publishing papers got results, but doing business didn’t.” It was only in 2020–21 that the company realized the dataset could train not only robots but all kinds of devices.

3. The 2021 Insight: What Was Missing Was Top-Tier Algorithms

  • In 2021, 群核 expanded its data team to 20–30 people and tried to turn it into a formal business. It soon discovered that customers wanted help “tuning the algorithms,” a task that required top-tier algorithm engineers; ordinary engineers “didn’t solve the problem.” 群核 could not do it at the time, and the customers themselves lacked the capability.
  • The solution was to bring back former classmate 周子涵 from overseas to build an AI lab. The fit was straightforward: 周子涵 was working on 3D vision at a university but lacked data, while 群核 “happened to have both GPUs and data, and was missing algorithms.” 黄晓煌 recalls that he “didn’t think that far ahead—I just thought this was an interesting new field.”

4. The GPU Cognitive Revolution and the “Innovation Army”

  • A CUDA veteran, 黄晓煌 acknowledged his own blind spot: “If you told the company you needed 1,000 GPUs to train something, whoever came in to request that would get yelled at and probably fired. I’d worked with GPUs for so many years and had never realized this was even a thing.” He understood that more data improved results, but had never imagined building GPU clusters at that scale.
  • The transformation was not a replacement of the old organization but an upgrade. The legacy business continues to provide steady cash flow, while the new era requires “a batch of new talent and a new organizational form.” The dividing line is clear: if the steps needed to hit a goal can be specified, use a process team; if even the method is unclear and KPIs cannot be formulated, use an innovation team. Innovation teams account for roughly one-tenth of the organization. They are expensive, “but compared with the opportunity cost, the cost of people is actually manageable.”

5. Shareholders: Ask for Money, Not My Life

  • Strategic shifts inevitably bring one or two dissenting shareholders—for example, investors arguing that 群核 should stop doing R&D in China and move it to Silicon Valley. 黄晓煌’s principle is not to change a strategy he believes is right just to accommodate shareholder opinion: “A shareholder just wants your money, not your life.”
  • His approach to shareholders who do not share the company’s direction is unusually decisive: “Don’t drag it out. Whatever the cost, get them out as fast as possible. Any price today is still the smallest price compared with the future.”

6. The Methodology of Riding Hot Trends: Wherever There’s a Trend, Nvidia Is There

  • 黄晓煌 directly reverses the usual narrative: “There’s nothing shameful about riding a hot trend—you can see Nvidia doing it too.” He strongly disapproved of Nvidia’s move into crypto mining, although the host argued that the money made in that cycle gave Nvidia the resources and confidence to increase its AI investment. 黄晓煌’s rule is that companies should ride trends but create value for them rather than follow blindly. Smaller companies have to ride trends; only once a company is large enough can it choose its own trends or create them.
  • The ROI logic is blunt: a hot topic can generate hundreds of thousands of reads, while a cold one may get only a few dozen. “If you insist on pushing some obscure industry, that’s painful for me too.” The filter is values: he would still avoid mining, would try the metaverse—“a very cool idea; if it didn’t work out, that happens”—and never entered Web3. “I still want to work on things that create real value.”
  • The deeper insight is that trends do not appear from nowhere. Either companies or governments are pushing them, and they need others to help. Going with the flow means understanding where that support is coming from and contributing to it.

7. Swept Into the Water: Giants Disrupting the Market and Employees He Couldn’t Convince

  • Asked whether a wave had ever knocked him flat, 黄晓煌 corrected the wording: “You can’t call it being knocked down; I went into the water.” During the Kujiale 1.0 era, the core business had become strong, only for a giant to enter and disrupt the market. Many employees left because they believed, “How could a company this small compete with a giant?” He has seen enough storms to regard all of them as part of life.
  • He says maintaining employees’ confidence is harder than resetting his own mental state. During the 2023 push into spatial intelligence, some employees asked, “Why is the boss riding another hot trend? What does AI have to do with you?” One even came to his office and said, “Why are you doing this fluffy stuff? Can’t you just focus on the supply chain?” His response is to listen actively but not try to persuade: “A lot of things can’t be explained just by reasoning. People either trust you or they don’t.”
  • He ends with a self-deprecating line from Ma Yun: “If you can’t even change yourself, how are you going to change other people?” He tries to listen to even the harshest voices but does not change direction because of them. The 3 founders share similar educational and professional backgrounds, and “in all these years, there hasn’t been a directional disagreement.”

8. Rebuilding the Business Model in 2023: From Annual Plans to Token Pricing

  • ChatGPT’s first impression on 黄晓煌 was not AGI but business-model design. OpenAI charged by compute and by token, while 群核’s annual and monthly SaaS plans meant that “a lot of very useful but very compute-intensive capabilities couldn’t be launched; every time someone used them, I lost money.” The company had developed video generation early, but after realizing that a single video could consume several GPUs for an hour, “the CFO said it couldn’t be released. If we launched it, I’d lose money and the financial statements would be impossible to look at.”
  • The underlying philosophy flipped. SaaS had been measured by “users multiplied by the time they spend each day,” with the ideal user keeping the software open for 12 hours a day. “In the future, it may be that the less time I occupy, the more value I create.” New businesses are already priced by token; the legacy business is still turning slowly because “the ship is too big.”

9. The Four-Part Spatial Intelligence Stack and the “Turn the Light Off” Problem

  • 群核 takes 李飞飞’s three-part framework—understanding, reasoning, and action—and adds generation, creating 4 components: spatial understanding, reasoning, generation, and action. The core problem is illustrated by a simple example: many current robots are still doing image recognition, so “if you turn off the light, they think they’re in a different place.” If multiple robots hold inconsistent views of the same room, they cannot work together. “We need to build a system that gives robots this memory and this ability to understand.”
  • Generalization follows from that definition. Without generalization, a robot stops recognizing an object as soon as someone turns on or off a light. With generalization, it recognizes the same thing under different lighting and other conditions. 李飞飞’s papers, including Thinking in Space, have clearly helped 群核 organize its direction.

10. Necessity Driving Change: The Difficulty of 2022 and the Door Opened by ChatGPT

  • The timeline is candid. In 2021, the legacy business was booming, and “you don’t think about changing when things are going well; you change when you’re poor.” The capital markets were strong in 2020–21, 群核 prepared for a US IPO and hired aggressively, then the market structure, capital environment, and a series of major events forced a reset. The arrival of AI offered “a very good glimmer of hope” for choosing the next direction.
  • Asked whether he was grateful to ChatGPT, 黄晓煌 kept the edge in his answer: “Grateful is too strong. It suddenly opened some new doors for us.” Waves will always come; the question is whether to choose a ripple or a tidal wave. The large-model wave “has already lasted 3 years,” unlike the metaverse, which went quiet after 1 year.

11. Resolving Founder Disagreements: Turn Emotional Questions Into Math

  • The 3 founders have disagreed, but rarely argued. The questions were concrete: how many resources to commit, whether fresh graduates could do the work, and whether to buy GPUs in bulk or just enough. On new graduates, the conclusion was clear: “Of course they can. Everyone with experience has already been taken by ByteDance and Alibaba; you wouldn’t get them anyway. Play the cards you have. Don’t complain about your luck.”
  • Under intense pressure from the CFO, the company chose “just enough” GPUs. In retrospect, 黄晓煌 regrets it: “We could actually have bought more. Limited GPUs definitely constrained our imagination; some models could have been trained much larger.” But the decision framework remains: look at whether GPU utilization is continuously rising or falling, and “turn the discussion into a math problem, because math problems have answers. Emotional problems don’t.”

12. Why Physical AI: A Data-Led Market Under Compute Constraints

  • The strategic reasoning was direct. Large language models opened the door to “more data, more intelligence,” but “compute was constrained in China, so internally we didn’t pursue large models.” 群核 instead moved into a market where data is less constrained but scarce, and compute requirements are less extreme. Physical-AI data is harder to obtain, while the compute burden is comparatively manageable, so “we strategically shifted in this direction.”
  • The other hard constraint was not starting an unrelated second company. The new models had to “not only expand into new business but also empower the old business.” Physical parameters and material information had previously been handled manually; now a large model can infer those parameters from a photograph, creating a major efficiency gain for the legacy business.
  • By 2023, the company had identified 3 initial scenarios: robotics—then described as “intelligent devices,” with no decision yet on whether they would be humanoid—e-commerce, and Industry 4.0. Film, games, advertising, and “many other sectors” could also use the technology.

13. Academia Leads by 2 or 3 Years: Jensen Huang Is Following the Wave, Not Building the Future

  • 黄晓煌’s demystification of Jensen Huang deserves separate mention: “I don’t think he’s creating the future. Those things were already proposed by academia. You can’t see what academia proposes; you can see it only when a commercial company talks about it.” Robotics was hot in academia in 2022–23 but reached industry in 2024–25. “If you don’t read papers, all you can see is what’s in the public eye, and that’s 2 or 3 years, even 3 or 4 years behind.”
  • He has turned that lag into a commercial discipline. The host mentioned an anecdote that Hinton and his students once wrote to Jensen Huang asking for chips and were allegedly rejected. 黄晓煌’s response was, “I assume that’s what happened,” followed by the rule: “If scholars haven’t solved it, don’t go messing around with it yourself.”

14. Choosing the Base Model for SpatialLM: Asking 梁文锋 for a Small Model

  • SpatialLM was trained in 2 versions, based on Tongyi and Llama. The reason for not using DeepSeek was edge compute: “The base model is too large to run; even its smallest model is several dozen B, which is too big for robots.” 黄晓煌 directly asked 梁文锋 whether DeepSeek could release a small model, and was told that DeepSeek’s main mission was breaking the frontier, while small models were more about building an ecosystem. According to 周子涵, Tongyi performs best at small parameter counts; 黄晓煌 also recognizes Llama’s overseas ecosystem advantage: “The data we generate has to be understood by other models.”
  • Open source was a prerequisite for this work. Before large language models emerged, 群核 had tried to generate scripts but could not solve the problem: “You’re saying I should train a large language model just to train this little thing? I’d have to be insane.” The company is therefore both a beneficiary of open source and an active advocate of it. The host said SpatialLM quickly ranked 3rd on Hugging Face after open-sourcing, behind only Tongyi and DeepSeek. 黄晓煌 said there were “definitely no direct commercial returns for now.”

15. The World-Model Map: Video, 3D, and Academic Approaches

  • The key distinction is by route. Google’s Genie 3 is “mostly trained on video,” while 群核 enters through 3D, combining 3D with video models. “Our model size definitely can’t compare with theirs; we’re just providing some additional capability for a video foundation model.” On Nvidia Cosmos’s claim of 90 quadrillion tokens and 20 million hours of human-interaction data, 黄晓煌 said Cosmos was “very well done,” but focused more on engines and infrastructure—ultimately “selling more GPUs.” 群核 is also working with Nvidia.
  • His view of 李飞飞’s World Labs is politely skeptical: it is “closer to an academic result from a laboratory,” and he has not seen an obvious path to commercialization—“though that may just reflect the limits of my understanding.” Tencent’s Hunyuan 3D is built mainly on game data and a game aesthetic. 群核’s Tech Day demo used a single image of a real photography studio awaiting demolition to generate a realistic room. “We want to do things with technical foresight that can still serve real life.”
  • The 3 concepts can be separated this way: Jensen Huang’s physical AI emphasizes connection to the physical world; 李飞飞’s spatial intelligence focuses more on understanding space in the digital world; and world models emphasize video models that obey physical laws. In practice, they all point to the same problem: current large models understand the world very differently from humans.

16. The Blue Ocean Was Chosen After Giants Came Calling

  • The host noted that every major tech company is working on spatial intelligence. 黄晓煌’s correction is central: “For now, the major companies are still mainly exploring with very small teams. It’s not like large language models, where the entire company is mobilized.” The host summarized the market as a blue ocean rather than a red one.
  • The preference was learned through scars: “I’ve been hit by giants myself—I’ve been running from them ever since. Why would you go looking for a market packed with giants? That’s suicide?” The filter is to find, within your field of vision, something that a giant has little reason to pursue but that fits your own capabilities. 群核 could also build an image-generation model, but “if the giants are definitely going to do it, I won’t.”

17. The Flywheel: Tools Create Data, Data Trains Models, Models Improve Tools

  • 群核’s spatial-intelligence system consists of tools, data, and large models. “You can’t leave out any one of them. It’s a flywheel.” Training data cannot simply be scraped from the internet; it has to be generated by tools. Kujiale and virtual studios are such tools. The data trains large models, and the models make the tools more efficient.
  • The data goldmine came from reversing existing business workflows. 群核 accumulated data with physical parameters through businesses such as whole-home customization and Industry 4.0. In Industry 4.0, it had already built simulation systems to ensure that production matched design. The question behind InteriorNet was: “Industry 4.0 turns digital things back into physical things. If you see a physical thing, can AI turn it back into a digital thing?” The company had an initial proof of feasibility in 2021; after learning about scaling laws in 2023, it realized that larger training runs could produce greater intelligence and stronger generalization—something it had not anticipated.
  • The company is also trying to move beyond its existing data stockpile. SpatialGen, which generates an entire space from a single image—for example, using a picture of Hawaii instead of building a 3D model—is itself a data-generation tool and has been open-sourced. The longer-term goal is to let agents explore on their own and generate their own data.

18. The Commercial Reality of SpatialVerse: NDAs, Repeat Purchases, and Sim-to-Real

  • SpatialVerse’s one-line positioning is “synthetic data for spatial intelligence.” It originally served 群核’s internal model teams, then became a business line parallel to Kujiale once outside demand grew. Customers include embodied-AI companies, AIGC players such as Kuaishou’s Kling and Pika, and VR/AR companies. The host said he had heard that nearly every embodied-AI “brain” company was a partner; 黄晓煌 replied that some companies provide feedback and some do not. “If Google makes it public, we can talk about it. Everyone else is under NDA.” In the early days, the proof that the data was useful was straightforward: after a product launched, “you could tell at a glance that some products had been built using the synthetic data we provided.”
  • The revenue answer was disarmingly candid: “The revenue from robotics companies is generally pretty ordinary. Being able to get enough revenue from them to cover our investment is already very good.” The key metric is whether customers continue buying more data, and repeat purchases are “pretty decent.” The legacy business serves nearly 50,000 customers, while “having 1,000 customers in any industry today would already be good,” so the company wants the technology to generalize across markets.
  • The technical answer is equally direct: sim-to-real always has a gap. “You definitely can’t use data from only one source; you have to mix in real data during training for robots to generalize. Some people train a model in a virtual environment and find it unusable in reality because the dataset contains no real-world data.” 群核 is not working on VLA for now because it is heavily tied to hardware; even different humanoid robots may not be able to run the same VLA. “Why would I train that?”

19. Video-Generation Companies: Complement or Direct Competition?

  • 群核’s differentiation is spatial consistency. “Sometimes a video-generation company changes the camera angle and the whole room changes, which violates physical laws. We can keep the room stable.” The host pointed out that this was not complementarity but something 群核 could do that the video companies could not yet. 黄晓煌 agreed, while maintaining that the future should be cooperation rather than competition. Video models are more efficient and cheaper for fantastical animated stories; 群核 itself calls models such as Kling and Hunyuan. “Use whatever you like. Everyone has different data, so the capabilities they train will be different.”
  • Virtual studios are an early commercial example. In e-commerce, the product stays the same while the environment and background change. Tech Day demos simulated dusk lighting and the tracking light of a flat advertising shot, prompting applause; e-commerce has provided the strongest industry feedback so far. The underlying philosophy is simple: “The real world is not computable. You have to move it into the virtual world before you can process and optimize it.”

20. Talent: 9x More Resumes, Interviews Until Dawn, and Fresh-Graduate Leaders

  • The Six Little Dragons effect is visible in the pipeline: resumes from C9 universities are up roughly 9x this year, and overseas-returnee resumes are 20x last year’s level. During the property downturn, “you barely dared to ask for the salary they were offering,” and many projects were delayed indefinitely because there were not enough experts to go around. This year, the results in the new directions “have all been delivered by newcomers.” 黄晓煌 spends more time recruiting than meeting customers: “In the end, everything that stays is created by people.”
  • His interview process is distinctive. He does not focus on resumes; he gives candidates a paper to read, asks them to explain it, and then assigns a problem to implement on the spot. The exercise is open-ended and untimed. One candidate worked from 9 a.m. until 1 or 2 a.m., and “I stayed with him”; that attitude counts in the candidate’s favor. His AI-era talent map is polarized: highly experienced people remain valuable, while experience that can be found by searching DeepSeek is under greater pressure. Those in the vague middle are in the most awkward position.
  • 群核 borrowed its hiring model from Nvidia. The leader of the CUDA team may have been a Stanford fresh graduate, and “one-third or even half” of the team may have come from Stanford. “Would you use a fresh graduate to head a new business? It’s hard to imagine.” 群核 is now trying to let people only a few years out of school lead important projects, and has begun experimenting with it.

21. The Riskiest Year: 2021 and Never Running Out of Ammunition

  • The company came close to a major failure: “In the first half of 2021, everything seemed to be exceeding expectations. We hired like crazy—nearly 3,000 people at one point—and growth was rapid. Then everything suddenly changed.” Property, the pandemic, and capital markets all turned at once. “I thought that if I had been any more reckless then, I probably would have been finished.” The rule afterward was a break-even floor: “Even when I burn money, I don’t want to lose money.” He saw other companies burn through their ammunition and enter an unrecoverable situation.
  • The lesson still governs today. Even while highly bullish on embodied AI, “I still won’t burn money recklessly.” When investors push him to spend more for faster growth, he says, “We’ll grow steadily. I’m not burning.” He personally reviews every role and “will not let standards slip.” Delegating hiring in 2020 produced lasting regret; hiring the wrong person harms both sides, because someone may spend 2 or 3 years trying to build the business before being told they are not a fit.

22. INTJ, the Trap of Digging In, and the Comfort of a Makeshift World

  • 黄晓煌 describes himself as INTJ, with the motto “Life is tough, but we are tougher.” He admired Bill Gates when he was young; now he thinks Jensen Huang is more impressive. Nvidia’s underlying judgment—that GPUs would replace CPUs—stayed consistent for years, while its target industries kept changing. “At first I was half-convinced. After watching him rise all the way, I realized the method worked.”
  • His rare self-criticism is that toughness makes him prone to digging in, which conflicts with going with the flow. After the government changed course on the property business, he kept pushing that line and only later realized he was spending his energy on the wrong things, working every day on low-ROI projects with little meaning. He asks his partners to remind him whenever he gets stuck: “Working harder doesn’t guarantee that you’ll get it done.”
  • He recommends Breaking Twitter for its outsider’s view and demystification of leaders: “Even the greatest leader in the world will mess things up sometimes.” There is no need to fear the behavior of a makeshift operation; taking action is always better than taking no action. Asked how many billions 群核 is worth, 黄晓煌 replied only, “It probably is worth something.” The host added, “Quite a lot,” and left it there.