Pioneers Insight Method Research Author
139: ICCV Best Paper, Light Year Beyond, Sand.ai: Cao Yue’s Decade in AI, from Researcher to CEO
Back to Episodes

139: ICCV Best Paper, Light Year Beyond, Sand.ai: Cao Yue’s Decade in AI, from Researcher to CEO

Summary

  • Cao Yue sums up OpenAI’s methodology as designing a scalable system to “squeeze compute to the maximum,” rather than piling on tricks around paper novelty. That requires an organization to evolve from a small author team into a system where data, training, infra, evaluation and product work together; his experience with Swin Transformer also convinced him that spotting an opportunity is only half the job—having the nose to detect it and the organizational ability to actually seize it are both essential.

  • He relayed Wang Huiwen’s “three-character” explanation for why China historically lacked an OpenAI-style organization: “not rich enough.” In the catch-up phase, the highest-ROI opportunities were efficiency innovation, business-model innovation and copy to China; once near the frontier, the direction becomes less clear, requiring capital, entrepreneurs and society to recalibrate their tolerance for failure around original exploration. Cao Yue believes DeepSeek, Unitree, Black Myth: Wukong and Ne Zha between 2023 and 2025 have already signaled this transition.

  • Sand.ai paid a technology-first tuition fee with Magi-1: a 30- to 40-person early team bet on an autoregressive video route with no reference implementation, and only released the model roughly 13–14 months later. The model drew feedback from the tech community, offline discussions during ICLR and GitHub stars, and brought in hundreds of thousands of registered users, but produced no clear commercial metrics; Cao Yue’s review is that a startup must first establish PMF and its first curve—“overemphasizing technology can make it very difficult to achieve fit between technology and demand.”

  • The core judgment behind Sand.ai’s next-generation model is not to invent another architecture, but to address video’s most rigid gap first: people speaking and performing. Existing models are good at empty establishing shots, transitions and other B-roll, but struggle with A-roll dominated by people; post-production lip sync cannot get characters across the uncanny valley, and the most common feedback from short-drama customers is “the people look fake; the actors aren’t performing.” Sand.ai is therefore focusing on people speaking and performing with synchronized audio and video, while treating consistency in appearance, voice and related attributes as a major capability problem.

  • Sora 2 triggered a key update in Cao Yue’s thinking: storytelling and shot composition within 10 seconds may not require an agent pipeline and can instead be produced end to end by the model. Synchronized audio and video turns a silent clip from a half-finished asset into directly consumable content, Cameo embeds identity, voice and relationships into the product, and narrative ability improves distribution; what truly surprised him was not any isolated feature, but the vertically integrated mechanism through which product demand feeds back into the model team.

  • The near-term commercial opportunity in AI video still looks more like a productivity tool; the real consumer-platform opportunity remains unproven. An early live-action short drama ran roughly 60–100 minutes and cost about RMB100,000, before costs had risen to around RMB300,000–400,000 around 2023; AI production costs about RMB2,000–5,000 per minute, or roughly RMB200,000 for a full title, but cannot yet compete on quality. The Sora App still relies on existing distribution channels such as WeChat Moments, Xiaohongshu and Douyin, and Cao Yue believes a platform needs either a new content form or a new distribution chain to truly exist; retention matters more than launch-period heat.

  • Low resolution, speed and cost may matter more than model leaderboards in determining whether a consumer product can scale. The interview noted that 360p may already be sufficient for consumer use, while Sora still takes roughly 1–3 minutes to generate a video and its free allowance has fallen from 100 clips per day to 30; Cao Yue believes the experience would be completely different if generation could be compressed to 10 seconds. Synchronized audio and video also changes the quality trade-off: dropping the image from 1080p to 360p does not cause sound quality to fall in tandem, making immersion more resilient to compression than in a purely visual model.

  • Cao Yue now sees timing and organizational readiness as the hardest variables—and the ones with the greatest investment significance. The broad direction from AI video as a tool toward consumer use is relatively easy to identify; the real non-consensus question is “when, and what opportunity will appear?” Even if an opportunity is correctly identified, an organization that has not prepared the ability to make joint decisions across models, product and operations will fail to capture the window. “The opportunity appeared, but the organization lacked the ability to seize it” is itself a failure of timing judgment.

Deep dive

1. Betting on deep learning in 2014 was Cao Yue’s first attempt to follow a faint signal against consensus

  • Cao Yue’s first sense of an AI moment came in his senior year of college in 2014, when he decided to work on machine learning and caught the deep-learning wave. At the time, senior professors in China still wrote on their homepages that they “did not do neural networks with more than two layers,” while his mentor visited Berkeley and found that Silicon Valley was already talking about deep learning everywhere.

  • His research group became one of the earlier teams in China to buy GPUs and begin training deep networks, even though “there were very few cards” at the time. In 2017–2018, Cao Yue first interned with and then joined MSRA’s computer-vision group, continuing to work on computer vision, multimodality and pretraining; the group had previously given rise to ResNet, influencing him not only technically but also in how to choose problems and build organizations.

2. The best problems must be both important and far from convergence

  • The first principle Cao Yue learned at MSRA was to “work on the most watched topic”; attention alone was not enough, however—the problem should also be among the field’s most important and have clearly substantial room for progress.

  • Manqi pointed out the tension: when a direction is already hot, that often means the opportunity has moved later in the cycle. Cao Yue’s answer was that judgment comes from intuition formed through long immersion—what looks externally like “a very subtle foundational signal” may already be “an extremely strong signal” to someone who truly understands the industry.

  • This is not simple trend-chasing: researchers must identify structural opportunities and judge whether they have already converged. Whether with Swin Transformer, his move to the Beijing Academy of Artificial Intelligence or his decision to pursue AI video, Cao Yue has consistently applied the filter of high attention, a high ceiling and an early stage.

3. Transformer entered vision by first fixing convolution’s limitations, then replacing the entire backbone

  • After Transformer appeared in 2017, the first phase of exploration in vision remained centered on convolutional networks: researchers embedded Attention Blocks into CNNs to expand the receptive field constrained by local convolutions at relatively low cost.

  • Cao Yue’s team also tried in 2018–2019 to let attention directly replace convolution, but the macro-architecture still followed ResNet; what changed was the local operator, not the network paradigm that computer vision had built over many years.

  • ImageGPT in June 2020 took a more “brute-force” route, applying self-attention directly to pixels, with very low computational efficiency and unremarkable results. Cao Yue recalls, with the caveat “if I remember correctly,” that many people read the paper without truly understanding why OpenAI had taken that approach.

  • ViT in October 2020 changed pixels into roughly 16×16 patches and delivered good results on ImageNet classification. For Cao Yue, the core technical change was turning pixels into patches; more importantly, it showed that vision research “should no longer be constrained by the structures iterated out of convolutional neural networks.”

4. Swin Transformer targeted not a classification title, but a replacement for the entire visual foundation

  • ViT had shown that Transformer could handle image classification. Cao Yue’s team then asked whether it was possible to design a macro-architecture suited to most vision tasks, covering classification, object detection and semantic segmentation as well as fine-grained tasks.

  • This was not point optimization of a single metric, but a replacement for the basic unit of representation learning. If a new backbone performed well enough across task categories, it could replace the entire convolutional-network structure represented by ResNet.

  • Cao Yue clearly felt that this was a “big opportunity”: network architecture was one of deep learning’s most watched topics and a foundational unit capable of affecting every vision task; ViT had already opened a conceptual gap, while the team happened to understand classification, detection, pretraining and Transformer.

5. Swin’s outcome depended not only on the idea, but on putting the entire organization behind it

  • Cao Yue’s team was not the only one to see the opportunity. What created separation was their decision, once they judged the opportunity large enough, to bring in everyone in the group who could contribute and push the work “to the extreme” on task coverage, metrics and completeness, producing an outcome of a different order of quality.

  • The experience left Cao Yue with a dual conclusion: “you need the ability to sniff out the opportunity,” but also “the organizational ability to actually seize it”; attention without solid execution cannot stand the “test of time.”

  • Manqi noted that the two abilities can conflict: people who constantly scan for new directions may not want to go deep, while those who are extremely rigorous may find that “the wave has already passed” by the time their work is complete. Cao Yue has also seen friends whose work was consistently high quality but who repeatedly missed windows by digging too deeply.

6. After winning ICCV Best Paper, Cao Yue saw the ceiling of the paper-driven paradigm

  • Around 2021, Swin Transformer won ICCV Best Paper. During this period, Cao Yue experienced a larger mindset shift: even if one could produce a top-tier paper, the ceiling of that methodology was still clear; DeepMind could build AlphaFold, and OpenAI could build systems of an entirely different scale.

  • DALL·E 1 and CLIP in early 2021 shook him in particular. Many people around him responded that the compute was inaccessible and the work was “untouchable,” and stopped studying it; Cao Yue believed that if one could “put aside one’s ego” and examine the work seriously, it was clear that its objectives, organizational form and way of working differed from those of a traditional paper team.

  • His central question shifted from “what should the next paper be?” to “why can they make something this impressive, and where exactly are we falling short?” That directly pushed him to leave MSRA and join the Beijing Academy of Artificial Intelligence.

7. Paper-driven organizations optimize author order; scalable systems optimize total output

  • Cao Yue believes Chinese research groups at the time were generally driven by papers. The author list naturally introduced a distribution of interests among first, second and third authors, which meant less encouragement for cross-role collaboration; MSRA was already trying to adjust, but organizational inertia could not disappear quickly.

  • Reviewers often asked about novelty, so teams optimized for “can we propose something new and tell a fresher story?” Many OpenAI projects were technically simple by comparison because they did not begin by searching for tricks; they began by building a system that could scale with compute.

  • Cao Yue compressed the methodology into one sentence: “How do you design a scalable system so that it can squeeze compute to the maximum?” It differs from the traditional research route of adding human priors around a single task.

  • Training a large system also requires gathering and cleaning data, training, infra, evaluation and even final PR; members must care about the system together rather than paper credit. Cao Yue only later realized that this organizational form was “essentially a startup.”

8. The pandemic widened the cognition gap, while BAAI offered China’s closest approximation to an OpenAI testbed

  • From 2020 until ChatGPT’s release, the pandemic reduced China–US technical exchange, while academic conferences and cross-team communication shifted online for an extended period. OpenAI’s momentum was already building, but people without on-the-ground information were more likely to remain trapped in their own “knowledge bubbles.”

  • Cao Yue judged BAAI to be one of the earliest institutions in China to embrace large models and OpenAI’s methodology: it did not need to make paper publication its core metric, could select people with an entrepreneurial mindset and jointly build complete systems similar to DALL·E and CLIP.

  • By mid-2022, BAAI had a cluster of roughly 1,500 A100s connected together; at the time, clusters with more than 1,000 GPUs in China were “extremely, extremely rare.” Its loose research environment, relatively large compute base and open-source, open-access mission made it, in Cao Yue’s view, “the organization in China most like OpenAI,” though the resemblance remained limited.

9. ChatGPT released accumulated technical momentum in an instant and rewrote the entry conditions for entrepreneurship

  • After ChatGPT appeared, the OpenAI momentum that had not been fully understood because of communication barriers was suddenly released. Sparks of AGI in early 2023 showed people the initial sparks of AGI while also confirming that the models still had substantial problems and room to improve.

  • Cao Yue does not believe the opportunity belonged only to traditional NLP practitioners: OpenAI’s methodology had already crossed the old boundaries between computer vision, language and multimodality. The key was whether one could build a scalable system, not which task one had worked on before.

  • Participating in the competition required far more than knowing how to train models: a company also had to solve frontier exploration, data, compute and financing, productization, strategy and organization. Cao Yue’s self-assessment at the time was that he could build a team to train models, but still lacked many of the capabilities entrepreneurship required.

10. Wang Huiwen reached talent through referrals, while Liang Wenfeng also proactively sought out Cao Yue

  • After ChatGPT appeared, Wang Huiwen issued a “call for heroes,” which Cao Yue believes rapidly heated up the domestic landscape. Wang used a snowballing referral process to meet talent: after speaking with one person, he would ask that person to recommend 2 or 3 of the most worthwhile people to meet in the relevant field.

  • From the first meeting to ultimately joining Light Year Beyond, Cao Yue took only a few weeks. His first impression of Wang Huiwen was that “you could clearly feel this person was extremely strong,” with both extensive practical experience and the ability to distill a methodology worth repeated thought.

  • Cao Yue recalls that in March 2023 he spoke only with Wang Huiwen and Liang Wenfeng. When Liang approached him, Cao Yue had already agreed to join Wang, so they did not continue working together.

  • Manqi asked whether Wang Huiwen had also met Liang Wenfeng at the time. Cao Yue said only “I think so,” explicitly retaining uncertainty; he also remembered that Liang wanted to build a research organization in China that would not face strong commercialization pressure over the long term.

11. “Not rich enough” explained the catch-up era and pointed to a system change for the originality era

  • Cao Yue asked Wang Huiwen, “Why has China never produced an organization like OpenAI?” Wang quickly answered that, given China’s internet companies and stage of development, “we were not rich enough.”

  • Cao Yue later understood the phrase to mean that the catch-up phase offered a clear target: companies only needed to approach the leader faster, with high certainty and the highest ROI in efficiency innovation, business-model innovation and copy to China; the farther behind the frontier, the stronger the sense of direction.

  • As companies move closer to the frontier, the leaders themselves are exploring, and the sense of direction weakens. Capital’s willingness to fund original ideas, entrepreneurs’ ability to withstand repeated failure, society’s tolerance for failure and the exit process for failed companies all need to change across the entire chain.

  • Looking again in 2025, Cao Yue believes representative moments such as Ne Zha, Black Myth: Wukong, DeepSeek and Unitree are beginning to appear. They do not prove that the transition is complete, but they make him certain that China is in transition and that similar events may become more frequent.

12. Light Year Beyond’s core assets were talent judgment and a visceral sense of CEO pressure

  • Light Year Beyond valued talent who had graduated within 3 to 5 years or were about to finish a PhD: still on the front line, at a capability peak and learning quickly. The team cared little whether candidates had previously worked in NLP, computer vision or another field, because language models could be brought to a good level within a few months.

  • Cao Yue remembers that DeepSeek was the fiercest competitor for this talent profile. The same profile later carried over to Sand.ai; he believes some organizations did not begin hiring with a similar mindset until 2024 or even later.

  • Light Year Beyond’s external ending was sudden, but not so for Cao Yue. He acknowledged feeling disappointed that it could not continue, but he and Yuan Laoshi first had to fulfill their responsibilities and help employees and the organization complete a “comparatively smooth transition”; the work itself kept his emotions in check.

  • Another lesson was that “the pressure on a CEO is extremely, extremely high, so you have to take care of your body.” Cao Yue uses thoughts of death to shrink everyday anxiety; after DeepSeek exploded in early 2025, he saw Liang Wenfeng remain calm, then returned and reduced noise channels such as Moments, putting his attention back on fundamental questions.

13. Entrepreneurship was the first time Cao Yue felt that his capability profile “all fit”

  • After having relatively more free time in August 2023, Cao Yue broadly compared joining another company with starting independently. He realized he was not the typical researcher willing to work alone until he had exhausted a narrow problem; instead, he had long been interested in fields, organizations, people and opportunities.

  • He summarized his deepest self-awareness as ambitious—not simply wanting to win, but wanting to move ever closer to something capable of having a major impact on the world. Borrowing Munger’s formulation, the path to extraordinary achievement is to “make yourself worthy of the achievement.”

  • When he saw CLIP and DALL·E 1, his first reaction was not to dismiss the work because he lacked compute, but to ask: “Why can’t we do this?” The answer would not simply be “you can’t” or “you’re not good enough”; it required finding a better way to organize and execute.

  • Entrepreneurship requires relatively comprehensive capabilities, has a high ceiling and offers many things one might build, while representing “hell mode” for the individual. This made Cao Yue feel that his previously scattered tendencies had finally aligned. As for why he did not join OpenAI earlier, he admitted that his understanding was insufficient at the time; after ChatGPT, joining would no longer have been the same non-consensus choice.

14. AI video in August 2023 combined an early stage, a high ceiling and commercializability

  • Cao Yue compared continuing to train language models, agents, Character AI-style directions and AI video. Before Sora appeared, AI video was still embryonic, with weak model capability but clear technical upside, matching his consistent search for something highly watched with substantial room remaining.

  • The commercial ceiling was equally high: each newly unlocked model capability could unlock new creative scenarios and demand, while the iteration cycle was long enough. Competition was only one evaluation dimension, not the main axis of the startup decision.

  • Cao Yue believed that starting another general-purpose language-model company in August 2023 would already be late. The other companies he encountered were difficult to surpass Light Year Beyond in overall completeness across models, infra, product, commercialization and financing, making entrepreneurship more attractive than joining an “incomplete setup.”

15. Magi-1’s autoregressive bet had intuitive appeal, but execution exceeded the early organization’s capacity

  • Early autoregressive image and video research sought to unify language, images and video in one model, but exploration was shallow and results were weak. Sand.ai’s starting point was that video naturally plays over time, much as language is read sequentially, so the most effective way to compress information might likewise be to predict the next token.

  • Cao Yue still believes the intuition “should be broadly right.” The real problem was the lack of a good reference implementation. Algorithm, infra and code design all had to be built from scratch, concentrating the innovation burden in a newly founded company.

  • In a 30- to 40-person team, nearly all of the key people were absorbed by the autoregressive route, while data, training workflows, evaluation and post-training also required sufficient investment. To realize the route’s advantages, every dimension had to be good enough, which an early team struggled to cover simultaneously.

  • The company was founded in early 2024 and released Magi-1 around April 2025, a process of roughly 13–14 months. Cao Yue’s review was that he had underestimated “how difficult it is for a newly formed organization to do work that is overly innovative”; if the initial mandate had required a usable model within 6 months, the team’s mindset might have been different.

16. Magi-1 won technical attention without establishing the company’s first curve

  • After its release, Magi-1 received substantial praise in technical circles, generated considerable offline discussion during ICLR and accumulated many GitHub stars; the technical spread also brought in hundreds of thousands of registered users.

  • But the model did not produce clear commercialization metrics. When Manqi asked whether this had diverged from the original expectation, Cao Yue acknowledged that he lacked a sufficiently deep, closed-loop understanding of the business at the time: he wanted commercial value but had not truly connected users, demand and monetization.

  • The feedback pushed the company toward a clearer PMF orientation: a startup must first establish its first curve, while overemphasizing technological advancement can prevent technology from fitting demand; letting immediate demand drive everything, however, can leave the technology behind.

  • Cao Yue defines the challenge as balance: maintain the necessary frontier edge in the model while ensuring that each iteration has business meaning in the present. People with product and commercial backgrounds must learn model know-how, while people with model backgrounds must learn product, operations and industry.

17. Pure visual generation is nearing convergence, while synchronized audio and video has reopened the capability curve

  • Cao Yue believes pure video generation with a single asset and a single shot has entered a relatively converged phase; that does not mean video models have no window, but the marginal differences in old capabilities are increasingly hard to turn into new scenarios.

  • The new variable this year is synchronized audio and video, along with real people, performance and narrative. In the past, a silent 5-second clip was only a half-finished asset for ordinary users and was difficult to consume directly, requiring scripts, voiceover, music and editing to complete it.

  • Once the model generates image and sound together, the individual clip itself becomes consumable; person ID also expands from preserving appearance alone to preserving both appearance and voice. This is the common foundation for Sand.ai’s new model and its potential on the consumer side.

18. Sora 2’s real increment is forming a consumable narrative within 10 seconds

  • Cao Yue breaks Sora 2’s capabilities into three layers: synchronized audio and video; preserving person ID given a photo and voice; and forming basic narrative through shot changes within roughly 10 seconds.

  • Synchronized audio and video was not entirely new—Google Veo 3 had similar capability in May that year—and character consistency had long been studied by image and video models. What impressed Cao Yue most was the third point, because multiple shots no longer merely cut between one another but form a perceptible narrative relationship.

  • Complete person ID therefore acquired a new meaning: previously it preserved facial features and appearance; with synchronized audio and video, it must also preserve voice. Manqi added that height, body type and the relative proportions between two people also affect whether someone “looks like” the intended person, which Cao Yue agreed remained a dimension requiring alignment.

  • Earlier models could also cut between multiple shots within 10 seconds, but struggled to control the relationships between shots, let alone make ordinary people want to consume the result. Sora 2 optimized existing shot-switching ability into narrative ability, producing a clear quality jump.

19. Moving narrative from an agent pipeline into a single model was a clear update to Cao Yue’s thinking

  • The more natural approach in the past was to have a language model write the script, a storyboard model design the shots, an image-generation model create the visuals, an image-to-video model render them, and then add voiceover and music; agents could replace the human labor in parts of that process.

  • OpenAI’s question, however, was: “Why can’t the model produce this narrative capability directly end to end?” Cao Yue acknowledged that before seeing Sora 2, he had not realized narrative should, at the current stage, be built directly into the video model.

  • The route has a prerequisite: output from a silent model remains production material, while synchronized audio and video turns it into content consumers can watch directly. Once the capability boundary changes, the old anchor that “the model is only one step in the workflow” must also change.

  • Cao Yue believes expanding from 10 seconds to 20 or 30 seconds “may not be technically that difficult,” but he retained another possibility: the system may still have a language model write the script internally. The fact that prompts become worse when users specify too much also suggests that the model may be better off improvising.

20. OpenAI’s advantage is allowing product demand to feed directly back into the model team

  • Cao Yue sees Sora 2 as the result of vertical integration: product first defines the need for a “10-second narrative short,” after which the model team jointly aligns on what narrative means, establishes a benchmark and delivers the capability through data, training and model optimization.

  • Traditional product development would prioritize combining existing models. Cao Yue speculates that OpenAI’s organization may instead be more inclined to first ask whether a problem can enter one model end to end, though he is unsure whether this was a founding principle of OpenAI. If the direction is feasible, the organization can preserve exploratory branches and release them to the product once mature.

  • Similar capability appeared after GPT-3: the foundation model’s capability boundary was visible, and the next step was to use InstructGPT, SFT, RL and post-training to align the model with an interface ordinary people could use easily. Cao Yue believes that likewise reflected vertical integration between product sense and model capability.

  • He declines to draw conclusions about Sora 2’s specific architecture. The reports are more ambiguous than for Sora 1; it is relatively clear that the base model “should be diffusion,” but there is no external consensus on whether it is bidirectional or unidirectional diffusion, or whether it is an autoregressive model. Narrative may not depend on some “earth-shattering” secret idea either.

21. Cameo turned model capability into a social loop, while Meta Vibes looks more like capability stitching

  • Cao Yue describes Meta Vibes as “seeming to have everything,” but more like vertically stitching an organization’s existing capabilities into a product. It lacks a distinct value proposition and a sell point users will remember.

  • The Sora App’s chain is more organic: synchronized audio and video makes content consumable, person ID lowers the creation barrier, and Cameo connects relationships with friends, co-creation, @ mentions, invite codes and remix within one product experience. The invite code naturally creates a following relationship between friends.

  • Manqi noted that early content would likely feature public characters such as Ultraman Zero, with real friends brought in later. Cao Yue believes OpenAI had concerns about deepfake risk but made trade-offs: some celebrity likenesses were not comprehensively banned, and this type of content is genuinely highly distributable.

  • Cao Yue speculates that model training inevitably used some types of copyrighted data, while multimodal labeling taught the model to recognize characters; this is his explanation, not a confirmed fact. Whatever the source, low-friction ID, 10-second narrative and decent quality together created a distribution event.

22. The next-generation model started with people by working backward from user pain points

  • In setting the direction for its next-generation model, Sand.ai chose synchronized audio and video, but the priority was not to cover every sound-related scenario; it was to make people speaking and performing sufficiently realistic.

  • Cao Yue categorizes the old model’s strengths as B-roll: empty establishing shots, transitions and visual material. A-roll in narrative content, by contrast, is made up largely of people, who usually need to speak and express emotion; generating the image first and adding lip sync later looks “very strange and very fake.”

  • In conversations with short-drama and AI-narrative-content practitioners, the most common primary bottlenecks were “the people look fake” and “the actors aren’t performing.” Cao Yue estimates that people occupy more than half of the frame in ordinary video, making this the most frequent demand and the gap most worth solving first.

  • The result also shows that this is not an industry-wide consensus: not every team prioritizes character performance. Cao Yue believes the difference comes from whether a team starts with demand and then combines technical judgment to find fit between demand and model, rather than continuing to optimize only against model leaderboards.

23. AI short dramas have not yet achieved cost dominance; the quality bar matters more than unit price

  • Cao Yue’s rough industry figures are that an early live-action short drama cost around RMB100,000 for a 60–100-minute title; around 2023, costs had at times risen to roughly RMB300,000–400,000, possibly higher, before falling back somewhat recently.

  • AI short dramas currently cost about RMB2,000–5,000 per minute, with a full title potentially around RMB200,000. They have not created an order-of-magnitude cost difference from live action, while generation quality still cannot truly compete, so AI short dramas with a live-action bias have not crossed the consumer threshold.

  • Sand.ai hopes to lower model costs by focusing on scenarios while improving character performance first. If a short drama has 60 or 100 episodes, the lead characters’ appearance, voice and body shape must also remain consistent over time; these high-frequency hard needs determine the priority of model features.

24. The same character capability can serve professional creators and become a “video battle” tool

  • The first target group is creators of short dramas, narrative content, advertisements, promotional films and performance-marketing assets. If characters become more expressive, the capability is not limited to cinematic narrative; any segment featuring people, including explainers and marketing, could benefit.

  • The second group is ordinary users. Sand.ai found internally that a single image, a line of dialogue and an emotion description are enough to generate a performed video; in group chats, it behaves like a customizable animated sticker, upgrading users from “image battles” to “video battles.”

  • The company has also held internal video-battle competitions, with categories such as “Most Cinematic” and “Definitely Going Viral.” Cao Yue therefore believes the interaction barrier is already low enough for natural social distribution, rather than being a concept assembled at the last minute after Sora 2 launched.

  • The first release on October 11, 2025 was a Web product, allowing more people to experience the model’s capabilities; the more complete consumer form will likely be an App, but its positioning and model–product design still need refinement and cannot simply copy the surface form of the Sora App.

25. Consumer video should first optimize resolution, speed and budget—not professional image quality

  • Cao Yue noted that for a long period users could generate only roughly 360p video. That was unacceptable to professional creators requiring at least 1080p, but might be enough to support casual mobile use. Sand.ai’s internal discussions also concluded that 270p or 360p could work for WeChat sharing.

  • Synchronized audio and video changes the cost of degradation: dropping from 1080p to 360p sacrifices only visual clarity, while sound does not deteriorate in parallel. If sound is an important part of immersion, low-resolution content may remain consumable.

  • Sora takes roughly 1–3 minutes to generate a video, which Manqi considers slow. Cao Yue agreed that if generation took 10 seconds, user behavior would be completely different from waiting a minute, and creation frequency might rise by an order of magnitude.

  • The compute constraints of a free product are already visible: the individual daily allowance has fallen from an initial 100 clips to 30. Cao Yue believes the cost is painful but may not be the biggest problem; if retention holds, many large consumer products can find a monetization path later.

26. The Sora App currently looks more like a tool; retention must prove platform status

  • Even if large companies assign a new platform only a 10%–20% chance, they will invest because the cost of missing it is too high; their cost of capital is low, allowing them to “charge when they see an opportunity.” Startups, by contrast, must judge more precisely whether this is truly a platform window.

  • Cao Yue gives two direct conditions: whether a new content form appears, and whether a new distribution channel appears. Today, Sora’s 10-second videos are still frequently reposted to WeChat Moments, Xiaohongshu, Douyin, WeChat Channels and Kuaishou, showing that existing platforms still control the consumption chain.

  • Without new content or new channels, Sora’s value is more tool-like. It has the beginnings of a community but lacks mature creator incentives; interaction remains mainly likes and remix notifications, and is constrained by the generation budget.

  • Cao Yue watches retention most closely: a short-term influx of users can help the team find the right audience faster, but cannot prove a long-term hard need. His conclusion on whether Sora 2 can become a new consumer platform is that “no one has the answer right now.”

27. Sand.ai’s vertical integration begins with rebuilding context

  • This year Cao Yue shifted substantial time from pure algorithms to product and operations, gradually filling in the product-development, operations and model systems. The organization still has only a few dozen people, so the change lies less in complex reporting structures than in communication density.

  • The company arranges one-on-ones and joint discussions among key model and product members, with Cao Yue serving as the information-distribution hub. The aim is not for product to “submit requirements” to algorithms, but for both sides to share demand, model capability and capability boundaries.

  • Product may simultaneously have requirements A, B, C, D and E. The model team must judge which are solvable now, which require only a small iteration and which must wait for the next-generation model. Only with aligned context can priorities avoid being distorted at departmental interfaces.

  • Cao Yue defines vertical integration as model people increasingly understanding product and operations, while product people understand the model’s current state and development trajectory. Freer information exchange allows different functions to work together more closely.

28. A model-level product does not remove product from the picture; it lets model and product amplify each other

  • Cao Yue believes “model-level product” was initially understood as telling product and operations not to over-polish, but to maximize the display of model capability. That made sense during rapid foundation-model evolution, but is not the endpoint of maturity.

  • Cameo is a fuller example: the model supplies consumable content and character insertion, while product design around relationships with friends, invite codes, @ mentions and co-creation lowers the barrier. Product does not overpower the model; it amplifies the model’s most distinctive feature.

  • Sand.ai’s new model likewise no longer aims to “propose some new technology,” but to make “people speaking and performing sufficiently realistic”; consistency in appearance, voice, body shape and long-form content all emerged as priorities from the collision between product demand and the model team.

  • Cao Yue contrasts Claude Code with Cursor: when model and product are controlled by the same organization, capability, cost and experience can be adjusted for a specific feature, even creating interactions that other products cannot replicate; a product that merely calls an API cannot achieve the same depth.

29. The new model’s launch will not presuppose a winner; real retention will identify the user

  • The Web product will initially be opened as broadly as possible, then observe whether the users who remain are professional creators, AI short-drama teams, advertising users or consumer meme-makers. Cao Yue does not want to assume a single persona in advance, because the true fit may lie somewhere between A and B.

  • Direct feedback from small-scale internal testing is that the realism of people speaking and performing is very high. Cao Yue believes this dimension is at least comparable to Sora and has a chance to perform better; this remains a judgment from the company’s internal testing stage.

  • Compared with Sora 2, he phrases the claim cautiously: synchronized audio and video for people is “at least about the same” and “has a chance to perform better”; the clear gap is the lack of 10-second narrative and automatic shot-composition capability. The program preserved both the advantage claim and the capability boundary.

30. A professional CEO’s job is to break industry trends into an executable cadence

  • Cao Yue drew on an around-2020 presentation by Li Auto titled “How to Be a Professional CEO”: first define the industry, trend and stage, then identify current pain points, solutions, target users, required financing, timeline and company strategy.

  • The framework matters because a CEO cannot simply “charge ahead” on intuition, enthusiasm or existing capabilities. Faced with an event such as Sora 2, the CEO must first decide whether it is noise or structural change, write down the implications, ask different people for their judgments and then decide what the company should do.

  • His highest priority now is to understand the biggest opportunity in AI video over the next period and turn Sand.ai into an organization capable of capturing it. The broad trend from productivity tools toward stronger consumer capabilities is relatively consensual; timing is the real non-consensus question.

  • “The opportunity appeared, but the organization lacked the ability to seize it” means the cadence itself was mishandled. A professional CEO must judge when to prepare the model, when to add product capabilities and when to establish channels, rather than forming a team only after market signals become completely clear.

31. Gemini 2.5 Pro has become a research partner and the organization’s “context translator”

  • Cao Yue again felt the power of top models in May and June this year: during a team discussion of an algorithmic problem, Gemini 2.5 Pro filled in a part researchers had missed and proposed a solution that was “actually fairly credible.” The team jokingly calls this experience “web research.”

  • He found the model especially good at breaking down analogies: people associate event A with event B, while the model can systematically lay out their similarities, differences and reasons. Cao Yue also feeds in scattered thoughts all at once and asks the model to structure them, sometimes discussing them for 1 or 2 hours a day on average.

  • This does not mean he trusts models more than people. A person’s few dozen words often compress a large amount of context, and alignment may take half an hour; a model can combine the supplied background with broad knowledge to quickly infer the reasons behind a view, sharply reducing communication friction.

  • Sand.ai employees send screenshots of difficult cross-department messages to “Teacher Gemini” and ask it to explain what the other side is really saying. Cao Yue believes this use as a bridge across context gaps between algorithms, product and operations may be closer to an underlying organizational transformation than simply asking for answers.

32. “The final form of all content is narrative,” connecting technology with the business model

  • A veteran film-and-television professional told Cao Yue: “The final form of all content is narrative.” Short video evolved from “recording beautiful life” to fully optimizing the 15-second viewing experience and then to short dramas; even creator content on Bilibili has its own narrative structure.

  • The feedback made him place greater weight on characters and performance, and explains why Sora 2’s end-to-end narrative shook him so much: the model is not merely lowering the cost of production assets, but moving closer to the core organizing principle of content.

  • When Cao Yue decided to make AI video, Wang Huiwen offered only one brief suggestion: “You could study Pixar.” Cao Yue later understood that Pixar uses graphics technology to make films and generate box office, while retaining character IP in-house for long-term monetization through merchandise and other channels.

  • Cao Yue’s understanding is that after a live-action film becomes a hit, the character IP is often taken away by the actor, while Pixar can retain the characters if it creates them itself and build long-term revenue. Cao Yue used to ask, “What would Ilya think?” After becoming an entrepreneur, he increasingly looks to operators such as Wang Huiwen, Zhang Yiming, Li Xiang, Lei Jun and Duan Yongping.

33. Cao Yue’s next question is where humanity will stand before an ASI with an IQ of 1,000

  • Cao Yue treats ASI as an open question discussed internally at the company: if the intelligence of language models rises every year and hypothetically reaches an “IQ of 1,000” in 5 years, existing testing systems may be incapable of measuring it, leaving humans without a reference point for understanding the number.

  • He uses the gap between adults and children, and between humans and monkeys, as an analogy: when an agent with an IQ of 500 or 1,000 appears, it is difficult today to infer what impact it will have on the world from a linear improvement in capability.

  • Many Silicon Valley companies believe building superintelligence requires scale, but Cao Yue thinks there may be different views on whether scale is truly necessary. He offers no definitive answer, only the view that continuing to reason in this direction exposes many questions that current organizational and product discussions have not yet reached.