Mootion’s Tong Chao on Putting Down the Hammer and Making Its First $1M
Summary
- With roughly 20 people, Mootion reached about 2 million overseas users and approximately $1.2M in revenue by early 2025, leading Tong Chao to conclude that AI applications do not need to wait for foundation models to mature before creating commercial value. The project had been operating for two and a half years and raised roughly $7M–$8M; apart from one finance and one operations hire, the team was almost entirely R&D. Tong Chao’s view: “The model is at 50 points, but users want an 80-point product,” and the product company’s job is to close that 30-point gap.
- Mootion does not treat viral hits or technology early adopters as its long-term customer base; it focuses on people who create continuously and are willing to renew. Tong Chao acknowledges that virality is a powerful cold-start lever, but believes revenue and conversion are stronger evidence of a real pain point than traffic; the roughly 3% of users at the front of the adoption curve may have a lifecycle of only one month, with almost no switching costs. The company’s choice is: “Are we serving a viral hit for one week, or a user for ten years?”
- Foundation-model upgrades will not automatically erase the application layer, especially when video is a multistep, multirole, long-chain workflow. Ronghui cited Claude’s ability to write a close-up shot of the eyes of “a terrified woman”; Tong Chao responded that the hard part is keeping the script, semantics, visuals, and video consistent across 64 consecutive scenes. Even if the model moves from 50 to 90 points, commercialization and rising user expectations will turn it into “a new 60,” leaving the remaining experience to the product.
- Mootion’s most expensive lesson came from taking its proprietary 3D motion-generation model to market and, after receiving feedback from 20,000 professional creators, still being unable to answer the question: “Then what?” The team once believed it had built one of the world’s first models of its kind, but the generated motion could not yet enter game or film production directly, and four or five months of work had not produced a complete value chain. The capability was eventually absorbed into the product, while the company truly abandoned the technology-startup path of “using a hammer to look for nails.”
- AI video applications currently fall into five categories—editing, talking heads, effects, end-to-end generation, and point tools—but the next phase will likely see generation and editing converge. Users do not care whether a piece of footage was shot, generated, or re-edited; they care only whether they can get the content they want. The biggest product opportunity, therefore, is to define the merged interaction model first. Tong Chao previewed that Mootion may launch such a product in the second half of the year or toward year-end.
- Globalization cannot be judged by user growth alone: Brazil is suited to cold starts, the Middle East depends on cultural entry points, the US is broadly as expected, while Japan and Taiwan are difficult to enter but offer high willingness to pay and strong loyalty. Brazilian users are eager to try new products and form communities organically, but conversion three months out must be addressed in advance; Oman, with fewer than 5 million people, already has roughly 70,000 Mootion users, and about 90% of growth across the Arab region is organic. Ramadan content at one point accounted for more than 10% of daily usage, showing how localized features can create powerful usage wedges.
- Tong Chao believes this generation of GenAI will not repeat the previous AI cycle’s collective failure to deliver, but model companies still face commercialization and competitive shocks. An older relative proactively asking about DeepSeek was, for him, a sign that “普惠” had arrived; application companies carry lighter burdens, while model companies could still face a sudden shock similar to DeepSeek. The corresponding startup method is not to predict the winner, but to “practice more, fail fast, and iterate at full capacity.”
Deep dive
1. A 20-Person Team Has Already Reached Approximately $1.2M in Revenue
Tong Chao’s operating snapshot: Mootion had been in business for two and a half years, raised roughly $7M–$8M, and by early 2025 had about 2 million overseas users and approximately $1.2M in annualized revenue. The product is not aimed at the AI industry; it is meant to “help people who have never used AI make their own videos and tell their own stories.”
The team has only about 20 people, including one finance hire and one operations hire; everyone else is in R&D, split roughly evenly between algorithms and engineering. That structure reflects Mootion’s two-pronged investment: developing and adapting generation capabilities while engineering long-form video workflows and driving down inference costs.
Tong Chao is 36, with a background in computer science and machine learning. He previously led AI and product at 360 and served as head of product at Innovent AI. Precisely because of his deep technical experience, he now repeatedly reminds himself to “throw away the technology” and first ask what problem the user actually has.
2. Mootion Breaks Video Generation into Three Sustainable Use Cases
The first user group is faceless short-video creators, who use generated content to replace the old process of sourcing footage and editing. One operations colleague spent just five minutes a day producing a Xiaohongshu video and accumulated roughly 13,000 followers in two months; a European religious-content creator now has more than 100,000 followers, with videos generating millions of views.
The second use case is education. Mootion gives teachers and students an entry point for using generative storytelling in the classroom, including bilingual content. One teacher turned lesson points into a two-minute video and found that classroom engagement jumped: “Everyone’s eyes were glued to the screen.”
The third is professional creators, who do not rely on Mootion for the entire production process but embed its underlying capabilities into their own workflows. A professional studio in Poland used Mootion to generate scenes for a 3D music video showing a character putting on glasses and entering a virtual world, then combined those outputs with its own production expertise.
3. Tong Chao Would Rather Serve a Ten-Year User Than Chase a One-Week Hit
The host’s challenge was direct: Pika, Vozo, Viggle, and other AI video products all achieved milestone growth through viral content. Does Mootion’s lack of an equivalent breakout mean it has not truly hit a pain point? Tong Chao acknowledged that virality is “a very good lever” for a cold start or a milestone jump, but said the team had not spent much effort pursuing it.
His counterquestion captures the product choice: “Are we serving a viral hit for one week, or are we serving a user for ten years?” Video’s value does not exist only in traffic channels such as TikTok and YouTube; if users keep paying real money and conversion is healthy, the pain point is real, even if it does not show up in public traffic.
The archetypal renewing customer is still someone running a faceless account. Tong Chao estimates that AI output may have improved from 20 points to 50, but a good story, expression, presentation, and account operation still require the user to contribute another 30 or 40 points—capabilities that “are not directly related to AI.”
Mootion therefore does not promise “generation equals virality.” The product provides the capacity for continuous production; users remain responsible for editorial judgment and operations. It is a slower value loop than an occasional viral hit, but one more likely to produce retention and revenue.
4. Better Foundation Models Raise the Application Ceiling but Do Not Automatically Consume Long Workflows
Faced with a Tom and Jerry video that was almost indistinguishable from human-made content, Tong Chao acknowledged that test-time training “is definitely a very promising direction,” and said the team was following the research. Such capabilities could push video generation to the next level, but he sees model upgrades as good news for applications, not an apocalypse.
The key difference is task length. A Jasper- or Copy-style text task may be a single jump—“I think, then I output”—whereas video involves many steps, a long process, and multiple stakeholders. The more steps there are, the harder it is for a single model upgrade to overturn the entire system, and the more the product must reorganize features, experience, and collaboration.
Ronghui raised the term “wrapper” that often appears during technology hype, arguing that people overestimate foundation-model capabilities and see everything outside the model as shallow packaging. Tong Chao said whether Mootion is a wrapper is beside the point; his constraint for founders is not to develop a purist obsession with doing only what others have not done, nor to imagine that the technology will become “superhuman” in two months.
5. As the Model Moves from 50 to 90, User Expectations Rise with It
The host cited another product thesis: users will not accept products at 20, 40, 60, or even 80 points; opportunity exists only between 90 and 100. Instead of struggling today to take a model from 50 to 80, should founders wait for the foundation model to reach 90 before starting?
Tong Chao admitted that this “may indeed be a problem,” but his answer is a moving benchmark. Once a 90-point model becomes a commodity, the market will rise with it, and relative to new demand it will become “a new 60.” The remaining 30 points still come from understanding video creation, specific use cases, and users—not from a general-purpose model automatically.
His conclusion is unambiguous: “Starting early is definitely better than starting late.” Closing the 30-point gap early is not just about accumulating engineering code; it also builds user intuition, scenario knowledge, and product judgment, all of which may remain valuable through model upgrades.
6. The Hard Part in a Four-Step Workflow Is Keeping 64 Scenes Connected
Mootion compresses creation into four steps: explain what you want and what you have; generate the full content structure; choose effects, transitions, sound, music, and other additions; then assemble and share, including the description and hashtags—not merely export an isolated video.
The automation incorporates knowledge from a professional film company among its investors: what structure a script should use, why a frightened scene requires an eye close-up, why a joyful scene calls for an ultra-wide shot with many people, and how different narratives should map to shots. After testing, Tong Chao said the team found that large models genuinely did not understand these conventions—and still did not at the time.
Ronghui used Claude on the spot: given only “a terrified woman,” the model produced instructions for widened eyes, dilated pupils, and a camera pulling back from a facial close-up. Tong Chao’s response was that a single-point test “sometimes works,” but that does not amount to a deliverable system.
The real test is to take the user’s prompt and source material, organize 64 semantically continuous scenes, and make every image and video accurately express the intended meaning at that point. “You give it one input and it can do it; but when you turn it into a system, turn it into a network, it fails.”
7. Templates, Localization, and Inference Efficiency Together Drove Payment
The “Templates” feature launched last December and became a monetization inflection point. These are not visual styles but entry points for different use cases; each sits on top of a long AI workflow, almost like an agent chaining together multiple tasks. That allowed the product to go deeper into specific scenarios, and payment rose visibly.
At a user event in Tokyo, a director Tong Chao guessed was in his sixties challenged him: “You’re an Asian team. Why can’t you generate Asians when generating Asians?” Western-skewed training data meant Japanese and Chinese characters still looked conspicuously Western. Tong Chao returned to optimize the model and instruction following; he later learned that the speaker was Takashi Asai, a Japanese producer and director who had worked on Suzhou River.
A trip to the Middle East produced a more specific feature. As Tong Chao summarized the local religious rules, Allah and the 23 prophets cannot appear as human figures and may be represented only by light or a halo; he said no model could generate such content by default. After Mootion launched its Islamic Stories feature on March 1, usage surged and at one point accounted for more than 10% of daily volume.
Low pricing was another visible product capability. The team initially set out only to solve slow generation, then realized that latency fundamentally represented GPU time and cost. About six months after launch, architecture and inference optimizations improved gross-margin headroom by roughly 50%, turning “faster” and “cheaper” into the same engineering outcome.
8. Not Letting Users Choose the Model Means Not Treating the First 3% as the Long-Term Customer
Mootion does not let users freely choose among Flux, Kuaishou’s Kling, Vidu, and other models because its target customers want the result, not the model brand. Tong Chao paraphrased Zhao Benshan: “Don’t look at the ad; look at the outcome.” Showing too many names would force ordinary users to first learn what Kling, Hailuo, Pika, and Sora are.
The host asked whether Kling 2.0’s strong launch meant users who knew about the new model might go directly to Kling or switch to competing products that supported it—effectively leaving Mootion behind. Tong Chao explained that these users are technology early adopters at the front of the adoption curve, not the core group Mootion intends to serve over the long term.
He places technology early adopters in roughly the first 3% of the adoption curve: active and highly transmissive, but with product lifecycles that may last only a month before a newer tool appears. Their primary criterion is being first; they have no stable attachment to a particular use case, so switching costs are extremely low.
This is not a denial of their ability to spread a product, but a question of resource allocation. When technology is immature, users are poorly defined, and team resources are limited, no company can serve everyone at once. The more important job for a product lead is to identify “the most important problem.”
9. Fifty Cold Emails a Week Keep the Founder from Losing User Intuition
Tong Chao still sends cold emails every week to roughly 50 users. Usually only one or two respond, after which he schedules a roughly 20-minute interview. He has maintained at least one such session per week since launch because when both the technology and the point of user adoption are changing, “the only way to know is to ask users.”
The host cited Stripe CEO Patrick’s view that user research does not necessarily tell you which feature to build; it first shapes a way of thinking, which then guides the product. Tong Chao agreed and called it “user intuition”—continually putting yourself into the user’s scenario and circumstances so that you can think about the product through 2 or 3, or even 3 to 5, different identities.
His warning is that highly intelligent people are especially prone to losing user awareness: because they understand the technology, they assume their own judgment must be right, even when the people they are serving may belong to an entirely unfamiliar group. User interviews are not a wishlist; they prevent founders from mistaking their own cognition for the market.
10. The “Then What?” of the 3D Motion Model Consumed Four or Five Months
In its early days, Mootion was heavily focused on foundational technology and built a relatively small 3D motion-generation model in late 2023. In Tong Chao’s cautious wording, it “should have been the first in the world.” A user entered a text prompt, and the model generated movements for film, animation, or game characters.
More than 20,000 professional 3D creators quickly gave positive feedback after launch, but two months later the team encountered the more important question: “You generated this character animation—then what?” The capability still could not enter real game or film production directly; the model alone was neither a professional workflow nor a complete product.
The team quickly stopped pursuing it as a standalone product and embedded the model into Mootion to improve generation controllability. Tong Chao looks back on the four or five months of cost this way: “If you go out before you’ve figured out what problem you’re solving, you will definitely get sent back.” The more technological possibilities there are, the easier it is for founders to explore markets that do not exist.
11. The Previous AI Cycle’s Lesson: the Hammer Wasn’t Hard Enough, and There Were Only So Many Swings
Tong Chao understands the previous generation of AI companies, roughly from 2016 to 2022, against a backdrop of technology optimism generated by DeepMind. Breakthroughs such as Go convinced people that AI would quickly enter many industries, so large numbers of companies took their CV, NLP, and machine-learning “hammers” out looking for nails.
The first problem was that the hammer was not hard enough. Capabilities typically reached industrial performance only in narrow scenarios, with weak generalizability. Even when a customer was found, large teams and additional engineering were needed to complete delivery, producing a services-heavy model that failed to fulfill AI’s original promise of improving labor efficiency.
The second problem was that startups had only a limited number of swings. If they failed to hit enough large enough nails early on, the risk of each subsequent attempt rose and the team’s willingness to keep betting declined. Even with substantial funding, they could not absorb unlimited scenario exploration and custom delivery.
The host put Innovent AI’s seven funding rounds, backing from SoftBank and others, continued losses after listing, and roughly 80% share-price decline on the table. Tong Chao did not defend any individual outcome; he reduced it to three structural difficulties: industry is highly fragmented and demanding, finance requires model ownership and ongoing maintenance, and retail offers many opportunities but each individual scenario is too small.
12. GenAI’s Accessibility Changes the Hammer but Does Not Remove the Risk for Model Companies
Tong Chao believes this wave will not simply repeat the previous one because the biggest generational difference is not a particular architecture but “普惠”—broad accessibility. During a Lunar New Year visit home, an older distant relative proactively asked about “that something DeepSeek”; 8 or 9 years ago, it would have been hard to imagine an elderly person in a county town actively seeking out an AI product.
Tong Chao said the team built a roughly 1-billion-parameter model based on the Transformer architecture around 2022, but he does not think that fact matters. What matters is that ordinary people can now access the capability directly, making it more likely that the hammer in a founder’s hands will actually hit a nail.
He also calls the current state “the first step in a long march”: architectures are more transparent, data can be scaled, and the probability that technology will continue along an established path is far higher than in the previous generation. In the past, CV teams took turns competing for the top benchmark score without necessarily being able to explain how they got there; today, at least, the path to the target is easier to see.
But this optimism applies mainly to applications. Model companies still carry similar risks: models may rapidly commoditize, or a new entrant may suddenly appear, as DeepSeek did, and shock every incumbent. Application companies carry much lighter baggage by comparison.
13. Ilya Did Not Know Before Launch Either; Speed of Practice Beats Prophet Narratives
When Mootion was getting started, Tong Chao’s partner met Ilya and Greg in Silicon Valley through Kai-Fu Lee, roughly one week before ChatGPT launched. The only message he brought back was that OpenAI seemed to be building an application—and Ilya himself said, “I don’t know how this thing will turn out.”
ChatGPT subsequently became a world-changing product. The contrast made Tong Chao realize that during a rapid technology breakout, “no one has a higher level of understanding than anyone else,” especially when it comes to the market: top researchers and founders are all exploring what user value the technology can actually produce.
The host added YouTube co-founder Steve Chen’s framing: outsiders imagine great founders as gods or prophets who knew the answer long ago and then executed it, but entrepreneurship does not work that way. Tong Chao’s operating principle is: “Practice more, fail fast, and iterate at full capacity.”
14. The Five AI Video Tracks Will Eventually Converge on “Generation Plus Editing”
Tong Chao defines the first category as AI editing: using AI to recreate existing editing capabilities, with Descript, OpusClip, and Captions as examples. The second is talking heads, centered on real or digital humans, with HeyGen, Hedra, D-ID, and Synthesia among the examples; the latter, he believes, has just crossed $100M in ARR.
The third is generated video effects, such as Pika and Viggle, which use generation to replace traditional post-production effects. The fourth is end-to-end applications such as Mootion, Fliki, and InVideo, whose goal is not a single video clip but a structurally complete, publishable piece of content.
The fifth category solves only a single point in the workflow, such as video dubbing, video translation, or face swap. Tong Chao’s summary is that these companies “are all shovels”: some provide one part of the final content, while others connect a single step in the process.
The next opportunity lies in combining generation and editing. Users do not care whether footage was shot, generated, or re-edited; they care only whether they can get the highlights or complete content they want. Whoever defines the merged product form first may control the new entry point. Mootion may offer its own answer in the second half of the year or toward year-end.
15. Descript, VEED, and HeyGen Show That Product Iteration Can Outlast a Technology Label
Tong Chao especially admires Descript’s effort to rebuild cumbersome editing interactions around natural language. The product still has many small issues, but it sends a clear signal: AI in video creation should not merely generate a clip; it can also become an interaction layer running through the editing process.
VEED.io is a more direct example of “growing with the user.” When it launched in 2019, it did not have many users, but the founders continued to document the startup journey publicly on Twitter and other platforms, staying in the same environment as users and adding AI capabilities in real time. VEED also launched the first VideoGPT inside ChatGPT.
According to Tong Chao’s figures, VEED reached roughly $1M in ARR in two years and then grew to more than $20M in ARR over approximately three more years. Its value lies not in correctly betting on one model, but in an agile team that keeps responding to users and upgrading with the technology cycle.
Among Chinese AI video founders, he particularly admires Joshua of HeyGen. The reason is not that the product never changed direction; it is that the team stayed close to users and the market after its pivot, rapidly completing product design and iteration. That, Tong Chao believes, is also an area where Chinese founders have an advantage.
16. Kling Leads, but the Video-Model Landscape Has Not Settled; Jimeng Is Betting on Content Consumption
Tong Chao believes the Kling team is “genuinely excellent.” When Sora appeared, its R&D effort was in a period of quiet development, but the team did not lose its footing. It continued investing along its established route, reached SOTA relatively quickly, then productized the research and attracted sustained use from creators in China and abroad.
He believes Kling bet on a technical route similar to Sora, with persistence from the research team as the core strength. The product layer is comparatively thin, but it globalized quickly. He even said Kling’s own internationalization “may have been done better than Kuaishou’s own overseas expansion.”
For teams that have yet to publish sufficient results, Tong Chao is watching 3AI and Cao Yue’s team, partly because their technical route may differ from DiT. Video foundation models are still in a divergent phase, so startups such as ShengShu Technology and PixVerse, along with major companies continuing to invest heavily, may coexist for some time. The eventual winner cannot yet be predicted.
The host asked about ByteDance’s pressure, citing Jianying, CapCut, and Pippit AI reaching the top of Product Hunt. Tong Chao is more bullish on Jimeng’s bold experiment: putting a TikTok-like feed on the first screen so AI content is not only produced but also consumed continuously, extending user time. It could become “the new TikTok,” he said, while stressing that “there is still a long way to go.”
17. Globalization Must Be Decomposed by Market; Local Users Beat Local Representatives
Mootion’s top three user regions are Brazil, the Arab region, and the US. Brazil is large, quick to embrace new things, and prone to spontaneous user-led group formation and sharing, making it suitable for a cold start. But payment habits resemble China’s market 5 to 10 years ago, so the team must think ahead about conversion “three months from now,” rather than watching new-user growth alone.
The Arab market was an unexpected growth driver. Oman, with fewer than 5 million people, already has roughly 70,000 Mootion users. Invited by the local investment authority and education department, Tong Chao visited public and private schools and saw teachers and students using the product in class. About 90% of local growth is organic; the team has produced only 2 or 3 influencer pieces.
The region has a dual character: religious constraints are strict and users are cautious about new things, yet that makes them especially curious about tools with an AI angle. Once a narrow entry point is found, organic sharing can scale quickly; after conversion, user lifecycles are long, but payment efforts should prioritize markets with better conditions, such as the six Gulf states.
Japan and Taiwan also rely on organic sharing, followed by influencer support once users reach the tens of thousands. Together, the two markets have roughly 150 million people; they are difficult to enter but offer long lifecycles, high loyalty, and high conversion. The US is broadly as expected.
18. Three Advantages for Chinese Teams and a Different Way to Globalize
Tong Chao says the first advantage of Chinese AI founders is being “grounded”: even when the technology is immature, they can use multiple approaches to find product-market fit. The second comes from a fiercely competitive mobile-internet environment, where growth and operations capabilities can create an asymmetric advantage overseas, especially in cold starts and scaling.
The third is the simultaneous availability of research and engineering talent. China has both a large pool of young, imaginative researchers and strong full-stack engineering teams that can deliver new technology to users quickly. This combination does not require everyone to be a top researcher; it shortens the distance between a research result and product feedback.
Mootion does not plan to appoint representatives in every country. Tong Chao wants to rely on 2 forces: let AI communicate with users worldwide like an employee or intern, then let local users become “the best representatives,” driving sharing, feedback, and product iteration themselves.