119: How Do AI Video Products Go Viral? PixVerse’s Answer at 60 Million Users
Summary
PixVerse has validated rare user and revenue growth with consumer AI video: more than 60M global users, over 16M mobile MAU, and monthly revenue above RMB10M. The mobile app surpassed 10M MAU in its first month after launching at the end of 2024; revenue has grown roughly 10x over the past few months, all from overseas subscriptions. More importantly, excluding model R&D, subscriptions already cover serving costs and leave room for profit: “This isn’t a demo.”
AIsphere does not position video models as “all-purpose AI directors,” but as a new kind of “AI camera,” with the larger product opportunity in helping ordinary people express themselves through short video while retaining professional creative services on the web. 谢旭璋 estimates that AI’s participation in complete audiovisual works is currently often only 10%-20%; scripts, shot transitions, sound and narrative still depend heavily on humans. On the consumer side, however, there is a more direct end-to-end need: helping people who cannot shoot or edit create videos that are “fun, personal and worth sharing” for the first time.
Templates are PixVerse’s growth engine because they eliminate both the prompt barrier and the cost of “drawing lottery tickets.” 谢旭璋 estimates that more than 95% of ordinary users cannot write long prompts well; after model fine-tuning and template constraints, users only need to upload photos of themselves, friends or pets to obtain publishable results with high probability. The Venom transformation attracted millions of users worldwide in a single month, and related videos generated more than 1B views—key validation that an AI video generation model could be widely used and shared by ordinary people.
AIsphere’s model moat is measured not only by absolute quality, but also by iteration efficiency and generation speed per unit of resources. With a team of more than 50 people today, the company trained seven generations of models in 2 years, iterating from V1 to V4.5. 谢旭璋 says its GPU usage may be only one-tenth or even one-twentieth that of many peers, while its high-quality online model can generate a roughly 5-second shot in 5-10 seconds, versus the 1-2 minutes required by most competitors. “Throwing money and resources at it, and hiring enough people,” does not automatically solve video generation.
The next model upgrade does not mean blindly pursuing longer videos or stronger text-to-video, both of which 谢旭璋 sees as “false needs” inflated by industry narratives. He estimates that more than 95% of shots in traditional works are under 8 seconds, and film professionals may actually prefer controllable 1-2-second shots. The mainstream creative entry point remains image-to-video or “image plus text.” Beyond its existing DiT architecture, AIsphere is exploring a new architecture combining autoregression and Diffusion while seeking to preserve near-real-time speed.
The competitive landscape remains open: 谢旭璋 puts Google, Kling and Hailuo in the first tier, but says Sora’s actual delivery is “10 pixels” behind its demos. He does not deny the overlap between TikTok, Kuaishou and template-based products, but stresses that video is not a winner-take-all industry: cinemas, long-form video, short video and livestreaming can coexist for decades. The real organizational challenge is that professional productivity and mass-market entertainment require different products, operations and growth systems, making it difficult to do both equally well in one product.
Domestic launch and the goal of reaching hundreds of millions of users are the main milestones for the next phase, but 谢旭璋 continues to place earnings quality alongside scale. Paiwo AI launched on June 6, after Chinese users had already been paying a few yuan to RMB20 on Xianyu for outsourced Venom videos. The company’s full-year target has narrowed from “at least 2x to 3x growth” to “hundreds of millions of total users”; over 3-5 years, it hopes to reach 10M DAU, tens of millions to 100M MAU, and achieve scaled profitability: “While helping users create, we also need to make the money back.”
Deep dive
1. The Twilight of Mobile Internet Pushed 谢旭璋 Toward Generative Visuals
In 2022, 谢旭璋 saw a clear inflection point at Lightcone Capital: many internet companies had “extremely strong data and teams,” yet were beginning to hit resistance in both fundraising and business execution. He described the feeling as entering “the twilight of mobile internet” and began searching for the next major trend.
Midjourney and Stable Diffusion affected him through personal experience rather than abstract technical judgment. He typed “谢旭璋” into Midjourney, and the model generated a beautifully composed image whose imagery seemed somehow related to his name. Someone who described himself as having “no artistic cells,” who could not draw and had never made short videos, felt for the first time that he could participate in content creation.
In the second half of 2022, he traveled from Shanghai to Hangzhou every week, exploring with programmers and researchers in teahouse sessions resembling hackathons. Lensa’s breakout, and ControlNet and LoRA’s improvements in controllability and character consistency, convinced him that visual generation would become enormous. The team tried a Chinese Lensa-style app; although it never launched, it produced one key conclusion: a “part-time interest group” could not capture the core opportunity.
After ChatGPT launched, he spoke with nearly every Chinese generative-AI company he could reach at the time. Most of the “six little tigers” had not yet been founded, and video-generation entrepreneurs were even rarer. After Chinese New Year in 2023, he resigned before a specific product had been decided, judging that only going “all in” on visual generation could turn his firsthand experience into a real investment thesis.
2. Complementary Founders and a Small Team Set AIsphere’s Early Resource Discipline
王长虎 had spent more than 20 years working on computer vision and video technology at Microsoft and ByteDance, and had helped build multiple large-scale products from zero to one. 谢旭璋 handled capital, fundraising, external partnerships and overall direction; 王长虎 supplied technical and frontier product judgment. After the two developed their partnership through a mutual friend, they founded AIsphere in April 2023.
The first fundraising round was not held back by video generation’s lack of popularity. Within days of beginning fundraising, the team secured support from an industry investor, while Lightcone Capital, the founders and the team also invested. 谢旭璋’s review was that there were few visual-generation startups at the time, making the team’s backgrounds and direction easier to differentiate.
Resource constraints were an actively accepted premise from day one. The company had only 20-30 people for much of its first year and still has just over 50 today, remaining small overall. Unable to cover multiple markets and product forms simultaneously, AIsphere started overseas, launching through Discord, the web and other formats at the end of 2023, while reserving China for a later stage when its models, products and organization were more mature.
3. Seven Model Generations and 60 Million Users Formed Two Curves of Validation
AIsphere trained and released seven generations of models in just over a year, iterating from V1 to V4.5. 谢旭璋 says the models rank near the global frontier by internal experience, subjective and objective evaluations, and blind tests. He also acknowledges that video-model evaluation is highly subjective, so no single leaderboard can provide the final answer.
The user numbers need to be separated. The program opened with more than 60M global users and over 16M mobile MAU. In the interview, 谢旭璋 summarized combined mobile and web MAU as “more than 10M, close to 20M,” and said PixVerse was among the earliest video-generation platforms to surpass 10M MAU.
The mobile app launched around November or December 2024 and exceeded 10M MAU in its first month. Even after seeing a large number of internet products, 谢旭璋 still regarded “more than 10M people rushing in within a month” as extremely rare. After the traditional traffic dividend had ended, only new technology and new application possibilities could recreate that pace of growth.
He does not mythologize MAU. The product still resembles a paid tool: many free users cannot access the core experience, and it remains some distance from a mature community used daily. The team therefore tracks DAU, MAU, paid users, repurchase and retention. But while the industry is young and usage frequency has not stabilized, monthly metrics are temporarily easier to compare across products.
4. Professional Film and Television Still Cannot Be Delivered End to End; AI Video Solves Only the Asset Layer
Runway gave AIsphere a clear reference point. When 谢旭璋 asked about the company’s vision for the next 2-3 years, the answer was “to win an Oscar.” That positioned its models and products naturally toward film, design and professional production, and companies entering the field later could easily follow the same route.
谢旭璋’s estimate of current professional work is more restrained. In so-called AI videos seen by the public, AI’s real participation is often only 10%-20%. Humans still handle how the script progresses and how different segments connect; dubbing, music, sound effects, language logic and multimodal composition are far from forming a reliable integrated production chain.
AI video is therefore currently more like an asset generator than a deliverer of complete audiovisual works. “It isn’t an all-purpose AI director; it’s a new AI camera.” The camera itself matters, but whether it is placed into a film, short drama or short video determines how much value the model can ultimately unlock.
5. The Mass-Consumer Opportunity Comes From a Huge Viewing-to-Expression Gap
Billions of people watch short videos every day, but very few can create videos that attract viewers and spread. 谢旭璋 initially estimated that fewer than 5%-10% of consumers might be capable creators; the host pushed back that this was too high. The shared conclusion was that consumption scale and effective creative supply differ by an order of magnitude.
The ubiquity of smartphone cameras has not automatically removed the barriers. Ordinary people still need to shoot usable footage, record voiceovers and edit, while bearing the risk of being mocked or receiving no traffic after publishing. Video has become an important medium through which young people acquire and transmit information, yet existing tools have not met the expressive needs of most people.
AIsphere therefore defines the problem as how to help someone who has never made a short video produce something “fun, personal and worth sharing” for the first time. Posting to TikTok or another public platform is only one outcome; sending it to friends, parents, group chats or private messages works just as well. The key is helping users take the first creative step.
6. Templates Remove Both Prompt Friction and Generation-Lottery Friction
Around the third model generation, the team concluded that image-to-video would become the core mobile entry point for ordinary users. It therefore changed the model and product in parallel, packaging complex capabilities into templates. A template is not merely a UI wrapper; it combines model capabilities, fine-tuning, prompts and expected outcomes into a repeatable experience.
The first barrier is prompting. 谢旭璋 estimates that more than 95% of users cannot write prompts well, especially long prompts; it may even be harder than simply taking a photo or video. Templates let users upload photos or videos without first learning how to describe motion, style and camera movement in words.
The second barrier is the “draw rate”: open-ended generation is unstable, and ordinary users find it difficult to keep paying while waiting for an occasional good result. Template constraints can sharply raise the success probability. What users actually care about is not “whether it is AI,” but whether they can “simply make something interesting, fun and shareable.”
7. The Venom Transformation Turned Model Capability Into 1 Billion Views
“We Are Venom” lets users turn themselves, friends or pets into Venom with one click; both the transformation process and final image are generated by AI. During the Venom movie’s theatrical run, millions of people worldwide used the template within a single month. The program said related videos accumulated more than 1B views—“more than the number of people who watched the Venom movie.”
谢旭璋 sees this as more important than traffic alone: it may have been the first time an AI video-generation model was directly used and shared at scale by ordinary people. Users did not need to understand the model; they only had to place something from their lives into it and receive a result with personal relevance, visual impact and social currency.
A breakout hit cannot be replicated simply by copying an idea. The base model must first ensure that the transformation process and result match expectations; the team must also fine-tune the model, add engineering fixes and anticipate cultural trends. Venom’s heat might last only several weeks to a month. If followers cannot match the iteration speed, “by the time they make it, it’s already no longer hot.”
8. The Community’s First Step Is Not to Copy TikTok, but to Keep Creation Moving
PixVerse has added a feed that lets users keep generated results inside the product, but 谢旭璋 repeatedly emphasizes that the community remains very early. Traditional platforms share revenue with creators and encourage submissions, while AI products incur inference costs and users often have to pay for the core experience. Mature mobile-internet formulas therefore cannot be applied directly.
The team is currently focused on two loops. First, help people who have never made short videos complete their first generation and publish. Second, let creative users develop formats that others can reproduce “with one click.” The full content-consumption community will continue to evolve, but his attitude is: “Take it step by step; moving too far too fast” risks losing focus.
Some templates online already come from user experimentation, mainly entering the system through submissions. A complete self-service publishing flow does not yet exist. 谢旭璋 acknowledges that internal creativity is limited, while “users’ creativity and ideas are endless.” The long-term goal is to build the stage and “let everyone come up and perform.”
Template creators currently receive no revenue share, and the team is considering incentive mechanisms. Unlike Jianying’s professional effects creators, PixVerse wants to attract a new generation of creators who previously did not know how to shoot or edit video. In the first stage, the most effective reward may not be cash but “being seen,” having an idea recognized and building influence itself.
9. Monthly Revenue Above RMB10M Has Moved Beyond a Pure Traffic Story
PixVerse’s monthly revenue has grown roughly 10x over the past few months and now exceeds RMB10M, all from overseas subscriptions. 谢旭璋 says revenue quality and gross margin are both high. The company has not disclosed its paid conversion rate, but paid users, renewals and repurchase are now core operating metrics.
He refuses to quote ARR directly because many companies simply multiply one day’s revenue by 365 or one month’s revenue by 12, which may not satisfy the meaning of recurring. “What exactly is recurring?” In his view, ARR fits SaaS better. Consumer products should be judged by whether revenue persists, how many people pay, conversion, ARPU and repurchase—not by mechanically annualizing one-off top-ups.
Overseas subscriptions come in monthly tiers of $10, $30 and $60. Higher tiers provide more usage at a lower unit price. Ordinary users tend to choose the lowest tier, while professional users are more dispersed. 谢旭璋 only says renewals are “much better than expected,” possibly no worse than professional creative tools, without disclosing specific rates.
The more meaningful metric is unit economics. Excluding model R&D costs, subscription revenue fully covers model serving costs, leaving room for profit on inference itself. The company has no mature second revenue stream yet, but is testing an API and considering letting users exchange ads for generations in the future to lower the barrier to core experiences.
10. The Multimodal Workflow Opportunity Is Large, but Adding Defects Does Not Automatically Produce a Product
The host mentioned products that aggregate video, voice, music and large language models to create visuals, voiceovers, music and subtitles in one pass. 谢旭璋 agrees with the direction but says the modalities remain fragmented: when video, sound effects, music, voice and plot logic each have problems, simple combination usually produces “an even bigger problem.”
He says the simplest way to identify what a product truly wants to solve is to look at which feature it promotes on the first screen after entry. A long feature list may only show that a team is experimenting; it does not mean the core need has been established. AIsphere is willing to provide APIs to partners exploring end-to-end workflows, but does not equate demo completeness with PMF.
In his view, the API is a new business model for standardized model capabilities, not necessarily a distraction. Feedback from individual users and enterprise customers can both expose defects in the general model and feed back into training. The market does contain customization, distribution and solution businesses, but AIsphere prefers partners to connect to the standard API and serve end users themselves.
11. Among Three Key Decisions, Resource Discipline Mattered as Much as Direction
谢旭璋 groups AIsphere’s key decisions into three. First, choosing video generation in 2023. Second, improving the efficiency of each generation under limited resources rather than spending excessively on model training. Third, shifting productization toward To C and low-barrier template creation. The first determined the track; the latter two determined whether the company could survive long enough to build differentiation.
By his estimate, the number of GPUs AIsphere uses to train its current model may be only one-tenth, or even one-twentieth, that of many peers. The company is not short of money; it has deliberately maintained ample cash. But it continues to operate under resource constraints, improving organizational understanding and training efficiency.
The host asked whether the company regretted not using more money and GPUs to train Sora-like capabilities earlier, after 王长虎 had said a year before that additional resources could have accelerated that work. 谢旭璋 said no. The company could not have obtained 10x the resources at the time, nor did it want to build an organization that only knew how to spend money. In China, it is difficult to replicate the capital supply available to U.S. companies.
He also does not present restraint as the only correct path. Kling has committed significantly more resources and people, and has done very well. But many major domestic and overseas companies have still encountered training bottlenecks after massive investment, showing that video generation “is not something you solve simply by throwing money and resources at it and hiring enough people.” Team efficiency and technical judgment remain hard constraints.
12. Sora’s Delivery Gap Shattered the Assumption That the U.S. Was Naturally Ahead
谢旭璋’s leading models include Google overseas and Kling, Hailuo and AIsphere domestically. By user count, he believes the world’s 3 largest video-generation platforms are all Chinese companies: Kling, MiniMax and PixVerse. The view is subjective and shaped by his industry position, but he does not include Sora’s current delivered version.
On Sora, he used an exaggerated formulation: the production model and the early demo were “10 pixels apart.” He also said people overseas had tried to reproduce the demonstration from a year earlier using the current Sora but failed. The problem was therefore not leaderboard fluctuation, but a huge gap between theoretical demonstration and usable product.
The more damaging spillover is that the industry has begun to normalize showing a demo first and then announcing what has supposedly been achieved, without delivering a real model or product. 谢旭璋 is not entirely pleased to see competitors slow down: OpenAI is an industry trendsetter, and a poorly executed video product could weaken confidence and expectations across the entire market.
13. China’s Video Industry Base Is the Underlying Reason Local Models Reached the Front Rank
When the host met 谢旭璋 in November 2024, he seemed visibly anxious; by February 2025, user and revenue data had risen. The host linked the change in mood to the impression left by Sora’s formal release in December 2024 and the confidence boost from DeepSeek during Chinese New Year. 谢旭璋 replied that the OpenAI episode had also hurt the industry, but it validated that Chinese or Chinese-speaking teams could build the technology well.
He rejects claims that China might take 10 or 20 years to produce something like Sora. In recent years, new video platforms including Kuaishou, TikTok and Douyin were first scaled by Chinese or Chinese-speaking teams, training large pools of visual, video and algorithm talent. Kling, Hailuo and AIsphere reaching the front rank was not an accidental catch-up.
When Kling launched around June or July 2024, it invited industry peers to attend. AIsphere was unusually public in showing support. 谢旭璋 already believed Kling might be the first case in large-model or foundation-model development where a Chinese team moved faster than overseas competitors. It was both a competitor and collective validation of domestic technical reserves.
14. New Architectures Are Worth Backing, but Near-Real-Time Generation Matters More Directly to the Product
谢旭璋 speaks highly of GPT-4o Image Generation’s handling of complex semantics and controllable generation, and recognizes the autoregressive approach behind it. AIsphere’s current online model uses DiT, but the company is actively testing architectures that combine autoregression and Diffusion. It hopes to become one of the first teams globally—and perhaps the fastest—to deliver a usable high-quality model.
His reservation is that a model with stronger language understanding and better responses to long prompts may not be most popular with ordinary users. Professional creators may gain more control, but mass-market users “may not know how to write, or how to write prompts.” The model will ultimately need to be accessed through templates or more natural interactions.
The variable that matters more directly to consumer experience is waiting time. 谢旭璋 says PixVerse’s high-quality online model needs only 5-10 seconds to generate a shot; at base resolution, the fastest mode can generate a 5-second video in roughly 5 seconds. Most closed- or open-source high-quality peer models still need at least 1-2 minutes.
AIsphere’s next step is not simply to chase a speed metric, but to preserve this speed with stronger models. “This isn’t a demo; everyone can use it online.” Near-real-time feedback reduces queueing and generation-lottery costs, making repeated template experimentation feel closer to an ordinary mobile product than an offline rendering job.
15. Longer Videos and Text-to-Video Are Both Needs Inflated by Industry Narratives
谢旭璋 calls “the longer the generation, the better” a counterintuitive false need. He counted shots in films and online videos himself and estimates that more than 95% are under 8 seconds. Long takes are rare; the ability to generate a 1-minute video does not mean creators need a 1-minute shot.
PixVerse’s common output is 5-8 seconds. It has also tested longer shots online, but users rarely choose them. More counterintuitively, a professional at a leading film company wanted 1-2-second shots because shorter clips are more controllable and easier to organize in editing. Sora’s emphasis on “1-minute generation” misled the industry, while its demonstrations looked more like stitched-together shots.
The second false need is treating text-to-video as the biggest use case. The gap between language and vision remains too large: ordinary people cannot write good prompts, while professional creators often use storyboards to control style and composition. For both mass consumers and professionals, the more realistic mainstream entry point remains image-to-video, or video generation from “image plus text.”
16. Large Platforms Will Overlap, but Video Has Never Been Winner-Take-All
Faced with AI effects built into Jimeng, Douyin and Kuaishou, AIsphere is first betting on foundation-model iteration: generation quality, speed and the ability to support new formats quickly are the foundation for template products. Without a sufficiently good model, many ideas cannot be executed; without real use cases, model potential cannot be unlocked.
谢旭璋’s industry analogy is that television stations and cinemas were not eliminated by Youku, iQiyi or Netflix, nor were long-form video platforms completely replaced by Douyin, Kuaishou or TikTok. Video is not a narrow track but a continuously expanding commercial ecosystem. Long-form video, short video, livestreaming and different distribution channels can coexist for the long term.
AIsphere therefore does not define its goal as “PK’ing” TikTok. It wants to serve people who already watch videos but have never made one. AI-generated supply could expand the entire pie by bringing more ordinary people into creation and forming content and distribution relationships that previously did not exist.
There is genuine overlap with Douyin and Kuaishou’s internal templates, but mature short-video platforms and professional creative tools both face conflicts over what they can support. Kuaishou creating Kling as a separate product itself suggests that new production methods may require new products. A product matrix can resolve some conflicts, but creates new problems around team capacity, user positioning and organizational focus.
17. Private Messages and Group Chats Mean AI Video Need Not Depend on Public Algorithms
PixVerse supports one-click sharing to TikTok, YouTube and Instagram, but also sees substantial content flowing to Telegram and other messaging and private social networks. Videos involving oneself, one’s pets or one’s friends may not be suitable for public posting; private messages, group chats and Moments can be more natural consumption settings.
谢旭璋 says he “makes videos for everyone in different group chats every day,” and colleagues also circulate them in small groups. This suggests that product value does not depend entirely on distribution from any single platform. As long as content creates surprise and response, users’ own social graphs can carry the sharing.
Public platforms still provide stronger positive feedback. On TikTok, the team has seen ordinary people who had never made popular content use simple formats such as dancing pets to receive their first videos with 10K, 100K or even 1M likes. Having “your own life and your own creation seen by people” is a more durable incentive than the AI label itself.
18. Domestic Demand Already Existed; Paiwo AI Simply Added the Official Entry Point
AIsphere did not enter China early mainly because it lacked people and bandwidth, not because it believed demand was different. Once its models, products and team matured, China—as one of the world’s largest markets—had to be added. Paiwo AI launched on June 6, initially with features broadly similar to the overseas version.
The Venom template had already undergone a form of gray-market validation. 谢旭璋 estimates that several million Chinese users may have used it that month. On Xianyu, people offered to make the videos for a few yuan each, with prices as high as RMB20. Users in Douyin comments also pleaded, “Please, RMB6 to make one for me,” showing real willingness to pay for one-off entertainment creation.
“Paiwo” is a homophone of PixVerse and also reflects the idea that everyone can become the director of their own life. The team wants to combine shooting with AI generation, letting users create around themselves and their surroundings and then share the results with friends and family, rather than initially defining the domestic version as a professional production tool.
China will retain the usage-based pricing logic behind the overseas $10, $30 and $60 tiers, but specific pricing and products will be adjusted according to launch feedback. The team will start with a cold launch, watch the data and iterate. It acknowledges that overseas remains the larger market and that it does not yet have the capacity to operate every country with high granularity.
19. Globalization Is Not Country-by-Country Replication; Local Culture Can Still Create Spikes
谢旭璋 believes video is more universal than text. A visually entertaining template can spread simultaneously in the U.S., China, Brazil, Thailand and Europe. PixVerse has reached relatively high positions in app rankings across roughly 80% of countries worldwide, with key markets including the U.S., Brazil, Russia and Europe.
Indonesia also has a substantial user base, while India periodically produces spikes in popularity. He attributes these results more to population, smartphone ownership and mobile-internet penetration than to highly granular local insight by the team. The product supports multiple languages, but operations remain close to a “run the global version first and see where it grows naturally” model.
A pre-Christmas “Cyber Jesus hugging believers” demonstration showed another side of cultural adaptation. Religious classics are mostly text, yet users wanted to visualize scenes from their faith. The template took PixVerse to No. 1 on the overall App Store rankings in Germany, Spain and Italy. The team even contacted an assistant to the former pope through friends and found that serious religious institutions did not show strong resistance.
20. Model and Product Are Not Either-Or; They Form a Loop Between General Capability and Specific Use Cases
The language-model industry often debates whether “the model is the product” or applications matter more. 谢旭璋 believes video generation cannot simply import that framing. The AI camera still has to solve general problems such as resolution, clarity, physical laws, motion range and generation speed, but it must enter concrete use cases before it can create revenue and user value.
AIsphere therefore retains two product logics. The web product serves professional video creators and offers a fuller feature set; the mobile app targets ordinary people and deliberately removes most professional editing capabilities. The underlying model has not abandoned professional capability; different endpoints simply expose the layer useful to their target users.
“Start and end frames” show why the uses of a capability cannot be fixed in advance. The feature was initially designed for professional frame interpolation: users input the first and last frames, and the model generates the motion between them. Ordinary users later uploaded childhood and adult photos to watch themselves grow up over several seconds, or transitioned a portrait into a travel photo, turning a professional function into a highly shareable format.
This is also why API and To C products can share underlying investment. The model expands the capability boundary first; product teams and users discover new uses; real-world feedback then returns to training. 谢旭璋 refuses to assign fixed weights to model and product: “Both things are important.”
21. Productivity and Entertainment Both Work, but One Organization Rarely Excels at Both
谢旭璋 does not reject productivity. PixVerse’s web product already has many professional users. His question is how much it actually solves. If AI video contributes only 10% of the assets in a complete work, its productivity value is discounted accordingly; the model cannot claim credit for the entire delivery simply because AI was used somewhere in the work.
Mass-market entertainment has a different PMF. The goal is not to make an existing professional workflow more efficient, but to help people who could never create produce their first piece of content. The two require different interfaces, operating models, growth channels, team collaboration and model presentation, making it difficult for one product to serve both well.
Adobe’s example captures the conflict best. Photoshop has a professional ecosystem built around complex editing capabilities; suddenly placing low-barrier text-to-image generation at the center could damage the existing experience. Professional tools often need a separate product to serve ordinary users. The hardest part is “doing one thing with focus while also doing the other thing well.”
22. A Nontraditional Product Founder Makes Up for Missing Credentials With Industry Context
In the early years, 王长虎 focused on core model R&D, while 谢旭璋 handled external affairs, fundraising, recruiting key people and the work he knew best. As the team expanded, 王长虎 shifted more attention to core R&D and product. The company remains relatively flat, does not emphasize titles and now has dedicated product staff handling day-to-day work.
Some early investors questioned whether 王长虎 could build a young, fancy AI product. The host observed that PixVerse ultimately seemed to have younger users and fancier formats. 谢旭璋 believes 王长虎’s experience building mass-market products at ByteDance gave him sufficient context to take an internet product from zero to one—more important than whether he had been a standard product manager.
Neither founder came from a typical product background, but AIsphere has insisted since inception that model and product matter equally and has run extensive experiments over 2 years. Entering video generation earlier also gave the team information that those who joined only after Sora did not yet possess.
When 谢旭璋 spoke with overseas growth experts in 2023, he found that almost no one had truly built global growth for generative-AI products on Discord or the web. Traditional mobile-internet experience could only be partially reused. “Put simply, we were entering unmapped territory.” Product methods had to be discovered through real launches and feedback, not obtained automatically by hiring someone with the right résumé.
23. The Value of Investing Experience Is Seeing Enough Patterns to Bet Against Consensus
While studying at Peking University’s Guanghua School of Management, 谢旭璋 organized Silicon Valley study trips and saw peers discussing startups and “changing the world,” in contrast to the prevailing choices of consulting and investment banking among Chinese students. He later did odd jobs at an institution on Beijing’s Zhongguancun startup street, watching founders pitch in speed-dating formats and investors place bets. The seed of entrepreneurship was planted there.
After joining Lightcone Capital in 2016, he spent more than 6 years covering text, graphics, video, livestreaming, music, audio, short drama, social, community and overseas tools. He estimates that among Chinese apps reaching 1M DAU between 2016 and 2023, he had spoken with or participated in more than half. There may have been fewer than 5 cases throughout his career that broke through to 10M MAU as quickly as PixVerse.
The period from late 2021 to 2022 became the clear inflection point. The capital environment changed, the industry entered mature consolidation, and giants such as ByteDance filled traffic gaps with products including Tomato Novel and Red Fruit Short Drama. His investment experience ultimately gave him no formula for success, but the belief that “things depend on people” and “there is no fixed template.” When he sees a non-consensus opportunity, he is willing to bet—and once he bets, he goes all in.
The period after Sora launched and before AIsphere’s next model was trained was the company’s closest approach to a trough. The market was full of noise about who would launch next month and whether China could ever build the technology. The team had no special morale technique; it continued R&D through shared internal conviction about the direction and its capabilities. His advice follows from that experience: people entering AI should use the tools themselves over the long term, understand what the technology can solve and then form firsthand judgments others do not have.
24. Beyond 100 Million Users, Scaled Profitability Is the Three-to-Five-Year Destination
谢旭璋 initially said the user base should grow by at least 2x to 3x within the year. When the host projected nearly 200M users from the current 60M-plus base, he narrowed the wording to “hoping for hundreds of millions of total users.” The revision preserves the ambition while showing that the company emphasizes crossing the 100M threshold rather than committing publicly to a precise year-end figure.
He further defines “scaled user volume” over 3-5 years as 10M DAU and tens of millions to 100M MAU. The product should grow steadily, solve a real need, have people playing every day and spread through word of mouth rather than simply accumulate one-time registrations.
“Scaled profitability” means more than large revenue; it means the company can actually make money. 谢旭璋 believes most AI companies may remain deeply loss-making for a long time, while video is closer to real usage, traffic and monetization. Using Midjourney’s annual subscription revenue of “several hundred million dollars, possibly more than $300M” as a reference, he judges that the potential user base for universal video creation should be larger.
Subscriptions currently cover inference. The next step is to expand both free experiences and business models: ads-for-video, APIs and additional user tiers are all candidates. The ultimate goal is not to trade cash burn for scale, but to make the technology sufficiently accessible: “While helping users create, we also need to make the money back.”