Pioneers Insight Method Research Author
Interview with ONE2X Founder 王冠: Generative Systems and Platform Power
Back to Episodes

Interview with ONE2X Founder 王冠: Generative Systems and Platform Power

Summary

  • 王冠 identifies data as the first-principles variable in the current AI cycle, and on that basis divides the industry into three stages: public-domain data, domain data, and product-native data. With algorithms, compute and data broadly equal, public-domain data has a fixed boundary, so foundation models “will not have a generational gap”; domain data favors large companies with businesses, channels and accumulated digitization; the real application window is to design, from day one, data that previously did not exist, then train it back into your own system. More data is not necessarily better: less than “one-thousandth” of current FSD data may be usable for training, and a system should learn only from samples that exceed its own level.

  • ONE2X is betting on video not merely because video commands a higher price per unit, but because it sees video as the starting point for content in the AI era. 王冠 observes that roughly 20-30 video-processing SaaS companies in the US may have ARR in the tens of millions of dollars; a single capability can generate meaningful revenue if it is good enough. Technically, video can also be decomposed into a finite set of “atomic capabilities,” making it suitable for a closed domain. The long-term goal is not to get more people to learn editing, but to “turn video from a form of creation into a form of expression.”

  • The winners in applications will be determined not only by the model, but by who can build a higher-quality, lower-cost System 2 around the same System 1. Reasoning, workflows, Agents, domain libraries, memory and DSLs are all, at root, ways of producing effective tokens; 王冠 summarizes this as “Context is everything, everything is context.” Calling an application a “shell” is not pejorative. The real questions are how thick the shell is, whether the same tokens can deliver better results, and whether the system can generate proprietary data.

  • 王冠 sees current multimodality as closer to the GPT-2 or “Nokia era” than to a multimodal ChatGPT or iPhone moment. Sora 2 and Veo 3 are only the first attempts to bring isolated tasks such as audio-visual synchronization and storyboarding into more unified models; the real inflection point requires end-to-end input from any modality, reasoning in video or audio, and several orders-of-magnitude improvements in inference speed and cost. ONE2X is therefore starting with production tools closest to model change: “Don’t build apps in the Nokia era.”

  • The boundary between model companies and application companies will eventually disappear in both directions: model companies will build products, while application companies that become large businesses will have to build models. Vertical applications will eventually hit the capability ceiling of general models and, at scale, demand more controllable costs and margins; once upstream companies enter the same scenarios, the two sides become competitors rather than merely upstream and downstream partners. The only temporary advantage of application companies is speed: they must establish a lead on the map and keep widening it.

  • Generative systems may push the information industry from a “distribution economy” toward a “production economy,” shifting power further from platforms to consumers. In the software era, producers determined supply; the internet handed power to distribution platforms such as search, recommendation and e-commerce; on-demand generation sends demand directly to production: “No middleman taking a cut.” Creators will not simply disappear, but will shift from producing content item by item to supplying recipes, taste and incremental intelligence above the system baseline. More ordinary people will become “prosumers,” consuming and producing at the same time.

  • ONE2X has not formally released a complete product, but it has already collected several commercial signals worth tracking. 王冠 used an early version of Mydou to make videos and generated more than 2 million views and several hundred yuan in revenue share on WeChat Channels in just over a month; a leading AI content creator uses the half-finished product launched quietly in May every day, borrowing Google accounts from everyone around him and maxing out every points package; another customer used the product to “buy out” the AI videos related to a certain trendy toy on Xiaohongshu. These cases validate the possibility of revenue, but are not yet evidence of a scaled business.

  • The company is trying to make its product architecture, organizational architecture and capital path serve one proposition: the system should amplify the strongest people. At the time of the interview, ONE2X had roughly 30 people, worked remotely, had no dedicated managers, and about half its members had been founders or co-founders; its north star was not simply user count, but “system intelligence,” represented by content quality divided by token consumption. The company began in early 2024 and was not formally incorporated until Q2, after a period in which it prepared to bootstrap; 王冠 believes DeepSeek reignited confidence in models, Manus more directly rekindled capital’s enthusiasm for applications, and the final judgment is: “AI and AGI are a long-chain thing, so I bet China.”

Deep dive

1. ONE2X Is First a Product Studio, Not a Traditional Company

  • 王冠 defines ONE2X as a “product studio for the AI era”: AI will change how information goods are produced, distributed and supplied, while the team keeps products at the center and continuously explores experiences and business models different from those of the software and internet eras.

  • “Studio” emphasizes experimentation and the character of the work itself. The priority order leans toward interest, product quality and outcomes rather than simply obeying short-term commercial metrics; 王冠 nevertheless makes clear that a business must still clear the basic threshold of commercial value if the organization is to survive.

  • The studio wants every person to operate independently, which it calls being an “AI full-stack engineer”: each member should occupy a unique position, have a clear personal objective, and overlap their interests with the content of their work as much as possible.

2. 王冠’s Decade in AI Spanned Three Domestic Cycles

  • During the big-data and traditional machine-learning phase, he worked at Baidu on user profiles, multidimensional user tags, and differentiated pricing and subsidies based on those profiles; at the time, most AI product managers served the back end of recommendation systems.

  • During the CV and deep-learning phase, he built an algorithm open platform that delivered model capabilities to developers through APIs. He also worked on Baidu’s PaddlePaddle, then joined Megvii to build algorithm productivity tools and study how to produce algorithms faster and more cheaply.

  • The shock of GPT-3 in 2020 was that one prompt plus few-shot examples seemed able to temporarily turn a general model into different APIs for translation, summarization and other tasks. This extended the thread he had always cared about: how to give ordinary people access to algorithmic capabilities at a lower threshold.

3. Unstructured Data Let AI Fit a Continuous World for the First Time

  • 王冠 does not divide the cycles into “AI 1.0, 2.0 and 3.0,” but by whether models can fit unstructured data. Traditional CV relies on definite labels such as coordinates and categories, and is still fundamentally fitting structures that humans have discretized.

  • Language, images and video express the world more richly and continuously. Only when generative models fit these unstructured distributions can they express relationships between points in reality rather than merely complete predefined tasks.

  • This also changed the position of the AI product manager. In the past, much of the work consisted of labeling, strategy and capability supply, which was “quite boring”; today, the model can move to the front stage and become an independent product whose subject is AI, as with ChatGPT and Manus.

4. “Model as Product” Expanded Rather Than Weakened the Product Manager’s Role

  • 王冠 compares the model to a person’s System 1: a large amount of information has already been compressed into it, allowing fast, instinctive reactions to inputs. “Model as product” ultimately asks what capabilities that System 1 should have.

  • A model’s capabilities come from its data distribution and can therefore be designed. A model product manager must define the desired effects and capabilities, ensure those capabilities can actually be trained, and ultimately make users perceive them in reality rather than leaving them at the level of benchmarks or technical narratives.

  • He treats evals as part of product work and recalls discussing evaluation continuously at a relatively early stage. Later, product leaders at OpenAI and Anthropic also began emphasizing that product people should write evals and define model capabilities, gradually turning the practice into industry consensus.

5. System 2 Turns Model Capability into Monetizable Product Value

  • Next-token prediction means that subsequent output depends on existing tokens. Reasoning models add effective tokens by unfolding their own reasoning; workflows, Agents and domain libraries add effective context from outside the model.

  • 王冠 sees the latter as “completely a product problem.” The technology behind Agent frameworks, business processes and knowledge bases may not be complicated; the hard part is understanding industry know-how and converting it into context that the model can absorb and that makes economic sense.

  • The terminology has shifted from prompt engineering to context engineering, but the underlying task has not. Product managers both help define System 1 capability and lead the release of System 2 value; their role has expanded rather than narrowed.

6. 澜舟科技 Showed Him That “Small Models” Were Not a Backward Path

  • Before ChatGPT, 王冠 joined 澜舟科技, founded by 周明, and formally entered the pretrained-model industry. 澜舟’s philosophy was that models should ultimately be genuinely used rather than become huge systems that no one could operate or afford, so the endpoint should be “smaller and smaller.”

  • 张小珺 asked whether a company should first build large and then make the model small. 王冠’s answer was that there is no absolute right or wrong path. Small models can be used to validate a technical prototype before scaling, while large models must also be compressed and reduced in size when brought into concrete commercial scenarios.

  • Pretraining technology had not yet converged, and the community had multiple architectures. The team selected directions for Chinese-language reproduction, lightweight implementation and commercialization, experimenting with text generation and text-to-image generation before Stable Diffusion and ChatGPT.

7. After Three Collisions with OpenAI’s Roadmap, He Stopped Building on a “Void Foundation”

  • The first attempt used GPT-3 for writing assistance, text processing and completion inside Notion. Just as the product worked, ChatGPT appeared and broadly subsumed capabilities like those of Jasper and Copy.ai behind a general conversational interface.

  • The second attempt shifted to code: users uploaded Excel files and described charts, while the system called Codex to generate polished visualizations automatically. The demo had just learned to draw charts when GPT-4 launched with much stronger coding capabilities, putting him once again on the extension of OpenAI’s iteration curve.

  • The third attempt used LangChain relatively early, turning models, data sources and prompt design into workflow nodes. The demo was just finished and the financing process was approaching an investment-committee meeting when OpenAI Plugins launched, showing that model companies were thinking about the same layer.

  • “Once and twice, but not three times.” 王冠 reflected that he did not know how model capabilities were produced, where they would go next, or how far the product boundary was from the capability boundary. That meant building a product “on a very void foundation,” so he gave up on entrepreneurship for the time being.

8. The Entry Point to Moonshot Was a Three-Hour Lesson in “Compression”

  • Around March or April 2023, former Megvii colleague Tim 周星宇 invited 王冠 to join Moonshot. Over a meal at Longrenju in Wudaokou, 周星宇 spent 3 hours explaining “compression is intelligence,” almost entirely through formulas.

  • 王冠’s exact words were: “I completely didn’t understand it … but I was deeply shocked.” He then studied the concept of compression through videos by Jack Rae at OpenAI, gradually piecing together different materials and observations from his work.

  • Moonshot did not yet have the aura of the “Six Little Dragons,” and other model companies may have had more resources and attention. He chose Moonshot because he believed it was better suited to answering, at close range, where model capabilities came from, where they were headed and how far products were from them.

9. Compression Connects Discrete Data, Appearing as Emergence, Generalization and Even Hallucination

  • 王冠 explains that data is originally a discrete representation of objects. After compression, points that were separate in the distribution acquire continuous relationships; externally, that continuity appears as intelligence and may also be called emergence, generalization or hallucination.

  • NLP offers the clearest example. Training data may contain “Chinese-to-English translation” and “Chinese summarization” but no “English summarization.” Once the task is unified as natural-language input and output, the model may nevertheless learn English summarization, because continuity has formed between the tasks.

  • Language can express the world richly while costing less to train than images and video, making it the best compression medium available today. 王冠 later became even more comfortable calling the framework “language is intelligence.”

10. Moonshot’s Most Valuable Traits Were Its Pure Goal and Low Organizational Friction

  • 王冠 compares the atmosphere of early Moonshot to the film The Birth of a Nation: “Build AGI and stand up straight.” The team shared a pure, unified goal; smart people pushed forward proactively from different directions, and the pieces often came together naturally at some point.

  • This collaboration did not depend on endless meetings, coordination or alignment. He instead found the working state relaxed: people could invest thought according to their own understanding, and “most of the output would probably be useful.”

  • He spent about a year at Moonshot and was “the company’s first person to leave, and the first person to leave to start a company”; the company’s resignation process was also established at that time. He did not leave because the experience was poor, but because his preparations for entrepreneurship were complete.

  • The real condition for leaving was that he believed he had answered 3 questions: where model capabilities come from, how they will develop, and how applications should relate to general models. With that theoretical foundation, he again believed he could build something with both product and commercial value.

11. “As Much Human Labor as You Have, So Much Intelligence You Have” Remains a Foundation Fact

  • The AI industry’s self-deprecating line—“as much human labor as you have, so much intelligence you have”—is not outdated in 王冠’s view. “Human labor” is essentially data, and how much a model resembles a person is still bounded by its training data.

  • When algorithms cannot reach a business metric no matter what, the common conclusion is that “the product manager should go get some more data.” Open-domain data represents an understanding of the problem, especially when it contains industry know-how; acquiring and defining the data is itself product work.

  • Mathematics, code and other closed domains can use explicit rules to verify results and trade compute for data through RL. But many open problems lack an automatic, objective quality judgment, so human business understanding must first be converted into learnable data.

12. Data, Compute and Algorithms Determine the Boundary, Speed and Emergence of Intelligence

  • 王冠 uses a large circle to describe data: its boundary is the intelligence space represented by existing data. Compute determines the speed of approaching the boundary; the more compute available and the higher its utilization, the sooner the system reaches the capability ceiling supported by the data.

  • The algorithm is like a small circle continually moving inside the large one. As it approaches the boundary, part of it may protrude beyond the original range and draw a new boundary. The part beyond the original data is what he calls “emergence.”

  • Algorithms therefore determine how much additional generalization the same data can produce, while compute determines how long it takes to get there. Data remains the more fundamental, first-principles factor. He later used this framework to derive the industry’s 3 stages.

13. The Public-Domain Data Stage Is a Race to Reach the Same Destination

  • The first stage revolves around public-domain data accumulated through the internet and historical digitization: “You have it, I have it too.” Since everyone’s data circle is broadly the same, the core competition is who can clean and train faster and approach the common boundary sooner.

  • This stage favors foundation-model companies with high talent density, ample compute, fast decision-making and low organizational overhead. Models may win or lose on different tasks because of architecture and data distribution, but 王冠 believes “there will be no so-called generational gap.”

  • He goes as far as saying that Chinese and US foundation models are “already flat today.” Overseas advantages come more from first-mover status, compute and scenario data accumulated by serving users earlier than from an unbridgeable gap in underlying capability.

  • As Chinese models gain more users, usage data that is effectively filtered and fed back into training may further erase the gap. But the key future differentiator will be “effective filtering,” not call volume itself.

14. The Domain-Data Stage Naturally Favors Large Companies and Digitally Mature Industries

  • Once the public-domain boundary is nearly reached, the next stage is “domain data that I have and you do not.” Internal business records, channels, user relationships and industry processes begin to create clear divergence in model capability.

  • This stage favors large companies that already have scenarios, product channels and accumulated digitization, as well as traditional industries with enough digital maturity to absorb technology and talent spillovers.

  • 王冠 therefore believes the first 2 stages are not typical windows for application entrepreneurship. The first belongs to foundation models; the second is more like large companies re-modeling their own data.

15. The Application Window Is to Create a Third Pool of Product-Native Data

  • The third pool is neither existing internet data nor existing domain data. It is data that did not exist at all before the product appeared, and must be generated jointly by a new product form, new interaction and the model behind it.

  • ChatGPT is the core example. Before it, history contained no such data for solving many different problems through natural-language conversation; only after the product launched did these interactions gradually become an accumulable, trainable distribution.

  • 王冠’s conclusion is that an application company should design a new dataset from day one and ensure that it can eventually train back into its own model or system. The product’s value and moat are not the interface, but the data that “exists because you exist.”

  • This is also the way for an application to maintain a safe distance from general models: do not bet that the foundation model will forever be unable to perform a function; make your own product continuously generate learning material that the foundation model did not previously have.

16. A Data Flywheel Works Only When It Selects “Higher Intelligence”

  • 王冠 is not entirely comfortable with the term “data flywheel,” because it can imply that all usage data has value. If every conversation is fed back into the model without filtering, capability may converge toward the average user level rather than continually improve.

  • He gives 2 possible explanations for the period when ChatGPT was perceived as becoming “dumber”: one is cost optimization or switching to a smaller model; the other is only speculation—that a large volume of low-quality user data changed the model’s distribution. He consistently preserves the qualifier “possibly.”

  • Effective data is not merely high quality; it must be “above the model’s current level.” The model must learn from smarter samples. In the FSD example he cites, less than “one-thousandth” of currently available data may be usable for continued training, and the standard will keep rising.

  • Different industries require different filtering mechanisms, but the principle is the same: identify the portion of existing behavior that adds genuine intelligence to the system, then turn that judgment into an evaluation, reward or data process.

17. Applications Form a Different Path Through Goals, Position and Speed, Not by Hiding from Giants

  • 王冠 rejects the competitive premise that “only I can do this.” Any valuable direction may attract large companies and OpenAI; if they are not even thinking about the problem, that may instead suggest the direction has no value.

  • The first difference comes from the goal. Products that look similar may pursue entirely different endpoints, and the goal determines resources, path and subsequent feedback, which in turn changes where the company ultimately arrives.

  • The second variable is the starting point. Starting by designing proprietary data is fundamentally different from starting by solving a predetermined software function; even if the interfaces look similar in the short term, the underlying accumulation will differ.

  • The third variable is speed. Early chat products all looked like chatbots, and coding products all looked like Cursor, because methods had temporarily converged. The real question is who can maintain the lead and widen the distance from large companies over time.

18. ONE2X Designed a Video Language Before Drawing an Editor

  • ONE2X’s first step was not to design video-editing software, but to define “why video looks this way and how it is made,” then express the image and production process in a clear structure.

  • This DSL sits between natural language and code. It has a fixed format that humans may not be able to read directly, and abstracts video into objects, attributes, values and syntax, creating a domain dataset that did not previously exist.

  • From the DSL, the team derived data storage, engineering architecture, Agents, workflows, the software conceptual model and front-end functions. Every button and user instruction should be traceable back to the underlying language system.

  • 王冠 emphasizes that this starts from a different place than finding a commercial problem first and then adding software functions. The former defines what the system should eventually learn; the latter often defines only what users can click today.

19. The Environment Must Accommodate Actions by Both Humans and AI

  • ONE2X calls its software interface an environment because it contains 2 types of actors: humans and intelligent agents. A capability must first exist in the environment before either actor can take the corresponding action and leave data behind.

  • Why a video uses shot sequence A rather than B, how shots are combined and which steps are modified can all be recorded structurally during creation. The agent’s own actions generate data in the same way.

  • Ordinary SaaS products also generate large volumes of online logs. The key is not whether something was “recorded,” but whether the data can be learned from, mapped to a clear task and method, and ultimately used to improve the next output.

20. Expert Labeling Matters More Than Learning User Behavior Indiscriminately

  • 王冠 does not plan to use all user behavior as a basis for iteration in the early stage, because average user behavior may not represent effective data. For a long period, labeling should be done by people who genuinely know how to make good videos.

  • He compares the product to an internal labeling platform: professional roles use the same environment and leave video-production know-how inside the system, much as model companies invite high-level domain experts to provide training samples.

  • Once the data is internalized, users may feel that “the video I made yesterday and the video I made today are different” even if the front-end software has added no features. Product upgrades therefore do not happen only at the level of interfaces and buttons.

  • 张小珺 wanted AI to record and perform his fixed writing workflow. 王冠 pointed out that this is precisely method data that has “not yet been digitized,” and one of the clearest sources of the third dataset.

21. Video Is First a Business Where a Single Capability Can Make Money

  • The team “started from humble beginnings” with limited capital, so the early priority was to generate revenue quickly. Video has a higher unit value than text, images and audio, making it better suited to early commercialization.

  • In his research on the US market, 王冠 found that, apart from CapCut’s runaway lead, there may be 20-30 video-processing SaaS companies with ARR in the tens of millions of dollars. This is a classic “ant-tool market.”

  • The implication is not winner-take-all. Any single video capability can generate meaningful revenue if it is useful enough. For a resource-constrained team, this is a more realistic starting point than lower-value modalities.

22. Video Production Can Be Compressed into a Finite Set of Atomic Capabilities

  • Technically, ONE2X needs to select problems that can be designed as a closed domain. Like Go, with its fixed board and rules, a closed domain means the space of next actions can be computed.

  • Video processing can be broken into a finite set of “atomic capabilities”: a particular effect, a stylized text treatment and so on. Production is then these capabilities arranged and combined over time.

  • A minimal set of atoms may therefore express a large number of videos, making it suitable for a DSL and for allowing humans and agents to act, evaluate and generate training data in the same environment.

23. Once AI Flattens Production Barriers Across Modalities, Video Becomes the Starting Point for Content

  • The PC and mobile internet moved first through text, then images, audio and video—not because that was the order of value, but because it was the order of production difficulty. The internet solved connection, distribution and consumption, but not production, so modalities that humans could produce more easily developed first.

  • As AI gradually flattens the barriers to writing articles, making images, composing music and producing video, modalities with higher value and greater information density will take a larger share. 王冠 therefore says: “Video is a starting point for content in the AI era.”

  • “Starting point” means that higher-dimensional forms such as software and games will follow. When he made the choice in early 2024, video generation had begun to look commercially viable, while high-consumption-value generative games were still further away.

  • Text, images and audio will not disappear, but will increasingly become components of video, software and games. Their old boundaries reflected the fact that humans had to specialize; they do not mean machines will preserve the same divisions once they produce content.

24. The Creator Moat Will Shift from Skill Combinations to Taste and Recipe

  • Once generative capability becomes public supply, creators will still determine how it is used. 王冠 calls this difference taste, though internally he prefers recipe: a repeatable, transferable method of production.

  • The cooking analogy preserves the boundary. Give the same recipe to different people and the flavor will still differ, but kung pao chicken does not thereby become twice-cooked pork. A recipe locks in the basic type and method of content while leaving room for personal variation.

  • Creators will therefore resemble product managers in the AI era. Product managers design production capability itself; creators control how that capability is invoked, combined and expressed.

25. The Entry Point Is a Tool Because Generative Technology First Rebuilds Production

  • 王冠 equates “generation” directly with “production.” Since this cycle first brings a productivity revolution, the product form closest to technological change is naturally a production tool, not a mature community, distribution platform or consumer application.

  • Directly producing content supply is not impossible; the issue is timing. Multimodal System 1 is still changing rapidly, and a content company designed around the performance of a particular model today may soon have its costs and processes rewritten by the next capability leap.

  • Tools can sense new capabilities immediately while building System 2, data and user workflows in parallel. Once production capability stabilizes and costs fall, those accumulations can also migrate to distribution and consumption.

26. Current Multimodality Looks More Like a GPT-2 Moment Than a ChatGPT Moment

  • 王冠 sees the progress from the Will Smith videos to Sora 2 and Veo 3 as early integration: isolated tasks such as audio-visual synchronization and storyboarding, previously handled by multiple models, have for the first time been brought together in a more unified system.

  • But he is explicit that this is not the ChatGPT moment for multimodality; it is closer to the GPT-2 moment. The industry still needs to scale data and parameters, lower costs, improve generation and inference speed, and complete the models’ capabilities.

  • To use an existing video as a reference for a new one, the current workflow often asks Gemini to “shot-list” the source and convert it into a language script before handing it to a generative model. That is not end-to-end multimodality; it is language serving as an intermediary between modalities.

  • The ideal model should directly accept arbitrary inputs such as video, music and text, and understand the different modalities internally. In the future, it may also reason in video or audio itself, just as language models reason in language.

27. “Don’t Build Apps in the Nokia Era” Defines ONE2X’s Product Clock

  • ONE2X’s internal reminder is: “Don’t build apps in the Nokia era.” Nokia had calculators and simple games, but they were not the same product opportunity as apps in the iPhone era.

  • 王冠 believes current multimodality may still be in the Nokia era or even the mainframe era. The better opportunity is to build tools close to production that continuously adapt to models, rather than assume the underlying capability is stable and prematurely lock in the final consumer form.

  • Possible signals of an iPhone moment include end-to-end multimodality with the input and reasoning capabilities described above, a relatively stable System 1, and another several-orders-of-magnitude reduction in inference speed and cost.

28. The “Shell” Is Not the Problem; the Quality of Its Context Is

  • 王冠 accepts the description of an application as a “shell”: “The word shell is not pejorative at all.” System 2 is fundamentally a shell built around System 1, although it can be thin and simple or thick and constitute a complete domain system.

  • “Context is everything” means that on top of the same foundation model, product performance is ultimately determined by external context. The video DSL is the foundation layer for all subsequent context.

  • Competition in context engineering includes Agent architecture, workflows, domain libraries and memory. A ReAct-style free loop may consume huge volumes of tokens; for determinate problems, predefining the steps can reduce newly generated tokens.

  • “Everything is context” points toward the terminal model: music, images, text and even another video should be able to act as native context that affects the output, rather than first being translated into language.

29. Model and Application Companies Will Converge on the Same Company from Opposite Ends

  • Model companies used to be viewed as technology companies that did not build products, but today every model company is building products. Application companies, correspondingly, should not be defined as software companies that will never build models, but as “companies that have not yet started building their own models.”

  • 王冠 uses Cursor as an example: it has been continuously trying to build its own models. Once its coding business, ARR and financing scale become large enough, it can pursue the training and compute path previously taken by foundation-model companies.

  • Owning a model matters both for performance and margins. Calling someone else’s model may be cheaper in the short term, but an application cannot optimize its capability distribution; only by building its own can it reduce costs for its scenarios and retain control of the cost structure.

30. Vertical Applications Will Eventually Be Forced to Train by the Boundaries of General Models

  • General models have finite parameters, which must be allocated across different tasks. Some capabilities reinforce one another, while others may conflict. A model optimized for generality and an application optimized for a specific commercial outcome cannot have identical capability distributions forever.

  • Once a vertical team reaches the foundation model’s boundary, it can only keep tuning prompts, lower its target or abandon the desired result. The host suggested that an application company could turn its own scenarios into evals and seed data for domestic model makers; 王冠 regards those model makers as “its own model department.”

  • Once the data enters a shared model, it may also serve others. 王冠 accepts this trade-off: at least the model will fit his scenario better, and he should be the person who best understands how to invoke that capability.

  • Once the business is large enough, a proprietary model becomes necessary for performance, cost and competition. If the upstream company enters the same product category, the parties shift from partners to opponents, making it even more dangerous for the application to hand its fate entirely to an external API.

31. The Only Temporary Advantage Applications Have Against Giants Is Speed

  • As goals, resources and technology gradually converge, the outcome will come down to talent density, decision speed and what 张一鸣 calls “a difference in understanding of one particular thing.”

  • 王冠 acknowledges that AI startups consume more compute, capital and GPUs than internet startups. An application company must at least lead large companies on its own map: its start date, business state and data-design approach should create a positional gap.

  • The question later is not whether the giant will arrive, but whether the gap is widening or shrinking. A small team’s advantage is lower organizational overhead and greater freedom of action; if it starts behind, the word “vertical” alone will not protect it.

32. No “Everything Bursting into Life” in 2025 Does Not Mean the Agent Path Has Failed

  • 王冠 disagrees with 张小珺’s criticism that progress is below expectations from the start of the year. Coding Agents, Manus and Lovable have already shown that building Agents, workflows or shells around models can solve commercial problems and make money.

  • The market feels progress is below expectations largely because there are not yet enough examples and the scene has not become one of “vitality everywhere, everything bursting into life.” But many vertical fields are in a state of “deep currents beneath still water,” and the speed at which they acquire the required model capabilities, know-how and domain libraries naturally differs.

  • 张小珺 cites Harvey and OpenEvidence in law and healthcare, arguing that they perform better through stronger System 2, domain libraries and data quality. 王冠 responds that model readiness, know-how, team conditions and resources differ by field, but the path has been validated.

  • For video, asking whether “the model is good enough” has limited value. The team can only add value within current capabilities; the real commercial question is whether the cost of System 2’s contribution is below the price users are willing to pay.

33. Agents Will Eventually Disappear into Product Names, Just as the Internet Did

  • 王冠 dislikes calling companies “Agent companies,” because an Agent is a technology that will ultimately be internalized into every product. People will not forever describe Douyin as a “mobile-internet short-video app.”

  • He uses the phrase “the moon shines on every river” to describe vertical systems: each river has its own data, intelligence and commercial value, and a small product may contain many Agents, tools and complete chains of work.

  • A general Agent is like “a sky stretching for thousands of miles beneath a cloudless horizon.” If everything claims to be general, in the long run there may seem to be only one sky. But that is merely a transitional state and does not mean vertical and general systems will remain permanently separate.

  • Vertical systems will expand their capability boundaries. Video might even be “downgraded” into a PPT with image-text structure, while general systems must also go deeper on high-value tasks. The final questions remain simple: what is produced, how good is it, and how much does the user pay?

34. A Generative System Is a Method, Not a Synonym for “Generating Video”

  • 王冠 uses recommendation systems as an analogy. Recommendation is a general method: recommending articles became Toutiao, recommending video became Douyin, and recommending jokes became another product. Generative systems can likewise serve video and migrate to other information goods.

  • From its first day, ONE2X treated System 2 as a generative system and divided it into 3 modules: a DSL defining objects and methods, context that constructs effective tokens, and an environment where humans and agents act together.

  • The debate between Agents and workflows is not important within this framework. The Agent side criticizes workflows as unintelligent, while the workflow side criticizes Agents as uncontrollable and expensive; in substance, both are simply methods for generating effective context.

35. Context’s Direct Task Is to Reduce Both Intent Entropy and Action Entropy

  • A user’s request—“Give me a report”—contains enormous uncertainty. Early models would write poorly or repeatedly ask questions; reasoning models add context by unfolding reasoning, essentially reducing the entropy of the user’s intent.

  • An intelligent agent must also plan, select tools and execute steps, so it faces action entropy. If context cannot constrain the next step, the system may invoke the wrong capability, take an unnecessarily long path or incur uncontrollable costs.

  • A generative system must turn vague intent into precise action and make the mapping stable at both ends. Controllable performance does not require users to write longer prompts; it requires the system to perform the necessary clarification on the user’s behalf.

36. The Real Value of an Environment Is That It Can Transfer Recipes Without Loss

  • The environment’s GUI will still have buttons, editors and operations, but intelligent agents may gradually become the main users. Humans will use the same interface for validation, correction and fine-tuning rather than performing every step themselves.

  • 王冠 uses the phrase “a pinch of salt” in Chinese recipes to illustrate information loss. If an environment can record the minute a chef adds oil, the angle of the arm, the oil temperature and the pan temperature, and let another system understand and execute them, the recipe can be reproduced with near-zero loss.

  • An environment therefore cannot merely save logs. It must correspond one-to-one with the DSL, tools and methods, and include a reward function that selects the parts of activity data worth reinforcing.

37. Good Video Is an Open-Domain Problem That Requires Layered Evaluation

  • If a user accepts an Agent’s output and makes few changes, that suggests the method may have worked for that user and can become a signal to reinforce. The same output may not work for a user with a different level or preference.

  • ONE2X places greater weight on internal expert labeling, with product managers and people who possess video aesthetics and production skills evaluating together, rather than treating all user behavior as ground truth. 王冠 compares this to the internal artist role at Midjourney.

  • The evaluation system must ultimately allow the system to distinguish between making a 60-point video and making a 70-point video, while continuously retaining samples that genuinely raise overall capability. This is the observable foundation of what he calls “system intelligence.”

38. ONE2X Is Targeting the World of Ideas, Not More Camera Records

  • 王冠 broadly divides video into the physical world and the world of ideas. Smartphones and cameras everywhere have already made the supply of the physical world abundant, and most content in today’s short-video ecosystem comes from reality that can be filmed directly.

  • The world of ideas includes knowledge, thought, formulas, imagined images and brand stories, none of which can be obtained through cameras alone. Turning an article into animation, motion graphics and visual structure is closer to “video-izing ideas” than a talking head.

  • Internally, ONE2X wants to build a “library, opera house and cathedral” in the video world, corresponding respectively to knowledge, art and spiritual content. The existing ecosystem is more like “nightclubs, public squares and supermarkets,” corresponding to dopamine, socializing and selling.

39. The Opportunity in Idea-Driven Video Comes from the Cost Curve, Not a Sudden Desire to Study

  • 张小珺 questioned why the company should pursue spiritual and knowledge content when entertainment has already demonstrated commercial value. 王冠 did not answer that people would suddenly love learning; he emphasized that the supply of physical-world video is already abundant, while idea-driven video remains constrained by high production costs.

  • One customer produced videos for a trendy-toy brand and, using ONE2X, “bought out” the AI videos related to that toy on Xiaohongshu. Trendy toys naturally lend themselves to stories and imagined settings, but when production took weeks and cost thousands or even tens of thousands of yuan, the business did not work.

  • A brand advertisement at the level of Apple or Nike might once have cost hundreds of thousands or even millions of dollars per spot, putting the same form of expression out of reach for ordinary products. Generative systems may not replicate top-tier quality, but by lowering costs they can make previously nonexistent content categories commercially viable.

40. The Half-Finished Product Already Shows Real Signals of Payment and Revenue

  • At the time of the interview, the formal version had not launched. In May, the team quietly released a product that was “half finished” because early customer tests had already shown strong commercial value, even though the underlying DSL and context were incomplete.

  • A leading AI content creator with substantial views on Bilibili and WeChat Channels has used the product for an extended period, with individual works potentially generating hundreds of thousands to more than 1 million views. He borrowed Google accounts from people around him, maxed out every points package, and eventually contacted the team because the points still were not enough.

  • 王冠 himself used an earlier version of Mydou to make WeChat Channel videos for more than a month, generating more than 2 million cumulative views. One day the platform notified him that a traffic revenue share had arrived; the amount was only several hundred yuan, but it was the first time he confirmed that “this product really can make money.”

  • The December version mentioned by the host is not merely a feature upgrade. It reconnects the DSL, context, Agent and environment into a complete generative system. The earlier commercial validation proved local capabilities, not the final product.

41. AI Will Move Power Further from Production and Distribution to Consumption

  • The software era was like a supply-and-marketing cooperative: producers made whatever they chose, and consumers used it, concentrating power on the production side. The internet took over distribution and consumption through search, recommendation and e-commerce platforms, shifting power to platforms.

  • Following the same trajectory, AI should hand power further to consumers. The information goods users see can be generated instantly based on their user profile, environment, and even their mental and physical state, producing personalization more extreme than recommendation.

  • 王冠 uses Infinite Tsukuyomi from Naruto to describe the distant endpoint: every person lives in a content world tailored to their own condition. He also emphasizes that this will take a long time and is not on the product roadmap today.

  • Along this long chain, 20 years of internet content have inadvertently become training data for large models. The new data defined by application companies today may likewise become one piece of a future larger model, rather than AGI being completed by a single lab alone.

42. AGI May Be a State or May Emerge Point by Point in Commercial Domains

  • 王冠’s broad definition of AGI is a model that truly knows what it does not know. Humans can distinguish between knowledge and ignorance, while a model’s refusal today may simply be a trained pattern; neither the model nor the user can confirm that it actually recognizes a capability gap.

  • Once a model can identify what it does not know and already has information-access and tool-use capabilities, it can actively learn to fill the gap. Even without all knowledge in place, reaching this state would mean that subsequent learning could be rapidly automated. He acknowledges that this is currently difficult to evaluate technically.

  • A narrower definition occurs in a specific commercial domain: AI can earn money or resources on its own, then buy data, compute and GPUs, complete evals, optimize itself and earn even more. Stock trading is the thought experiment he gives.

  • This would be gradual. Language, coding and other domains would be “lit up” one by one as humans gradually leave the optimization loop. Video remains much earlier because end-to-end multimodality is still immature, putting it further from local AGI.

43. The Core Challenge Generative Systems Pose to Recommendation Platforms Is “No Middleman Taking a Cut”

  • 王冠 calls search, recommendation, e-commerce and nearly every internet platform a middleman: they do not directly produce information goods, but control how content is distributed and seen.

  • He uses Newton and Leibniz publishing calculus at the same time as a thought experiment. Even if the content, timing, location and creator profile were identical, the platform could still deliver different levels of traffic, showing that the same goods do not receive the same distribution.

  • Generative systems do not eliminate matching altogether; they internalize matching into production. Consumer demand goes directly into production, where the system generates and delivers the result instantly instead of first producing inventory and then letting an independent platform decide what to display.

  • The same logic may extend to physical goods, connecting demand directly to programmable production lines. 王冠 only says he believes teams are already exploring this; he does not claim a mature general-purpose goods system exists.

44. Creators Will Not Go to Zero, but Will Split into “Artists” and “Prosumers”

  • The first group consists of people who can consistently provide incremental intelligence above the system baseline. They may become employees or partners of generative platforms, contributing better recipes, taste and training data, while the system copies the strongest person’s capability 10,000 or 100,000 times.

  • The second group is the prosumer, where production itself is consumption. Retirees practice calligraphy, users ask ChatGPT to research topics they care about, and fans generate images of niche characters; the work creates value for its maker first.

  • AI images have already eliminated some niche fan-art package businesses because users can directly generate the characters they want, but more people are participating in creation. 王冠 therefore believes generative systems will expand the scope of UGC rather than simply eliminate all creators.

  • 张小珺’s rebuttal is worth preserving: very few people can continuously create new recipes, and once machines learn them, the original creators may no longer be needed. 王冠 does not deny the shock; he only says the transition may be very long and, logically, difficult to avoid.

45. Recommendation Algorithms Are Training Creators to Become People Who “Hit the Database”

  • 王冠 gives the example of a 2- or 3-second video that first shows a distant mountain and then suddenly zooms in without ever providing a real answer. Viewers replay it because they did not see clearly, raising both completion and repeat-play rates.

  • “The golden 3 seconds,” delaying the answer and repeatedly setting hooks are all ways creators adapt to recommendation mechanics. The production objective becomes not the work itself, but hitting the platform’s rules for calculating attention.

  • This is the opposite of the prosumer. A prosumer only needs the content to have value for themselves, while a platform creator competes for returns in a public traffic pool. 王冠 uses the example to show how distribution power can reshape and even distort supply.

46. The Endpoint of AI Democratization Is Not Everyone Becoming a Creator, but Creation Becoming Expression

  • Language is available to almost everyone. Text has a higher threshold; 王冠 believes that, across the world’s population, only roughly 70% of people may be able to read and write. Images, music and video are progressively harder to use as personal expression tools, but higher modalities also transmit more information.

  • When writing required brushes, ink, paper and inkstones and professional creation, books were extremely precious. Only after text became a low-cost form of expression could WeChat emerge. 王冠 expects video to undergo the same transformation rather than merely making today’s editing software faster.

  • Love letters and wedding videos are his concrete examples. Today, a groom who makes a video of the journey from meeting to marriage is praised as “very thoughtful” precisely because the work remains expensive to create. In the future, it should become an everyday form of expression like speaking or sending a message.

  • Mydou currently charges in a tool model, with copyright belonging to the person who makes the content. At least under the current economic structure, the person who uses the tool first should receive greater efficiency and better returns.

47. Once Content Supply Shifts from Individual Items to IP, the Currency of the Economy Moves from Attention to Trust

  • Internet creators produce discrete pieces of content that enter a public pool for algorithmic distribution. Producers in a generative system are more likely to create recipes that continuously generate content with shared attributes, approximating an IP.

  • Once stable-quality content is no longer scarce, scarcity will not disappear; it will move. 王冠 believes it will shift from individual pieces and short-term attention toward long-term trust in an IP, author or channel.

  • Substack, Medium, OnlyFans, and public-account writers directing users to Zhishixingqiu or private communities are early forms of this model. Public traffic may be limited, but high-value audiences are willing to pay continuously because they trust the source.

  • The new system can still accommodate search, recommendation and consumption by others. The change is that self-production and self-consumption become the base value, while external traffic is no longer the only way to monetize. Once everyone has a high-quality baseline, what remains scarce is taste, selection and credible continuity.

48. ONE2X Uses “System Intelligence” to Unify Product Metrics, Organization Design and Its Long-Term Bet

  • 王冠 is not attached to a single AI-native form. Poe-style capability aggregation, tools with editing capabilities, Agents and workflows have all demonstrated value. ONE2X simply chose the slower, heavier DSL path because it better fits the team’s DNA.

  • Commercial value is the baseline; the north star is system intelligence: raising output from 60 points to 70 points for the same input, or using fewer tokens for the same quality. That is why “3 users contributing 1 million” may be better than 100,000 users contributing the same revenue—the generative system is meant to scale the strongest people.

  • He writes product intelligence as “organizational intelligence × conversion rate.” At the time of the interview, the team had roughly 30 people, worked remotely, and held one in-person day per week proposed by members themselves; there were no dedicated managers, and all 3 co-founders remained hands-on. Hiring emphasized Building in Public, open-source projects, work outside the job and passion for specific things. Loneliness and trust problems from remote work were offset through a “warm and trusted program” and an internal Feishu social feed.

  • The company began in early 2024 and was not formally incorporated until Q2, having prepared to bootstrap in its early days. DeepSeek reignited confidence in models, while Manus more directly rekindled capital’s enthusiasm for applications, prompting 王冠 to joke that “everyone should bow down to Manus.” The team is drawn mainly from first- and second-degree networks, and roughly half have been founders or co-founders. His final macro judgment is: “AI or AGI, it’s a long-chain thing—so I bet China.”