Vol. 48 The AI Industry in 2025: Whether It’s Native or Not May No Longer Matter — with 三五环刘飞
Summary
AI’s main battleground in 2025 may shift from model races to application delivery, with product definition, operating speed, and organizational coordination becoming scarcer assets as model capabilities converge. ByteDance, Alibaba, and Tencent are all separating or restructuring their model and application teams, reflecting a shared view that “models are models, applications are applications.” OpenAI’s jump from GPT models to ChatGPT was enormous, but the market now wants products to go “another ten or twenty steps” beyond that.
Plunging inference costs are expanding demand while potentially continuing to compress the revenue pool for the model layer and cloud API market. 刘飞 relayed 修寒’s rough estimate that leading third-party applications could each generate daily traffic on the order of tens of billions of tokens, with perhaps hundreds of such applications, implying a market of only several billion yuan to RMB10B. Per-token costs fell roughly 100x in 2024; if total costs fall another two orders of magnitude in 2025, maintaining the existing market size would require a 100x increase in usage. 刘飞’s social business currently consumes several million tokens a day, and he doubts such usage KPIs are realistic for most applications.
“Applications” and “Agents” are not separate tracks: most AI products classified today as scenario applications can be understood as small Agents with controlled boundaries. Perplexity breaks down questions, calls information sources, and recombines the answers; AI coding, AI PPT, and file summarization similarly package general-purpose models into defined workflows. The key is not whether a product carries the Agent label, but whether it can find the “key” that fits a specific user need and keep refining every tooth.
The advantage of products such as 豆包 and 夸克 comes not only from their underlying models, but from the organizational ability to rapidly replicate validated use cases and turn them into low-friction experiences. 庄明浩 chooses products based on whether the team can quickly follow new breakout use cases; “this is completely not a technical reason, but an organizational and operational one.” 夸克 has evolved from a browser and “Swiss Army knife” toolkit into “an AI all-purpose assistant for 200 million people,” while 豆包 positions itself as a “super assistant.” Both are competing for the next unified entry point.
The AI Native label is losing decision-making value: users do not care whether a feature is powered by a foundation model, rules, or human judgment; they care only whether it solves the problem. 刘飞’s summary is that “AI is not a scenario at all,” just as no one calls themselves a C or Java product manager. A knowledge-base founder once removed language such as “AI food analysis” from the product description, suggesting that technology narratives aimed at investors are giving way to real usage and retention.
AI glasses are one of the strongest areas of hardware consensus in China and the US, but the opportunity comes with supply-chain, inventory, and form-factor risks. The market could split between products using Qualcomm chips priced at roughly RMB1,500–2,000 and domestic-chip products priced below RMB1,000; glasses can also capture first-person, personalized data, creating new inputs for generation and Agents. But compute, battery life, weight, thermal management, and whether to use a split-device design remain unresolved. “All those pitfalls are still right there.”
Near-term automation value will appear first in tasks with verifiable outputs, while subjective content production remains clearly behind coding. The path from Cursor to Devin has made “hand-writing code becoming performance art” imaginable, but AI PPT and content summarization still miss what humans consider important. The most practical way to participate is not to “learn AI” in the abstract, but to use products such as 豆包 and 夸克 against everyday tasks and search for product opportunities in real-world friction.
Deep dive
1. The AI Industry’s Time Unit Has Shrunk to Two or Three Months
庄明浩’s strongest impression of 2024 was that everything was “accelerating insanely.” He began preparing his annual review at the end of November; by December 11 or 13, it had reached 132 pages. When he presented it roughly a month later, the deck had grown to 143 pages, while the original material continued to become obsolete.
This shift is not limited to foundation models. Every hot subcategory can turn over completely “once every two or three months”; news, conclusions, and product perceptions from three to five months earlier quickly lose their validity, while industry attention moves in waves, with one wave piling onto the next.
刘飞’s observation is that after ChatGPT 3.5 appeared, the initial shock centered on underlying technology. By 2024, the application and middleware layers had begun to diverge and evolve rapidly even when they shared the same foundation-model logic.
2. Model Curves Are Converging, While Cheaper Tokens Create a Growth Paradox
By Q3 2024, it was increasingly becoming consensus that pretraining was nearing its ceiling, shifting the discussion toward post-training, distributed inference, and Agents. The question of when the internet’s existing data would be fully “consumed” also emerged. Whether in benchmark rankings or the distance between open- and closed-source models, 庄明浩 saw “all the curves converging toward roughly the same place.”
刘飞 added that compute-card prices have become more rational after the initial bubble, and some companies increasingly feel that pouring more resources into giant models may not produce results. This does not mean technology has stopped advancing; teams are recalculating the return on investment of the model race.
刘飞 relayed 修寒’s rough estimate: looking only at China’s leading third-party applications, daily usage could be on the order of tens of billions of tokens. There may be hundreds of such leading applications, implying a market of roughly several billion yuan to RMB10B on a cloud-services basis. The estimate depends on assumptions, but it shows that the model-calling market may not be as large as its traffic numbers suggest.
The sharper issue is pricing. From the beginning to the end of 2024, per-token inference costs may have differed by 100x; even if the 2025 decline is less than 100x, costs could still fall another 10-20x. Combined with stronger multimodal capabilities, total costs could drop another two orders of magnitude. 刘飞’s social business already consumes several million tokens a day, but he says that raising its usage KPI 100x simply to preserve industry revenue would put “enormous pressure on the business team.”
3. The So-Called Agent Is Often Just an Application with Clearer Boundaries
Responding to the idea that “2024 was the year of applications and 2025 will be the year of Agents,” 刘飞 asked where the difference actually lies. 庄明浩 used Perplexity as an example: it breaks a search question into modules and reasoning steps, calls familiar information sources to collect and organize material, and presents the result in a unified format. By a broad definition, that is already an Agent.
Under this definition, most AI products classified as scenario applications can be viewed as “small Agents.” The issue is not whether they carry the Agent label, but whether they can decompose tasks, select tools, constrain workflows, and deliver results users can evaluate.
The difficulty in 2023 was that both the technical frontier and user needs were changing rapidly, leaving product teams unable to find a stable point between two moving curves. By 2024, what models could do had become clearer, allowing teams to define a “small boundary” and package the fit between demand and technical capability into concrete features.
4. The Moat for Big-Tech Applications Is Organizational Speed in Following Use Cases
By Q3 and Q4 2024, products such as 豆包 and 夸克 had reached larger user bases and gradually become industry consensus. 庄明浩 explained that when he chooses an application, he is not paying for differences in underlying technology; he is betting that its product and operations teams can quickly follow any new use case that proves itself.
That means the time lag between users and frontier capabilities will remain short even when a new application breakout appears, and new features will arrive through familiar interfaces and operating methods. 庄明浩 stressed that this is “completely not a technical reason, but an organizational and operational one.”
刘飞 used 豆包’s evolution to make the point. It initially resembled a general chat tool following the virtual-avatar craze, then rapidly improved its visual design, interaction, and details. He also noted that 庄明浩 had previously pointed out abnormal product responses at 极客公园, and the team was able to respond quickly. That density of response is a luxury for a small team with limited resources.
5. “Models Are Models, Applications Are Applications” Has Reached the Org Chart
ByteDance’s 豆包 model team was already separate from its product and operations teams. Alibaba subsequently had 夸克 take on the team formerly responsible for the Tongyi App product, while keeping model-layer R&D within the model team. Tencent kept the 元宝 model team within its technical organization and merged the application team with the Tencent Meeting team. The common direction is clear: “models are models, applications are applications.”
庄明浩 believes that from 2023 through the first half of 2024, “the model is the application” was the more popular idea. But as competition intensified, scenarios, operations, user experience, and technical implementation all became more granular and complex; product experimentation led solely by R&D personnel was “starting to fall behind.”
OpenAI’s leap from GPT models to ChatGPT was enormous, but ChatGPT initially felt like something that had suddenly escaped from an internal Demo Day to become a core product. Only after the addition of GPTs, Agent services, and Calendar integration did it begin to look like more than a technology demonstration. Users today are not asking products to take one more step; they want them to go “another ten or twenty steps” outward around a clear need.
At OpenAI, most employees used to be focused on technical R&D; now perhaps half are focused on product and operations. 庄明浩 sees the same shift taking place in China and the US: model capability still matters, but application delivery has become a separate professional discipline.
6. The Foundation Model Is the Sun; Products Are the Conduits That Deliver Its Energy
刘飞 acknowledges that his view changed. When ChatGPT first appeared, he thought that if a foundation model could handle NLP sub-tasks such as tokenization and machine translation in a unified way, perhaps a chat window could cover every product form. For a time, he did not understand why Perplexity or standalone AI coding products were still necessary.
Later use convinced him that a foundation model’s ability to cover a scenario does not mean users can directly obtain the result. Prompting still has a learning curve, and every specific task needs additional workflow, context, interaction, and result-presentation layers.
庄明浩’s analogy is that a foundation model is a “sun” with enormous light and heat, but people cannot easily turn the sun directly into usable energy. A product is a conduit with a predefined length and width that delivers exactly enough capability to the point of demand. “The sun itself is still extremely powerful, the source of everything,” but once the sun becomes stable, more value is generated by the surrounding conversion layer.
7. 夸克 Shows How AI Can Amplify the Capabilities of an Earlier-Generation Product
庄明浩 was already a 夸克 user before the AI wave. He valued its simplicity, lack of advertising, cloud-drive integration, and small tools such as college-application planning and PDF processing, which together formed a “Swiss Army knife.” In an era when the browser landscape was largely settled, those experiences were enough to attract some users, but not necessarily enough to unlock the mass market.
AI added greater leverage to that accumulated base. Search, files, plugins, and various generative capabilities could be recombined, while the original product experience could continue to improve. The real difficulty, beyond simply adding a few icons, was deciding where features belonged, how to connect them to legacy workflows, and how to navigate organizational and page-level control inside a large company.
庄明浩 initially used 夸克 mainly as a Perplexity substitute. Only later did he notice scenarios for younger users, such as photo translation, scanning exam papers, and organizing wrong answers. The underlying technology is not mysterious on its own, but once packaged, it changes the experience qualitatively: users no longer need to search for instructions or hand-write complex prompts.
Selection matters just as much. Which scenarios enter the homepage’s handful of core slots, which are hidden in the search box or a secondary page, and which are built now versus later all determine whether an all-in-one product becomes a useful system or a chaotic toolbox. The existence of good products in specialized fields does not mean a platform should place every feature in the most prominent position.
8. A Universal Input Box Cannot Replace Well-Designed Buttons and Workflows
OpenAI’s desktop client looks like “just one box,” close to the old ideal of “box computing” and all-in-one functionality. 刘飞 points out that this approach may not suit mainstream users in China. People can express some needs in natural language, but more functions still need to be buttonized, foolproof, and executable in as few steps as possible.
AI PPT illustrates the divergence within the same underlying need. Some tools emphasize summarization, organization, and outlines; some rely on extensive template libraries; some excel at visualization; others focus on logical structure. Before recommending a product to a friend, 庄明浩 first asks about the purpose, scenario, reference material, text-to-image ratio, and level of formality. There is no single generic answer.
For most white-collar users in China, the common denominator may still be selectable, editable templates that match the content. As users move from “summarize this for me” to specifying structures, the pyramid principle, and complex output formats, “every step up is a massive source of user churn,” because every step requires users to understand the task better.
刘飞 therefore believes the industry is entering another period in which product managers gain value. They must preprocess complex prompts into scenarios and buttons. 庄明浩 adds that a good product manager in this generation must also understand model boundaries precisely; otherwise, they cannot truly connect technology with the human tendency toward “laziness.”
9. The Next-Generation Browser May No Longer Be Called a Browser
庄明浩 cited the view that “your next browser may not be a browser.” Chinese browsers had already built in a large number of tools before AI, but a mature market made it difficult to reactivate users. AI gives the search box a chance to become a unified entry point for calling search, applications, and even Agents.
This also revives the logic behind Chrome’s earlier attempt to become an operating system: a computer needs only one entry point, through which shopping, socializing, content, and tools are all completed. That vision ran into limitations after several years, while mobile internet further compressed the browser’s use cases. AI may now reconcentrate mid-tail and long-tail scenarios into one box or a new interface on PCs and even phones.
Browser plugins are suitable for early, small-scale PMF validation, but are inherently constrained by the browser itself. Once a product wants to extend further, it cannot remain satisfied with the plugin form. Whether OpenAI, Anthropic, Chinese startups, or big tech, “no one will be satisfied with building a plugin.”
10. The All-Purpose Assistant Is Absorbing the Old Boundaries of Search, Browsers, and Input Methods
夸克 has updated its slogan to “the AI all-purpose assistant for 200 million people” and no longer emphasizes search or the browser. 庄明浩 cited a 七麦数据 report saying its downloads rank first in the category. 豆包 now says on its homepage, “Your super assistant is online.” Products with different origins are beginning to look increasingly alike.
This move toward all-purpose functionality is expensive. Beyond basic conversation, products must continuously add images, music, video, and other modalities, along with translation, Agents, PPT, file summarization, mind maps, and an endless range of extensions. Most teams cannot keep up over the long term; only a few giants with the determination to commit can sustain the investment.
庄明浩 reinterprets 王小川’s “three-stage rocket” of search engine, input method, and browser. In the AI era, all three are fundamentally entry points for users to interact with models and may eventually be unified. For now, they are still experimenting along their old product boundaries, but leading AI assistants on PCs already “look like browsers” in the UI while operating on an entirely different underlying logic.
11. What AI Native Cannot Answer, User Value Will Answer Directly
On the definition that only products that would not exist without AI qualify as AI Native, 庄明浩 believes there is still no definitive answer. Even in games, it remains unclear what an ideal AI Native game would be. AI is already more diverse and complex than the previous generation of technology; adding another Native layer only makes the boundary harder to draw.
The mobile internet offers counterexamples. Pinduoduo’s long-standing decision not to build for PC and Web does not automatically make it a Mobile Native e-commerce company. Didi and Meituan, by using location data, are closer to narrowly defined mobile-native capabilities. Distribution form, technical characteristics, and product value are not the same concept.
刘飞 believes the AI Native discourse had a historical rationale. After ChatGPT triggered a financing bubble, the market needed to distinguish teams genuinely built around the new technology from companies that added a chatbot and called themselves AI. Once competition moved to the application layer, judging products by technical purity lost its meaning.
A knowledge-base founder once removed language such as “AI food analysis” from the product description, and the results improved because “users really don’t care that much whether it’s AI.” 刘飞 takes the point further: “AI is not a scenario at all.” No one calls themselves a C or Java product manager. A product can combine foundation models, rules, human judgment, safety, and compliance; if the result is better, it works all the same.
12. The Real Moats in AI Products Often Hide in Rules, Costs, and Local User Habits
庄明浩 used 小宇宙’s AI summary to pose an implementation question with no obvious answer: after a popular show has been summarized for the first subscriber, should the second person who clicks receive the existing result or trigger a new generation? That small choice involves cost, prioritization, and dynamic balancing, showing that an application is never just a one-time model call.
刘飞 recalls that the industry spent much of the past year tracking Transformer paper authors, technical teams, and Altman’s background. Entering 2025, the more important questions are what products companies actually build, what needs they meet, and what value they deliver. This was not previously irrelevant; technology and the market simply had not reached this stage.
庄明浩 believes US teams move more naturally into To B and enterprise services because VC or customer funding is easier to secure. China, meanwhile, has a huge internet population and 20 years of accumulated To C product and operations know-how across the internet and mobile internet. For many boundary definitions and scenario packages, “you really can’t expect a US company to build them.”
刘飞 now uses 豆包 and 夸克 more often than Perplexity and ChatGPT, and believes domestic product experiences have improved markedly from a year ago. TikTok’s competition also reminds him that the capabilities of consumer applications include more than the interface: they also encompass operations, supporting infrastructure, and complex execution systems.
13. AI Hardware Opportunities Are Clear, but None of the Previous Generation’s Pitfalls Have Disappeared
In his annual outlook, 庄明浩 is bullish on AI hardware in both China and the US. The recurring presence at CES of glasses, rings, companion robots, humanoid robots, and machines that clean pools and lawns suggests that several directions have become “an absolute consensus.”
Consensus does not mean easy. Hardware must absorb real costs involving supply chains, tooling, minimum production runs, channels, and inventory. AI brings new capabilities, but “all those pitfalls are still right there”; a mistaken product boundary is far more expensive to test and revise than in pure software.
AI glasses alone could split into two tiers. One would use Qualcomm chips and sell for roughly RMB1,500–2,000, competing with the Ray-Ban model. The other would use domestic chips and potentially fall below RMB1,000, with some experience gap but possibly faster cost reductions.
The form factor has not converged. Weight, battery life, thermal management, and power consumption may force compute or power into a separate device, but it is still unclear whether AI glasses need another device at all. Watches and rings may eventually handle sensing, gestures, and control; how Apple, Meta, and other ecosystems combine glasses with existing devices remains open to extensive experimentation.
14. Glasses and Companion Robots Must Move from Passive Q&A to Active Perception
Glasses can continuously capture first-person, personalized images and environmental data—inputs unavailable from public internet corpora and potentially valuable for personalized generation and Agents. 刘飞 therefore sees them as a genuinely exciting consumer scenario of the kind that has been rare in recent years.
火山引擎 once demonstrated a business-trip workflow: an employee filmed with a phone and asked where they were, which vehicle to take, when their flight was, and what the itinerary looked like. 庄明浩 thought the video felt awkward because “take out the phone—open the camera—film—then start talking” involved too many steps. With glasses, the entire workflow would connect naturally.
Companion devices expose the same problem. 庄明浩 has a children’s walkie-talkie with a built-in foundation model and a first-generation AI plush toy at home. But after the novelty wore off, his son, who was ten and nearing eleven, and his five-year-old daughter “didn’t know what to say to the toy.” Companionship cannot depend only on children initiating questions.
Next-generation products need visual scanning, memories of family members, proactive greetings, and coordination with devices such as children’s watches. Internet interaction once moved from manually typed URLs to portals, search, and personalized recommendations. 刘飞 believes the past two years of AI entrepreneurship have also produced a cohort that understands models, the internet, and hardware simultaneously, raising expectations for hybrid founders in 2025.
15. Coding Is Near the Automation Threshold; PPT and Content Summaries Remain Unreliable
PPT remains 庄明浩’s most discouraging use case. Existing tools cannot satisfy demanding users. If one day even those users can confidently hand the work to AI, manual layout will become a historical artifact. Coding, by contrast, is already much closer to that threshold, from Cursor to Devin: “hand-writing code will eventually become something done by historical relics.”
The difficulty in content summarization is not that the prompt is too short; it is that the task contains subjective judgment. Even with a highly complete framework as input, 庄明浩 worries that the model will omit the core point. By contrast, when he sees a summary in the style of “蓝心一言,” he feels confident that important content has not been lost and can even skip the original video or podcast.
庄明浩 therefore distinguishes deterministic from non-deterministic tasks. Explaining or searching for problems, removing watermarks, generating Excel formulas, and writing code all have relatively clear inputs, outputs, and validation standards, so they will be automated faster. Content summarization, analysis, and creation require judgments about “what matters,” so their automation curve will not fully replicate coding.
16. The Most Effective Way to Learn AI Is to Use It Against Real Tasks; the Best Startup Opportunity Is to Refine One Key
刘飞 does not recommend “learning AI” in the abstract. Reading papers and studying company histories do not automatically translate into capability. A more direct approach is to start with domestic products such as 豆包 and 夸克. If PPT is central to one’s work, keep testing whether AI PPT can cover the real workflow; let the task, rather than novelty, drive the learning.
庄明浩 has also moved past his most anxious phase. Even if Chinese foundation models still lag the US and it remains difficult for startups to train large models independently, the conclusion is not to stop: “The work still has to get done, and life still has to go on.” There are already more than enough product problems on the immediate agenda.
刘飞 compares application exploration to “keys and locks.” A table with 100 keys may initially make them look similar, but a team must keep adjusting the number, height, sharpness, and slope of the teeth until one matches a lock’s cylinder. There will inevitably be failed attempts, but without repeated refinement, no final product-market fit will emerge.
AI has expanded the available code, interaction, and generative tools by more than one order of magnitude, lowering the barrier to turning ideas into products. 刘飞 cites 花生, the creator of “Kitten Fill Light”: someone from product operations with no technical background who was able to teach himself development over several months and build a small, elegant product. In a market without standard answers, any effective attempt may earn “more positive feedback than this thing” has so far received.