Pioneers Insight Method Research Author
2025 New-Year Conversation: AI’s Pivotal Year, the First Year of Agents | A Conversation with ZhenFund’s 戴雨森
Back to Episodes

2025 New-Year Conversation: AI’s Pivotal Year, the First Year of Agents | A Conversation with ZhenFund’s 戴雨森

Summary

  • 戴雨森 believes the decisive variable in AI in 2024 was the speed of capability gains, not the number of version releases. SWE-bench rose from 2.8% for GPT-4 at the start of the year to about 50% for Claude 3.5 Sonnet and 71.7% for o3; o3 also scored 25 on FrontierMath. Meanwhile, Kimi reached 40M monthly active users roughly a year after launch, while Sora—which stunned everyone early in the year—was facing usable or even free alternatives such as 可灵, 混元, and Veo 2 by year-end. “Something everyone found astonishing a year ago may now seem merely ordinary.”
  • AI has crossed the threshold of being able to “do a job” in a few areas such as coding, but the industry as a whole is still nowhere near covering its costs. Cursor’s ARR is approaching $100M; Bolt.new surpassed $20M ARR in 2 months, while Bolt.new and Lovable each reached $4M ARR within 4 weeks; Higgsfield grew from about $1M to nearly $50M, and Monica also surpassed $10M. 戴雨森’s qualification remains unchanged: “In some areas, it can already start doing a job,” but ChatGPT appeared only 2 years ago, and overall commercial revenue remains far below costs.
  • The technical frontier has shifted from simply piling on pre-training to a combination of RL, inference scaling, long context, code generation, and Computer Use. Ilya compared internet text to already-mined “fossil fuels”; 戴雨森 takes this to mean that the intelligence embedded in existing text has already been compressed quite thoroughly, so the next phase will depend on post-training, tool use, and models generating new knowledge. At the same time, the cost of a given level of intelligence falls to roughly one-tenth every year, and advanced capabilities can fit into smaller models—“brute force” is no longer the only scaling law. The host added that function call and structured output could also give Agents more precise instruction-following capabilities.
  • Devin’s significance is not better code completion, but that it was the first to show that money and compute could approximately buy asynchronous work directly. It can plan tasks, execute them in its own virtual machine, be corrected midstream, accumulate organizational knowledge, and return to a human only when finished or genuinely stuck; $500 buys 250 ACUs, each lasting about 15 minutes, implying roughly $8/hour—half of the $16 California minimum wage cited by 戴雨森. “Programmers like Cursor; bosses like Devin,” because the latter demonstrates the “scaling law of work.”
  • The applications most likely to find PMF in 2025 will either make customers money directly or improve important tasks by more than 10x. Of Midjourney’s hundreds of millions of dollars in annualized revenue, 戴雨森 estimates that about half comes from commercial image-making such as advertising; Higgsfield is focused mainly on marketing, while Cursor, Devin, and Perplexity compress the costs of coding and information gathering. The avoid list is equally clear: be cautious about “killing time” categories already dominated by giants such as Douyin, physical-world operations that have yet to converge, replacement hardware that overlaps heavily with smartphones, and enterprise products that require major workflow redesign without delivering overwhelming ROI.
  • 戴雨森 sees “go global or die” as an overheated consensus, not a universal commandment for Chinese founders. Higher wages, stronger willingness to pay for tools and subscriptions, and access to more capable models do make productivity software easier to monetize in Europe and the US; but Chinese teams going abroad typically operate with a “low buff,” and enterprise services must make up for gaps in local customer understanding and Go-to-Market. Growth also depends more heavily on SEO, social media, and virality. Conversely, Chinese companies may be unwilling to pay for tools yet willing to buy AI-delivered work outcomes at one-tenth the price.
  • 戴雨森’s core opportunity set for 2025 is Agents, scalable personalization, superhuman research capabilities, and cross-modal transformation, but he still defines the present as the “BlackBerry era.” Products will move from selling tools to selling work, while software, content, and education may be generated instantly for each individual; o3 moves benchmarks upward from ordinary people and expert humans toward superhuman performance. The basis for optimism is not that current products are mature, but that the loop in which technological progress unlocks applications and applications in turn activate models is accelerating.

Deep dive

1. The Most Important Fact of 2024 Was That Capabilities Grew Faster Than Public Standards Were Reset

  • 戴雨森 uses SWE-bench to mark the year: at the start of the year, GPT-4 could solve only about 2.8% of common GitHub tasks; by year-end, Claude 3.5 Sonnet had reached about 50%, while o3’s initial eval reached 71.7%. His optimistic extrapolation is that if this pace continues, AI could cover the vast majority of individual GitHub tasks in 2025.
  • The contrast in mathematical ability was just as dramatic: early in the year, ChatGPT could still get 3-digit multiplication wrong; by year-end, o3 could handle IMO-level problems and scored 25 on FrontierMath, a benchmark endorsed by Terence Tao that extends from Olympiad mathematics toward frontier research difficulty.
  • Product adoption accelerated in parallel. Kimi went from launching on October 9, 2023 to reaching about 40M monthly active users by the end of 2024; Sora’s promotional video stunned users in February, but by year-end they could access models such as 可灵, 混元, and Veo 2, which may have been better than Sora at the time and were even free.
  • The host captured the mismatch in perception: the public often sees only GPT-4o, o1, o3, Claude 3.5, and Gemini 2.0 cycling through version numbers, while in reality “something everyone found astonishing a year ago may now seem merely ordinary.”

2. In the Early Industry, the Norm Was Not Consensus Being Vindicated but Predictions Being Repeatedly Proven Wrong

  • 戴雨森 recalls that early in 2024 the market predicted a “100 C battle” in China, with numerous teams trying to build a Chinese Character.AI; by August, Character.AI had announced its acquisition by Google, also exposing how difficult it is for companion products to break into the mainstream.
  • When Cognition released only a Devin demo in March, many people said it was “a con,” or “maybe even a scam,” and dedicated debunking efforts appeared; after the formal product opened in December, sentiment flipped to “wait, this is actually real.”
  • OpenAI’s trajectory was equally unexpected: in late 2023, employees flooded social media with “OpenAI is nothing without its people,” but by the end of 2024 early employees and core researchers such as Alec Radford had departed one after another. The market waited a year for GPT-5, and even 4.5, but got o1 and o3—and the inference scaling they represented—instead.
  • 戴雨森 has therefore stopped trying to be perpetually right in his predictions: “In early-stage investing, especially in early technology, being proven wrong is normal. Only if you’re not afraid of being proven wrong can you keep learning and growing.”

3. Coding Has Crossed the Work Threshold, but Commercialization Has Only Been Partially Validated

  • Six months earlier, 戴雨森 called large models “elementary-school students,” not to deny commercialization but to emphasize that technological revolutions usually move from infrastructure and research investment to product deployment and finally revenue realization. He now acknowledges that coding has been the first field to cross the threshold of being able to “do a job.”
  • Once Claude 3.5 Sonnet could solve roughly half of SWE-bench, Cursor, Windsurf, and Devin could genuinely help programmers solve many problems and improve productivity, rather than merely showcase demos; this is a textbook case of higher model capability directly unlocking products.
  • The revenue signals are already strong: Cursor’s ARR is approaching $100M; Bolt.new surpassed $20M ARR in 2 months and once reached $4M ARR in 4 weeks; Lovable also reached $4M ARR in 4 weeks. Higgsfield grew from about $1M to nearly $50M, while Monica also surpassed $10M.
  • But 戴雨森 preserves the most important qualification: “Overall, the revenue it generates is still far below the cost.” Model improvement, scenario unlocks, value creation, and commercialization are still unfolding layer by layer; the industry needs patience.

4. The Endgame for AI Tools Is Not Just Making Experts More Efficient, but Giving Ordinary People Expert Capabilities

  • The host pointed out that Cursor, Bolt.new, and Higgsfield have no traditional network effects, yet spread rapidly through a group of passionate early users; he urged listeners not to “watch from across the river,” but to spend some time and money experiencing the capability frontier themselves.
  • 戴雨森 borrowed Gibson’s line: “The future is already here—it’s just not evenly distributed.” To ordinary people, AI may still be no more than news; for programmers and digital artists, it is already an indispensable part of the production process.
  • The host went further: the larger significance of AI coding and digital creation is not merely serving programmers and artists, but enabling ordinary people to create things that previously only those professional groups could produce. “This is actually hugely relevant to everyone.”

5. More Than 1,000 Projects Show That Agent Deployment Requires Three Capabilities to Mature Together

  • The ZhenFund team saw more than 1,000 AI application projects over the year, while 戴雨森 personally spoke with roughly 100 to nearly 200 founders; his overall impression is that application deployment is indeed accelerating.
  • The first unlock is reasoning: GPT-4o and o1 reduced hallucinations, allowing models to plan and complete more complex tasks. The second is coding, because many problems in the digital world can be translated into coding problems.
  • The third is Computer Use, which Anthropic pushed first: models can use browsers and existing software, turning the software systems accumulated by humans into their own tools. Reasoning, code writing, and tool use must stack together before Agents can move from prototypes to actual execution.
  • The host added that function call and structured output could give Agents more precise instructions, representing another possible technical unlock.
  • 戴雨森 expects every industry to experiment with Agents in 2025. “Many of these are still at a fairly primitive stage,” but Devin has already turned an abstract idea into a product paradigm that can be imitated and adapted.

6. The Divergence in Chinese and US Startup Directions Ultimately Reflects Different Prices Put on Productivity Value

  • In China, enterprise-service deployment is difficult, so founders have leaned toward To C “killing time” products such as emotional companionship and AI chat; US teams more often enter vertical industries directly, replacing some labor and delivering cost reduction and efficiency gains.
  • Another domestic boom is robotics and embodied intelligence, with large numbers of companies being founded and funded; the host thought parts of the sector looked overheated, and 戴雨森 later said he was cautious about general-purpose humanoid robot platforms.
  • The conversation also touched on generational turnover: after the post-80s generation benefited from the internet boom, the post-00s generation briefly felt there was “nothing left to do” on the internet. AI is now reopening a technology platform and startup window for younger founders.

7. The New Generation of Founders Is More Global and More AI Native, but Still Needs to Make Up for Gaps in Growth and Commercialization

  • The information cycle has compressed from 3-6 months in the internet era to the same day; overseas models and products are reported, translated, and discussed as soon as they appear. Models’ built-in multilingual capabilities also allow many teams to target domestic and overseas markets from day one.
  • The strongest young teams 戴雨森 sees are generally more AI native, with members who have AI research or engineering experience and can therefore identify and execute on new opportunities earlier.
  • Their weakness is lack of experience: teams that did not live through the full internet-era business process are often unfamiliar with promotion, monetization, and organizational building. “Old hands” with internet-growth experience, such as Monica, have a temporary advantage, but these capabilities can be recovered through learning, hiring, and team complementarity.

8. The Single-Track Pre-Training Route Has Peaked, and Model Scale No Longer Equals Incremental Intelligence

  • 戴雨森’s biggest shift in thinking is that he no longer believes pre-training can deliver unlimited “brute-force miracles.” At NeurIPS, Ilya called internet text “fossil fuels”: text accumulated by humans over many years and easy to compress has already been absorbed into models in large quantities.
  • The next shortage is new knowledge—including knowledge still inside human brains and not yet expressed, as well as knowledge discovered and generated by AI itself. Its supply will not grow as quickly as scraped internet text.
  • He has also revised his view that models must move from 7B and 70B to 700B: models of the same 70B size can continue to improve, and the same level of intelligence can be compressed into smaller models. Truly huge models may increasingly serve as teacher models and alignment systems.
  • 戴雨森 uses CPUs as an analogy: once single-core clock speeds reached roughly 3GHz, they stopped rising in a straight line, and the industry shifted toward architecture and energy efficiency. Models may likewise learn more knowledge and skills at roughly similar sizes, while the cost of a given level of intelligence falls to about one-tenth every year.

9. RL, Inference Scaling, and Long Context Form the New Axes of Capability Growth

  • At the start of 2024, reinforcement learning was still discussed by only a small number of researchers; after the release of o1 and o3, the continued improvement of capabilities through post-training and RL began to become an industry consensus.
  • The key to the inference scaling law is not merely making a model “think longer,” but having it plan, check, call tools, and continue working—moving from real-time one-question-one-answer interaction toward System 2-style deliberation.
  • The third variable to be revalued is context. Models have already compressed a great deal of intelligence, but with only a single user prompt, even a highly intelligent model struggles to understand the task accurately. Cursor can read an entire codebase, while Devin can combine Slack conversations with organizational records, greatly increasing product value.
  • 戴雨森 therefore judges ChatGPT-style Q&A to be “a very, very primitive way” of interacting. The next generation of products will compete on making it painless for users to hand models the context of their personal lives, organizations, and current tasks.

10. Screens, Cursors, and Cameras Are Turning AI from a Pen Pal into an On-Site Assistant

  • The host used the Mac version of ChatGPT as an example: it can not only read screenshots, but also access content inside a window that is not currently visible and requires scrolling, while combining the cursor or selected text to understand what currently holds the user’s attention.
  • 戴雨森’s analogy is that ChatGPT used to resemble a pen pal who could only send and receive emails. If that pen pal “stood behind your computer,” or even lived inside the computer and could see organizational information beyond the screen, it would obviously be much more useful.
  • Gemini 2.0’s camera understanding extends context into the physical environment: the host pointed the camera at a film-festival poster on a wall and asked directly for the festival’s name and edition. An interaction that once belonged in science fiction can now return an answer quickly at an acceptable cost.
  • The capability has not yet been fully packaged into a mature C-end product, but both participants saw hands-on use as the best way to understand the pace of technological progress.

11. AI Coding Has Completed a Four-Level Shift from “You Ask, I Answer” to “You Ask, I Do”

  • The first stage was ChatGPT: the user stated a need, and AI produced code without knowing the objective, execution environment, or result. When an error occurred, the user still had to run it locally, copy the error, and send it back to the model.
  • The second stage was GitHub Copilot: the organization’s codebase became part of the context, so AI was no longer completely blind, but users still had to place, run, and debug the code themselves inside an IDE.
  • The third stage is represented by Cursor and Windsurf: AI predicts what the user will write next, creates or modifies files, and executes command-line and deployment operations. 戴雨森 describes this as AI moving from “writing code on a piece of paper” to working directly on the user’s computer.
  • The fourth stage is a virtual employee such as Devin: a planner lets it advance asynchronously, a virtual machine lets it run and debug websites itself, and users can insert new instructions during execution. In short, the interaction has advanced from “I ask, you answer” and “I ask, you write” to “I ask, you do.”

12. Going Overseas Offers Better Revenue Soil, but It Is Not a Natural Buff for Chinese Teams

  • Average wages are higher in Europe and the US, and users are more willing to pay for productivity tools and subscriptions; software that saves labor and is priced in dollars can build meaningful revenue more easily.
  • Going overseas also provides access to stronger models such as Claude 3.5 Sonnet and GPT-4o. Multilingual capabilities lower the technical barrier to global products, and together these factors explain why AI applications commercialize faster overseas.
  • But 戴雨森 warned: “When every VC is telling founders to go overseas, that often means the market is too hot.” For most Chinese teams, going overseas is an away game with a “low buff,” requiring additional understanding of language, culture, channels, and unfamiliar customers.
  • The Chinese market may also scale first and monetize later. Using early Taobao’s free model versus eBay’s commissions as an example, he stressed that different markets develop different paths; not every team should copy the European and US subscription model.

13. Success in Overseas Enterprise Services Depends on Defining Demand and Acquiring Customers Cheaply, Not Just on Engineering Execution

  • 戴雨森 believes Chinese teams often overestimate the overseas advantage created by “strong engineers and fast execution.” The real difficulty is defining the key problem; enterprise services in particular require direct contact with customers and the addition of local Go-to-Market expertise.
  • C-end products with relatively universal demand, such as Monica, can be operated more remotely, but sales-driven enterprise products require people to go out in person. The greater the language and geographic distance, the less user research can rely on imagination.
  • The common thread among overseas products such as Higgsfield and Monica is strong SEO, overseas social media, quality content, and viral distribution—not reliance on expensive paid acquisition. Unless monetization is strong enough to clear the ROI hurdle, indiscriminate buying of traffic is difficult to sustain.
  • For Chinese teams, product execution is usually not the bottleneck. “What to build” and “how to promote it” are the two questions most likely to create differentiation.

14. AI Hardware Appears to Leverage China’s Supply-Chain Advantage, but Its Expansion Speed May Not Match Software

  • 戴雨森 has reviewed many AI hardware projects but remains cautious overall: “Hardware looks beautiful on paper, but it may not actually be that easy to deploy.” The approach Chinese teams have executed well in the past is to make overseas prototypes faster, cheaper, and smaller; he cites PLAUD as an example of an inventive product but believes hardware typically scales more slowly than software.
  • He remained skeptical of products such as Rabbit and Humane that attempted to rebuild the primary interface from the moment they launched.
  • ZhenFund is not completely unwilling to back hardware founders, but it has not concentrated its bets as some funds have, because software remains the more direct vehicle for spreading AI capabilities at this stage.

15. Devin’s Most Convincing Task Was Handling Incomplete Information Retrieval Like an Intern

  • 戴雨森 asked Devin to research the value statements of leading US venture capital firms. It first used sources such as PitchBook and CB Insights to identify 10 top firms, then searched each one for materials variously named manifesto, ethos, about, or philosophy.
  • When Accel’s website had no directly corresponding page, Devin did not mechanically report “not found.” It continued searching the news section and ultimately extracted the closest approximation of Accel’s values framework from a 2023 article. 戴雨森 saw this as evidence of junior-employee-style task comprehension and creative gap-filling.
  • It also cut corners: when asked for full text, several firms returned only a summary, forcing the user to explicitly request the “exact full text.” 戴雨森 did not hide the flaw, instead treating Devin as an intern who needs review, instruction, and correction.
  • This non-coding task showed him a broader boundary: if work can be completed by a person sitting at a computer, browsing the internet, and using software, there may eventually be corresponding Agents for finance, law, research, and other industries.

16. Devin Defines a Third Category of Tool: It Handles Non-Repetitive Problems Without Needing Constant Supervision

  • 戴雨森 divides human tools into two categories. Hammers, power drills, keyboards, and mice require sustained attention; washing machines, vending machines, and assembly lines do not, but can execute only mechanical, repetitive processes.
  • Humans have always lacked a third category of tool: “something that doesn’t require my continuous attention, but can also plan and solve problems on its own.” In the physical world, Waymo comes close to this definition; in software, AutoGPT was previously still only a prototype.
  • Devin’s key advantage is not getting everything right every time, but allowing tasks to progress asynchronously: after being @-mentioned in Slack, it works independently and returns only when genuinely stuck or finished. Users can therefore assign work to multiple Agents at once and reserve their attention for more important judgments.

17. Virtual Machines, Accumulated Knowledge, and Work-Based Pricing Make Agents More Like Employees Than Software Licenses

  • Devin has its own virtual machine in the cloud, where it can operate browsers, deploy, and debug without occupying the user’s computer like RPA, Cursor, or Windsurf. It is equivalent to “hiring an intern and having to equip them with a computer.”
  • It can also accumulate organization-specific knowledge: after completing a task, it proactively tells the user what it learned, and after confirmation, stores that knowledge as a rule for the future. This resembles an employee writing a work summary and receiving review, rather than a static tool that starts from zero every time it opens.
  • A $500 plan includes 250 ACUs, each lasting about 15 minutes, implying roughly $8/hour. 戴雨森 uses California’s approximately $16 minimum wage as a reference and argues that the correct comparison is not Cursor’s $20 subscription, but the labor cost of producing equivalent output.
  • Add 7×24-hour availability, no office space, and no personnel management, and once an Agent can independently complete an intern’s tasks, companies will seriously compare “hire an intern or use a Devin Agent.”

18. Shared Tasks and Parallel Execution Give “the Intern in the Office” Product Reality for the First Time

  • In a shared team account, the host saw Devin build a VC-manifesto website for 戴雨森. The first draft was poor, so the host took over, provided a reference website and an image-generation API document and key, and asked it to create illustrations for 10 manifestos.
  • The feeling was not of operating a tool, but of giving additional guidance to a shared intern after a colleague had gone downstairs for lunch: “What 戴雨森 actually wants is that—go improve it.” The task and context could be handed off between different managers.
  • Another task required collecting LinkedIn information. Devin had no account, so it asked the user to enter the username and password inside the virtual machine, then continued using the logged-in environment. The interaction—“Boss, could you enter the account details?”—reinforced the sense of working with an employee.
  • Because each Agent uses its own machine, work can proceed in parallel. Users no longer need to commit their full attention to a single tool; they only need to allocate work, review deliverables, and correct course at key points.

19. “Everyone Is a CEO” Depends on Everyone First Learning to Make Executable Requests

  • The host revised the old slogan “everyone is a product manager” into “everyone is a CEO”: when using Devin, the main actions become issuing instructions, checking the work, and providing inspiration and direction at a higher level.
  • 戴雨森 immediately added a limitation: if a boss says only “build me Taobao,” a human team cannot deliver it either. AI still “stupidly” accepts tasks beyond its capabilities, ultimately disappointing both sides.
  • The scarce capability therefore becomes knowing “what I want to do” and organizing that intent into a clearly structured, easy-to-understand task. Everyone used to hate bosses who asked for “a colorful black”; now every user must avoid becoming that kind of AI boss.
  • Agents will not eliminate management. They will distribute management capabilities more broadly: task decomposition, prioritization, context, review criteria, and feedback quality will all directly determine output.

20. Agents Will Turn Open-Source Wheels into Productive Assets Ordinary People Can Call Upon

  • The host quoted a reminder from a friend: vast amounts of human intelligence exist as repositories on GitHub and Hugging Face, but ordinary people do not know where the wheels are, let alone how to download, deploy, and connect them to their workflows.
  • Take a chess application: implementing the rules may require hundreds or thousands of lines of code, while a search can return hundreds of pages of results. Devin can asynchronously find an existing repository, identify community-validated best practices, and integrate them into the project.
  • 戴雨森 believes AI is best at finding an existing solution and then assembling it with “glue.” The first jobs to see major productivity gains or even replacement will be copy-and-paste work such as junior graphic design, front-end development, and code porting.
  • Education will therefore need to move from repeatedly training execution toward understanding principles, asking the right questions, and creatively solving unsolved problems—just as people stopped spending most of their energy on hand calculation after calculators became widespread.

21. The Scaling Law of Work Creates a More Direct Mapping Between Capital, Compute, and Organizational Output

  • 戴雨森’s blunt definition of a scaling law is: “I can spend more money to buy more productivity.” After a traditional company raises money, it still has to hire people, build an organization, and manage collaboration; capital cannot automatically become execution.
  • Asynchronous Agents can accept many tasks simultaneously, while a product-manager-style Agent can break down requirements and direct multiple coding Agents, ultimately forming a virtual organization. The main inputs shift partly from headcount and attention toward objectives, compute, and electricity.
  • The host used the logic of The Leadership Pipeline: once promoted into management, a person’s output is no longer the work they complete by hand, but the output of the entire team. Agents extend that leverage to ordinary people who could never have managed 1,000 people.
  • At one end, wealthy companies use compute to complete more work; at the other, people with ideas but no programmers can validate products at lower cost. Once execution is no longer as scarce, “what to do” will become the core constraint on entrepreneurship.

22. Most Criticism of Devin Is Valid, but It Does Not Disprove the Paradigm It Points Toward

  • 戴雨森 believes directly comparing Devin’s $500 price with Cursor’s monthly fee confuses two types of products: Cursor improves the productivity of the programmer’s own time, while Devin attempts to independently deliver a piece of work for the user.
  • For an experienced programmer, today’s Devin resembles a “clumsy intern”: explaining the task, waiting, reviewing, and cleaning up after it may be less efficient than fixing the bug directly in an IDE. Cursor fits existing workflows, making it more comfortable in the short term.
  • Organizations have also not yet developed the expectation that an “AI employee” will learn and improve. When a human makes a mistake, managers understand that training may work; when expensive software makes a mistake, users are more likely to think, “I bought a tool—how can it also have problems?”
  • 戴雨森 does not claim Devin will be the ultimate winner. Cursor is incremental innovation; Devin is disruptive innovation, requiring new onboarding, management habits, and organizational processes. What matters is that Devin was the first to show clearly what an Agent product might look like.

23. 2025 Is Still the BlackBerry Era, but Models Will Keep Activating Products That Previously Could Not Be Built

  • 戴雨森 does not agree that ChatGPT’s release means AI has entered the iPhone era. He would rather call it the “BlackBerry era”: technology is expensive, directions are fragmented, standards have not converged, and many ideas can be seen but not built.
  • AutoGPT was the archetype of “building Douyin in the BlackBerry era”: the concepts of planning, execution, and checking were correct, but the models hallucinated too much, could not use tools reliably, and could not browse the web consistently, so the product naturally failed to run.
  • Cursor already existed in 2023, but it took Claude 3.5 Sonnet to make next-action prediction and high-quality coding work in practice. Cursor, in turn, helped make Claude 3.5 Sonnet the model favored by coding users. Products and models activate each other.
  • Devin may likewise still be waiting for stronger o1, o3, or new Anthropic models. 戴雨森’s optimism is not based on the product being mature, but on the fact that in 2 years the industry has moved from “you ask, I answer” to “you ask, I do,” far faster than traditional platform iteration.

24. PMF Favors Making Money and 10x Productivity Gains, and Dislikes Weak Benefits and High Migration Friction

  • 戴雨森 first asks whether a product can make customers money. Midjourney already has hundreds of millions of dollars in annualized revenue, and he estimates that about half comes from commercial images such as advertising. Higgsfield also focuses mainly on marketing videos because customers are willing to learn and pay for tools that directly generate revenue.
  • The second category is 10x productivity gains on important tasks. Cursor and Devin sharply compress the time spent finding libraries, writing code, and debugging; Perplexity turns reading 10 or 20 search results and summarizing them independently into a single synthesized answer.
  • He is more cautious about ordinary “killing time” applications: mobile internet turned fragments of time that previously had no internet access from zero into one, but those moments are now occupied by mature platforms such as Douyin. Without strong differentiation, AI companionship is more likely to gain a foothold only in NSFW or niche demand.
  • Three other categories face similar difficulty: physical operations that are still unlikely to scale within 3-5 years, replacement devices whose capabilities heavily overlap with smartphones, and Agents that require large companies to rebuild permissions, privacy, and workflows while delivering only a 50% rather than 10x improvement.

25. The 2025 Themes Are Selling Work, Scaling Personalization, and Moving AI into the Superhuman Research Zone

  • The commercial expression of Agents will shift from SaaS to “Service as a Software,” or “Sell Work not the Tool”: asynchronous, plannable, tool-using, and priced by workload. 戴雨森 expects many attempts and many failures, but work outcomes rather than software seats will become the key pricing unit.
  • The second theme is scalable personalization. Content distribution is moving from one-size-fits-all portals, through keyword personalization in search and user personalization in recommendation engines, toward generation on demand. Bolt.new and WebSim already let users generate websites temporarily with prompts, while NotebookLM demonstrates personalized podcasts.
  • Personalization may also enter software and education: users will not have to accept the same WeChat tabs or a standardized curriculum, but can receive applications and instruction suited to their habits, level, and goals. ZhenFund’s investments in AI education and AI coding for personalized application generation follow this thesis.
  • The third theme is the superhuman benchmark. o3 reached about 2700 on Codeforces, corresponding to the top 0.01% of humans; historically, only roughly 130-plus people have reached that level. GPQA has surpassed 70 points and exceeded human PhD performance, AIME has surpassed the level of elite US high-school students, SWE-bench is close to being solved, and FrontierMath pushes evaluation toward frontier research.
  • 戴雨森 is not troubled by the expense of o3’s high-compute mode because it is supposed to solve humanity’s most difficult exploration problems. Future models may split into a Sheldon-style research model, an o3-mini-style workhorse, and a cheap model for routine requests such as checking the weather.
  • Multimodality will fill in AI’s understanding of the world and open up recombination across modalities: text-to-podcast is not read-aloud conversion, but adaptation into content suited to listening; video understanding can overlay operating instructions on footage of a coffee machine, while AI glasses may provide real-time guidance on tennis posture.
  • Truly AI-native giants will still have to wait for technology to diffuse. The internet first brought email, portals, and self-operated e-commerce, followed by search, social media, and platform e-commerce; only after mobile internet became widespread did Douyin, Meituan, and Didi emerge. Once individuals and enterprises broadly have AI assistants and Agents collaborating with one another, organizational software and advertising models will be redesigned.
  • 戴雨森’s final non-consensus view is that the next ByteDance may not look like ByteDance; general-purpose humanoid robot platforms and data scaling have not converged, so ZhenFund is more cautious and prefers upstream components such as dexterous hands and motors. Chinese companies may not buy tools, but they may buy work outcomes directly delivered by AI outsourcing at one-tenth the price. “There are so, so many reasons to remain optimistic and try to break through.”