
戴雨森
Frontier Insights
Frontier Thesis: AI has crossed the chasm via inference scaling, tool use, and reasoning, pushing every vertical toward an inevitable “Lee Sedol moment.” While models trend toward small-scale knowledge discovery, true defensibility lives at the application layer through context, persistent memory, and bespoke execution environments.
Strategic Playbook: Back resilient founders with extreme adaptability (betting on people, not fixed tracks). Capitalize on read-only agents delivering instant PMF, anticipating a 100x–1000x surge in inference compute demand as agentic workflows proliferate.
Key Risks: Hyper-burn, valuation dilution, permission boundary failures in write-enabled agents, and the persistent bottleneck of robust online learning.
Key Views & Dialogues
142. 戴雨森’s Venture Capital Observations, Episode 2: Harness, the Next ByteDance, the Big Opportunities of 2026, and Stanley Druckenmiller
- 🗓️ Date:
2026-05-27| 🎙️ Show:张小珺Jùn|商业访谈录
Harnesses may be the durable layer as users keep their context while swapping models, with Cursor’s Composer showing how wrapper data can strengthen the product. The larger catalyst is agent-to-agent commerce, but Anthropic’s return problem remains unresolved: customers must turn rising token spending into products, savings and profits.
View Dialogue Notes & Key Takeaways
戴雨森 revisits how his “year of R” call got humbled. He was right about OpenAI’s consumer ceiling: the $20 subscription is hard to raise, 50M paying users have already filtered out everyone willing to pay for a chatbot, and “everyone was too optimistic about putting ads in ChatGPT”; he was wrong about the threshold discontinuity in coding from Claude 4.5 to 4.6. His explanatory framework is worth remembering: “You only get a steam engine once the water boils; at 99 degrees, you still don’t” — intelligence creates value through discontinuities, not gradual improvement, and even Anthropic calling the model 4.5 rather than 5 suggests the company itself did not anticipate the jump.
The return problem has not been solved; it has merely been pushed down to Anthropic’s customers. Tokens are customers’ inputs, not their output; the chain is input → output → result, and “the stagnation of mobile internet was not caused by a shortage of programmers, but by people not knowing what to build.” The numbers are now too big to ignore: AI hardware profits will reach $700B this year, Samsung and SK Hynix are approaching Nvidia’s profit level, and Anthropic’s year-end AR expectation of $100B implies selling $10B worth of tokens every month. The valuation framework has 3 parts: in the short term—3 months—both OpenAI and Anthropic are undervalued, and a public listing “could be hyped to $3T”; over a 1-2 year horizon, both are overvalued; over 10 years, the winners could be $5T or $10T companies.
Harnesses are the OS; models are the CPU. Users are loyal to the harness, not the model: people keep rescuing the memory in OpenClaw while casually swapping in Kimi K2.5, “90-point performance at a 20-point price”; Cursor training Composer on data generated by its wrapper proves that “without the wrapper, there would be no such model.” “The model is the product” and “the wrapper is worth more” can both be true.
His most exciting new thesis is network effects between agents. Once a harness has accumulated personalized context, different people’s agents will produce different results on the same task: “My agent hires 张小珺’s agent. $1,000 is the value of the tokens; $9,000 is the proprietary knowledge accumulated by your agent.” That proprietary information may never be distilled into the model, and agent-to-agent marketplace companies are already emerging.
The next ByteDance will not look like ByteDance. AI feeds are about “wrapping new technology in the shell ByteDance is best at, but beating yourself under your own game rules is extremely difficult.” Consumer entertainment has to be more fun than Douyin and RedNote from day one, which is a hard problem; productivity is the “easy problem.” The opportunities are hiding in attributes that incumbents dismiss: wrappers, open source, and insecurity. “If you think wrapper value is zero, what if you’re wrong?”
His positioning and investing approach: after exiting his positions before the Lunar New Year, he added back “chokepoint” hardware exposure in storage, optical components, CPUs and related areas, but remains unwilling to be aggressive. His investing idol is Stanley Druckenmiller; he trades like a trader, with a “strong opinion weekly held” mindset. In venture, he continues to back people rather than chase themes: 刘松明 and 丁宁’s embodied-AI companies have risen from valuations of roughly RMB200M to several billion renminbi. His consumer-electronics hot take: AI hardware will repeat the fate of new consumer brands, and wearable recording devices are a “need invented by VCs.”
At the individual and social level, he warns against outsourcing thought to AI: “If you get the answer directly, your brain has not changed; your weights have not been updated.” People need to build a “gym for thought.” Education should prioritize agency, not taste—taste may no longer be uniquely human. His closing principle: “Being humbled frequently is a blessing. When someone is never proven wrong, it is probably not because they are always right, but because they have stopped progressing.”
🔗 Original source & video: 142. 戴雨森’s Venture Capital Observations, Episode 2: Harness, the Next ByteDance, the Big Opportunities of 2026, and Stanley Druckenmiller
2026 AI New-Year Conversation: the year of R | A Conversation with ZhenFund’s 戴雨森
- 🗓️ Date:
2025-12-14| 🎙️ Show:十字路口Crossing
2025 AI is framed as a shift from the BlackBerry era to the iPhone era, while agents remain an early market. 2026 turns to return: data-center, talent, and chip spending must prove economic payback as subscriptions, advertising, coding, and enterprise AI face deflation or slow deployment. Research and Remember may reshape paradigms and memory-based assistants, while Chinese open-source models reaching 80%–90% of SOTA in about six months could keep pressuring model rents.
View Dialogue Notes & Key Takeaways
戴雨森 believes 2025 marked the shift from AI’s “BlackBerry era” into the application-driven “iPhone era.” O1 introduced thinking-time scaling; Claude 3.5 Sonnet brought coding capability; OpenAI o1 added reasoning; and Anthropic enabled long-horizon planning. Together, they pushed GPQA, SWE-bench and other capabilities past the usability threshold, triggering the breakout of Cursor, Claude Code, Codex, Manus and Genspark. The coding-agent market went from zero to potentially around $1B in ARR within a year. “Progress in model capability unlocks application opportunities” remains the conversation’s central thesis.
The Agent thesis has been validated, but product maturity remains at the early-market stage; renaming a workflow does not make it an Agent. A real Agent derives from agency: independently decomposing goals, selecting and calling tools, and changing course based on feedback. Its core output is “saving people time.” Devin, Manus, Claude Code and 豆包手机助手 are all only beginnings. 戴雨森 agrees this is the “first year of Agents,” while stressing that it will be the “decade of agents”: moving from L2 assistance to L3 responsibility could still take years.
戴雨森 defines 2026 as the Year of R, with the first R standing for Return: markets spent the past 3 years trading the I in ROI, and must now validate the return. Tens and hundreds of billions of dollars in data-center investment, $100M annual talent offers, and the rallies in Nvidia, optical components and memory are all driven by the potential return from AGI and application profits. But moving frontier models from 80 to 90 requires far more investment while capability gains are slowing, and Chinese open-source models can reach 80%–90% of SOTA within roughly 6 months. The gap between “high expectations and slow deployment” may become a key variable for market sentiment and financing conditions in 2026.
All 4 mainstream monetization paths can grow, but none automatically delivers the grand valuations previously assigned to them. Chatbot subscriptions at $20 or $200 face token deflation and free competition; advertising and e-commerce may need years to find native formats and will partly redistribute the existing pie of Google, Meta and ByteDance; replacing a $100K programmer does not mean the model can capture $100K, and may instead rapidly commoditize the task; and enterprise AI, despite companies reaching $100M–$200M ARR, still has to cross the chasm of slow large-enterprise deployment. Investors are therefore shifting from pure growth toward gross margin, retention and cash flow, asking whether $1 of tokens can be value-added by an application into $2.
The second R is Research: as the marginal return on scaling declines, the next level of capability will require new paradigms, organizations and measurement methods. Ilya describes the current phase as a return from scaling to research, while Demis expects AGI to remain 5–10 years away and require 1 or 2 breakthroughs. Self-play, continual learning and world models are candidate directions; New Labs such as SSI, Thinking Machines Lab and Reflection are seeking alternative paths in lower-KPI environments. Meanwhile, MMLU and SWE-bench are nearing saturation. Gemini 3 Pro’s roughly 78 and GPT-4.5’s low-80s do not reflect the gap users feel in practice: without new benchmarks, it is difficult to know whether training is heading in the right direction.
The third R is Remember: memory may become an application differentiator and move passive assistants toward proactive agents. Existing memory is still like “retrieval with a large notebook”; the next step is using online learning to form a model belonging to each user and understand preferences rather than merely retrieving exact words. 戴雨森 cites ChatGPT recommending Japan’s Yakushima, a remote destination aligned with his niche preferences, as a personal test case. Once AI can combine long-term memory with real-time context to prepare meeting materials and anticipate needs instead of waiting for a prompt, 戴雨森 believes it could open a “10x opportunity.”
Chinese open-source models are the structural force Koji believes remains underappreciated, and they have changed the table for application founders. Before DeepSeek R1, founders worried they could not access closed-source SOTA. Ten months later, Chinese open-source models had rapidly narrowed the gap, and DeepSeek Math-V2 reached the IMO-gold-medal level of general models roughly 6 months later. Open source acts like a “nuclear weapon,” destroying the short-term lead of expensive closed models while gaining advantages through low cost, transparency and visibility for talent. Application teams should start globally, keep “turning over cards,” and use industry expertise, proprietary data, distribution or integrated hardware and software to avoid being swallowed directly by the models.
In an environment of rapid model commoditization, the best teams to back are those that can see technical thresholds 6–12 months early and use human agency to amplify creativity. 戴雨森 emphasizes learning, leadership, innovation, willpower and the resilience that is becoming scarcer in the AI era. He also stresses product thinking and marketing: “The saddest thing is not saying the wrong thing; it is working hard to build something that nobody knows about.” AI can reproduce Picasso from a data distribution, but still struggles to create a new OOD style, an original joke or a company. The strongest near-term combination is therefore not AI creating alone, but humans adding the finishing touch and AI supplying 100x or 1,000x leverage.
🔗 Original source & video: 2026 AI New-Year Conversation: the year of R | A Conversation with ZhenFund’s 戴雨森
124. Yusen’s Venture Capital Observations, Episode 1: 2026 Expectations, The Year of R, Pullbacks, and How We Bet
- 🗓️ Date:
2025-12-13| 🎙️ Show:张小珺Jùn|商业访谈录
戴雨森 frames 2026 as the “Year of R”: compute investment is accelerating while SOTA gains slow, making returns decisive and leaving US equities vulnerable to a second-half pullback. Model commoditization, weak pricing power, slow enterprise adoption and margin pressure will reshuffle winners, while China’s valuation discount, research, memory, multimodality and voice offer optionality.
View Dialogue Notes & Key Takeaways
2026’s keyword is “Year of R.” Return, Research, and Remember. 戴雨森’s core logic: the progress of the new SOTA over the previous generation is slowing (think GPT-4 versus GPT-3, compared with GPT-5 versus GPT-4), while investment is growing exponentially—and roughly 50% of data-center investment is compute that will be obsolete within 4-6 years. “As everyone invests more, the focus on returns will only grow.”
The clear market call: US equities could see a meaningful pullback next year, potentially in the second half. The trigger would be a softening US labor market combined with AI returns falling short of expectations; “OpenAI’s own user and revenue growth will be a very important trigger for this broader concern.” He has essentially cleared out his secondary-market equity positions, but does not short: “Shorting is wanting other people to become miserable.” He only goes long, and can reduce exposure.
The bubble debate directly counters 朱啸虎’s claim that “there will be no bubble for three years; that’s pure nonsense.” “Every technological revolution in human history has brought a bubble, almost without exception.” AI is the most important revolution, so “it is perfectly natural that it could bring the biggest bubble in human history”—but we are not at the top yet. The peak comes when bad and fraudulent companies are also expensive; Nvidia’s short-term valuation is even low. The risk lies in a mismatch between the long and short ends: once long-term AGI expectations turn bearish, “even excellent short-term results may no longer be the main trading factor.” On the token thesis: usage rising 10x a year is inevitable, but “rising token volume does not necessarily guarantee the arrival of returns.”
Models have no secrets, and selling APIs is not a monopoly business. Silicon Valley has no non-competes and talent moves quickly, so “there are no secrets that can truly be kept for the long term.” Chinese open-source models have compressed the gap to 6-12 months at one-tenth to one-fifth the cost; “it is basically the age of Chinese models now.” First-party products are becoming increasingly important to model companies (Claude Code reached several hundred million dollars of ARR within months, and Dario was willing to compete directly with its largest customer, Cursor), while thin-shell applications where “the model gives you code and you give the user code” will face sustained pressure.
All four monetization paths are slower than consensus expects. Subscriptions are hard to reprice (Netflix has barely raised prices in 20 years; “the model you bought for $200 last year can now only be sold for $20”); advertising and e-commerce are existing pools of spending (Google took 4 years to find AdWords and Facebook 6 years to develop feed ads); replacing programmers is value deflation, not wage transfer (Jevons paradox); and enterprise adoption is slower than expected (Office Copilot remains below expectations). Sequoia’s David’s $200B question has inflated into a trillion-dollar question—the elephant in the room, with no good answer yet.
The China-US AI valuation gap is the biggest option in global allocation. Thinking Machines’ $50B valuation “is greater than the combined valuation of all Chinese AI startups”; Mistral at $14B versus Kimi below $4B; high-growth US applications receive 30-100x ARR, versus roughly 10x in China. The allocation is a barbell: OpenAI/Google/Anthropic at one end for stability, Chinese companies at the other for maximum optionality and potentially 100x returns. “I think the gap has already passed its widest point.”
For founders, the era of growth through negative-margin token resale is over. Silicon Valley now expects applications to have SaaS-like gross margins above 50% and retention; founders need a high-quality view of the SOTA 6-12 months out (杨植麟 on long context, Manus on agents), must target global markets from day one, and must “stay at the table and keep turning over cards”—Manus only worked on its third attempt. Memory, multimodal generation, and voice will be the key battlegrounds in 2026.
🔗 Original source & video: 124. Yusen’s Venture Capital Observations, Episode 1: 2026 Expectations, The Year of R, Pullbacks, and How We Bet
Vol.189 Dai Yusen Crossed the Hype and the Bubbles into the AI Era: A Conversation with Yusen on His Entrepreneurial and Investing Journey from Jumei to ZhenFund
- 🗓️ Date:
2025-10-20| 🎙️ Show:高能量
Dai Yusen sees angel investing as backing people and exercising patience, because direction changes while learning ability and creative drive can endure for a decade. ChatGPT convinced him on November 30, 2022 that AI had crossed from research into productivity, while scaling could also fuel the biggest bubble in human history. Application moats depend on the distance from model output to a finished deliverable, with Manus’s roughly $90M annualized revenue offering an early validation point.
View Dialogue Notes & Key Takeaways
Dai Yusen reduces the essence of angel investing to “backing people” and “patience,” because everything around an early-stage company can change, while a person’s ability to learn and desire to create may be what carries them across ten years. His first investment was VR player Skybox; after the sector cooled, founder 罗子雄 pivoted into games and spent four years polishing Party Animals. That completely unforeseeable path taught Dai what it means to “see because you believe.”
More important than chasing the next hot sector is finding the person without whom the thing would not exist. “Test-taking founders” often buy a ticket after a successful template has emerged, then face overcompetition from large companies, peers and capital. Dai therefore looks for founders who can articulate a vision, have shown a history of creating independently, and can withstand an extreme stress test of “why start a company?”
ChatGPT confirmed for Dai on November 30, 2022 that AI had crossed the chasm from research into an enabling technology for almost every form of knowledge work. He used it until 4 a.m. that night, warned his team the next day that it was significant, organized a discussion three days later, and invested heavily in projects including 光年之外 and Kimi in early 2023. His allocation analogy: “If you went back to 2010, three years after the iPhone launched, it would not be excessive to put 80% of your money and energy into mobile internet.”
He also believes AI will create “the biggest bubble in human history,” but that a massive bubble may be exactly what pays for real infrastructure and innovation. Scaling law will push the race for compute, infrastructure and talent to extremes, while most projects in the hype cycle will still fail. Early investors cannot avoid bubbles; they must look for “the beer beneath the foam” and distinguish good bubbles that leave behind infrastructure and innovative capacity from bad bubbles that merely inflate the price of existing assets.
Models and applications are not mutually exclusive; an application’s value depends on how much processing lies between model output and the deliverable the user actually wants. Dai compares model output to sashimi and a finished application to boiled fish: if the output is already the product, the wrapper will be thin; if the application must use multiple tools and steps to deliver a PDF, file, video or completed task, it accumulates context, environment, interface, data, trust and habits. Manus disclosed annualized revenue of roughly $90M within months of launch, a factual rebuttal to “applications will inevitably be swallowed by models,” though Dai still acknowledges that time will decide.
Heavy spending by a frontier-model company is not automatically a reason to pass; the key question is whether the leading model can control a massive general-purpose entry point. On the program’s figures, ChatGPT had roughly 400M DAU and close to 1B MAU; OpenAI was valued at about $500B, Anthropic at about $175B, and Kimi’s last round at about $3B. Dai acknowledges that model investments are unlikely to produce 10,000x returns and face heavy dilution, but says that if he could do it again, he would still invest—and “should have put more money into Kimi.”
China’s venture market is recovering from a trough, but a return of heat requires higher standards, not another rush into consensus trades. More than half the founders in ZhenFund’s latest fund are under 25; Dai sees investment activity in AI, robotics and hardware approaching 2020–2021 levels and points to large valuation gaps between Chinese companies and overseas peers. His portfolio stance is “long China,” angel-only, believe in young people during winter, and believe in the cycle while maintaining high standards in summer.
🔗 Original source & video: Vol.189 Dai Yusen Crossed the Hype and the Bubbles into the AI Era: A Conversation with Yusen on His Entrepreneurial and Investing Journey from Jumei to ZhenFund
127: 25 AI Mid-Year Check-In with ZhenFund’s 戴雨森: OpenAI’s IMO Gold-Medal Performance, Kimi K2’s Comeback, Agent Adoption, and the Talent War
- 🗓️ Date:
2025-07-21| 🎙️ Show:晚点聊 LateTalk
OpenAI’s unreleased general-purpose model solved five of six 2025 IMO problems at gold-medal level, strengthening the cases for inference scaling and AI-assisted discovery, though IMO has not officially certified the result. Model progress will not automatically erase application value: Agents still need context, memory, tools, and execution environments, while adoption hinges on raising one-shot success from roughly 20% to 70%–80%.
View Dialogue Notes & Key Takeaways
OpenAI’s unreleased general-purpose large language model solved five of the six problems at the 2025 IMO, reaching gold-medal-level performance and marking the episode’s most important capability jump. It had no internet access, no math-specific optimization, and did not use Code Interpreter; despite Google’s objection that the result was not officially certified and 陶哲轩’s reminder that proof scoring can vary, it still suggests LLMs are beginning to crack tasks that are hard to produce and hard to verify. The episode relayed one researcher describing it as a “moon landing moment for AI”: if AGI was once a train smoking in the distance, “we can hear it now.”
The breakthrough strengthens three technical theses at once: inference scaling, general-purpose generalization, and scientific discovery. More thinking time can continue to improve performance, while the model’s base model was reportedly the same as GPT-4o, implying that much of the gain may have come from post-training and inference; if true, substantial room for optimization remains even if pre-training hits a wall. IMO proofs and unproved minor theorems are structurally similar, leading 戴雨森 to argue that AI may be close to discovering “small new knowledge,” with potentially greater productivity value than coding tasks that mainly copy, adapt, and assemble existing code.
The value of both model capability and the application-layer “shell” is being underestimated. In the same weekend, the model delivered gold-medal-level IMO performance, while ChatGPT Agent’s actual outputs, including PPTs, lagged behind some online results from Manus, Genspark, Kimi, and MiniMax; model progress does not automatically eliminate applications. Agents need applications to provide organizational and personal context, cross-session memory, tools, and execution environments, and “the same model understands me better” could become a durable moat.
By the first half of 2025, coding and reasoning had crossed the chasm, while Agents entered the early-adopter phase of mass adoption. Cursor, Claude Code, o3, Deep Research, and Kimi Researcher have demonstrated clear productivity value, while Manus and Genspark are beginning to handle goal decomposition, tool selection, execution, and review. True L3 is not a chat window with a few extra buttons; it is “humans stop doing the work, AI does the work,” with users shifting from operating tools themselves to learning how to be AI bosses.
Kimi K2 is a case study in the market underestimating a strong team. As of the recording date, July 15, 戴雨森 called K2 “the best open-source model in the world, period,” with particular enthusiasm for its coding, agentic workflow, and Chinese writing; its OpenRouter coding call volume climbed from No. 13 to No. 10 within days, offering a more direct user vote than benchmarks. Behind it are a stable core team, a long-standing bet on long context, renewed investment in pre-training, and continued exploration of RL, tool use, and open source; K2 was still a non-reasoning model, with reasoning and multimodal versions yet to come.
Agent-driven compute demand could dwarf the Chatbot era, and Nvidia’s move past $4T may not be the end of the road. 戴雨森 observed that Agent applications can consume more than 1,000x the tokens of ordinary Chatbots, just as dial-up-era assumptions that everyone would only chat on QQ could not have forecast the bandwidth required for 4K video. Productivity demand is also not directly bounded by individual leisure time: AI can cover 50 stocks in parallel and process multiple earnings calls at once. In his view, the conversion of tokens into productivity has only just begun.
The binding constraints have shifted from whether AI can land to talent, organization, safety, and the distribution of gains. Meta’s “disruptive” compensation has reset the industry’s cost base, but whether assembling a large group of star researchers creates a coherent team remains unproven; embodied AI is the opposite case, with lower Optimus production expectations underscoring that manipulation and productization cannot be rushed. By year-end, the key tests will be whether Agents can raise one-shot delivery success from roughly 20% to 70%-80%, whether memory can generate true compounding returns, and how society governs AI code no one can fully review, content that is difficult to distinguish from reality, and widening gaps between individuals.
🔗 Original source & video: 127: 25 AI Mid-Year Check-In with ZhenFund’s 戴雨森: OpenAI’s IMO Gold-Medal Performance, Kimi K2’s Comeback, Agent Adoption, and the Talent War
106: A Long Conversation with ZhenFund’s Dai Yusen on Agents: Every Industry Will Face Its “Lee Sedol Moment,” Attention Is Not All You Need | Agent #1
- 🗓️ Date:
2025-03-09| 🎙️ Show:晚点聊 LateTalk
Agent adoption depends on reasoning, coding, and tool use crossing usable thresholds together, with O3 showing frontier-level performance. Read-only research is the clearest early PMF, while write Agents should deploy more slowly because failures carry greater operational risk. Agent usage could amplify inference demand and support compute growth, while Nvidia’s high share leaves room for ASICs, TPUs, and Ascend.
View Dialogue Notes & Key Takeaways
Dai Yusen believes the inflection point for Agents in 2025 will come when reasoning, coding, and tool use all cross the usability threshold at once. GPQA has risen from single-digit and low-teens scores for frontier models early in the year to above 70 for O3; SWE-bench has climbed from the single digits for GPT-4o to 70–80 for O3 Mini/O3; and Codeforces is around 2703, placing O3 among the world’s top human competitors. Once reinforcement learning enters specific industries, more “Lee Sedol moments” could follow.
O1’s significance was validating post-training RL and test-time compute as two new scaling laws; R1’s was opening that path to the entire industry. R1-Zero showed that running RL directly on the V3 base model, without SFT, could steadily lengthen outputs and improve intelligence. GRPO, failed MCTS approaches, and similar “one-bit” findings reduce duplicated trial and error while giving WeChat, Baidu, and application developers practical tools.
The real change Agents bring to the internet is a new growth equation: tools with agency no longer need to continuously occupy human attention. Dai Yusen summarizes the shift as “Attention is not all you need”: money can be approximately converted into compute, and compute into work output, creating a “scaling law for work” in which productivity no longer depends solely on hiring, training, and expanding organizations.
The first clear PMF is read-only Agents, followed by higher-risk write Agents. Deep Research can already deliver, within minutes, reports approaching the work of strong white-collar employees over one or two years. Operator, MCP, and products such as Monica are beginning to operate websites, send emails, and call tools, but privilege escalation and collateral damage mean “write” will deploy more slowly than “read.”
Lower costs will not reduce compute demand; Agents could instead amplify inference by 100x or even 1000x. GPT Pro costs $200 per month for roughly 100 Deep Research runs, or about $2 each; US Agents charging $6–8 per hour are already below California’s roughly $16 minimum wage. Total demand for the compute stack should keep rising, but Nvidia’s more than 90% share and high gross margins leave room for ASICs, TPUs, Ascend, and other specialized solutions.
Model companies and application companies will coexist, but the real danger is being trapped by the first PMF and a huge user base. BlackBerry was trapped by keyboards, Yahoo by portals, and Chatbots may be trapped by fragmented conversations. DeepSeek looks like an open-source ecosystem’s “Android moment,” while Cursor, Perplexity, and Kimi show that the application layer can still capture value through model combinations, user mindshare, accuracy, and scenario understanding.
The 2025 investment thesis is not only continued gains in model capability; it is also the beginning of a shift in productivity and wealth distribution. An individual who can direct Agents could become a “super individual,” while large companies and the wealthy may use capital to mobilize thousands of Agents at once. Dai Yusen supports broad access to technology and faster innovation, but warns that resources could become even more concentrated, emphasizing that “use first, fix problems later” does not mean ignoring safety.
🔗 Original source & video: 106: A Long Conversation with ZhenFund’s Dai Yusen on Agents: Every Industry Will Face Its “Lee Sedol Moment,” Attention Is Not All You Need | Agent #1
2025 New-Year Conversation: AI’s Pivotal Year, the First Year of Agents | A Conversation with ZhenFund’s 戴雨森
- 🗓️ Date:
2025-01-05| 🎙️ Show:十字路口Crossing
AI’s 2024 inflection was capability growth, not model-version churn: SWE-bench rose from 2.8% for GPT-4 to about 50% for Claude 3.5 Sonnet and 71.7% for o3. Devin’s roughly $8/hour asynchronous execution suggests Agents could sell work rather than software, but industry revenue remains far below costs and PMF should favor direct revenue or 10x productivity.
View Dialogue Notes & Key Takeaways
戴雨森 believes the decisive variable in AI in 2024 was the speed of capability gains, not the number of version releases. SWE-bench rose from 2.8% for GPT-4 at the start of the year to about 50% for Claude 3.5 Sonnet and 71.7% for o3; o3 also scored 25 on FrontierMath. Meanwhile, Kimi reached 40M monthly active users roughly a year after launch, while Sora—which stunned everyone early in the year—was facing usable or even free alternatives such as 可灵, 混元, and Veo 2 by year-end. “Something everyone found astonishing a year ago may now seem merely ordinary.”
AI has crossed the threshold of being able to “do a job” in a few areas such as coding, but the industry as a whole is still nowhere near covering its costs. Cursor’s ARR is approaching $100M; Bolt.new surpassed $20M ARR in 2 months, while Bolt.new and Lovable each reached $4M ARR within 4 weeks; Higgsfield grew from about $1M to nearly $50M, and Monica also surpassed $10M. 戴雨森’s qualification remains unchanged: “In some areas, it can already start doing a job,” but ChatGPT appeared only 2 years ago, and overall commercial revenue remains far below costs.
The technical frontier has shifted from simply piling on pre-training to a combination of RL, inference scaling, long context, code generation, and Computer Use. Ilya compared internet text to already-mined “fossil fuels”; 戴雨森 takes this to mean that the intelligence embedded in existing text has already been compressed quite thoroughly, so the next phase will depend on post-training, tool use, and models generating new knowledge. At the same time, the cost of a given level of intelligence falls to roughly one-tenth every year, and advanced capabilities can fit into smaller models—“brute force” is no longer the only scaling law. The host added that function call and structured output could also give Agents more precise instruction-following capabilities.
Devin’s significance is not better code completion, but that it was the first to show that money and compute could approximately buy asynchronous work directly. It can plan tasks, execute them in its own virtual machine, be corrected midstream, accumulate organizational knowledge, and return to a human only when finished or genuinely stuck; $500 buys 250 ACUs, each lasting about 15 minutes, implying roughly $8/hour—half of the $16 California minimum wage cited by 戴雨森. “Programmers like Cursor; bosses like Devin,” because the latter demonstrates the “scaling law of work.”
The applications most likely to find PMF in 2025 will either make customers money directly or improve important tasks by more than 10x. Of Midjourney’s hundreds of millions of dollars in annualized revenue, 戴雨森 estimates that about half comes from commercial image-making such as advertising; Higgsfield is focused mainly on marketing, while Cursor, Devin, and Perplexity compress the costs of coding and information gathering. The avoid list is equally clear: be cautious about “killing time” categories already dominated by giants such as Douyin, physical-world operations that have yet to converge, replacement hardware that overlaps heavily with smartphones, and enterprise products that require major workflow redesign without delivering overwhelming ROI.
戴雨森 sees “go global or die” as an overheated consensus, not a universal commandment for Chinese founders. Higher wages, stronger willingness to pay for tools and subscriptions, and access to more capable models do make productivity software easier to monetize in Europe and the US; but Chinese teams going abroad typically operate with a “low buff,” and enterprise services must make up for gaps in local customer understanding and Go-to-Market. Growth also depends more heavily on SEO, social media, and virality. Conversely, Chinese companies may be unwilling to pay for tools yet willing to buy AI-delivered work outcomes at one-tenth the price.
戴雨森’s core opportunity set for 2025 is Agents, scalable personalization, superhuman research capabilities, and cross-modal transformation, but he still defines the present as the “BlackBerry era.” Products will move from selling tools to selling work, while software, content, and education may be generated instantly for each individual; o3 moves benchmarks upward from ordinary people and expert humans toward superhuman performance. The basis for optimism is not that current products are mature, but that the loop in which technological progress unlocks applications and applications in turn activate models is accelerating.
🔗 Original source & video: 2025 New-Year Conversation: AI’s Pivotal Year, the First Year of Agents | A Conversation with ZhenFund’s 戴雨森