Pioneers Insight Method Research Author
114: Metaso’s Min Kerui 2: “I’m Not an Actor”
Back to Episodes

114: Metaso’s Min Kerui 2: “I’m Not an Actor”

Summary

  • “What Should I Learn Today?” was called by 冯骥 “the best user interface of any AI tool I’ve used to date,” but 闵可锐 gives the current version only 70 points. He believes it may be the first product to connect source materials, scripts, slides, visualizations and course delivery end to end, but the interface, writing quality, aesthetics and feature depth all need substantial work. The unexpected K12 demand is mainly blocked by product workflow rather than model capability, and the team is preparing to add exam-paper explanations; 曼琪 imagines it could generate courses better than most teachers, but that is not 闵可锐’s explicit promise.

  • Metaso’s product method is not systematic research, but a founder using technical judgment to compress the decision cycle and then finding the next product in execution details. 闵可锐 says, “More than 80% of product-strategy decisions are made on a gut call by me”; a large company needs research, reviews, technical validation and scheduling, while he may decide within 10 minutes. “What Should I Learn Today?” came from an anti-summary judgment: AI search’s 300–800-word answers work for precise fact checks but discard the value of deep material; the real opportunity is helping users absorb a 100-page document, not compressing it further.

  • 闵可锐 believes general-purpose Agents had still not crossed the reliability threshold for delivery in the first half of 2025, and flashy demos cannot substitute for success rates. His rough estimate is that perhaps only 3 out of 100 tasks worked in 2023, rising to 30 in 2025, while users expect at least 95 to succeed and meet expectations; vertical Agents work only when the single-point result is valuable enough. In China’s B2B market, the more viable path is less “assistive productivity” than “pure replacement”—for example, automating some online law-firm cases end to end, theoretically reaching 100x the efficiency of traditional lawyers.

  • In 2025, Metaso shifted its weight from users toward revenue because it predicted Tencent, Alibaba and ByteDance would each invest at least several billion yuan in the AI To C market. 闵可锐’s view by the end of 2024 was that Tencent Yuanbao would probably invest RMB5B in 2025, Alibaba might invest no less than RMB5B, and ByteDance clearly would invest no less than RMB5B. Metaso AI Search is nearing 1M DAU, but general search is naturally treated as free and current payments do not cover inference costs; 曼琪 suggested that “What Should I Learn Today?” might generate $1M in monthly recurring revenue overseas with tens of thousands of subscribers. 闵可锐’s underlying principle is: “First, you need the ability to survive.”

  • 闵可锐 refuses to substitute fundraising size for operating proof because “I’m not a good actor,” and he cannot explain why some “very clumsy stories” win investors over. After the 2023 boom, Metaso raised only a little over RMB100M and also turned down several acquisition opportunities in 2024; his extreme downside calculation is to sell the search data, architecture and engineering reserves “as scrap metal,” repay investors and still retain a surplus, whereas raising $75M might actually make that safety line harder to defend.

  • The danger boundary for an application company is not a particular feature, but a product that relies mainly on foundation-model capability, is delivered as mostly pure software, and could reach 10M DAU. 闵可锐 believes model companies and large tech groups will both want such directions, leaving startups to build moats in sales, services, professional workflows and other areas while continually searching for the next product; Metaso’s position is that “we’re feeling our way across the river, while the big companies are feeling their way across on top of us.”

  • 闵可锐 attributes DeepSeek’s advantage to talent organization, quantitative-engineering DNA and the taste of the top decision-maker—not to a mythical low-cost structure. The $5.57M figure covers only the GPU-hours cost of the final training run; total investment “may need two more zeroes.” Quant engineers convert 0.1 milliseconds directly into $10M, allowing them to push hardware and engineering optimization to depths few internet teams reach. He believes DeepSeek can treat the pursuit of AGI as consumption rather than investment, and that the China–US rivalry gives it access to other resources without requiring a return commitment.

  • Metaso’s technical moat is a system that dynamically recombines as models advance, not a bet on one proprietary or external model. Beyond integrating DeepSeek, 闵可锐 says all of Metaso’s earlier models were trained in-house; the actual architecture may include more than four models, using specialized models in exchange for 10x speed or cost advantages. The company has around 60 people, with 闵可锐 simultaneously leading product and R&D, participating in model training, sketching UI concepts and breaking projects into 100 tasks—founder leverage that makes decisions fast but is also the organization’s clearest bottleneck.

Deep dive

1. “What Should I Learn Today?” works end to end, but 闵可锐 gives it only 70 points

  • 冯骥 posted about it unsolicited despite having no personal relationship with the team, calling it “the best user interface of any AI tool I’ve used to date” and saying it was the first time he could clearly see how AI might transform education.

  • 闵可锐 considers that assessment “a bit overpraised”: the current version gets at most 70 points, with substantial room to optimize and polish the interface, script quality, slide aesthetics and feature depth.

  • What he credits is not a local feature breakthrough, but the possibility that this is the first product to connect material processing, course organization, visual presentation and explanation into a complete loop, producing content and an experience that are “pretty good.”

2. Parents pushed an adult-learning tool toward K12 problem solving

  • The product was originally designed for adults and lifelong learners: users find relevant material, then Metaso restructures it into a friendlier course and presents it in the user’s preferred style.

  • After launch, more than one parent wanted to upload an exam paper and have AI provide the answers, organize the key points, difficult points and knowledge points, and actually teach the child. The request exceeded the original definition but exposed a direct need.

  • 闵可锐’s view is that the K12 knowledge range is largely within current model limits; the main gap is stitching the full problem-solving workflow into a good user experience.

  • The team is therefore preparing an upgrade, which 闵可锐 expects could launch as early as next week. He joked that if listeners treat it as a business opportunity, “it may already be too late” by the time the episode airs.

3. A software form tries to bypass learning-device price and experience constraints

  • 闵可锐 observes that major education companies typically put AI capabilities into learning machines and study-practice devices priced at several thousand yuan, rather than offering the same service as pure software.

  • He gives two reasons to reject that model: hardware prices limit the number of people who can use the service, while dedicated learning devices struggle to match the computing power, animation and visual quality of the best phones and PCs.

  • 曼琪 imagined that a user could upload one exam paper and receive a one-hour course explaining it from start to finish, with teaching quality broadly on par with a teacher and perhaps better than most teachers. That was the host’s scenario judgment, not 闵可锐’s promise about the final product.

  • 闵可锐 did not define the future too narrowly, saying it is still too early to tell whether the product will become an education product or a content platform three to six months from now.

4. “What Should I Learn Today?” came from AI search’s long-content problem

  • After more than a year building search, Metaso found that AI made information retrieval easier but typically produced a 300–800-word “short essay”; that may suffice for precise fact checks but is nowhere near enough for broad topics.

  • For the latter type of need, search should recommend the relevant PDF, paper or original text. Even a 100-page context may be the material the user genuinely needs to learn.

  • “What Should I Learn Today?” was not born from a flash of inspiration. Metaso had already built document-library and academic products, forming a judgment that the original material itself carries substantial value; the problem is that ordinary people struggle to digest it.

5. Metaso chose to summarize in reverse: less compression, more absorption

  • 闵可锐 estimates that roughly a year after ChatGPT launched, more than 100 audio, video, podcast and public-account summary products may have appeared; many platforms also added summarization as a feature.

  • His central objection is that the information loss “may be very large”: different readers care about different things, and the deepest valuable information often cannot be extracted into a summary.

  • He uses One Hundred Years of Solitude as an example: readers may lose track of the characters by page 50, while academic papers assume prerequisite knowledge and follow academic conventions. Metaso’s goal is not to shorten 800 pages again, but to lower the barrier to entering the original work.

6. A good course must fill in the layer of knowledge experts omit

  • After 王与桐’s 2023 interview was imported into the product, the course immediately generated diagrams of Transformer, BERT, RNN, encoder and decoder structures. These were neither fully explained in the original conversation nor taken from prebuilt material.

  • 闵可锐 believes that “the best person in the industry” is not necessarily the best teacher because they are too far removed from ordinary users and assume too much background knowledge; a generated course must actively add annotations, metaphors and analogies.

  • A typical addition is that LSTM processes one step after another, with each later step depending on the previous one, making execution slower, while Transformer is better suited to parallelism. That background materially lowers the barrier to understanding.

7. Product decisions are compressed into the founder’s 10-minute window

  • Outsiders assume that a product-oriented Metaso must have a powerful product team. 闵可锐’s correction is blunt: “More than 80% of product-strategy decisions are made on a gut call by me.”

  • He reads extensive feedback but conducts fewer and fewer formal interviews or research projects. Repeated demand still will not be built if it does not make sense; a one-off comment can be immediately prioritized if it allows him to derive a coherent case for its importance.

  • At a large company, a product manager must use research data to prove the problem, hold review meetings, discuss technical feasibility and then schedule the work. 闵可锐 tries to compress the entire process into “a 10-minute consideration cycle.”

  • The method is better suited to consumer products. For professional or B2B scenarios, he admits that the team must deeply understand the relevant group; simply asking whether users “need it” usually produces a yes and may not generate useful communication.

8. Metaso did not build an Agent just because Agents were hot

  • 闵可锐’s counter-question is direct: “Shouldn’t you work on something valuable in the market that nobody has done? Why do something everyone else is doing?”

  • When Metaso AI Search was launched in late December 2024, few people in China viewed search as the main track. “What Should I Learn Today?” likewise came from its own chain of questions rather than a financing trend.

  • He has not developed an intuition for which Agent direction will be especially good. Imagining demand is not scarce; what is scarce is reaching the model stage at which most people agree that the product works.

9. General-purpose Agent success rose from 3% to 30%, still far short of 95%

  • 闵可锐 tested AutoGPT for roughly 15 minutes in 2023 and concluded that “it won’t work now,” even though it may have accumulated tens of thousands of GitHub stars within a week.

  • His rough scale is that perhaps 3 of 100 tasks could be completed in 2023, rising to 30 that performed well in 2025. The improvement is significant, but still below the user expectation that at least 95 tasks succeed and meet expectations.

  • He sees “I have a computer and want to make $100M—give me a plan” as treating an Agent like Aladdin’s lamp. The boundary between a reasonable task and pure wish fulfillment is precisely what makes “general-purpose” so difficult to define.

10. Vertical Agents work only when they deliver the result

  • If 2 of 20 tasks are legal tasks such as contract review, a specialized team can abandon the rest of the capability set and optimize that one point until it becomes valuable enough.

  • 闵可锐’s pessimistic conclusion on Chinese legal tools is that merely “helping lawyers improve efficiency” has no opportunity because the economics do not work; what might work in B2B is “pure replacement.”

  • Metaso seriously studied online law firms: law firms acquire large volumes of cases through advertising on Douyin and Kuaishou, and some cases could theoretically be automated end to end, with theoretical efficiency reaching 100x that of traditional lawyers.

  • The project was not pursued because Metaso is not good at front-end customer acquisition. A partnership with a third party could instead fail because the two sides did not understand each other’s needs. 曼琪 concluded that if the model is really to work, Metaso may need to build a law firm itself.

11. China’s professional market is too small, forcing Metaso from vertical to general

  • When Metaso was founded in 2018, it started with legal translation and served law firms directly. Writing Cat then expanded to writers, while AI Search and “What Should I Learn Today?” moved further toward the broad consumer market.

  • 闵可锐 says the core reason is that China’s To B and professional markets are too small and difficult to monetize. The decline in law-firm business over the past few years also directly hit its original customers.

  • In a parallel universe in the US, Metaso might have continued focusing on professional services: US legal products could reach $100M in annual revenue, and the growth of companies such as Harvey shows that such a market may exist.

12. Product selection first follows the principle that the company must survive

  • 闵可锐 reduces the goal of a domestic product to users or revenue: “You have to aim for at least one of the two.” Revenue’s first meaning is not profit maximization, but independent survival.

  • Many companies that were once famous have disappeared or contracted sharply over the past seven years. Metaso has therefore always valued real users, payments and cash flow rather than relying solely on the next funding round.

  • He does not require every product to hit both metrics. 曼琪 suggested that some products might need only tens of thousands of subscribers, strong conversion and solid delivery to reach $1M in monthly recurring revenue; this was a hypothetical example of product trade-offs.

13. Long-term value fits Metaso’s bets better than a two-week viral spike

  • 闵可锐 does not want to spend three to six months for only two weeks of distribution. AI filters, lip-syncing and “Little Show”-style formats look more like fads, and fads quickly peak and fade.

  • Better search and lower barriers to teaching can create value over time. A product need not go viral, but users should still have a reason to keep using it after the buzz fades.

  • Distribution is also difficult to control: making the DeepSeek model good was within the team’s control, while its later mass-market fame was impossible to predict. Metaso sometimes spends on marketing, but cannot match the distribution resources of the largest companies.

  • 闵可锐 estimates that a giant such as Yuanbao could spend tens of millions of yuan a day. Internal resources such as WeChat’s nine-grid placement are difficult even to price and may be equivalent to Metaso’s entire annual budget.

14. Application-driven development makes room for specialized models

  • 闵可锐 chose application-driven development as early as 2023: without an application pulling the model team forward, it can become “good at everything” while remaining insufficiently usable in every direction.

  • Model competition is also a very flat world. A code model is not judged only on a “China leaderboard”; it faces Claude 3.7 directly. Being perpetually slightly worse than the leader makes it difficult to justify a product.

  • Metaso therefore does not pursue a “hexagonal warrior” that maxes out language, math and code. It allows models to specialize around user tasks in exchange for speed and cost advantages.

15. DeepSeek’s talent strategy was “how it should have been done”

  • When 闵可锐 first saw DeepSeek’s talent organization, he thought, “Wasn’t this how it should have been done all along?” He found it awkward when large companies later reverse-engineered the reasons for its success and imitated them.

  • 曼琪 asked why Metaso did not organize itself the same way. His answer was unvarnished: “Because I don’t have the money. If I had the money, I would do the same.” One DeepSeek researcher may cost as much to develop as 10 or more people at Metaso.

  • He admires 梁文锋’s “taste”: when smart people face equally brilliant competitors, they should not imagine they can handle 8 products and 2 R&D lines at once. They should suppress one major opportunity, concentrate resources and do it as well as possible, even at the expense of products and some users.

16. DeepSeek can treat AGI as consumption rather than calculate the return first

  • Asked how DeepSeek will create commercial value over the long term, 闵可锐 asks: “Why must it create commercial value?” Someone with enough money can treat the pursuit of AGI as consumption, like buying a luxury car.

  • This explains the difference he sees between DeepSeek and Alibaba, Tencent and ByteDance: the latter have more talent, chips and money, but cannot pursue AGI as pure consumption in the same way.

  • DeepSeek’s entry into the China–US rivalry changed its resource position again. 闵可锐 believes it can obtain more funding through other channels without needing to promise a return of the investment type.

17. Alibaba has not been separated from DeepSeek by a chasm; ByteDance’s problem is organization

  • 闵可锐 sees the gap between Alibaba and DeepSeek as closer to A versus A-, not a huge chasm. He believes Alibaba teams were probably already using reinforcement learning, producing better data and studying reasoning internally before R1 appeared.

  • Alibaba may have lost the chance for a nationwide communications moment rather than having fallen completely behind. Open-source models also let application companies learn from and fine-tune them while emphasizing controllability.

  • On ByteDance, he sees no problem with talent density or resources, but “given the money it spent,” the result was not A-level. Whether strong individuals can perform depends heavily on how the top decision-maker allocates core resources.

  • He wrote in a WeChat post while reviewing Llama 4: “First-rate resources and a second-rate team cannot beat second-rate resources and a first-rate team.” Strong résumés do not mean an organization can make those people work as one.

18. Quantitative-engineering DNA turns 0.1 milliseconds into calculable money

  • 闵可锐 believes quantitative engineers squeeze hardware utilization to the limit because “0.1 milliseconds means $10M.”

  • An ordinary internet engineer may not know how many milliseconds it takes to read a file from disk or how many times a simple for loop can execute per second. Quant engineers understand those orders of magnitude instinctively.

  • The dense engineering innovations in DeepSeek’s papers are therefore not accidental. Even if 梁文锋 did not come from a conventional long-term AI research background, he came from a group with unusually strong engineering practice and optimization experience. 曼琪 added that the team later brought in people with strong algorithmic and mathematical abilities, creating the combination.

19. “I’m not an actor” determined Metaso’s fundraising pace

  • 闵可锐 says that “even when I tell the truth, people don’t believe me,” so he is even less able to invent a story he does not believe. He questions the weaknesses in advance and knows he may not answer every follow-up question.

  • He cannot understand why some “seemingly very clumsy stories” win broad buy-in, nor does he know how to make a flawed story accepted by the market: “That is beyond my ability boundary, so I can’t do this job.”

  • The AI primary market can also reverse 180 degrees every two or three months. By the time a story wins approval, investor sentiment may have flipped before diligence ends. Even if he were willing to perform, the mechanics would be extremely difficult.

20. More money is not a free option; it changes the company’s constraints

  • After the 2023 boom, Metaso disclosed only one financing round of a little over RMB100M. 闵可锐 says taking more money at the time probably would not have been a problem, but the funds would not have been enough to compete with leading companies on the same model, while an application might not need that much upfront investment.

  • He rejects the logic of “if you have RMB500M on the balance sheet, spend RMB400M first and shrink later with RMB100M left.” Additional funding brings dilution and introduces the preferences of different investors into product direction.

  • Some investors want a more aggressive consumer strategy; others want a stronger B2B business. The more investors there are, the harder it is to form a unified expectation. In an uncertain market, 闵可锐 would rather bootstrap with users and revenue first.

21. Metaso rejected acquisition because its downside value was still defensible

  • In 2024, some parties discussed investment while others proposed a share swap, merger or outright acquisition. 闵可锐 rejected them because it was “too early,” not because the offers were necessarily too low.

  • His logic was that Metaso’s valuation was not high and that, as long as the products continued to improve, it would remain an acquisition target: “There’s no need to rush. We haven’t reached the point where people can’t afford to buy you.”

  • His most pessimistic calculation is to treat the search data, architecture and engineering reserves “as scrap metal,” repay every investor and still retain a small surplus. Raising $75M might instead make that safety line disappear.

  • There is no answer for the upside. Metaso could become a company with broad users and excellent products, or build a DJI-like technical moat that cannot be copied with a few hundred million dollars. 闵可锐 does not try to summarize its mission, vision or endpoint.

22. “What Should I Learn Today?” is better suited than search to carry overseas revenue

  • Metaso is considering taking “What Should I Learn Today?” overseas. Even in the US, where willingness to pay is stronger, general search is often viewed as free; 闵可锐 does not expect overseas revenue to materialize immediately.

  • 曼琪 suggested that the value delivered by a learning product is more concentrated, and that tens of thousands of subscribers in Europe and the US could support $1M in monthly recurring revenue. That was a commercial-scale hypothesis; 闵可锐 only explicitly said the product is better suited to overseas expansion.

  • The team has no current plan to move staff overseas. Metaso AI Search is nearing 1M DAU; investor expectations pushed it toward user growth in 2024, while in 2025 it plans to put more weight on revenue.

23. After three giants each bet RMB5B, a small company stops fighting for the top three

  • By the end of 2024, 闵可锐 had already judged that Metaso’s 2025 competitors would no longer be companies worth tens or hundreds of billions of yuan, but trillion- and 10-trillion-yuan enterprises: Tencent Yuanbao would probably invest RMB5B, Alibaba might invest no less than RMB5B, and ByteDance clearly would invest no less than RMB5B.

  • Once well-resourced players effectively lock up the top three positions, Metaso’s question changes from whether it can reach 10M DAU to whether it should keep pursuing 1M–2M DAU or first solve revenue and survival.

  • The unexpected variable was DeepSeek reaching the top spot—and still holding first place at the time of the interview. 闵可锐 will not comfort himself by assuming that large companies might not act this way; he insists that the worst case will happen.

24. Writing Cat showed how models swallow application features

  • Writing Cat launched its AI-writing feature around October 2022, roughly two weeks before ChatGPT launched and formally went live at the end of November. The window to accumulate resources was extremely short.

  • AI writing subsequently became standard in free models, and Writing Cat’s long-term value shifted toward assistance such as checking and proofreading. Subscription use remained relatively stable, though 闵可锐 said he would need to check the exact data and that usage is heavily affected by thesis and exam seasons.

  • 闵可锐 refuses to draw a permanent boundary between models and applications. His case-by-case rule is simple: after a new model launches, will the product gain more users or lose more users?

25. 10M DAU sends an application into the model companies’ hunting ground

  • 闵可锐 proposes a more direct boundary: whenever a relatively general model capability could support an application with 10M DAU, model companies will be interested.

  • If a product combines a large user base, online delivery and heavy dependence on foundation models, it is difficult for a startup to explain why a model company would not build it itself. In his view, this is “an objective prediction, not a pessimistic prediction.”

  • Defensible value must lie elsewhere—in a long sales cycle, professional services or proprietary workflows—not merely in the lightest possible model wrapper that captures the most users.

  • On Manus, 闵可锐 values the team’s “extreme flexibility and lack of baggage.” Compared with teams arriving with a big-tech or star-founder halo, it can change direction more easily and does not have to endure questions about “abandoning its beliefs.”

26. The market can push Jasper straight from the altar to “dead”

  • 闵可锐 uses Jasper to illustrate swings in sentiment: from being celebrated to a lower valuation, layoffs and then “everyone taking a turn kicking it,” the entire sequence may have taken only a year.

  • But by his recollection, Jasper still had roughly $100M in ARR in 2025. A three-year-old company reaching that scale has already achieved something 99.99% of startups never can.

  • His conclusion is not that Jasper escaped the shock, but that the market often moves from one extreme to the other within three months, and both extreme judgments may ultimately be wrong.

27. DeepSeek’s $5.57M is not the full cost

  • 闵可锐 believes the most common misunderstanding is that “DeepSeek was very cheap.” The $5.57M covers only the GPU-hours cost of the final training run and does not represent the absence of long-term R&D investment.

  • He estimates the full investment “may need two more zeroes.” DeepSeek was already one of China’s wealthiest companies with one of its deepest talent pools; it did not prove that a company without money can build a strong foundation model.

  • The opposite misunderstanding about Metaso is that it merely “integrated some model.” 闵可锐 says that beyond this DeepSeek integration, every model Metaso used previously was trained in-house; it simply never packaged them for external API sales.

28. Metaso’s model architecture is not a two-piece “large model plus small model” stack

  • The public explanation that “DeepSeek and R1 handle deep reasoning, while Metaso’s proprietary LLM handles search and integration” is deliberately simplified. The actual system may contain more than 2 or 3 models, perhaps even more than 4.

  • The architecture keeps changing with new models, new open-source capabilities and business conditions. 闵可锐 does not believe in one fixed route; he continually adjusts based on the latest resources and information available.

  • The value of in-house models is that they can specialize. Sacrificing irrelevant capabilities in exchange for a 10x speed or cost advantage can create a real edge in application competition.

  • If the next generation of models erases those differences, Metaso’s task is not to defend an old moat but to quickly find “the most worthwhile thing to do” in the new cycle and optimize it as well as possible.

29. Model capability accumulates linearly, but outsiders perceive steps

  • 闵可锐 relays the view of a frontline model executive: internal capability improvements over the past two years have been more linear, while outsiders see steps only at version releases because teams optimize many things simultaneously.

  • He considers o1 a substantive breakthrough, but one that mainly steepened the improvement curve for relatively verifiable tasks such as math and logic. Reasoning-generated data can also feed back into the base model’s fast-thinking capabilities.

  • Companies such as Anthropic have also discussed not strictly separating reasoning and non-reasoning models. The best model will likely absorb strengths from multiple approaches rather than remain permanently split into two tracks.

30. “Does it have intelligence?” is harder to verify than “Can it get the answer right?”

  • 曼琪 relays 马毅’s counterexample: a model can solve competition-math problems seen during training but still get elementary-school questions wrong; if it truly understood math, the inversion in capability would seem impossible.

  • 闵可锐 responds that reasoning and intelligence themselves are difficult to define. Humans can also make mistakes in a three-step instruction, so one cannot simply conclude that they lack intelligence.

  • He compares it to a mathematician solving an elementary problem: the mathematician may not know the trick for the chickens-and-rabbits puzzle but can set up equations or use integration to reach the answer. Performance through different paths alone cannot prove or disprove understanding.

  • 曼琪 believes ordinary users have found it increasingly difficult to perceive improvements in language ability since GPT-4; multimodal advances such as GPT-4o are more obvious. 王与桐 gives a more gradual example: in multi-turn ambiguous contexts, models now proactively add tables and diagrams.

31. Benchmarks can be gamed upward without matching real use

  • 闵可锐 believes that whenever a task can be measured, teams can design methods on the data and algorithm sides to keep raising the score. The question is whether the benchmark actually corresponds to a real need.

  • When he sees people comparing o3 with o1, the test often looks more like searching for weaknesses and proving “I’m still better than it” than identifying the boundaries of tasks where humans and models can work well together.

  • The real difficulty is establishing evaluation standards for real-world tasks and then improving models against them. That may also be the core problem OpenAI and other model companies need to solve.

32. Reinforcement learning cannot conjure an answer from a base model with no solution

  • 闵可锐 believes many people underestimate pretraining. In the R1 paper’s GRPO example, the model randomly generates 64 or 128 paths and reinforces the paths that get the answer right.

  • If a hard problem produces no usable solution after 128 attempts or even 10,000, the method does not work. Reinforcement learning can reduce “guessing 100 times” to 2 attempts or 1, but it cannot create a capability the base model has never reached.

  • 闵可锐 sees Grok 3 as a success: top engineers paired with the best hardware caught up with first-tier model capability in a relatively short time, proving that “brute force can still produce miracles.”

  • Llama 4 shows that having the resources, code, data and GPUs in place does not guarantee successful linear extrapolation. Algorithms, data processing and team execution still determine whether the model can convert resources into results.

33. Hitting a data wall does not mean technical iteration has ended

  • The original scaling law was based mainly on text tokens. The more realistic bottleneck today may be high-quality data rather than compute alone. Multimodal data could help, but it no longer fully matches the definition and initial conditions of the original scaling law.

  • 闵可锐 compares the future to semiconductor process technology: each apparent wall is followed by top teams finding indirect paths such as immersion lithography. Those breakthroughs are difficult and cannot be predicted linearly in advance.

  • For Metaso, the ultimate technological endpoint is less important. The direct question is which capabilities the next-generation model unlocks from “doesn’t work” to “works well,” and whether those capabilities can be connected to user demand immediately.

34. Coding and web generation are the most usable new capabilities of the past six months

  • 闵可锐 sees the code capabilities of Claude 3.5 through 3.7 as a major unlock: with one sentence, the model can not only generate a website but also produce something with “pretty good aesthetics” even when the user has not described the design rules.

  • Metaso AI Search launched interactive web-page generation in March 2025 precisely in response to this capability shift. “What Should I Learn Today?” also uses the now-improved code and HTML generation capabilities.

  • Agents benefit from longer planning horizons and self-correction as well, with success rates moving from single digits to several tens of percentage points. But it is important to identify who provides the capability: the polished result may be primarily Claude’s contribution.

35. Not betting on GPT-3 applications in 2020 was a reasonable judgment, not an obvious mistake

  • GPT-3 in 2020 was a 175B dense model. 闵可锐 estimates that DeepSeek’s reasoning model in 2025 actually activates roughly 30B parameters, with extreme optimization potentially reducing its compute cost to one-fifth of GPT-3’s.

  • Even with several generations of updated GPUs, Metaso still could not serve GPT-3 at its original scale today. Given a startup’s costs and monetization prospects, viewing it as a laboratory product was entirely reasonable.

  • OpenAI’s special capability was to use a model to explain scaling and AGI three to five years out, then secure enough resources to keep pushing forward. 闵可锐 sees that as beyond the capability of ordinary founders, not an operating path that can simply be copied.

36. Cost optimization can save a company without necessarily winning users

  • Metaso optimizes per-query cost much more aggressively than large companies, sometimes reducing resource use to one-tenth or even one-hundredth of a rival’s. The trade-off is greater conservatism when applying models.

  • Users will not forgive a 90-versus-95-point gap simply because the product consumes one-tenth the resources. 闵可锐 admits that cost engineering is highly valuable internally but “not that important” from the outside.

  • AI Search has done almost no paid conversion, and a small number of B2B partnerships still cannot cover inference costs. 曼琪’s example of Manus spending $1.44M on Claude tokens in 14 days shows another kind of growth pressure.

  • 闵可锐 believes it would of course be better to avoid the giants and find a sufficiently large niche, but the real difficulty is figuring out how. Metaso can only “feel its way across the river,” while “the big companies can feel their way across on top of us.”

37. When the previous product is copied, the next one must already exist

  • 闵可锐 does not deny the speed of big-tech imitation: once a product shows potential, a copycat appearing within three months is a highly predictable environmental variable.

  • A small team’s response is not to hope competitors overlook it, but to know in advance “where exactly your next one is” when the other side begins copying the previous product.

  • He compares this again to lithography machines: nobody can predict the full path 10 years out, but at minimum one should find the next step. The next good idea is often hidden in the execution details of the current product.

38. A 60-person organization’s flexibility comes from the founder planning things himself

  • Metaso has around 60 people, only modestly above the 40-plus it had at the beginning of 2023. The organization appears flexible mainly because 闵可锐 does not need to persuade many people through formal mechanisms.

  • When building “What Should I Learn Today?”, he broke the project into 100 pieces and assigned each person 2 or 3. The host compared him to the planning model in an Agent, with the rest of the team handling execution.

  • The ideal state would be for him to break the project into 10 pieces, have 10 people break those into 200, and generate results beyond his expectations through interaction. 闵可锐 admits that this state “has never happened.”

  • Good ideas are not frequent in reality: “The vast majority of ideas from the vast majority of people in the market are not good enough.” If good ideas were that dense, large companies would not need to watch what startups are doing.

39. Hiring is harder because the people Metaso wants are also wanted by the giants

  • The AI boom did not expand Metaso’s available talent pool; it made recruiting harder. The company wants people with “a little AGI faith, but not too much,” rather than the type directly competing with DeepSeek or ByteDance.

  • 闵可锐 wants colleagues who “have ideas, have ability and do not have too much ego.” More precisely, ability and ego should be matched, avoiding people with many ideas but insufficient delivery capability.

  • Metaso even encountered an intern from a certain “Little Tiger” company. After Metaso issued an operations offer, the intern went back, received a higher offer and was hired back by the original company. In the end, competition often degenerates into “everyone bidding with money.”

  • He recognizes an advantage in professor-led teams: teachers and students spend 2 years together, learn each other’s capability boundaries and build trust. In the open market, both sides talk for only 2 hours, so candidates can mainly compare whose pitch and compensation are better.

40. 闵可锐 “had to become a product manager” because AI product judgment requires technical feel

  • He leads product and R&D simultaneously and also participates in model training. 王与桐 believes a good AI product manager needs technical understanding, while someone who understands foundation models deeply enough is unlikely to switch into product work.

  • 闵可锐’s distinction is that when the best model fails on a task, he asks why and whether it can be made to work, rather than treating “this model does not work” as a hard capability boundary.

  • He does not take notes and often builds products in his head. He sketched the “What Should I Learn Today?” interface with a book on the left and an outline on the right, and may decide the number of post-course interactive questions within a minute.

  • A new colleague cannot realistically take over 50% of his details within a month. The more viable path is long-term collaboration, gradually transferring the work of further decomposition to people whose baseline is already solid.

41. As the market accelerates, the boss’s patience also gets compressed

  • During the AI downcycle, Metaso could polish a product until it was ready before launching, without worrying that a comparable product would appear halfway through. Today the market changes constantly, and launch and iteration cycles have accelerated sharply.

  • 闵可锐 admits that early colleagues received more detailed explanations and more setup. Now he may give them the result directly and ask the team to flesh it out and move quickly.

  • But he does not believe a world in which 20 models launch every month can continue for another two or three years. Competition may not last as long as everyone thinks; the timing is simply not yet clear.

42. Steady state means revenue, spending and product accumulation begin to match

  • 闵可锐 defines steady state as users and revenue continuing to rise even without steep growth, revenue and spending moving in sync, and product capabilities continuing to compound.

  • Legal translation and Writing Cat could each be considered break-even on a standalone basis, but their growth is insufficient. Ideally, a product should keep doubling through its first 2 years and still grow by several tens of percentage points in its third year.

  • Before ChatGPT, Metaso had achieved company-wide break-even. If it targeted profitability immediately, it could reach it relatively quickly, but only by sacrificing certain things.

  • The current strategy remains to “put out quickly whatever needs to be put out quickly,” while judging revenue and profitability dynamically. If the result misses expectations, the team goes back to identify which assumption failed and optimizes the corresponding link.

43. 闵可锐 may have overinvested in detail and underinvested in fundraising

  • Looking back over 2 years, he is certain that Metaso invested heavily in product and R&D and suspects he overinvested in details. He is not equally certain about what received too little investment; fundraising is one example 曼琪 raised.

  • When Metaso became popular in 2024, he did not accept many investor meeting invitations. Spending several months on fundraising might have produced better terms and valuation, but he still asked: “Was this really worth it?”

  • Primary-market communication depends heavily on meeting in person and “smelling the scent”; it cannot be handled solely through an Agent. That means fundraising competes with product R&D for the founder’s time.

44. Founder growth is not about becoming better at storytelling, but “getting better at everything”

  • 闵可锐 summarizes his growth over the past 2 years as “I’ve become better at it”: engineering, product, details and judgment all became more precise because he had to handle them himself.

  • The satisfaction he enjoys does not come from immersing himself in one role. It comes from building a product puzzle in half an hour, spending 3 months making it real, and adding new pieces during execution.

  • AI amplifies that personal engine. He can handle models, product, design and technical experiments while the team executes the plan. The leverage works, but it also makes the company highly dependent on the founder’s single-point output.

  • He even believes that without Metaso, he could train people with basic programming experience over 2 or 3 months and give them a chance to enter a first-tier foundation-model team, perhaps competing for annual salaries of RMB1M after a short period of training. But the audience is too narrow and the approach could create counterproductive effects, so it is not suitable as the current main product.

45. The 2025 goal remains to do the work at hand well and faster

  • 闵可锐 offers no grand annual narrative: first make the existing products “better and a little faster,” then help the younger people on the team become “better at it” rather than merely executing.

  • He wants people with ability, ideas and organizational fit to join, so the founder no longer has to supply unlimited output. That matters more than immediately scaling to a larger size.

  • Fundraising is not on the explicit goal list; the answer remains “let it happen naturally.” Metaso continues operating not because it is certain about the endpoint, but because each stage is better than the last while preserving a defensible base of value.