Pioneers Insight Method Research Author
The Second Half of AI: A Deep Dive into Benchmark and Evaluation | A Conversation with 丁丁, Former Kimi Product Manager
Back to Episodes

The Second Half of AI: A Deep Dive into Benchmark and Evaluation | A Conversation with 丁丁, Former Kimi Product Manager

Summary

  • 丁丁 fully agrees that AI has entered its “second half”: the center of competition is shifting from gaming general-purpose benchmarks to defining problems and evaluations that mirror real business. 曲凯 cited models reaching “graduate student” or “PhD student” levels on leaderboards, yet in real-world deployment “not even quite at the level of an intern”; 丁丁 explained that the final experience is the product of the base model, system prompt, search API, knowledge base, interfaces and other components working together—not any single score.
  • 曲凯 sees DeepSeek as a wake-up call that brought attention back to intelligence through products and DAU. 丁丁’s clear point is that RL only works when built on a strong base model and pre-training. As models improve, complex prompts will matter less, but prompts will not disappear; they will return to being “simple, clear and explicit,” while the dedicated “prompt engineer” may disappear.
  • DAU is a proxy metric that capital markets and organizations readily adopt, but it cannot directly represent progress in intelligence. User scale can generate proprietary data, but an input from a Kuaishou user is not equivalent in training value to 50 consecutive rounds of context culminating in a research report; the key is not more data, but whether high-quality data aligns with the capabilities the target model is meant to improve.
  • 曲凯 likens benchmarks to tools that determine where a model invests its finite “skill points”; 丁丁 sees them as core assets of model companies. A good test set must be realistic, have a difficulty gradient and strong discrimination, and continuously retire old questions and add new ones as models evolve; she may retain a hidden benchmark unknown even to the algorithm team, preventing models from answering to the test or being hacked.
  • Benchmarks must be positively correlated with real user metrics; otherwise the evaluation itself should be rebuilt. Taking Manus as an example, 丁丁 suggests a more meaningful metric might be “the rate at which results are completed in the fewest steps and then downloaded or cited”; auto-eval, human eval and end-to-end experience must also be continuously cross-checked.
  • Evaluation is relatively well-defined for productivity products, while emotional companionship exposes the absence of a universal answer to “what good looks like.” When a user says, “I’ve just been dumped,” some want a solution, some want follow-up questions and memory callbacks, some just want a hug, and others enjoy being challenged; personalization still has to be abstracted into model capabilities, while being “smart enough to play dumb” is itself instruction following.
  • The opportunity in vertical AI is not betting that base models will stagnate, but combining domain know-how, proprietary data, feedback mechanisms and engineering into a closed loop. A strong AI product manager therefore needs to be more full-stack: use the best models and APIs frequently, read papers, build demos firsthand, and retain the classical product manager’s judgment on structure, tone and the overall experience, rather than relying on A/B tests to make local metrics look positive.

Deep dive

1. The Contradiction in AI’s Second Half: Leaderboard Intelligence Has Not Become Real Utility

  • After joining Kimi in early 2024, helping build the app in its early stages and working there for more than a year, 丁丁 says she “fully agrees” with the view that AI has entered its “second half.” She first draws a distinction: evaluation is the process of analyzing model performance, while a benchmark is a set of questions used for testing.

  • 曲凯 points out that models can appear to perform at a graduate or PhD level when climbing leaderboards, yet in real business applications are “at most not even quite at the level of an intern.” 丁丁 adds that what becomes scarcer in AI’s second half is product-manager-style problem definition, with the focus on actual experience and utility.

  • 丁丁 warns that end-to-end products cannot test only the base model; the system prompt, search API, knowledge base, interfaces and the entire workflow jointly determine the result. General-purpose test questions have a different distribution from real user inputs, so high leaderboard scores and poor product usability can naturally coexist.

2. The First Half Established the Base Model as the Priority; Prompt Engineering Is Reverting from a Job to a Tool

  • 曲凯 broadly places China’s first half of AI from early 2023 through late 2024 or early 2025. 丁丁 summarizes the core task as improving the base model and extracting more from pre-training. Kimi focused on RL early on, but RL’s impact must be built on a strong base model and pre-training—a point DeepSeek’s success validated again.

  • Over the past year, companies also tried to package existing capabilities as productivity products: Kimi built an app, ByteDance built 豆包, and MiniMax pursued directions including Talkie and 星野. When base-model capabilities were insufficient or the training paradigm had not yet shifted, prompt engineering was pushed into a position of outsized importance.

  • As models become stronger, users may need only to describe their goals “more simply, clearly and explicitly” to get results that surpass what once required complex prompts. 丁丁’s interpretation of Sam Altman’s comment is that dedicated prompt engineers may cease to exist, but prompts themselves “will definitely exist.”

3. DAU Is a Necessary Feedback Gateway, but an Easily Misread Proxy Variable

  • 曲凯 observes that the industry once shifted from pursuing AGI to competing on product DAU, before DeepSeek pulled attention back to intelligence itself. 丁丁 believes the single-minded pursuit of DAU reflects the inertia of the mobile internet era and is also “a little lazy”: when model quality is difficult to express objectively, scale is what organizations, capital markets and users can understand most easily.

  • She does not dismiss DAU: “You need users in order to get feedback.” Benchmarks can come from general-purpose test sets, online feedback, manually constructed data and synthetic data; sufficient scale can indeed generate proprietary data. The problem is noise and misalignment with the target.

  • The sharpest comparison is this: an input from Kuaishou users does not have the same training value as 50 consecutive rounds of context ultimately requesting a research report. User data becomes fuel for progress in intelligence only after being filtered and aligned with the capability the model is meant to improve.

  • 曲凯 asks why Anthropic can still catch up despite OpenAI having substantially more data. 丁丁 does not dispute the data advantage; she stresses that the outcome is a chain of multiplicative factors: pre-training, SFT activation, RL, the training paradigm and the infrastructure supporting RL. Any one link can change the final experience.

4. Whether to Take on Massive User Volume Is First a Resource Constraint, Not a Product Creed

  • Asked whether, as 梁文锋 six months earlier, she would take on the DAU and data, 丁丁 answers: “If resources were sufficient, I would definitely want to.” She acknowledges that it is hard for a product manager to see a huge user base without wanting to absorb it, but the real prerequisite is whether resources are sufficient.

  • 丁丁 believes DeepSeek’s later handling looked more like letting events take their course, without deliberately taking on all the traffic. The conclusion is not that user data is useless, but a more practical ordering: data value must be assessed alongside resource constraints, the training pipeline and every model component.

  • User data can also extend the product manager’s own cognitive boundaries: people inside a company cannot be experts in every industry or exhaust every modality and use case. High-quality inputs, combined with expert interviews and research, are more likely to help a model company understand real demand and define new benchmarks.

5. Benchmarks Do More Than Measure; They Directly Set the Course of Model Evolution

  • After moving from search products to model products, 丁丁’s biggest shift in perspective was understanding the lifecycle of a test set. Search can reuse a relatively stable dataset over time; once a model solves a dimension, the old benchmark reaches “the end of its lifecycle,” requiring new dimensions and difficulty gradients to be defined continuously.

  • A deep-search question might be: “Find all of Tencent’s financial reports from the past 10 years, then predict how much its net profit will rise this year.” The input, the model’s output and the evaluation of that output together form the smallest unit of a benchmark.

  • The same “good answer” cannot be copied across businesses. Deep search should draw on all available data sources and be as realistic and comprehensive as possible; in emotional companionship, a response such as “Based on your current emotional state, I have the following suggestions: one… two…” may expose product failure despite its clear structure.

  • 丁丁’s core judgment is that “what ultimately differentiates model products or gives them their distinct character is precisely how differently they define the act of writing benchmark questions.” Question difficulty, the standard for a good answer and the trade-offs between old and new capabilities all pull training in different directions. She also warns against simply transferring what model A was trained to do onto model B; deciding which capabilities the next generation should continue to develop is itself a trade-off.

6. A Good Test Set Must Be Realistic, Layered and Dynamic—and Explain User Metrics

  • 丁丁’s principles are that questions should come from real needs; they should have difficulty and discrimination rather than all sitting at the same level; and they should evolve with the model, with solved questions exiting and questions for new capabilities entering. Bad benchmarks are usually “especially simple” or concentrated in a single dimension and difficulty level.

  • Benchmarks must ultimately map to product outcomes. If test performance improves but the relevant user metrics do not, the benchmark should be changed so the two remain aligned; otherwise, “your evaluation has no meaning.” The expected relationship should be at least positively correlated.

  • Traditional metrics such as conversation turns, frequency, duration and retention may not be optimal. Taking Manus as an example, 丁丁 speculates that a metric closer to task value might be the rate at which the model delivers a result in the fewest steps while users download or cite that result.

7. Four Hundred Questions Can Be a Hypothesis, but Top Queries Cannot Represent All Demand

  • 曲凯 uses 400 questions as an example when asking about the scale needed at the beginning of a startup. 丁丁 believes more questions are not necessarily better, as long as there are enough to measure product performance. Ranking user prompts and extracting the 400 most frequent ones is indeed similar to evaluating high-frequency queries in the search era, but noise and invalid inputs must be filtered out first.

  • 曲凯 proposes a 1M QV scenario: the top 400 queries might cover only 200K users, leaving the remaining 800K as long-tail demand. 丁丁 therefore opposes focusing only on top queries. The test set should track the full online distribution as closely as possible and should especially capture strong negative feedback, such as downvotes, to identify the product’s floor.

  • As the user base expands, demand will move from uniformity toward segmentation, and the evaluation dimensions must also be supplemented or adjusted. Benchmarks have no fixed update cycle; 丁丁 says “the faster, the better,” while emphasizing that the data still has to decide.

  • Organizational structure also affects quality. Large companies often have data teams produce labels and evaluation sets, then hand them to strategy, feature or client-side product teams, creating multiple breakpoints in the chain. Startup teams are smaller and communicate over shorter distances, which may let them iterate faster on their understanding of “what good looks like.”

8. Subjective Preferences Cannot Be Eliminated; They Can Only Be Modeled, Segmented and Personalized

  • Mathematics and code have ground truth, making them easier to evaluate and more likely to be adopted; language style and modes of expression lack 100% consensus. When DeepSeek broke out, it was seen as having “philosophical depth and elegance,” implying that the team had implicitly defined that style as good, while other companies may not previously have treated it as an important metric.

  • 曲凯 therefore proposes 2 possible futures: different models developing different personalities, or benchmarks being built for different groups of people and even individual users. 丁丁 notes that some preferences already exist—users may choose Claude for coding and o3 for deep search—but personalization does not necessarily require “1,000 users and 1,000 sets of questions”; it can also be abstracted into product or model capabilities such as memory.

  • On whether users who “just like dumb” models conflict with the push for smarter base models, 丁丁’s answer is direct: there is no conflict, because playing dumb is also instruction following. “You have to make it smart enough that it can play dumb.” 曲凯 sums it up as “great wisdom appearing foolish.”

9. Emotional Companionship Exposes the Hardest Layer of Evaluation: Human Values Have No Standard Answer

  • 丁丁 uses the statement “I got dumped today” to unpack Character.AI’s challenge. A weaker model might immediately list suggestions such as going for a run or seeing friends; a more human-like response might simply ask, “What happened?” With memory, it might follow up: “Didn’t you tell me last week that things were going well between you two?”

  • Another good reply might ask what happened, or simply say, “Sending you a hug. I’m right here with you.” From a therapist’s perspective, the model might first focus on the emotional change and let the user talk instead of rushing to offer solutions. 曲凯 points out that some people want a problem solved, some want comfort, and others would prefer to be challenged with “So what if you got dumped?”

  • A benchmark may therefore serve only around 80% of users, leaving the remaining demand to personalization, stronger base models and niche products. The deeper question is whether a model’s evaluation criteria are essentially a mapping of human values—and whether that mapping is actually correct. Neither person offers a definitive answer.

10. Benchmarks Will Become Secret Assets; Vertical Moats Come from Domain Feedback Loops

  • Productivity scenarios have relatively clear standards for correctness and delivery, making benchmarks easier to define. Companionship products may create niche opportunities but are harder to evaluate uniformly. 曲凯 therefore asks whether people will eventually “steal benchmarks” the way they steal code; 丁丁 believes benchmarks are core assets.

  • 丁丁 says she might maintain a hidden test set known only to herself, “which even the algorithm team should not know about.” Once trainers know the questions, they may unconsciously make the model answer them specifically, or the model may be hacked. She therefore prefers to retest before launch with a secret set.

  • On whether to build out strengths in a “skill tree” or develop evenly, 丁丁 separates the issue into 2 layers. The stronger the base model, the better it will generally perform in vertical domains as well: “A PhD student is simply smarter than an elementary school student.” But vertical products can still use engineering, proprietary data and business understanding to design more accurate feedback signals.

  • In sales, for example, the real moat is not just model capability, but a team understanding sales interactions, positive feedback and reward mechanisms better than anyone else. Deep domain know-how combined with model understanding lets startups integrate the two faster, accumulate users and develop industry insight.

11. AI Product Managers Must Be Full-Stack, but Classical Product Judgment Remains Irreplaceable

  • 丁丁 believes the product manager’s core capability is still “translation”: identifying problems and abstracting users and business scenarios. In the past, that translation led to interaction and structure; today it must also turn the complete business process into evaluation standards, observability metrics and feedback nodes. Data quality and the boundaries of model capability have become significantly more important.

  • She recommends continuously using the best models and APIs, checking how versions change every month or 2, and “trying everything you think you might want to build with AI first.” GPT-4o’s “say it and it appears” image generation and its ability to revise through multiple rounds of instructions are boundaries that can only be updated through firsthand experience.

  • She advocates abandoning the mindset of passing work from module to module and treating oneself simultaneously as product manager, designer and front-end engineer, while even trying the back end to close the entire loop firsthand. Reading papers is also essential, though product managers do not need to understand the underlying principles as deeply as algorithm engineers.

  • The principles from her time at WeChat still hold: define the product structure before building features. The 4 tabs, scan entry point and Moments integration for multi-avatar needs reflected the organic links between modules. A/B tests can run 8 or even 20 experiments, but they cannot generate WeChat’s restrained simplicity or Instagram’s fancy tone; local metrics can also be hacked, ultimately harming the whole product.

12. Hiring a Good AI Product Manager Requires End-to-End Experience, Work Samples and Real Usage Habits

  • 丁丁 would prioritize candidates who completed 0-to-1, end-to-end work at an early-stage model startup or small company, because that experience better demonstrates full-stack ability than responsibility for a single module in a mature process.

  • Candidates who have built a small demo or product in their spare time send a very strong signal. The work does not need to be exceptional, but it must show that the candidate has run the workflow firsthand, validated some ideas and encountered the model’s boundaries.

  • Even if neither of the first 2 conditions is met, one can ask directly: “Which model do you like most? Which one do you use most often? In what scenarios? Why?” Specific answers can reveal industry understanding, enthusiasm and focus.