Pioneers Insight Method Research Author
158: DeepSeek Before V4 Launch: Talent Competition, Organizational Traits, and a Distinctive AGI Goal | Solo
Back to Episodes

158: DeepSeek Before V4 Launch: Talent Competition, Organizational Traits, and a Distinctive AGI Goal | Solo

Summary

  • Ahead of V4, the real issue is not a “mass exodus” but the departure of multiple core authors from the research team. DeepSeek has fewer than 200 employees, including a little over 100 in R&D; 王炳轩 was recruited away by 姚舜宇 at Tencent, while 魏浩然 and 郭达雅 have since left, with the latter 2 at the time only rumored to be joining major tech companies. 程曼祺 stressed that 3 departures are not a large number for a company of this size, and that the rumored group exodus “did not happen.” The small-parameter V4 variant was handed to several open-source frameworks for adaptation around January 2026; it is expected to be open-sourced in April, with community information pointing to versions near 300B and 500–600B.
  • The AI talent bidding war and the wealth effect across the industry are turning DeepSeek’s opaque option pricing into a retention problem. A Tencent Tsinghua Yao Class internship posting shared in the podcast group listed RMB5,500/day pre-tax; another comment put Qwen at RMB4,200/day. The highest internship pay 程曼祺 had previously heard of was RMB4,000/day for select top interns at the Cursor team. If the RMB5,500 figure is accurate and the intern works every day for a full month, monthly pay would exceed RMB100K. Meanwhile, MiniMax and Zhipu saw their share prices rise 5–6x after listing, reaching market caps of RMB250B–300B, prompting 梁文锋 to establish a clearer valuation for the company.
  • 梁文锋’s AGI roadmap is not a single-minded bet on model performance; it gives equal weight to China’s domestic ecosystem and original research with no clear near-term payoff. DeepSeek is using UE8M0 FP8 for next-generation Chinese chips and TileLang in place of CUDA and Triton, while also pursuing Janus, Prover, OCR and learning mechanisms closer to the human brain. The strategy could create differentiation, but it clashes with some researchers’ desire to work on the strongest models, secure prominent technical-report authorship, access ample GPU capacity and receive immediate external feedback. The goal “is not simply to chase the strongest model or win a performance arms race.”
  • The Agent era is rewriting what “strongest” means; DeepSeek’s model foundation remains strong, but product reach and iteration speed are clear gaps. Since early 2025, Zhipu, MiniMax and Kimi have released 5, 4 and 3 model updates, respectively. DeepSeek V3.2 ranked 12th in OpenRouter consumption among OpenClaw models over the past 30 days, placing in the top 10 paid models after free models were excluded. The fact that an older model still ranks near the top is itself a sign of competitiveness, but DeepSeek has invested less in AI coding, general-purpose Agents and OpenClaw applications, leaving it short of the long-tail use cases and diverse data generated by product distribution.
  • DeepSeek’s hardest-to-replicate asset is its research organization: fewer than 200 people, limited overtime and cross-team collaboration once took it to the global top tier. 梁文锋 believes high-quality work is capped at 6–8 hours a day, while fatigue produces poor judgment and “wastes valuable compute resources.” The research team has only 2 levels—梁文锋 and the researchers—and 3–5 people can initiate a new direction across teams. Before V3 and R1, DeepSeek achieved breakthroughs with roughly one-tenth the headcount and one-third to one-half the per-person working hours of major tech companies, but the global compute arms race is putting its resource advantage under pressure.
  • DeepSeek has begun to change, but V4 should not be treated as another miracle that proves whether the company will succeed or fail. Since autumn 2025, 梁文锋 has spoken more about productization and commercialization; in mid-March, a recruitment notice named Claude Code, OpenClaw and Manus for the first time and sought a model-strategy product manager for the Agent track, while the company also tried to clarify valuation and option expectations. 程曼祺’s closing advice was to cool the temperature: “DeepSeek doesn’t have to be the hope of the whole village; DeepSeek is DeepSeek.” R1 was a miracle precisely because miracles are low-probability events.

Deep dive

1. Core authors are starting to peel away, but there has been no “mass exodus”

  • DeepSeek currently has fewer than 200 employees, including a little over 100 in R&D across data and infra, plus a product team of a few dozen. At the end of 2025, 王炳轩, a core author of the first-generation DeepSeek LLM, was recruited by 姚舜宇 at Tencent. Around the Spring Festival, 魏浩然, a core author of DeepSeek-OCR, left. More recently, 郭达雅, a core author of R1, formally resigned; the latter 2 were reportedly considering moves to major tech companies.

  • 阮冲 left in the first half of 2025 and, after taking some time off, announced in January 2026 that he had joined 元戎启行. He was a veteran of the High-Flyer days and a core contributor to Janus Pro and other multimodal work. 程曼祺’s calibration is that full-time departures had historically been rare, but the latest changes are still only “signs of loosening”: more people have chosen to stay, and the group exodus described outside the company did not happen.

  • V4 has made the personnel changes more sensitive. The small-parameter version was handed to several open-source frameworks for adaptation around January 2026, while the large-parameter version was initially expected, optimistically, around mid-February or before the Spring Festival. The information she currently has is that V4 “should be open-sourced in April,” with at least versions near 300B and 500–600B; other releases could follow later.

2. Talent bidding has turned opaque options into a real DeepSeek weakness

  • ByteDance Seed has about 1,500 people. Even after drawing criticism for team overlap and horse-race competition, it cannot simply be labeled bloated next to DeepMind’s more than 7,000 employees, approaching 8,000; more research directions require more researchers, as well as more compute. In the second half of 2025, Tencent asked 姚舜宇 to lead its AI research effort. A new chief building a team means current joiners have a better chance of securing core roles.

  • Talent pricing has reached interns. A Tencent internship posting for Tsinghua’s Yao Class shared in the podcast group listed RMB5,500/day pre-tax. Someone in the same discussion said a Tsinghua Yao Class student interning at Qwen was earning RMB4,200/day. The highest figure 程曼祺 had previously heard was RMB4,000/day, paid last year by the Cursor team to select top interns. If the RMB5,500 figure is accurate and the intern works every day for a full month, monthly pay would exceed RMB100K.

  • DeepSeek employees have signed option agreements, but the company has not set a clear valuation or price, leaving them unsure what those options are worth. MiniMax and Zhipu saw their share prices rise 5–6x after listing, with both reaching market caps of RMB250B–300B. After the Spring Festival, StepFun and Kimi were also rumored to be planning IPOs. The visible wealth effect among classmates, alumni and former colleagues has made the issue more urgent.

  • One of 梁文锋’s recent efforts to change course is to give the company and its members a clearer valuation and set of expectations.

3. 梁文锋 breaks AGI into performance, domestic ecosystem and original research

  • DeepSeek did not suddenly turn to AI in 2023. 梁文锋 founded High-Flyer in 2015, began using GPUs for deep-learning live trading in 2016, and had almost fully AI-ized its strategies by the end of 2017. Fire-Flyer One was established in 2019 with 1,100 GPUs; by 2021, High-Flyer had 10,000 GPUs. He said buying cards early was “about satisfying a kind of curiosity,” like buying a piano for the home: the family could afford it, and someone was eager to play.

  • Beyond pushing the ceiling of model intelligence, 梁文锋 treats China’s domestic ecosystem as a separate mission. V3.1 adopted UE8M0 FP8 designed for next-generation Chinese chips. When V3.2 was updated at the end of September 2025, its technical report showed that DeepSeek had also replaced the mainstream CUDA and Triton operator libraries with the Chinese open-source project TileLang. This is not just a performance question; it is about building large-model capabilities on a domestic compute stack.

  • The third track is original research that major tech companies and typical startups may be unwilling to fund: Janus attempts to unify multimodal understanding and generation; Prover studies formal proof; OCR converts text sequences into images before feeding them into the model, bringing it closer to how humans read paragraphs and hierarchy. In 2025, DeepSeek also recruited neuroscience and brain-science advisers to explore learning mechanisms more like those of the human brain. The common feature is that the near-term payoff is unclear.

4. The Agent era is redefining “strongest” and exposing a product-reach gap

  • The disagreement begins with the evaluation system. Young researchers want to keep working on the industry’s strongest models, earn authorship on closely watched technical reports and have abundant GPUs for experiments. 梁文锋, by contrast, is not preparing to enter a pure performance arms race. Original research is supposed to tolerate uncertainty, but in an environment trained to look at scores, leaderboards and external feedback, this strategy can feel “out of place.”

  • Benchmarks are becoming harder to equate with real capability, especially as the competition shifts toward agentic models. Models with similar scores can produce completely different user experiences. 程曼祺 believes the long-tail cases and diverse data generated through product reach are becoming more important, while DeepSeek’s long focus on models and limited product investment are relative weaknesses.

  • The iteration gap is now measurable. From early 2025 to the present, Zhipu, MiniMax and Kimi have released 5, 4 and 3 model generations, respectively. MiniMax also released M2.1, M2.5 and M2.7 between January and that point. DeepSeek V3.2 strengthened its Agent capabilities, but its overall update cadence remains materially slower. Even if the forthcoming V4 is likely to remain the strongest open-source model, it cannot cover every definition of what “strong” means across use cases.

  • In OpenRouter data from February 24 to March 26, DeepSeek V3.2 ranked 12th in OpenClaw model consumption. The No. 1 model, Step-3.5-Flash, and the No. 5 model, Trinity Large Preview, were both free during the period; excluding free models, V3.2 made the top 10 among paid models. The ranking suggests that its Agent share among small and midsize developers is relatively low. An older model still ranking near the top can also be read as evidence of competitiveness—but “it depends how you interpret it.”

5. A flat lab traded fewer people and shorter hours for research density

  • DeepSeek employees typically leave between 6 and 7 p.m. 梁文锋 believes it is difficult for anyone to produce high-quality work for more than 6–8 hours a day, and that poor judgment caused by sustained fatigue wastes compute. This is highly unusual among leading AI companies: even DeepMind has operated at high intensity for years, while some researchers at xAI work as many as 80 hours per week.

  • The research organization has only 2 levels—梁文锋 and the other researchers—with no second-in-command. It resembles a large laboratory more than a conventional company. 梁文锋 directly leads a foundation-model architecture team of dozens, participates in architecture sign-off, and links data with infra. He attends each team’s weekly meeting, tracks overall progress and bottlenecks, and acts as the organization’s “connector and glue.”

  • Weekly meetings are usually open to other groups, making cross-team learning and discussion the default. A new direction may begin with just 3–5 people—even researchers from different teams—who agree that an idea is worth testing. They run a small experiment first, and the company adds resources once the potential is confirmed. This organic division of labor avoids the problem at large companies where infra becomes an “internal vendor” and model researchers drift away from engineering.

  • A profile of 172 researchers working on LLM, V2, V3 and R1 found identifiable backgrounds for 84 of them: more than 70% were undergraduates or master’s students, and more than 70% were under 30. Before V3 and R1, DeepSeek reached the global top tier with roughly one-tenth the headcount of major tech companies and one-third to one-half the per-person working hours, demonstrating how much research output it could generate with fewer people and shorter hours.

6. DeepSeek is changing, but V4 should not carry “the hope of the whole village”

  • 梁文锋 long concentrated his time on model R&D. When he met investors in 2023, he wanted an arrangement with a fixed cap on returns modeled on early Microsoft and OpenAI investment agreements. VCs would not accept it, and the fundraise failed. After R1 exploded in 2025, he stopped meeting investors and declined requests from most institutions to establish contact. This was not a refusal to engage—he still speaks frequently with AI professionals—but a decision to “do only a few things.”

  • The need to change is now visible. The second half of the “buying a piano” analogy still holds: there are more people who want to play and more pieces to study. But the ability to “afford it” is changing as the global compute arms race intensifies. 梁文锋 is working to resolve valuation and option issues while increasing product investment. Recruitment notices in mid-March named Claude Code, OpenClaw and Manus for the first time and sought a model-strategy product manager for the Agent track.

  • 程曼祺 frames the final question as how to separate noise from signal during a period of anxiety: preserve a flexible, original research model while responding to pressure to adjust compute, products and incentives. One industry practitioner said of DeepSeek: “People who keep their heads down and do the work may not necessarily be the ones who finish on top in a turbulent, frothy market. But only if more DeepSeeks emerge can China’s tech sector move from replication to leadership.” R1’s miracle should not become the minimum bar for every release: “DeepSeek doesn’t have to be the hope of the whole village; DeepSeek is DeepSeek.”