Pioneers Insight Method Research Author
Vozo Founder 周昌印 on Going from Google X to $1M ARR in 6 Months
Back to Episodes

Vozo Founder 周昌印 on Going from Google X to $1M ARR in 6 Months

Summary

  • Vozo reached $1M ARR in 6 months through a “create curiosity first, then address real demand” strategy, with virtually no marketing budget and only two Product Hunt launches. 周昌印 treated the route itself as strategy: “You definitely don’t want people to think you’re a Me too.” The team first used a never-seen-before hacked feature, rewrite, to establish an innovative brand; after discovering that many users were using it for translation, they pivoted back to the real translate market and went deeper, then expanded into lip sync and photo lip sync. His view: “Me too teams cannot succeed in today’s Gen AI era. That’s my bias.”
  • His core investment judgment is that general-purpose models will never become moats. He built a visual foundation model himself in 2023, before Runway V2 launched, which gave him a feel for the pace of breakthroughs in controllability, consistency, and compute cost. He expected Sora’s lead to be quickly caught—“Google’s might come out in another 3 to 5 months”—and it happened. His conclusion is to “stay as far away from it as possible” in entrepreneurship: only build models with application-specific advantages, such as translation-specific voice cloning, tone preservation, and lip sync, while using foundation models wherever possible.
  • The business is unusually healthy: only $6M–$7M raised, a team of more than 40 with over 70% in R&D, and roughly $6M ARR from a domestic teleprompter app, called Blink overseas, bringing the company to cash-flow break-even. “Everything Vozo earns is our profit.” The two products will merge into one membership system under the single name Vozo.
  • His quantitative definition of PMF is worth copying: 80 of every 100 paying users should renew, leaving 20% for churn caused by their own businesses, alongside an absolute target of first reaching $1M ARR. Reaching the top 3 on Product Hunt is “an operations matter” with little to do with the product itself. Its real value is forcing the team to explain the product in one sentence and providing roughly 1,000 visitors—enough for PMF iteration. “Launching promotion a month earlier or later actually isn’t that important. Getting PMF right matters more.”
  • The most expensive lesson came in 2023, when the company pursued foundational research and applications in parallel: “There was no synergy… you couldn’t help me, and I couldn’t help you,” and the two sides spent a year pulling in different directions. The October 2023 PR for the HiveNet multimodal model was effectively an internal declaration that the company was stopping work on foundation models, even if the PR did not say so. From then on, every R&D project had to start from the product. The researchers did not leave; they cared more about seeing their work used by large numbers of people.
  • His philosophy in moving from Google X researcher to “down-to-earth” entrepreneur is distilled into one conservative rule: “Build the last missing link in the industry chain.” His first startup, VR teleportation, was wishful thinking. The lesson was not to build the first of five links and hope others complete the remaining four. He applies the same framework to Vision Pro: “Even a company as large as Apple would struggle to connect such a massive chain.”
  • Will models swallow the decorative craftsmanship of the application layer? “I think they definitely will.” His internal rule is: “If it’s a standard model, don’t touch it.” Build things close to the application and meaningfully different. Midjourney is the example: the broad text-to-image framework is broadly similar, but detailed tuning creates enormous commercial differences. “Don’t worry too much that one model will do everything… that’s impossible.”
  • Two observations about people: top performers are extremely focused, while ordinary people score everything too evenly. His PhD adviser is nearly 70 and still won two best papers this year; “he actually thinks about relatively few things.” If you rate three priorities at 80, 60, and 40, they are probably closer to 90, 20, and 10. Deliberately widen the variance, ask “What happens if I don’t do this?” and let the data answer.

Deep dive

1. Vozo Was “Hard-Born”: From Freedom of Video Expression to Voice Rewrite

  • 周昌胤’s team returned from the United States to China in 2021 and set an internal mission: “freedom of expression through video.” The aim was to let ordinary people who could not edit—teachers, product managers, and marketing managers—communicate through video. Starting in 2022, the team gathered user demand while pursuing generative AI R&D, incubated and killed several ideas internally in 2023, and did not launch Vozo until July 2024.
  • The first feature was deliberately a lower-dimensional problem: not generating a video from scratch, but voice rewrite. You already have a video and want to change its story. One use case was taking a movie scene that had already gone viral and using it to tell your brand story; another was turning a Thanksgiving promotion into a Christmas promotion with one click. The operation required only a single prompt: “Turn this video into Spanish,” or “Make this video more exciting.”

2. A Hacked Leo DiCaprio Drew the Crowd; Translation Captured the Demand

  • Koji’s first impression came from a flood of posts in the member group: Leo DiCaprio in The Wolf of Wall Street was still speaking with the same serious, impassioned delivery, but about trivial matters. The lip movements and tone changed while the visuals stayed the same. Titanic and Harry Potter were similarly “hacked.” This was the breakout moment during the first Product Hunt launch on July 20.
  • Vozo Translate launched in November and returned to Product Hunt, probably taking the No. 1 spot on the monthly leaderboard. The pivot was data-driven: “A large number of users were using Rewrite to do Translate.” The team interviewed translation users and spent a long time refining the product internally before launch; retention was very high. From there, it followed the demand outward: users who wanted lip sync but not translation led to deeper lip sync work; users who said, “I don’t want to lip-sync my video, I want to lip-sync my photo,” led to photo lip sync in January, which grew quickly. In March, the team previewed a “bigger thing”: “Unlike Rewrite, which was something that had never existed, this is a function people already need, but we’re doing it differently and making it easier to use.”

3. The Route Is the Strategy: “Me Too Teams Cannot Succeed in the Gen AI Era”

  • Koji’s question went to the heart of it: if the team had launched with Translate or Lip Sync from day one, might it have struggled to gain traction? 周昌印 fully agreed. “The route really matters.” The first feature determines the first impression. “In today’s generative AI era, innovation is actually the main form of promotion. You definitely don’t want people to think you’re a Me too.”
  • There is also a deeper organizational problem: Me too is difficult to explain internally, and an innovation team loses morale if it is reduced to copying others. “I think Me too teams cannot succeed in today’s Gen AI era. That’s my bias.”
  • He acknowledged the paradox: the product must capture demand while also having an innovative hook. The solution is sequencing. “Get the innovative brand out first” with Rewrite, then return to the real demand market with Translate and go deep, gradually expanding from there. “This route may not exist for every business, but AI video was lucky enough to have this path of continuously expanding the circle.” Some companies sit on one big launch and explode without revealing the route; “that also exists.” His own path began with user demand.

4. Product Hunt’s Real Value: Clarify the Product in One Sentence and Get a Free Cold Start

  • Vozo has spent nothing on marketing and reached $1M ARR through two Product Hunt launches. He first launched on Product Hunt in 2015 and believes its greatest value is not traffic but forcing you to answer: “What kind of product are you, and how can you explain it in one sentence?” That is the most useful part for product refinement.
  • On traffic, Product Hunt brought roughly 1,000 visitors. “Around 1,000 is enough for us to iterate on the product and PMF,” effectively providing a low-cost cold start.
  • He is frank about the leaderboard ecosystem. Getting to No. 1 “doesn’t have that much to do with the product. If you understand operations and are willing to push it… getting into the top 3 should be no problem.” Ranking first does not mean the product succeeded. “Whether the launch succeeds or fails is an operations matter. Whether it brings real commercial value afterward is a product matter.” He is not surprised that many teams reach No. 1 but never build a sustainable product.

5. Restraint on Promotion; Use Intercom to Keep Users in the Iteration Loop

  • After launch, the team had opportunities to promote the product but deliberately held back. Before user feedback had been addressed, more promotion “didn’t have that much meaning.” The right move was opening Intercom early and having several core team members talk continuously with users on the website to understand what they wanted, what they did not want, and what dissatisfied them. The team shipped 1 to 2 versions every week.
  • His conclusion comes with a hedge: “This may not necessarily be right. Some teams might grow faster if they promote earlier. But this is my view: launching promotion a month earlier or later actually isn’t that important. Getting PMF right matters more.”
  • He quantified the feel of PMF with two internal goals. In absolute terms, “reach an M first—$1M ARR—and then talk.” They hit it without promotion right on schedule. “It may have been luck, but it was roughly consistent with our original judgment.” On retention, 80 of every 100 paying users should remain; he expected 20 to churn because of their own business conditions. “Having clear goals gives you more energy to do the work.” Each phase should focus on 1 or 2 goals rather than opening 5 channels at once.

6. The Truth About Entering Late: Research and Applications Took a Wrong Turn for a Year

  • Why did the company launch only in July 2024? “I think the other reasons were more important. Our own reasons were more important.” The team had been working on AI video since 2021 and saw generative AI earlier than most companies in 2022. It even set up a joint lab with one of his former professors, a prominent scholar, to pursue foundational research. At the time, the company had almost no revenue. “It was an outrageous thing to do.”
  • The thesis looked elegant: build very practical applications on one side and very ambitious research on the other, hoping “they would converge one day.” In hindsight, “it was actually a fairly wrong idea.” By 2023, the product needed features that the foundation model could not provide, while the research team pushed forward on its own terms. The results were exciting but could not be productized, full of sampling quirks and strange behaviors. “There was no synergy, no combined force. Instead, both sides felt regret: you couldn’t help me, and I couldn’t help you.”
  • The application side ultimately won. In October 2023, the team issued a PR for its in-house multimodal model, HiveNet. “The meaning of that PR was that we were no longer pushing it forward, although the PR itself didn’t say that.” From then on, every R&D project had to start from the product, with 20% of researchers’ time theoretically reserved for free exploration.
  • Would the researchers leave? No. After more than a year of the previous experience, “a large number of researchers hoped their research could make it into the product.” As Vozo’s user base grows and feedback improves, “those researchers become very happy, and it turns into a fairly interesting loop.”

7. The Ledger: $6M–$7M Raised, Four Years, Positive Cash Flow

  • The company stopped before Series A. Funding came from Xinxing Capital, Sequoia Seed, and individual investors, totaling roughly $6M–$7M. Koji’s reaction was: “From 2021 to 2025, four years with only $6M—that is extremely high capital efficiency.”
  • The source of positive cash flow is another product line: a domestic teleprompter app, known overseas as Blink, built on the previous generation of CV/NLP technology and generating roughly $6M ARR. “That product basically guarantees that our cash flow is break-even, so everything Vozo earns now is our profit.” Reaching break-even has also been good for the team’s mindset.
  • The team has more than 40 people, with over 70% in R&D. “We’re doing very heavy research.”

8. Not an Application Factory: The Two Products Are Becoming One Vozo

  • Asked whether the company operated as an application factory, he replied, “Good question. We actually hadn’t figured it out at the beginning.” Everything centered on freedom of expression through video. The team first built an app to capture demand, discovered that traditional CV capabilities were limiting, and then moved into generative AI. For a long time, the two products ran in parallel. “That was also a very painful point for our team.”
  • The team has now found a way to merge them. The app leans consumer-facing, serving KOLs, KOCs, and a small number of SMBs; Vozo targets enterprise marketing departments and a small number of SMBs. Their users overlap by roughly 20%–30%. After the merger, the products will cross-promote, share functionality, and use one membership system. The unified product will be called Vozo and serve content creators, marketing managers, and e-commerce businesses that use video to tell stories.
  • The name was generated by GPT. “It really impressed me.” The requirements were that it be short and related to video or voice. The vision was that “one day everyone will have their own Zoom. You’ll tell many stories every day, just like writing a blog.” From Vozo to AI, there are only 6 letters, and it is easy to say.

9. Sora Has Little to Do with Vozo: General-Purpose Models Will Never Be Moats

  • The key context is that the team had built its own visual foundation model, released, in his recollection, after Runway’s first generation and before V2. “You only know where the bottlenecks in visual foundation models for video generation are after building one yourself.” Those bottlenecks included controllability, consistency, and compute cost. If a minute of video cost $200–$300, ordinary content creators could not accept it. The experience also gave him a sense of when those bottlenecks might be overcome.
  • When Sora appeared, it was “several streets ahead of everything else.” But he expected Google’s model to arrive in another 3 to 5 months because everyone would push in that direction. “In the end, it happened as expected.” Today, many Chinese companies can do it. His core judgment follows: “Whether it’s a large language model, an audio model, or a visual multimodal model, if it’s general-purpose, it will never become a moat,” because of open source and other ways to replicate the capability.
  • The strategic conclusion is: “Our entrepreneurship should stay as far away from it as possible.” Every model Vozo builds has application-specific requirements. Translation, for example, has different demands around tone preservation, so the team builds voice-cloning, speech, and lip-sync models specifically for translation. “We iterate around the real demand inside this hammer,” while using external foundation models whenever possible. He also recognizes his own limits: “I’m definitely not very good at fundraising. I probably couldn’t do the foundation-model, burn-money game.”

10. The Devil in Translation: Chinese-to-German Is an Optimization Problem

  • The concrete problem is a major length mismatch between Chinese and German. “You may speak Chinese for 5 seconds, while German takes 15 seconds.” With the visuals unchanged, the two tracks fall out of sync. “You can’t keep your mouth closed for 15 seconds.” The solution is to turn it into an optimization problem: find wording with the right length, stay close to the original tone and intonation, and keep the lip movements plausible.
  • Missing context is machine translation’s fatal weakness. Without background, a brand term can be translated directly and incorrectly. Emotional replication requires “one sentence corresponding to one sentence,” but “translation cannot be done one sentence for one sentence, or the translation will be bad.” It must incorporate context, preserve the corresponding structure, and copy the emotion. That is why the industry believes machine translation is not enough; quality-conscious companies hire teams at $50 or $100 per minute. His conditional forecast is that “in another 1 or 2 years, machine translation may be a little better than human experts.” But “if the translator is truly an expert, they will still do a better job.”
  • The product solution to the trust problem is clever. If Chinese is translated into Arabic, the user may have no idea whether it is correct. With a human translator, the user can sign a contract and hold someone accountable. “As a SaaS product, you can’t come after me later. So what do we do?” Vozo built back translation: translate it over, then translate it back and compare. “If it’s roughly the same as the original meaning, it must be right.” Koji compared it to the blind telephone game on a variety show. The team has recently started tackling harder short-form translation: how to preserve exaggerated expressions and the emotion of slapping the table. “We’re slowly taking on some harder problems.”

11. Three Layers of Solutions and the Technical Base: Product > Engineering > Model

  • His priority order is a methodology. Product solutions come first—a pop-up telling the user to click once “is actually the best solution.” Engineering comes next: write algorithms to align timing and optimize length. Model iteration comes last, such as rapidly reproducing tone from a single sentence. “When you discover a problem, which one do you use to solve it? Which are workarounds for now, and which are things that must be built in the future?”
  • Technical breakthroughs over the past 1 to 2 years have lifted Vozo’s ceiling. Lip generation moved from GAN-based approaches 4 or 5 years ago, which had low clarity and realism, to this wave of Transformer-based approaches, and recently to what he calls “Gaussian Splatting” (possibly Gaussian Splatting). Generation is faster and quality is better. “Fast one-sentence voice cloning, highly realistic lip and facial movements, and full-frame generation are all things that have gradually emerged over the past 1.5 to 2 years. Some may have emerged only in the past 6 months.” The strategy is to ride the industry’s momentum, “standing on the gold board.” He says Vozo’s lip sync is probably among the best in the industry, supported by accumulated data and close tracking of the latest technology.

12. Will Models Swallow Application-Layer Craftsmanship? “Definitely”—So Don’t Touch Standard Models

  • Koji raised the industry anxiety: Google had just released Veo 2, rendered as “Vivo Two” in the ASR, and the reviews were strong. Could model evolution swallow the craftsmanship in the feature layer? He did not dodge the question. “I think it definitely will. It’s a big vehicle.”
  • The countermeasure is explicit: “If it’s a standard model, don’t touch it. We want to build things that are close to the application and different.” Midjourney is the example. The broad text-to-image framework is “more or less the same,” but Midjourney has done a great deal of detailed technical tuning. “That tuning actually creates a very large difference.” The same applies to video. “Maybe in the future there will be a more usable visual model like DeepSeek, but when you apply it to your application, the difference will be enormous.”
  • His reassurance to application-layer technologists is that even in an era of the same underlying technology, one team can do it far better than another. “Don’t worry too much that one model will do everything, leaving no technical space at the edges. That’s impossible.”

13. An Idea That Tempted Him but Never Got Built: Google Glass Plus a Low-Latency LLM

  • Asked about opportunities he had seen but not pursued, he first qualified the answer: “I don’t dare say. You only know after actually doing the research.” But his personal interest was obvious. A “Google Glass” device—referred to in the source as “Google Introduction,” possibly Google Glass—combined with a low-latency LLM “would be very interesting, with a huge amount of room for imagination.” Many things he wanted to build for it in the past but could not are now possible.
  • The idea he still thinks about most is “Google Glass making you smarter.” That “may also be what Google Brain wanted to do”: someone asks you a question you cannot answer, and it tells you almost immediately—“so fast that I think I came up with the answer myself.” “For someone like me, I’d be willing to pay for it.” He immediately corrected himself: “This may be another return to my earlier mistake, because I’m extremely excited about it. To actually build it, you still need to do the commercial analysis.”

14. Google X in Retrospect: Exploring for Sergey Brin, Poaching A+ Talent Until Larry Got Angry

  • Before finishing his PhD at Columbia in 2011, a Stanford professor recruited him to Google X to form a new group. It started with 3 people and grew to 12. “Among our 12 people, we had probably won 4 Grammy Awards,” as he put it. The group brought in many of the strongest people in computer vision for photography. It “was essentially set up to satisfy many of Sergey Brin’s exploratory needs.”
  • The output was substantial. The core imaging and video-processing algorithm stack for what he calls “Google Glass,” possibly Google Glass, came from the team. The work is now present in image processing on “basically all Android phones.”
  • The level of extravagance was striking: “Anything under $10,000, I could just buy directly.” In recruiting, the group first poached the A+ people from other Google teams—the highest performers and smartest people—and “they generally came.” That continued until Larry Page, who ran the formal businesses, became angry: “You can’t keep poaching people from other departments at Google.” The group then turned to recruiting the strongest people in the industry from outside. “This wasn’t very good for the business departments. They were making money, and we were spending it.”
  • The reason he left was the downside of extreme freedom. He led 6 or 7 people on a project; after the demo, everyone said, “Wow, that’s so cool,” and then nothing happened. The project was hung on a wall as an exhibit. “It wasn’t much different from when I was doing my PhD. When freedom goes to the extreme, you can’t productize anything or create impact.”

15. The First Startup’s Lesson: Wishful Thinking and “The Last Missing Link”

  • His first startup, founded after leaving Google in 2015 in the United States, built immersive video around the idea of teleportation: allowing 2 people to meet face-to-face anytime, anywhere, with the highest-resolution video rendering available at the time. In hindsight, it was “still extremely research-driven.” “I thought I had identified user demand, but I hadn’t.” The market looked large in theory, but careful commercial analysis showed that the business scenario did not work. “It’s not enough for an idea to make logical sense and have a large market for you to build it.”
  • That experience produced what he calls “a very conservative business choice”: “I want to build the last missing link in the industry chain.” If a business requires 5 links and you build the first one while hoping others build the remaining 4, “that is actually extremely difficult.”
  • He applies the same framework to Vision Pro. The experience is excellent and spectacular, but the form factor, price, must-have rationale, content ecosystem, and supply chain are all missing pieces. “Even a company as large as Apple would struggle to connect such a massive chain. For a startup, you should stay as far away from it as possible.” One employee from his first company later went on to work on Vision Pro-related projects.

16. From “Begging People to Use It” to “Being Asked for It”:杭州 MCN Interviews Changed His Commercial View

  • In the VR era, the company served major customers such as AT&T, Verizon, China Mobile, and the Publicity Department. But “every time we iterated the product, we had to beg them to use it. A lot of the time they didn’t use it at all. They just paid and left it sitting there.” That was one reason the previous company never became large.
  • During the pandemic in 2021, stranded in Hangzhou, he spoke with the CEOs of more than a dozen MCNs. The contrast was stark. Every conversation on this side produced a long list of needs: “I want to make videos this way. I have this problem now.” On the other side, “I had built a lot of things and was begging them to use them. That feeling was so painful.” The conclusion became a new business principle: “I want to build something many people want, something they can use immediately after I finish it. That is what a good business experience feels like.”

17. The Livestreaming Camera Was Killed; the Teleprompter Became an 8-Million-User Cash Cow

  • The first attempt was a livestreaming camera. MCNs wanted buildings with hundreds of livestreaming rooms, but high-end livestreaming required multiple cameras and a director shouting through a headset to switch between them. The team built a head-sized livestreaming device that used AI to understand the scene and switch cameras automatically. “It was still the instinctive response of researchers: use AI to replace the director.” But it was “still not down-to-earth enough,” with many commercial problems, and was killed after 6 months. The company raised financing on the livestreaming-camera idea. One famous head of a dollar fund even asked him directly: “Why are you doing this? Can you do anything else?”
  • The team then built an AI teleprompter. It floats near the phone camera, “a bit like karaoke, but the subtitles follow your voice. When you stop, it stops; when you speak faster, it scrolls faster.” Unexpectedly, users not only needed it but were willing to pay, with a fairly high conversion rate. Launched in 2022, it has accumulated roughly 8 million users and nearly 100,000 users in private groups. “The domestic market is truly enormous. There are a huge number of people who need down-to-earth products.”

18. The Elite’s Hang-Up: A “Low” Teleprompter and a Self-Indulgent Lab

  • The psychological gap was real: “I thought the first feature I built was so low, even though everyone wanted it.” When speaking with former professors and classmates, he generally did not tell them what he was working on. But he also stressed that making a teleprompter work well was not easy. Noise, strong accents, and erratic speaking rhythms required “a lot of very dirty work.”
  • That emotion “needed somewhere to be released,” which directly led to the joint lab. Some people have poor on-camera presence, an unpleasant voice, or weak delivery. “Even if you give them the best teleprompter and write the entire script for them, they still can’t shoot the video.” Those core problems required research. He calls it “a little self-indulgent, but somehow lucky.” It was not necessarily the lab that made the breakthrough; the entire industry made major advances in 2022–2023, and the lab caught the wave. “It may have been a risky kind of luck.”
  • Asked whether he would build a lab again, he gave an honestly uncertain answer: “Fifty-fifty. I’m not that certain.” If he did not build a lab, he would definitely do something else that was fairly crazy. “If I were only doing something down-to-earth that could make money, I probably wouldn’t accept it.” His current compromise is that Vozo makes him feel at least a little proud and gives him an answer to himself. “If all I had left was the teleprompter, I wouldn’t be able to explain myself to myself.”
  • Fieldwork also produced a useful shock. A user complained that the teleprompter was unusable, insisting that the room was “extremely quiet and the lighting was excellent.” When the team visited, the lighting was extremely dim and traffic beside the room was particularly loud. “He wasn’t lying. That was genuinely how he saw it. What we call bright isn’t what he calls bright.”

19. Three Hurdles for Technical Founders: Wishful Thinking, Market Ignorance, and Ego

  • Asked how to approach problems from a commercial rather than technical perspective, he offered 3 points, while acknowledging that they were “not particularly systematic.” First, abandon wishful thinking: “Whether it will happen, you can actually find out by asking.” Second, recognize missing knowledge: the innovations you care deeply about may matter to only 1% of users. Third, remove ego. He left Koji with an assignment: “Maybe you can find a way to summarize this. It could help a lot of entrepreneurs.”
  • Ego is removed through “passive lessons.” “You don’t think you’re wrong. After being wrong several times, you know.” His behavior changed accordingly. Technical proposals he puts forward “may be rejected by the kids. They may not necessarily be right, but rejected is rejected. As long as it isn’t something critical, I let it go.” The calculation is probabilistic: “My solution might produce 70 points; theirs might produce 65. But because it’s their solution, they’ll execute it better, and the result might be better than mine.” He was not like this before. “I used to think I was the smartest. A performance drop from 99% to 98.9% was unacceptable.”
  • His latest understanding of entrepreneurship is subtraction. “Every day it’s this and that, so busy. My energy was actually very scattered, and I don’t think I made several important company decisions correctly. Gradually I realized there aren’t many important things.” Perhaps the difference between truly great founders and ordinary people like him is that the great ones can instantly see which thing matters more and which thing does not need to be done.

20. Top Performers Are Extremely Focused; Ordinary People Score Everything Too Evenly

  • He has worked closely with several high-profile people: his mentor at MSRA was a foreign member of the U.S. National Academy of Engineering, his Columbia adviser was one of the strongest professors in computational imaging, and he interacted with the Google Brain team at Google. He found one common trait: “They are extremely, extremely focused.” His PhD adviser is nearly 70 and still won 2 best papers this year. “He actually thinks about relatively few things. What is the most important problem in this field? What is the most important subproblem within that problem? Once he solves the most important thing, resources and people naturally gather around him. Sometimes you feel he has it pretty easy.” People who are not at the top “don’t have the luxury of doing only important things.”
  • His anti-mediocrity scoring method is simple: “When you think of 3 things, your instinct is to feel that all of them are fairly important. If you rate one at 80, one at 60, and one at 40, the odds are they are actually 90, 20, and 10. People are always too moderate. You will definitely underestimate the gap in importance between the middle items.” Deliberately widen the variance.
  • His own filter is: “What happens if I don’t do this?” Not merely whether failing to do it feels uncomfortable. Will revenue actually decline? Will users really leave? Will 2 users leave, or 20%? Once you roughly calculate it, many things turn out not to matter.

21. How He Learns: Learn from GPT, Then Find the Best Person You Can Reach

  • His current learning method is disarmingly direct: “Mainly by learning from GPT. I’m a hardcore ChatGPT fan. They must have lost a lot of money because I use it every day.” When o1 first launched, he exhausted his quota every few days and had to wait until the next day. “GPT is actually already smarter than people. Just learn from it.” The other rule is to find the best person you can reach in the relevant area and talk to them first.
  • His advice to younger people lowers the bar: “As long as you find the best person you can find around you, 80% of the work is already done.” You do not need the strongest person in the field. His own chain of evidence began when, as a graduate student at Fudan, he wanted to do computational research and cold-emailed a researcher at Microsoft Research Asia, “a very important benefactor later.” After a phone interview, he went to Beijing, was referred to the head of MSRA, and then was pushed toward a PhD at Columbia. After one academic talk, an “old man” in the audience asked a question; that man later became his boss at Google X and called to invite him to join. “You just need to focus on people you can reach around you. The network is very small.”
  • Ronghui added a footnote: another guest once said, “Treat every conversation as an interview.” 周昌印 extended the point from a manager’s perspective: “Colleagues in China are clearly less aware of this than colleagues in the United States. I think it’s a key skill. Maybe universities or high schools should teach it.” His broader life pattern is similar: business school as an undergraduate, Microsoft after graduation, then resigning to pursue graduate study and a PhD. “Sunk costs aren’t very important to me. I’m a probabilist.”

22. Chinese Version Coming Soon: Users from Overseas Short Dramas and E-Commerce Are Already at the Door

  • The Chinese version has been planned for a long time. Internally, the team debated whether and when to support the domestic market. But Chinese users are already numerous, especially in overseas short dramas and cross-border e-commerce. They use the product while complaining that there is no Chinese version and that Alipay and WeChat Pay are unavailable. “Some companies say they want to kick the Chinese market out. Our team has never thought that way. It’s only a question of ranking China—do we do Japan first or France?” The decision now is: “Support the domestic market first, no matter what.” PMF iteration is nearly complete and the team is moving into growth. It may soon open customized support for the domestic market and recruit growth and business-development staff, as well as AI video product, R&D, and engineering talent in China. “We can create roles around the people.”