Pioneers Insight Method Research Author
136: Sora’s New World & Lovart’s 4-Month Review | Talking with 陈冕 About Building a Vertical Agent | Agent #5
Back to Episodes

136: Sora’s New World & Lovart’s 4-Month Review | Talking with 陈冕 About Building a Vertical Agent | Agent #5

Summary

  • 陈冕’s core read on the Sora App is not “AI TikTok,” but an AI social product powered by Cameo. A nine-second video can now deliver shot language, Chinese voice, character consistency and audio-visual sync in one pass, but the real product flywheel comes from “shooting with friends and celebrities” and Remix via left-right swipes. His Aha moment was: “Isn’t this social? Wow, this is Instagram.”

  • Sora may be opening a virtual-social opportunity serving billions of users, while compressing the competitive window to three to six months. 陈冕 believes that if there is even a 10%–20% chance it becomes a Super App, model leaders and companies that already own Super Apps cannot afford to sit out. But even a mid-sized startup that gets there first would face facial-privacy, regulatory and enormous generation-cost pressures, followed by rapid replication from the giants; opportunity and capital are more likely to concentrate at the top.

  • Two months after its public launch, Lovart’s DAU had reached 150K–200K. During Beta, daily activity was roughly 10K–20K under invite-only access, although the original source does not specify the exact metric; after Beta, DAU first rose to 80K–100K, with Nano Banana driving another step-up. The US accounts for more than one-third of users, with a slightly higher but not dramatically higher share of revenue; on a full-year estimated-revenue basis, revenue has exceeded $30M. Usage costs are currently broadly covered, but the higher free-user mix and conversion lag mean the business will still lose money overall in its early phase.

  • Lovart’s vertical moat is not exclusive access to a foundation model, but a proprietary Interface, Context and industry experience. Chat Canvas simulates a client and designer pointing at visual work together around a table, while the product plans to accumulate enterprise historical assets, Reference and Preference; when every application can call similar models, “whoever collects more Context will deliver the better experience.”

  • 陈冕 summarizes AI application growth as drawing the product that should exist after models mature, then waiting for the technology wave to catch up. Once models such as GPT Image 1 and Nano Banana delivered on complex instructions, editing and consistency, the pre-built Chat Canvas could amplify those capabilities immediately; the next phase is to grow the early Agent user base, deepen Context Engineering, and keep making the trade-off between waiting for models and filling today’s gaps with engineering.

  • Productivity applications have entered a knife-fight, and 陈冕 believes new entrants must “already be at the table this year,” while the next major opportunity is To C. Veo 3 once cut prices by 70%, Nano Banana is cheaper than GPT Image 1, and GPT-5 is cheaper than Claude; as creation costs continue to fall, the variable will shift from “who produces content” to “what new content people consume.” But startup To C Super Apps are more likely to grow gradually from an overlooked corner than explode at launch.

  • Speed is both Lovart’s organizational advantage and its most visible operating risk. 陈冕 believes an AI product’s effective window may last only two years, so the team prioritizes people who understand AI, learn at high frequency and accept constant adjustment. The pace is grounded not in blind optimism, but in two still-testable beliefs—that AI will continue to improve rapidly and that AI will replace fictional-content creation—along with a contradictory warning to himself: “I’m a little afraid I’m too anxious, and I’m also afraid I’m not anxious enough.”

Deep dive

1. Sora’s first shock came from a frictionless nine-second creation loop

  • When 陈冕 received an invite code at 3 a.m., he had no great expectations: the launch had been abrupt and the online introduction was bare-bones. But the product Flow was smooth enough that he made his first video—“Sam and I walking on the street”—with almost no learning cost. It generated quickly, the camera work was strong, and the result made him think, “This actually has something.”

  • The second video showed the two of them dueling in a bamboo forest. For the third, he asked GPT to write a nine-second rooftop-duel script from Infernal Affairs, then pasted it into Sora. Only on the third attempt did he turn on sound. The Chinese adaptation, a voice close to his own, and audio-visual sync came together, making him feel that “the door to a new world had opened.”

  • What truly Shocked him was the combination: continuous camera language, the potential for cinematic expression, reasonably strong character consistency, and the fact that “you don’t need to roll the dice—it comes out in one shot.” The model did not meet professional-production standards on every dimension, but for the first time it compressed all of these capabilities into a smooth workflow that an ordinary user could complete.

2. Cameo turned a video generator into a playground for social relationships

  • 陈冕 quickly realized that, aside from Sam, he did not know anyone appearing on the creation page. If he was neither shooting with celebrities nor seeing friends inside the product, he did not know what to create. He immediately urged his US colleagues to exchange invite codes, transforming a solo model experience into a co-creation game among acquaintances.

  • When a colleague used his face to generate “I’m leaving work early today—I’m not working today,” the entertainment value came from the relationship for the first time, rather than from the video itself. 陈冕 shifted from a user’s perspective to a product manager’s and reached the core Aha moment: “Isn’t this social? Wow, this is Instagram.”

  • His conclusion was therefore decisive: at least for now, Sora is an “AI social product,” not an AI short-video product in the conventional sense. The model is the foundation, but without leveraging social relationships, a pure AI Feed is not enough to create the same level of noise; other products have already tried it.

3. Remix productized meme propagation as low-friction relay co-creation

  • 陈冕 had already been thinking about the core mechanism of AI creation: once the barrier falls, users should be able to see someone else’s work and immediately “continue the co-creation.” By placing Remix alongside traditional actions such as likes and comments, Sora showed that the team was not only building generation, but also designing how works would continue to spread and mutate.

  • 程曼祺 mapped this to Bilibili memes such as “Look in my eyes” and MVP, which users repeatedly apply, as well as Jianying’s “use the same edit.” These were already forms of Remix, but previously required manual recreation and passed from person to person; AI makes replacing characters, relationships and situations a native action.

  • Swipe-based generation initially produced weak results, and 陈冕 did not even fully understand the interaction. The output kept improving, but what he valued was the interaction choice: “The best interaction will not go beyond swiping and clicking.” At the same time, long Prompts could trigger Coding errors and other Bugs, showing that the first version prioritized polishing the critical path.

4. Meta Vibes’ cold start highlighted the difference made by real people on camera

  • Meta launched Vibes on September 25, four days before OpenAI released Sora on September 29, yet almost nobody discussed it. The program’s explanation was that Vibes was buried inside the Meta AI App, had no standalone App, and lacked a social core built around shooting with real people through Cameo.

  • 程曼祺’s own use also exposed the boundary: even when a specific line was written in the Prompt, Sora sometimes still failed to say it. 陈冕 acknowledged that the model had room to improve, but the more important point was that this product loop was never designed around professional creation and precise control. It was optimized for fast, playful co-creation.

  • If the goal were “AI TikTok,” the first priority would be to supply a large volume of high-quality content as early as possible. Sora’s current appeal instead comes from friends, celebrities and relationship networks. 陈冕’s summary was that it “built the form best suited to the moment when model capabilities became suitable,” and that the social market may be larger than short video alone.

5. Models, organization and capital must arrive together, making Sora difficult for startups to replicate

  • Asked whether a mid-sized company could catch fire by building the same interaction first, 陈冕 rejected the premise. This was not merely a product innovation; Sora’s model was also SOTA at that moment. A company that did not build foundation models would struggle to be first. The likely first mover would need both a top-tier model and an organization capable of producing consumer-product innovation.

  • Even if a mid-sized model company broke through, the follow-on challenges would be severe: recording real faces creates privacy, compliance and regulatory pressure; generating huge volumes of free nine-second videos requires “burning an enormous amount of money”; and once demand is proven, companies with more traffic, capital and compute will move immediately.

  • AI social therefore looks more like a battleground for the giants than a startup window defensible through one clever interaction. Startups can still enter adjacent To C needs, but directly challenging the same network effects quickly turns model and capital scale into the barrier.

6. Sora’s long-term vision is a “virtual-world WeChat or Instagram”

  • OpenAI offering Sora for free does not mean it lacks commercial intent. 陈冕 agreed with Sam’s answer to why the company would build a “time-killing product”: if a good product can also make money, that helps realize AGI. A Super App does not need to make money early; once it has a network, it can generate substantial revenue.

  • He described future social life as two parallel worlds: a real-world social circle and a virtual one, expressing “the false real” and “the real false.” Both contain elements of reality and fiction, and their boundary will gradually dissolve. The ultimate contest is for user time, and for which identities and relationship expressions prove more attractive.

  • If the vision is fully realized, this could become a Super App with “billions of users.” But he repeatedly added a constraint: this is not a description of the current product. The first version’s model, product, operations and positioning all remain exploratory. What is nearly certain is that the model will continue to improve.

7. Giants have three to six months, while startups should not worship overnight virality

  • 陈冕 believes that whether a giant assigns Sora a 10% or 20% chance of becoming a Super App, it cannot afford to miss. The cost of losing a top-tier entry point is far greater than the cost of entering the fight. Companies with existing Super Apps and leading model companies will all participate.

  • Social-network effects mean the follow-up window may be only “three to six months.” The first batch of products still has a chance; once friend relationships and the content ecosystem have accumulated, later entrants will struggle to move users simply by copying the interface. That is why he expects the competition to escalate quickly.

  • Startups will follow a different path. This launch will create more To C opportunities, but genuine startup Super Apps often grow slowly from an unremarkable corner. 陈冕 warned that “if you blow up in To C, be careful”: arriving too fast can also mean disappearing too fast. Douyin, Xiaohongshu, Kuaishou and Bilibili all accumulated for a long time.

8. Sora forced 陈冕 to revise his view that giants would not move so quickly into consumer applications

  • He first clarified his earlier position: video had never left the AGI route. Pursuing a world model necessarily requires video capabilities. But the main path to AGI remains language; video and other multimodal capabilities will gradually be incorporated rather than determine the progress of general intelligence first.

  • The actual mistake was underestimating the speed. 陈冕 had thought giants would not launch a To C Super App this quickly. He now admits, “Clearly, I was wrong about this.” OpenAI was more aggressive than expected: it was not only pursuing AGI, but also trying to build the strongest consumer products.

  • This cuts both ways. The giants’ entry validates that the field will become more prosperous and that attention and opportunity will increase. But the giants’ boundaries are broader than expected, so the challenge for startups will expand sharply unless they scale quickly. His requirement for Lovart is therefore not retreat, but to “get big faster.”

9. Capital is moving faster than physical infrastructure, creating a potential short-term bubble

  • 陈冕 considers the fact that “in human history, no company has ever rapidly grown into a company worth more than $500B within a few years of becoming wildly popular” an extreme example of how quickly capital can concentrate. Future space is being priced with excessive optimism, while global capital is sufficiently abundant to concentrate consensus and ammunition at unprecedented speed.

  • Compute, energy and infrastructure cannot accelerate at the same pace. There is not enough compute or energy, and construction takes real time. In the short term, Super AI could cost so much that “the money cannot be earned back and the math does not work.” It may make sense several years later, but the gap between capital expectations and real-world delivery could still produce a bubble and a collapse.

  • User habits and “the human heart” are another slow variable. Pioneers, skeptics and traditionalists will coexist for a long time, and the boundary between reality and virtuality cannot be accelerated to the limit. 陈冕 still considers the broad trend “beyond doubt,” but he does not describe adoption as a straight line.

10. 陈冕 uses a handful of “Wow Moments” to judge nonlinear consumer potential

  • On Sora’s first day, the most-liked Feed posts had only a dozen or so likes, and the content was not good. From launch day to now, both creativity and variety have improved materially. 陈冕 compared this with Kuaishou’s early period, when most people also found the product difficult to accept, arguing that a creation network’s terminal state cannot be judged from the first day’s “AI junk.”

  • He compresses his consumer-product intuition into a very small number of moments: “The last one was ChatGPT, and before that it was Douyin.” A strong To C product shows enormous magic on first use. Sora gave him a Wow Moment of the same class, so his view of its long-term potential is close to a “no brainer.”

11. Three months in North America shifted Lovart’s mission from a tool category to creative access for everyone

  • After Lovart’s May testing, 陈冕 stayed in the US through August, roughly three months in total. People in Silicon Valley were less likely to begin by asking about competitors and the endgame. They asked why the founder had to build the company, what the Vision was, and where its distinctiveness lay. Repeatedly answering those questions forced him to think more clearly about the product’s target user.

  • Lovart ultimately positioned itself as a “design agent for everyone who wants to create.” It is not an entertainment product that everyone will open every day. Instead, whenever someone genuinely wants to create, it aims to be the professional AI Design Agency they can call on, turning imagination and ideas into finished work.

  • This also defines the difference from general-purpose products such as Doubao. Lovart may eventually serve a very broad audience, but the use case itself remains Professional: it should deliver better work like a professional studio, rather than equating “available to everyone” with lower quality standards.

12. AI may reduce narrow designers while increasing broad-based creators

  • 陈冕 does not avoid the substitution question. Designers who only Copy and execute repetitive deliverables are more exposed, and the number of narrowly defined design jobs may decline. But design capability will become broadly accessible: people who previously lacked tools and budgets will be able to create, so the broader population of designers and creators may actually increase.

  • 程曼祺 used weddings as an example. At certain life milestones, people care deeply about deciding the visual expression themselves, and hiring a designer may not be better than articulating their own ideas. 陈冕 added that after Nano Banana became popular, people began uploading photos to generate complete sets of artistic portraits, showing that demand for professional-looking personal creation is not limited to professionals.

  • For non-professional users, the value of an AI studio is delivering usable results at lower communication and labor cost. For professional designers, it is more like an entry-level team that first explores assets, inspiration and drafts, after which the designer uses personal judgment to raise the ceiling.

  • His boundary is clear: “There will always be humans who can raise the ceiling of human content creation.” Lovart is therefore neither a service only for ordinary users nor a demand that professionals accept machine-made final work. It gives different levels of ability different leverage inside the same studio.

13. AI entrepreneurship in San Francisco is an “excitement-driven arms race”

  • Lovart’s office is around California Street in San Francisco. 陈冕 observed that generative-AI application companies and players such as Krea and Luma are more concentrated in downtown San Francisco, while Chinese-founded companies are more concentrated in traditional Silicon Valley areas such as the South Bay and Palo Alto.

  • He initially sensed the “collective mood” through events, cafés and conversations with Founders. AI was being discussed everywhere, global entrepreneurs were densely present, and everyone was working at extreme intensity, but fewer people were trapped in zero-sum competition narratives. They believed the opportunity was large enough for everyone to do what they wanted. This was an “excitement-driven arms race,” not a purely anxiety-driven one.

  • The US To B ecosystem also makes it easier to combine external services. At a reasonable price, companies will subscribe or connect to an API without first going through a long sales cycle and price negotiation. Model Hosting and API provider fal could win ten startup customers directly at one Party; Lovart could even co-host an event with Freepik, a company in the same space.

14. North American users buy a substitute for expensive labor; professional designers buy inspiration

  • For US SMBs, designers are expensive. When a café owner sees Lovart, the first thought is how much money could eventually be saved while still getting a Logo, menu, brand assets and multiple proposals. Design services are relatively inexpensive in China, so the excitement is less intense.

  • European and US professional designers place more emphasis on self-expression and rely more heavily on familiar traditional tools such as Photoshop. They may not want to use AI output directly, but may want AI to act like a “junior designer with strong taste” providing Inspiration. Midjourney was adopted early even when it lived only on Discord and lacked a complete editing Workflow, because its aesthetic quality was strong.

  • Chinese designers are more oriented toward efficiency and delivery. They embrace new tools more readily, may be more sophisticated users, and may even use ComfyUI to connect an end-to-end workflow. 陈冕 did not force the two markets into one feature set: Europe and the US emphasize inspiration and exploration, while Chinese users may prefer to keep more production steps inside the AI product.

15. Globalization starts with local narratives, content operations and cultural taste

  • In North America, Lovart prioritized building a Marketing team rather than a large sales or performance-marketing operation. 陈冕 believes mainstream European and US products rely less on Pay Ads and more on Branding, complete narratives and the spread of strong work. For a creative tool, the most convincing marketing content is the good work the product actually produces.

  • Design must also be localized. How an English font should be arranged depends on the cultural environment in which users have lived for years. Because of childhood computer classes, Chinese users may find Huawencaiyun “tacky,” while foreign designers may not understand the association. This kind of judgment cannot be reasoned out remotely by a Chinese team alone.

  • At the time, Lovart was not running large-scale pre-roll advertising. It used KOLs and brand communications to acquire users with relatively high quality and strong willingness to pay. 陈冕 planned to test Pay Ads only after the brand and category narrative became clearer; buying traffic too early could bring in large volumes of low-conversion users.

  • On fundraising, he insists that “fundraising is a means, not the goal.” When Lovart began scaling, the next round was already essentially set, and the team was not short of money during the US trip, so meeting US investors intensively was not a priority. Long term, it may diversify its sources of capital, but at that point the better use of time was strengthening the team and getting closer to users.

16. Lovart’s metrics kept stepping up, but 陈冕 still viewed it as a test-stage product

  • By the end of September, US users accounted for more than one-third of Lovart’s user base. The US share of subscription revenue was slightly higher than its user share, but not by much. Because growth was driven mainly by Organic distribution, the users were relatively well targeted and were not diluted by the low-paying traffic commonly produced by large-scale paid acquisition.

  • During the May-to-July Beta, invite restrictions kept daily activity at roughly 10K–20K, although the original source does not specify whether that figure referred to DAU or another metric. After Beta ended, DAU quickly rose to 80K–100K. Following the public launch on July 28, Nano Banana improved the Agent’s final output, pushing usage up another step; recently it has stabilized at 150K–200K DAU.

  • 陈冕 disclosed that revenue had exceeded $30M on a full-year estimated-revenue basis. The product generated almost no revenue during the testing period; the formal release and Nano Banana each produced a step-up, with the revenue curve rising in staircase fashion alongside DAU. The host subsequently referred to the number as ARR, but 陈冕’s original wording was “estimated revenue for the full year.”

  • But he gave the product a subjective score of “not yet 60 points”: the model, engineering and product all had large amounts of unfinished work. Users were staying not because the current output was perfect, but because they could see the future and, amid rapid iteration, could see the product solving more of their work problems.

17. Current usage costs are broadly covered, but early free users still create losses

  • Asked whether revenue could cover API usage costs, 陈冕 answered, “Right now, basically yes.” But because the free-user share is higher and each user receives a generation allowance, conversion from free usage to subscription takes time, so the business will definitely lose money overall in its early phase.

  • He rejects the idea that Agents will lack a business model in the long run: “Token prices will definitely get cheaper. That is a no brainer.” Subscription is already the basic business model; the current issue is simply that the unit economics have not fully balanced. He compares Tokens to traffic and electricity: broader adoption itself will continue pushing costs down.

  • Longer term, pricing can be tiered by delivery depth. Asking an Agent to think for 15 minutes, two hours or a day involves completely different levels of Context collection, Thinking Process and tool complexity, just as a designer rushing through a one-day job should not charge the same as one conducting a month of Research.

18. 陈冕 still divides the AI productivity opportunity into two major categories: “Office” and “Adobe”

  • When he started a company in 2023, the directions he saw included Office and Adobe on the production side, and information acquisition, search, social and broad entertainment on the consumer side. Looking back today, large productivity products still largely fall into text and information processing or image and video creation; the framework has not fundamentally changed.

  • Coding is this generation’s new Office. Code and text are both native LLM capabilities, making the productivity applications built on them the most direct, but also the most likely to overlap head-on with general-purpose model companies. General-purpose vendors remain strongest in text, information and Coding, and will absorb a large share of general use cases.

  • Adobe-like applications can benefit from multimodal-model progress while sitting farther from the language axis. 陈冕’s analogy is that foundation models are “creating a highly intelligent person,” while vertical companies are training a designer. A PhD graduate may have powerful general knowledge, but still needs industry experience, working methods and data to produce professional design.

19. Chat Canvas is modeled on clients and designers communicating around a table

  • When explaining Chat Canvas, 陈冕 did not start with a feature list. He started with real human communication. A normal interview requires only looking at the other person’s face, which corresponds to a Chatbot. Communicating with a designer or director requires both sides to face a screen, a desk and a visual artifact, pointing to specific locations and saying, “Change this here, change that there.”

  • A Chatbot may be the generally optimal interface, but it is not the complete interface for every vertical. Creation requires visual Alignment first. Canvas is the shared working environment where the client, Agent and output align; it is not an extra canvas added to make the product look professional.

  • In his view, vertical applications solve two problems: how this “person” works after entering an industry, which determines the Interface; and which data, Context and experience it should accumulate, which determines how it thinks and delivers. Foundation intelligence supplies general knowledge; vertical companies supply professionalization.

20. Application growth comes from anticipating models, not waiting for them to mature

  • When Lovart began designing, GPT Image 1 had not yet launched. The team did not know whether the model could understand the long instructions generated by an Agent, maintain complex semantic consistency, or “change exactly what you point to.” They nevertheless inferred from the image-model team’s R&D direction that these problems were being actively solved.

  • GPT Image 1 first performed relatively well on language instructions, consistency and instruction following. Nano Banana and ByteDance’s Seedream 4.0 then moved the bar up another step. Once the model was ready, the Chat Canvas built in advance could immediately become a complete experience.

  • 陈冕 summarizes the method as: “Draw in advance what will happen in the future, then wait for it to happen.” Application companies do not control the underlying innovation, so they must anticipate how model evolution will overturn the Interface and, when the model is Ready, “show it off aggressively,” driving growth through innovation rather than simply through paid acquisition.

  • ControlNet, which was popular in the Stable Diffusion era, offers a contrast. In the past, specialized control structures were needed to compensate for weak understanding. Now, as one-shot image-text understanding improves, many old interactions disappear naturally. Being close to model companies does not mean receiving secrets; it means learning earlier what the next technical problem will be.

21. Lovart treats “AI becoming more like a human” as a bet that could also destroy the product

  • 陈冕’s design philosophy is to reproduce human-to-human communication as closely as possible, while betting that ASI will not arrive so quickly that this relationship immediately becomes obsolete. If models improve slowly and never become human-like, a product designed around real-person interaction will struggle to reach PMF. If models continue approaching human capability, Lovart’s interface will become increasingly natural.

  • He acknowledges the boundary. Once AI surpasses humans, human-style communication may no longer be necessary, and users may not need to provide Context at all; today’s interface could be destroyed. But “how do you interact after surpassing humans?” is currently unimaginable. The product therefore has to close the open question inside a range that can still be designed.

  • The application-growth path follows from this: pre-draw the interaction paradigm, capture growth when models mature, accumulate users and data, then continue investing in Context Engineering, Prompt Engineering and engineering optimization. If open-source models keep pace, fine-tuning or RL can amplify the advantage of proprietary data.

  • He later stopped caring whether a model was open or closed source. Open source mainly affects cost and whether Post-train is possible; closed source leaves pricing power with the model company. But even without training, the engineering depth of Context and Prompt is large enough, while image and video models can still be treated uniformly as tools called by an Agent.

22. Lovart naturally extends from image to video, but is not betting on 3D yet

  • 陈冕 believes products that make images will probably enter video, because video creators usually start from storyboard images and existing image users also have substantial downstream demand. The reverse is less true: pure graphic designers may not make videos. Inside Lovart, image and video usage is already close to 50/50.

  • A café SMB can use the same Agent to make a Logo, menu and full Branding package, then generate a promotional video for social media. Lovart therefore looks more like a creative studio for producing Digital Content than a graphic tool serving one professional label.

  • 3D users and workflows are more distinctive, while most content is still consumed on flat terminals such as phones, computers, paper and screens. Spatial consumption has not truly taken off, so 3D will not be a near-term priority. Generating conversational podcast videos was merely an experiment that became popular when Veo 3 launched, not a strategic direction.

23. The next Context module will actively ask for information clients leave unsaid

  • 陈冕 used LatePost as an example. To create a series of posters for a media company, a designer must understand its history, audience, existing assets and brand tone, not just receive a single Prompt. Clients often possess large amounts of implicit Context without knowing which information they should proactively provide to a designer.

  • The team divides design Context into Reference and Preference. The former includes an enterprise’s private historical assets as well as public references in current fashion, such as Miyazaki and dopamine aesthetics. The latter consists of personal preferences formed through long-term interaction, such as consistently rejecting skeuomorphism and preferring minimalism.

  • The planned module will first use a small model to quickly understand the Prompt, then ask follow-up questions like a human designer to fill missing information. Users can drop in a website link, images or existing work. The Agent will abstract them into traits such as “professional, restrained, non-entertainment and serious,” then continue Research and design.

  • After repeated use, the system will build a user-specific asset library, Reference Pool and Preference library. In the future, users will not need to explain again who their company is or what it likes; the Agent can retrieve historical preferences first, confirm them with the user, and then generate.

24. Multimodal models are good enough, but consistency, text and the last mile remain unsolved

  • Post-training and RL for aesthetic Judgment still require designer labeling, and the North American market needs people immersed in local culture to judge what looks good and what looks tacky. The Context module, however, currently relies more on existing multimodal understanding and engineering strategy than on large amounts of specialized training.

  • 陈冕 acknowledges that model judgments of style are “definitely not as good as a human’s,” but many tasks are already good enough. When shown LatePost’s materials, the model at least will not mistake it for an entertainment publication. The team can also improve analysis and references by filtering for higher-quality search and information sources.

  • Core unmet needs include consistency, style response, non-standard dimensions and text inside images. Chinese and English have model companies working on these problems, but for other languages, “even Japanese has not been solved by everyone.” Small text also fails frequently.

  • Other edits should not be language-mediated at all. Adjusting brightness or contrast is more efficient with a direct slider than by asking AI to understand a sentence. Professional creative products still need to decide which steps should go to a generative model and which should remain in traditional tools.

25. Lovart is not building a community; its Inspiration Feed exists to improve input quality

  • 陈冕 distinguishes between “a tool being a tool and a community being a community.” Liblib began as a community where users contributed assets and then added tools, making it a special form. Lovart was a delivery studio from day one and currently has no plan to build an interactive community inside the same product.

  • The Inspiration channel lets users view other people’s designs and replay how they were created. Its goal is not likes or social relationships, but helping users understand “what they want and how to provide Context to AI.” It is more like a designer showing ten examples first and asking the client to choose the one closest to the brief.

  • If a community is developed in the future, it will be a new business launched by the organization as a separate product, rather than a community mechanism bolted onto Lovart.

26. Being close to users and being close to technology constantly creates the choice between patching and waiting

  • Product managers need to interview users regularly, analyze data, find drop-off points and observe representative users creating in practice. The company’s internal designers and employees who run content accounts also expose pain points every day. The issue is not whether requirements can be collected, but whether to patch immediately when the model cannot yet do something.

  • Text is the clearest unresolved case. An Agent can use traditional Coding methods to add text boxes, solving problems with small languages, small type and fonts. But if the underlying image is deliberately generated without text and traditional tools add it afterward, the overall aesthetic may suffer. If the next-generation model solves text, the entire patch may become obsolete.

  • 陈冕 himself does not read large volumes of Papers. His technical partner and the team read continuously and explain changes in Pre-train, Post-train, Reasoning and RL. His job is to translate technical developments into analogies about human learning and professional experience, then project them into product scenarios. The technical partner joined in early 2024 after already spending one or two years exploring large-model entrepreneurship.

27. The productivity window is closing; the next variable is To C content consumption

  • 陈冕 believes productivity tools have entered a knife-fight and are beginning to concentrate at the top, while To B products are highly homogeneous. A team founded only this year and then spending six months to a year building a product may already have missed the window. He distinguishes between “being founded now” and “already being at the table now”; only the latter qualifies to participate in the next breakout.

  • Multimodal applications include the previous generation of generative tools—Krea, Higgsfield and Freepik—which may turn toward Agents. Canva and Adobe will first address their own growth until a new form produces a sufficiently strong signal. 陈冕 guesses that traditional giants will make a major bet only after a vertical Agent reaches $100M ARR, because only then will it touch their vital organs.

  • General-purpose model companies are more likely in the short term to build general-purpose Agents and Coding products; professional design remains far from their Vision. Lovart is currently competing mainly with startups and its own speed. As it scales, competition will rise step by step to the previous generation of application giants and the general-purpose model companies.

  • To C is the next wave he sees. Veo 3 once cut prices by 70%, Nano Banana became much cheaper than GPT Image 1, and GPT-5 became cheaper than Claude. As production costs continue to fall, AI-native content will reshape consumption and social behavior. Lovart wants to participate, but at the time it “didn’t have the bandwidth,” and was only thinking ahead.

28. Timing determines whether you get to the table; a high-frequency organization determines whether you stay there

  • 陈冕 says “timing is everything,” and AI further magnifies the value of time. In mobile internet, capturing PMF once could sustain a company for 10 or 20 years. In AI, even if a product catches the window, its underlying construction logic may be rewritten by new technology two years later.

  • Being first may not matter as much as being in the first batch, but if you are not in that first batch, “you may lose everything.” His three-wave review is: in 2023, he judged that global image tools were already too late and first captured the Chinese market; the Agent wave led him back into globalization; whether he can capture the third wave depends on future To C products.

  • The organization therefore values learning speed, AI nativeness and high-frequency iteration more than simply accumulating functional experience from the previous era. Experienced hires who cling to old paradigms may just begin to contribute when the direction has already changed. The real distinction is not whether someone has experience, but how they treat it.

  • Team responsibilities and personnel changes will be frequent, and some people will complain, “It’s one way today and another way tomorrow.” 陈冕 believes the company should shift toward refinement once technology slows, but with limited resources it can only prioritize the highest-leverage, highest-ROI work. “It has been proven that people who are not anxious cannot do AI” is his interim organizational conclusion.

29. Resilience comes from falsifiable beliefs, while anxiety is an information source 陈冕 deliberately keeps open

  • Since founding companies, his three periods of intense challenge have been: finding PMF for the first product; the team, funding and business crisis from late 2023 to early 2024; and, after confirming that the previous-generation product was not the future, rapidly envisioning and launching Lovart. The sense of achievement grew with each crossing.

  • During the difficult period, several companies expressed acquisition motivations or interest, but none became a full offer because 陈冕 did not give them the opportunity to proceed. He did ask the team for its views, but believes that when a CEO believes in the direction, the CEO should ultimately transmit that conviction firmly rather than pass a signal the CEO may not believe personally.

  • His persistence has a clear premise: whether AI technology is still advancing rapidly, and whether AI will replace all fictional-content creation. As long as the answers remain Yes, a team already at the table has no reason to exit. If technology stalls and the goal depends on capabilities that are materially stronger, he also sees no reason to persist blindly.

  • 陈冕 compares entrepreneurship to a Soulslike game and extreme sports, and casts himself as Dota’s position-two player: he does not enjoy silently farming, preferring to find breakthrough points, judge when to attack or defend, and set the pace. He says he fears both being “too anxious” and “not anxious enough.” What he can control is whether he continues taking in external information. For now, these ten years are his period of being “free, fully committed, and crazily determined to make one thing happen.”