Pioneers Insight Method Research Author
Agent Neo launches: Derek and 拐子 on the ultimate AI creation tool
Back to Episodes

Agent Neo launches: Derek and 拐子 on the ultimate AI creation tool

Summary

  • Flowith has moved from early product validation into accelerated scaling: registered users have surpassed 200,000, DAU is volatile but growing rapidly, and Derek says February revenue was “several times” January’s while March was several times February’s. March ARR reached $1M, with current monthly revenue in the several-hundred-thousand-dollar range. The company is closing two consecutive funding rounds totaling more than $10M; the team has grown from 3 people a year ago to 15–16, but still plans to stay below 30 and remain a small-team operation.

  • Agent Neo’s core bet is not “building another Agent that can plan,” but pushing task steps, context and final deliverables toward infinity while executing continuously in the cloud. It can run for a week or a month, periodically tracking OpenAI, Google, Anthropic and their relevant executives, then continuously updating emails, documents or webpages. Internally, the team calls this “can stop”—meaning, in its own explanation, “it can’t stop.”

  • Flowith believes the main bottleneck for large models is no longer intelligence but agency: models know a great deal, yet fail to convert that intelligence into revenue, finished work or complex deliverables through multi-step execution, tool use and context management. Neo is therefore defined as an Agent that “makes you smarter while also making you lazier.” The host warned that it could keep “silently burning through my tokens and credit card” after the user has forgotten about it; Derek sees long task chains and long memory as necessary steps toward AGI.

  • The real product differentiation lies in autonomously iterating on the same piece of work: Neo can repeatedly revise an initial webpage from roughly 5,000 lines of code to 50,000 or even 500,000, rather than planning once, generating once and stopping. Users control depth by task: simple jobs may run for only 1–2 steps, while product-grade delivery can run 100–200 steps. That is Flowith’s key answer to the question of how it differs from Manus and Genspark.

  • The team uses 3 demanding examples to demonstrate the value of long chains: extending the 100-episode, 200,000-character Chinese drama Empresses in the Palace; tracing a dog-harness image to a brand story dated December 8, 2022 and answering “bacon”; and building a 3D billiards game that can actually be played. A separate Black Myth: Wukong case combined retrieval, character-image generation, long-form code, an interactive webpage and a YouTube video in one deliverable, demonstrating what the team calls multimodal, product-grade creation.

  • Flowith deliberately limits Neo to “AGI for the creative domain,” because creative work needs to be seen, intervened in and repeatedly refined, while generic tasks such as booking a hotel only require an outcome. Text, images, video and webpages can naturally convert into one another in its framework. The company has intentionally avoided phone calls, text messages and food delivery because they add limited value to creative work; the trade-offs are concentrated on the creative workflow.

  • On competition, Derek calls Manus a milestone that brought Agents to the mainstream, while acknowledging that its emergence accelerated Neo’s release from the original second-half or year-end plan to May. He expects model companies to have more room in general-purpose Agents and startups to be better suited to vertical Agents; the cloud is ideal for delegating work before bed and reviewing it in the morning. The host noted that local browsers retain more login state, context and browsing history. Derek acknowledged the local advantage and said better cloud–local integration may eventually be needed.

  • Flowith’s growth playbook is AI-native as well: it used the knowledge-trading properties of its “knowledge garden” to ride DeepSeek’s momentum with content around “making money with DeepSeek,” generating roughly 500,000 Xiaohongshu views; “I think I found the OnlyFans of the AI era” became a headline reused by the media. The conversation called this Vibe Marketing—letting AI generate directions, headlines and creative assets at scale while people focus on input and selection. The same discussion stressed the need to “post aggressively,” while acknowledging that no headline can save a product with no real hook.

Deep dive

1. Flowith Has Cleared Early Validation While Staying Lean

  • At the time of the May 2025 recording, Derek said Flowith had more than 200,000 registered users. DAU was volatile but had been growing rapidly in recent weeks. On revenue, he said February was “several times” January and March was “several times” February.

  • Revenue has also entered visible territory: Derek said March ARR reached $1M, while current monthly revenue was in the several-hundred-thousand-dollar range. He expects Neo’s launch to bring another step-change.

  • The company is closing two consecutive funding rounds totaling more than $10M. Before that, it had taken only one small angel check; earlier still, the founding team self-incubated the business with savings from previous ventures.

  • Derek founded technology-education brand Tech X Academy in 2016, later renamed X Academy, which is now in its 10th year. In 2018 and 2019, he also worked on interest-based social product Realm. During college, 拐子 organized China’s first electronic-music festival and built several million followers across public WeChat accounts.

  • The team has grown from 3 people a year ago to 15–16 today. More than half work in engineering and technology, with the rest covering product, design and marketing; over the long term, the company still wants to keep headcount below 30.

2. Neo Is Betting That the Bottleneck Is Agency, Not Model IQ

  • Derek’s starting point is a paradox: models are already highly intelligent and know enormous amounts, yet most people still have not made more money or produced better work with AI because they have “failed to bring out its agency.”

  • The role of an Agent is therefore not to swap in another chat interface, but to use multi-step execution, tool calls and context management to turn static intelligence into the ability to execute complex tasks.

  • During the Oracle period, this was still exploratory. With Neo, the team has explicitly put longer task chains, longer context and longer outputs at the center because it does not want the Agent to be a toy users try once and abandon.

3. “Infinite” Scope and Cloud Execution Redefine the Task Boundary

  • Neo’s architecture pushes simultaneously toward infinitely long task chains and infinitely many outputs. A task can continue for a week or a month without requiring the user to keep a browser tab open and watch it work.

  • 拐子 offered a concrete media example: check the Twitter activity of OpenAI, Google, Anthropic and relevant executives every 2 hours, then continuously update emails, documents or webpages as instructed rather than regenerating each round from scratch.

  • Internally, the team summarizes this state as “can stop.” As they explain it, “it can’t stop; it can keep working for you without interruption.”

  • The host raised a potential cost: the user may have forgotten the task was ever assigned while it continues consuming tokens and charging the credit card. Derek returned to the point that long task chains and sufficiently long memory may be necessary steps for Agents on the road to AGI.

4. Flowith Wants to Build AGI for the Creative Domain

  • Derek defines Neo as “AGI for the creative domain,” rather than a fully general assistant. Booking a hotel requires only the correct result; creative work requires people to inspect the process, understand the choices and sometimes intervene while the Agent is working.

  • Creation is not limited by modality. Text can become an image, video or webpage, and a user who initially asks for a long report may discover that a webpage is the better deliverable.

  • “Everyone can become a creator” is the broader implication: students can write papers and reports, while professionals can use AI to become better product designers, interaction designers or programmers.

5. A 200,000-Character Empresses in the Palace Shows How Long Tasks Are Decomposed

  • Derek imagines having Neo extend the 100-episode, 200,000-character Chinese drama Empresses in the Palace: first retrieve the ending of the original, then plan an outline covering all 100 episodes. The full job could require more than 100 steps.

  • The Agent would then build an empty webpage shell, write Chapters 1–5 in batches, generate illustrations for each chapter, and place the text and images into the page. It would repeat the cycle until the full work was complete.

  • The point of the example is not the word count of a single generation, but the closed loop formed by planning, generation, illustration, assembly and continuous rewrites to the same deliverable.

  • Derek’s quality bar is that the final webpage will be “very long and very detailed,” with content good enough to read like a novel. The team’s product claim is not merely that the task reaches completion.

6. Neo’s Structural Difference Is Repeatedly Refining the Same Work

  • A conventional AI might generate a Snake game in one step. Neo produces a first version, then adds features one by one, modifies the code and optimizes autonomously, turning a single delivery into a long iterative chain.

  • Derek says an initial webpage might contain roughly 5,000 lines of code before growing to 50,000 or even 500,000. Cursor, by contrast, still typically requires a person to supervise the AI step by step.

  • The host’s challenge was: “How is this different from Manus and Genspark?” Derek’s answer was not simply planning but granularity—Neo implements one feature per step and keeps modifying the same work, addressing the weakness of existing Agents in refining deliverables.

  • The system does not enforce a minimum step count. For the same “build a Snake game” prompt, users can choose 1–2 steps for a quick result or max out the depth and let it run for 100–200 steps to pursue product-grade output.

7. The “Bacon” Test Measures Dynamic Reasoning, Not Long Output

  • A GAIA Level 3 question presents images of 2 dogs and a harness, asking the Agent to identify the harness brand, find a story published on the brand’s website on December 8, 2022, and determine what meat the dog ate.

  • The standard answer is a single word: “bacon.” The difficulty is that the image offers no explicit brand cue and the article contains many distracting terms, forcing the Agent to dynamically adjust its search steps, models and tools.

  • 拐子 compared it with “a high-school student competing in a math Olympiad” versus “writing a Chinese composition”: the same brain is involved, but the required combination of capabilities is entirely different. In the team’s testing, the other general-purpose Agents in the comparison failed to answer correctly.

8. 3D Billiards and Black Myth: Wukong Push Creation Toward Product Delivery

  • The 3D billiards case required handling ball collisions, shot power, camera movement and perspective. The team tested Oracle, Manus, Genspark, Cursor and Lovable; 拐子 said only Neo completed a version that could actually take a playable shot.

  • It did not write the game once and stop. It first reasoned through billiards rules, built the table, and then added collision and hitting mechanics step by step. 拐子 compared the workflow to that of a product manager or developer and emphasized that Neo continuously improves its own results.

  • The Black Myth: Wukong website evolved from the limited images and introduction created during the Oracle period into a product that studied the original’s aesthetic, generated character concept art, used gold visuals and artistic fonts, embedded video and delivered an unusually long interactive webpage.

  • 拐子’s conclusion is that infinite context is not merely a question of text capacity; it allows text, images, code and video to remain together inside one continuously expanding product.

9. The Infinite Canvas Came From Real Workflows and Creates Real Friction

  • The host said Flowith had “a bit of a learning curve.” Many users were attracted by the flashy interface but washed out after trying it. Derek called that assessment “a bit one-sided,” while acknowledging that a free-form canvas naturally requires learning.

  • The product has shifted from free dragging closer to Figma or Photoshop toward a “flow layout.” New users can simply type into an input box, making the experience closer to ChatGPT; only later do they discover branching, model switching and side-by-side comparison.

  • The canvas initially solved a pain point for the founding team: after modifying a prompt, a single-threaded chat made it difficult to retain and compare historical versions. They moved prompts and responses into Figma; one investor even used Excel to turn a one-dimensional conversation into two dimensions.

  • Derek’s view is that once creators learn to expand horizontally and extend vertically, “many people can’t go back.” Several products later adopted similar interactions, which he describes as “a contrarian view turning into consensus.”

10. Manus Educated the Market and Accelerated Homogenization

  • Derek sees Manus as a “milestone” for the Agent and AI industries: it brought the Agent concept to a mass audience and helped users with no prior understanding of Agents accept a new form of AI product.

  • His criticism is equally clear: after Manus, many products added Agent functionality without bringing anything new. “The publicity is huge,” while the functional differences are limited; he believes the market may be developing fatigue.

  • Flowith has chosen to continue making high-risk bets. Derek believes that even when an experiment fails or does not work, it still moves Agent products forward.

11. The Cloud Is Built for Asynchronous Delegation; Local Browsers Retain Login-State Advantages

  • Fellow has argued that Agents should run in the local browser. Derek admitted he had not used Fellow in practice, but said Oracle once came close to a local model: users had to keep the webpage open and watch the task complete.

  • After Neo entered internal testing, a new habit emerged: assign work before bed—conduct a 200-step deep dive or stock research, build a website, or collect background on the people scheduled for the next day’s meetings. The task continues while the user sleeps, who can review the process and results on a phone after waking.

  • The host pointed out that local browsers contain more context and browsing history, and that users are typically already logged in, avoiding repeated authorization requests. Derek said a cloud Agent asked to post on Twitter could run into login barriers, while logging into a virtual machine creates privacy risks.

  • His conclusion remains open: the local route “clearly has its advantages,” and a better hybrid may emerge. The cloud route, meanwhile, is focused first on solving long-running asynchronous work.

12. “Vertical” Comes From Choosing Not to Do Things, Not From Lacking Generality

  • After trying Neo, the host found it “pretty general” and asked why a creative Agent would not be overwhelmed by general-purpose Agents. Derek acknowledged that if the end state of general-purpose Agents is AGI, vertical Agents and many human jobs will eventually be disrupted.

  • But he does not expect that end state to arrive overnight. Current general-purpose Agents can “do all kinds of things,” but cannot yet do any specific domain extremely well, leaving room to optimize vertical workflows.

  • Flowith’s trade-offs show up at the tool layer. Technically, it could add phone calls, text messages and food delivery, but those capabilities have limited value for creative work, so it has chosen not to build them and instead concentrate resources on capabilities aligned with its positioning and long-term vision.

  • Derek also acknowledges that the boundary between general and vertical is blurry; the restriction itself is a product choice. Another, more “science-fiction” technical update may arrive 3–4 months after the recording, but he did not elaborate in advance.

13. Manus Raised Capital Heat and Forced Flowith to Launch Early

  • Against the backdrop of Manus’s roughly $500M valuation, the host asked how the capital market had changed. Derek emphasized that Flowith had already been growing visibly by February, before Manus launched, so the earlier funding round was not materially affected.

  • He attributed later rounds partly to Manus heating up the Agent category. Flowith had originally planned to fully upgrade Oracle in the second half of the year or by year-end; after Manus appeared, it accelerated development and ultimately moved Neo’s launch to May.

  • The funding environment went from conservative to hot in a matter of months. Derek said the team spoke with roughly 20–30 institutions over the prior 2 months; after receiving a strong offer, leading firms began competing for the deal and the valuation changed rapidly.

  • Recent conversations have still been mainly with Chinese funds, but the product targets a global market and the team wants to continue approaching US investors. The company structure and the issues US institutions will care about in an overseas financing process need to be designed in advance rather than addressed during fundraising.

14. DeepSeek Connected Product Capability, Hot Topics and Virality Into a Growth Loop

  • Flowith launched version 2.0 and its “knowledge garden” in January. Beyond knowledge-base functionality, the product also had a knowledge-trading dimension, creating a product-relevant angle for later riding DeepSeek’s momentum.

  • When DeepSeek’s servers crashed on Lunar New Year’s Eve and interest surged, a team member said they did not merely discuss the model news. They put their own internet-native copy into the knowledge base and had DeepSeek inside Flowith generate 20–30 possible distribution angles.

  • The winning theme was “making money with DeepSeek.” The content emphasized that the knowledge base could be monetized; combined with product-driven virality, it generated organic traffic in China and abroad. The speaker described the February and March results as “a small, exponential-style growth.”

15. Vibe Marketing Turns AI From a Copywriter Into a Creative Partner

  • DeepSeek-related viral content generated roughly 500,000 views. When the knowledge garden launched, AI produced more than 10 directions, and “OnlyFans” showed the team a powerful analogy between knowledge trading and social distribution.

  • The final headline was “I think I found the OnlyFans of the AI era,” which other media later reused. What impressed the speaker was not the gimmick itself, but AI’s ability to identify a core keyword that people might not have thought of first.

  • The discussion called the team’s long-running practice Vibe Marketing: leave most decisions, directions, headlines and image selection to AI, while people mainly provide input, filter the output and reduce intervention. Compared with Vibe Coding, the speaker believes it may attract even more attention in marketing and distribution.

16. Xiaohongshu, an AI-Native Organization and the Long Road Ahead Form an Execution Advantage

  • Xiaohongshu was described as close to “equality of attention”: both new and established accounts can enter recommendation feeds through a single strong post, while image-and-text content is cheap to produce. In practice, teams should study the hottest content in a topic, reshape the headline and images, and embed conversion into the body copy.

  • The Xiaohongshu playbook is rougher: “You have to post aggressively.” If 5 posts do not work, post 20 or 50 until one generates positive feedback and turns the flywheel. But if the product is ordinary and lacks a real hook, even strong internet instincts cannot rescue it over the long term.

  • The development process has compressed from the traditional product–UX–UI–engineering–testing chain into product directly to engineering: UX has been folded into product, UI is delegated to AI operating within a design system, and engineers combine web coding with AI assistance. Derek then tests, edits code and polishes the output. He personally uses Cursor for roughly 3–4 hours a day.

  • Derek’s biggest retrospective is not a specific product mistake, but that the team was “not confident enough” in contrarian calls around the canvas and Oracle and failed to go all in more aggressively. The short-term task is to keep refining Neo, the knowledge base and the canvas; the long-term plan is to build the “ultimate AI creation tool” over 1–3 years while preparing for a prolonged battle whose shape cannot yet be predicted.