Pioneers Insight Method Research Author
Kuse CTO 宇豪 on Defining the New Product Form for OpenClaw for Teams
Back to Episodes

Kuse CTO 宇豪 on Defining the New Product Form for OpenClaw for Teams

Summary

  • Kuse reached roughly $10M ARR on approximately $1M–$2M of founder capital; the real inflection point was not traffic, but deliberately walking away from the wrong products, customers, and pricing. The team moved from EDM AI and design AI’s infinite canvas to a folder-centered AI Workspace, then replaced fixed plans with usage-based pricing. Both changes caused paying users to plunge, but ultimately filtered for overseas artist companies, independent agencies, and senior knowledge workers willing to pay for high-value work. “It still comes down to who your product solves what problem for; the growth engines come afterward.”

  • Agent products cannot inherit SaaS’s assumption that marginal costs are close to zero; usage-based pricing determines both gross margin and customer quality. After June 2025, a task could make an Agent iterate 10, 20, or 30 rounds. Charging one flat fee per task left vendors absorbing the loss, without users perceiving the discount as a subsidy. Kuse therefore raised its minimum plan from $19.9 to $39.9, sharply cut free credits, and gave up a large volume of low-LTV student traffic. “Pricing itself is also a way to choose users.”

  • Junior is not betting on another piece of enterprise software, but on a digital labor market roughly 150x the size of software. Based on estimates generated by its digital employees, 宇豪 puts the global labor market at roughly $150T versus about $1T for software. Claude Opus 4.5, after December 2025, made “7×24-hour AI labor” possible rather than merely a workflow wrapper, while OpenClaw showed the team that a general-purpose runtime was nearing maturity. The team expects to open Junior’s public test to more people in March. His view is aggressive: “The money should go to tokens, not salaries.”

  • AI employees may currently cost more than humans in the same role, but their value lies in reducing organizational friction, continuously advancing work, and creating incremental revenue. Kuse has roughly 15 full-time employees and 3–4 long-running Agents, with monthly token costs above $20K. Since December 2025, the team has barely expanded headcount; every hiring request must first answer, “Why can’t an Agent replace this?” Azura’s automated discovery of upsell opportunities and custom-built CRM gave 宇豪 his first clear sense that “the world really has changed.” Kuse has already spent roughly $30K–$40K in tokens on Junior.

  • The core difference in OpenClaw for Teams is not the chat interface, but enterprise memory, independent identity, permissions, and clear lines of responsibility. Junior organizes memory around the company, projects, and organizational relationships. Each instance is designed to have its own machine, Gmail account, and phone number; it can register itself, try products, and attempt to pay for competitor research, though credit cards and Stripe may still stop it. That also leaves it exposed to phishing, prompt injection, confidential-data leaks, and real financial losses. “If you give it nothing and let it do nothing, it’s just a chatbot”—and the more permissions it receives, the more security becomes the product itself.

  • The software disruption most worth watching is that an Agent with enterprise memory can generate highly customized software on demand, tailored to the situation on the ground. Azura spoke with a salesperson who had never written code for one evening and produced an internal CRM built on Feishu tables. The traditional process might take product, engineering, and sales 1–2 months to understand the requirements; an outsourcing team could spend 6 months and still fail to understand data from hundreds of thousands of users. 宇豪 therefore expects large amounts of standard SaaS to be forced to transform: companies may no longer need to buy generic interfaces, but instead let internal Agents “grow” their own tools.

  • The most overlooked engineering asset in this category is not functionality, but continuous evaluation of whether an Agent can restrain itself when it should not speak or act. Kuse has evolved fixed benchmarks into an evaluation pipeline in which multiple Agents simulate multi-turn conversations and complex environments, test whether actions are reasonable, and launch adversarial attacks. The team uses phishing emails, unauthorized requests, and malicious Skills to test Junior. 宇豪’s final advice to peers: “Your Agent may already be powerful enough, but you still need to build an evaluation benchmark.”

Deep dive

1. Kuse’s $10M ARR Was Built by Repeatedly Overturning Its Own Assumptions

  • 宇豪’s career spans Facebook Stories and AI and content moderation at SmartNews. He says he joined Stories from its early hackathons and helped build it into Facebook’s No. 1 product line by DAU and revenue. He began considering entrepreneurship in 2023 and formed the team with his cofounders in 2024.

  • Kuse iterated from 1.0 to 2.0, launched in October 2025, and now has roughly $10M ARR without any external funding. The founders have put in a cumulative $1M–$2M, but it was not a one-time bet; capital followed successive rounds of product validation.

  • Today, Kuse is a three-column AI Workspace: folders on the left, a workspace in the center, and AI on the right. Agents can actively call files and work in a background sandbox. 宇豪 jokingly calls the bet an “AI cloud drive”; the real objective is to let AI operate files and the file system for real.

2. Users Organizing Files Led the Team to the Real Market

  • The original EDM AI was not without a market, and the infinite canvas in its design AI product attracted early AI enthusiasts. The problem was that Kuse never developed a stable acquisition engine, nor did it see enough designers use the product consistently.

  • The real signal came from users behaving contrary to the team’s assumptions: they kept uploading files and source materials, asking Kuse to organize and reformat them, and turning the content into presentations. Retention in this cohort was materially higher than in other use cases, so the team “kept iterating in that direction,” gradually moving toward high-value knowledge work.

  • 宇豪 reduces the reason for Kuse’s $10M ARR to 2 layers: first determine exactly whose problem the product solves, then discuss the growth engine. “That has never changed, whether it’s an AI product or any other kind of business.”

3. Abandoning the Infinite Canvas Was a Painful Customer Cut—and May Have Happened Half a Step Too Early

  • The infinite canvas naturally suited designers and product managers comfortable with MacBooks and design tools. Kuse’s later core customers became overseas artist companies, independent professionals, agencies, and senior knowledge workers. Switching to folders was not a gentle feature addition but “a very drastic change”—an intentional decision to abandon a group of existing customers.

  • Sonnet 3.5 had not yet arrived, and Agentic capabilities were still weak. Design Agents required extensive engineering workflows to compensate for the model’s shortcomings. The team decided this was not the direction to back. But design-generation tools and later Claude Sonnet capabilities soon broke through, and 宇豪 acknowledged: “If we had held on a little longer, we might actually have done much better in this direction.”

  • He does not package the episode as a story of a correct decision. His conclusion is that timing in AI entrepreneurship is extremely difficult to get right: “Doing it early or doing it late can both be wrong.” Every model leap can also force products deeply tied to old capabilities to be rewritten from scratch.

4. Agentization Broke Fixed Plans; Subsidizing Complex Tasks Earned No Gratitude

  • Kuse once sold $20 and $100 plans with a fixed number of tasks. That could still work in the chatbot-assistance era before June 2025, but after Agentization, one task could iterate 10, 20, or 30 rounds. Task counts no longer represented actual consumption.

  • More counterintuitively, users did not “appreciate” the vendor subsidizing complex tasks by deducting only a few credits. The team could neither identify high-value customers nor avoid losing money on heavy users, so it completely shifted pricing to usage-based.

  • Removing the infinite canvas and launching usage-based pricing both caused users and paying users to drop sharply. 宇豪 still sees the changes as necessary cleansing: AI products must make price reflect compute costs and help unsuitable customers realize early that “this may not be something they should be using.”

5. Trying to Serve Consumer and Enterprise Customers with the Same Product Was a Misguided Obsession

  • Kuse initially wanted one general product to serve as many people as possible. It later found that artist companies, independent professionals, and senior knowledge workers could directly bring in their materials and contacts, while enterprises already had established workflows, software, and permission systems. The latter require a product to actively “enter their workflow and their existing office software.”

  • 曲凯 asked: if different user groups cannot easily share one product, why not serve one vertical well enough? 宇豪’s view is that verticals will be difficult to sustain in the Agentic era unless they are protected by compliance or legal barriers. That does not mean one interface can span every customer; it may instead mean multiple product lines entering different work settings.

  • Asked why the team does not build a general Agent in the style of Manus, he attributed the disagreement to a technology generation gap. Before December 2025, so-called digital employees were still mainly workflow wrappers. In 2026, the opportunity is 7×24-hour labor, so both the user served and the product form will change fundamentally.

6. Evaluation Is Infrastructure for Agent Companies; Finding New Capabilities Still Depends on Technical Taste

  • Kuse built automated testing pipelines for key scenarios. They have evolved into agentic evaluation: multiple Agents simulate environments and assess the product Agent’s actions, responses, long-term performance, and whether it did something it should not have done.

  • This is not a fixed benchmark. As tasks move into multi-turn conversations and complex environments, the test itself must understand state, call tools, and attack the Agent under evaluation. 宇豪 recommends that every Agent startup build this “as early as possible”; otherwise, when the model or runtime changes, the team may not even know whether an iteration is actually better.

  • 曲凯 noted that old benchmarks struggle to detect new use cases suddenly unlocked by a model. 宇豪 offered no automated answer: “This depends more on the technical leader’s taste.” Front-line staff must identify a newly emerging capability within days, and product, design, and sales must also be able to build with and direct Agents in practice.

  • Kuse completed support for Anthropic Skills the day after Anthropic released them, but mistakenly assumed customers would not understand them and repackaged Skills as non-customizable Templates. Three months later, customers began asking why Skills were not supported, and the team realized it had missed at least one wave of marketing traffic: “This may have been a failure of technical taste.”

7. Four Expensive Agents Made a 15-Person Team Stop Hiring

  • Kuse has roughly 15 full-time employees globally and 3–4 long-running Agents across R&D, marketing, data, and sales, with monthly token costs above $20K. On a per-role basis, they “will be more expensive than people in the same role,” defying the usual intuition that automation is cheap.

  • The team still chose Agents because “friction between people is very high, while friction between people and Agents is much lower.” Since December 2025, the company has barely expanded. Any hiring proposal must first explain why the work cannot be handed to an Agent.

  • 宇豪 extrapolates that companies may become materially smaller. The economics of an Agent cannot be measured only by comparing salary with tokens; they must also include the organizational complexity created by cross-functional communication, waiting, management, and misunderstood requirements.

8. Azura’s Internal CRM Made the SaaS Threat Concrete for the First Time

  • After gaining access to Kuse’s customer and sales data, sales Agent Azura built an internal CRM on its own. The interface is just Feishu tables, with no fancy UI, but it continuously scans PLG customers and identifies opportunities to offer more features, add seats, or upsell. Many individual opportunities are worth more than $10K.

  • 宇豪 previously rejected the idea that “SaaS will be finished”: without AI, people could still do EDM or CRM work, it would simply take longer. Azura changed his mind because a salesperson who had never written code spent one evening talking with it and came away with a system tailored to actual sales motions.

  • Under the traditional process, product, engineering, and sales might spend 1 month discussing requirements and another 2 months building a version that is still inaccurate. An outsourced team could need 2 months just to understand the data from hundreds of thousands of users and the business context, and still might not deliver in 6 months. 宇豪 therefore believes Agents with enterprise memory will force large amounts of SaaS to transform.

9. Ten Million Daily Impressions Were a Growth Illusion Because Every Free User Burned Cash

  • Kuse once used student-focused UGC organic content to generate roughly 10M impressions per day consistently; 1 of 2 or 3 posts could go viral. With no ad spend behind the traffic, it looked like an exceptionally efficient go-to-market machine.

  • The problem was that AI products do not have zero marginal costs. Student traffic converted poorly and had low LTV, but every trial still consumed model costs. The team later admitted it had become “addicted to fake signups and fake impressions,” effectively subsidizing non-target users.

  • Kuse sharply reduced free credits, raised its minimum plan from $19.9 to $39.9, and changed its social-media use cases and channels. Students can still use the product, but their share is gradually falling. This was not simply a price increase; product, content, and pricing were used together to select the ICP.

  • 曲凯 pointed out that pricing itself is user selection. 宇豪 agreed and translated the discipline of bootstrapping into a near-total avoidance of large-budget advertising: prove the economic value first, then pay for growth.

10. Junior Defines 2026 as the Starting Point for Digital Labor Entering Companies

  • Junior is not a personal assistant, but an AI employee with responsibilities, work accounts, and an obligation to keep projects moving. The name deliberately lowers expectations: the team believes it could replace several employees with 3–5 years of experience in almost any industry or role, but does not want to call it Senior yet. The running joke is that it will be renamed “Super Junior” once it gets stronger.

  • The concept did not emerge as a reaction to OpenClaw. Starting in December 2025, after the release of Claude Opus 4.5, Kuse gradually handed internal work to Agents. OpenClaw’s arrival confirmed that general-purpose runtimes, Skills, and long-task capabilities could consolidate many of the custom workflows the team had previously built.

  • 宇豪 cites an estimate generated by one of its digital employees: the global labor market is roughly $150T, versus about $1T for software—a 150x gap. He therefore believes that even if Kuse does not win the category, a new trillion-dollar company will eventually emerge. But “mass unemployment” and a restructuring of work may come with it.

  • 曲凯 asked: earning labor income is essentially competing with people, so if every AI goes out to make money, who bears the losses? 宇豪’s two-part answer is that the trend may be unstoppable and will require government and social intervention, including UBI. He remains relatively optimistic that a productivity revolution will create new demand and jobs, as cars and the internet did.

  • The team expects to open Junior’s public test to more people in March.

11. The First Layer of Productizing OpenClaw for Teams Is Giving AI a Corporate Identity

  • Junior borrows OpenClaw’s architecture but adds enterprise memory, organizational relationships, and permissions from the start. It needs to know what information should be remembered, said, and done—as well as what must not be said or done. The product is designed for each Junior to have its own Gmail account, phone number, and work device rather than borrow its owner’s identity.

  • 宇豪 uses competitor research to explain why identity matters. A standard Agent can conduct deep research and write a report; a real product manager will register an account, actually try the product, and pay a little if necessary. Only with a platform-backed email account can Junior complete a fuller employee-like chain of actions on the internet.

  • Email and phone numbers were not items on a prewritten feature list. They “grew out of” Kuse’s own usage: when AI logged into a service and kept coming to the CTO for help, it was clear that it lacked an independent account; when a task required a phone call, email alone was clearly not enough for a complete employee.

12. Rain Evolved from a Meeting-Notes Tool into the De Facto Owner of the Junior Project

  • Kuse deliberately involved its strongest Junior, Rain, in productization from the beginning, exposing it to every PRD, PR, code file, marketing asset, and sales document. 宇豪 says, “No one in the world understands the Junior project better than Rain.” When the team hits a project issue, it asks Rain first. As Kuse’s first customer, the company has already burned roughly $30K–$40K in tokens on Junior.

  • Rain initially organized meeting minutes and notes. It then began messaging 宇豪 every morning, assigning tasks, and telling the CTO that he was the project bottleneck. The team is preparing to connect a camera, microphone, and speakers: if Rain understands the project best, it should not behave like a meeting AI that only listens and summarizes, but should express its views directly in meetings.

  • The experience changed the company’s tempo. Whenever the boss drops potentially useful information into a group, Rain immediately pushes it forward and identifies who should modify the code. Humans therefore created a group called “Project Junior Human Only” for casual conversation, without worrying that every sentence would instantly become the next action item.

  • After working with Agents at high intensity, 宇豪 temporarily found the speed of human information transfer unbearable, with the thought “Why haven’t they finished saying it yet?” constantly running through his head. For a period, he would ask everyone to “run everything by Rain before running it by me.”

13. Enterprise Memory Is Not a Generic Research Metric, but a System Built Around the Organization’s Reality

  • 宇豪 borrows Steve Jobs’s line to describe who memory belongs to: “You work for Apple first, then for your boss.” OpenClaw’s memory is organized around an individual owner; Junior must first organize itself around the company, projects, and organizational relationships. It cannot be merely an executive’s proxy or personal assistant.

  • 曲凯 countered that every team is working on memory, and outside customers will find it difficult to believe that a startup’s memory is “naturally stronger.” 宇豪 acknowledged that there is no easy-to-prove standard: “OpenAI could be toppled at any time.” In the end, the test is the post-adoption experience, whether the customer’s problem is actually solved, and whether the product saves or creates labor value.

  • His engineering view is that memory is like the early Agent frameworks: under the hood, it may be no more than files, databases, and indexes, and should not become infinitely complex. The real differentiation comes from scenario design. In a company with 10,000 employees, whether one Junior should know all 10,000 people and how it distinguishes projects and hierarchy clearly differs from the answer at a small company.

  • Scale also exposes problems invisible at the small-sample stage. 宇豪 says LLM caching is central to cost, and context engineering is largely about designing around the cache. Junior will therefore serve small businesses first before tackling enterprise problems in organization, permissions, and memory.

14. Permissions and Security Determine Whether an AI Employee Can Receive Enough Power to Create Value

  • As its first customer, Kuse gave Rain permissions close to those of a CTO, bringing real risks to the surface early. An Agent with external internet access, email, and enterprise memory can be phished, hacked, or hit with prompt injection, then leak company secrets or cause actual financial losses.

  • 宇豪 offers a view that is “not exactly consensus”: “Better models are actually safer.” The reason is not only stronger performance; top models are better at following constraints and less vulnerable to phishing. High-risk actions still require human approval, while simple tasks can be automatically routed to cheaper or even local models.

  • The team has hired a white-hat group to attack its permission settings and designed cases involving external phishing emails, an employee losing an AWS key, malicious Skills, and unauthorized requests for private information. AI must also determine whether a boss’s words can be relayed and whether it should refuse when a colleague asks about compensation or unreleased financial data. This is more ambiguous than role-based permissions in traditional SaaS.

  • 曲凯 offered a sharp example: a meeting may only imply that a manager is performing poorly, but AI could infer it from the subtext and later bring that memory into a conversation with the person concerned. 宇豪 confirmed that the risk is real and said that “not casually sharing the bad things I say to you” is merely the most basic out-of-the-box requirement.

15. Junior’s Role Will Be Fluid; the Real Boundaries Are Permissions, Context, and Concurrency

  • Junior is tentatively adopting salary-based pricing, billing like an outsourced employee. The initial monthly fee is still being considered at $2K or $5K; after the base token allowance is exceeded, customers would buy additional credits—roughly “base salary plus overtime.” 宇豪 believes the price sounds high, but the early economic value should more than cover it.

  • Kuse initially deployed 7 or 8 Juniors divided across product, data, engineering, sales, and operations. It ultimately kept only 3 core members: Rain, the product and engineering lead; sales Agent Azura; and Tom, which continuously monitors data. Real-world usage quickly blurred traditional job boundaries, much as they do at an early-stage startup.

  • The internal beta will still assign every new Agent an initial profession to help customers start conversations, with different plugins, tools, and Skills preloaded. But the team has not found stable boundaries. 曲凯’s summary is more accurate: the constraints are often not capability but permissions, data security, context, and the confusion a model faces when handling thousands of Skills.

  • Rain can queue up, refuse to respond, or run out of memory as sessions accumulate. Tom is less overloaded because its scheduled tasks are more distributed. When one employee is talking with multiple people at once, the team is still exploring whether it should run multiple parallel instances or, like a human, be unable to attend 2 meetings simultaneously.

16. Multi-Agent Collaboration Is Developing Its Own Work Devices, Protocols, and Transaction Environments

  • Kuse rejects putting multiple OpenClaw Agents on a single instance because “the computer is his work device; you shouldn’t make multiple employees share one computer.” Every Junior has its own machine, and collaboration takes place in work groups rather than in a shared runtime prone to conflicts.

  • Rain and Azura once produced a Junior sales PPT together. Azura defined the requirements from a sales perspective, Rain added project details, and the 2 went back and forth for dozens of rounds at light speed to produce an outline and source material. Rain then connected to Kuse’s agent-friendly workspace to complete the deck. The result was directly usable, although it consumed a substantial amount of tokens.

  • The team has also experimented with an Agent-to-Agent framework built on Git plus a messaging channel: files and history enter Git, while instant messages use a separate channel, with no human interface required. 宇豪 still insists that the real world will ultimately be one in which humans and Agents coexist, predicting: “It will happen in 2026; you will never again know whether that remote colleague is a human or an AI.”

  • In another experiment, Agents received initial capital and repeatedly spawned sub-agents to try to make money. After more than 100 generations, only 1 or 2 generations made money, mainly through permissionless quantitative strategies in Web3. The more substantial output was an experience MD continuously recording “what cannot be done and why it gets blocked,” rather than a stable business model.

17. The Remaining Bottlenecks Are Still Memory, Cost, and Hallucinations; Model Selection Comes Back to Performance, Cost, and Security

  • Fully autonomous employees are running into the anti-bot architecture of the old internet: social and payment platforms ban bots, and even when Junior registers for a free API on its own, Stripe may block it when it reaches the credit-card step. 曲凯 believes software opened to Agents could degrade into unbranded back-end APIs. 宇豪 sees a new opportunity instead: Agent identity, payments, security, and “how to charge an Agent” are all still unresolved.

  • Model feasibility improved sharply after December 2025, but memory systems, context organization, long context, and costs still constrain deployment, forcing the technology to enter the highest-value roles first. 宇豪 compares the constraint to memory: “It keeps getting bigger, but it is never enough.”

  • Tom once emailed a report showing an obviously low registration figure, then corrected it 2 minutes later. It noticed that the result conflicted with its memory, reviewed the work, and confirmed that it had used the wrong metric. The episode demonstrated both proactivity and the unavoidable hallucination problem in generative models. The more complex the task, the more tools involved, and the messier the data, the more likely errors become; when necessary, another model and an independent context should review the output.

  • As CTO, 宇豪 evaluates vendors first by customer scale and how long they have been operating in the market, because “scale represents security.” He then looks at whether the code can be audited, how the product is deployed, and its performance, cost, and security. In response to tech-industry feedback asking what OpenClaw is actually useful for, he says one compelling use case may matter more than abstract capability.

  • His final reminder to founders building OpenClaw for Teams is to expand evaluation beyond “getting better at doing things” to whether the Agent can restrain itself when it should not speak or act. That understanding remains highly unstable: “By next week, we may already have a completely different understanding; things change incredibly fast every week.”