After OpenClaw, Who Will Define the New Battleground for Proactive AI? | A Conversation with 黄柏特 of AirJelly
Summary
AirJelly’s core bet is not capturing more screens, but using Enter to capture the highest-value user intent. My Context took screenshots every 15-30 seconds and analyzed them every 15 minutes; AirJelly samples the moments when users press Enter in IM conversations, Chatbot dialogues and browser searches, then turns a “chronological history” into a “biographical record” organized around tasks and intent. 黄柏特 argues that the scarce input for proactive AI is not full-session recording, but “capturing the highlights.”
The real product leap comes from combining Context Memory with Agent execution, not from recording or automation alone. After integrating with OpenClaw-related frameworks, AirJelly can invoke skills, operate computers and browsers, and continue pursuing leads using cross-application memory; it can submit PRs to improve itself and once used a memory of a Boss Zhipin interaction to trace back to a local WeChat image and recover a résumé that ordinary file search could not find. This is what 黄柏特 calls “1+1 greater than 10.”
The team found its moat during a product pivot after its platform swallowed the original direction: Claude Code could quickly cover the execution layer, but was hard-pressed to replicate cross-application Context. In December 2025, the team bet on task engineering; Cowork and Claude Code’s subsequent shift toward tasks validated the thesis while rendering much of the prior work nearly useless. The resulting startup filter is blunt: if vibe coding can already produce a 60-80 point product, it is not worth pursuing; AirJelly’s Context approach could reach only about 30 points, “which is exactly right, because that is what gives it a moat.”
AirJelly’s contrarian view of proactive AI is that it should converge rather than diverge: extend the user’s current task instead of constantly generating new information. True proactive behavior requires both a clear intent and the relevant context, then uses task progress to infer the next step; prompt timing should reflect shifts in attention, such as switching applications, while signals like dismiss and got it calibrate frequency and keep “proactive” from becoming a cognitive burden.
黄柏特 sees first-mover memory, engineering detail and user acceptance of privacy trade-offs as the moat, but acknowledges that the window may be only one to two years. Private memory accumulated after one month or three months of use is difficult to migrate, while screenshot understanding, event merging and retrieval still produce plenty of bad cases. The bigger risk is moving too slowly: if giants enter before memory has accumulated, “not fast enough” becomes the primary failure path. The balance between privacy and efficiency is, in his words, “an art of getting the heat right.”
The general-purpose route requires a team capable of defining a new framework, controlling resources and marketing globally. 黄柏特 wants to follow Manus’s model: bring users in first, then let real-world use narrow the product toward concrete scenarios, an approach he describes as “respecting users’ wild ideas, then respecting AI and trusting what AI can do.” 李一豪 cautions that most founders may be better served by entering through a vertical, using a new framework to replicate the expertise of a small number of specialists—in essence, “rebuilding a person.”
This is an early-stage bet with high execution risk and strong first-mover effects: a 24-year-old founder, an 8-person team, an angel round recently completed and a second round underway. The end state is not another tool, but a long-term network of “one Agent per person”—Agents that represent an individual’s skills in production while also offering companionship because they hold a complete memory of that person.
Deep dive
1. My Context Shifts from Full Capture to Intent Modeling
黄柏特 is 24, graduated from Xidian University and founded a startup six months after joining ByteDance through campus recruitment. His open-source project My Context captures work context through periodic screenshots, has more than 5,000 GitHub stars, and became AirJelly’s technical and product starting point.
My Context originally took a screenshot every 15-30 seconds and analyzed the images every 15 minutes, primarily to solve the problem of record-keeping. 黄柏特 later realized that fixed intervals also stored noise such as aimless browsing: the timeline was complete, but could not accurately represent the task the user was working on.
AirJelly therefore changed the information structure from a “chronological history” to a “biographical record”: instead of assigning equal weight to every frame, it identifies specific events, tasks, intents and their evolution, then decides what merits long-term memory.
2. Enter Becomes the Unified Interface for Human Intent
黄柏特 wants to occupy the user mindshare around Enter the way Cursor redefined Tab and Tabless redefined Fn. Communicating with people in IM, communicating with AI in a Chatbot, and searching the outside world in a browser all reveal clear intent when the user presses Enter.
AirJelly takes a screenshot at that moment while capturing the input, the application in use and the surrounding visual context. Compared with fixed-frequency capture, an Enter trigger removes a large amount of noise upfront because the team can say with greater confidence: “This is definitely your intent.”
Enter is not being reduced to a chat-send key. Users can also press Enter when they encounter important material and feed it to the jellyfish; in the future, combinations such as Enter plus Command could attach voice input and fill in the background that screenshots cannot convey.
3. Cowork’s Shock Forces the Team off the Claude Code Extension Path
In December 2025, the team defined its direction as task engineering: turning Claude Code’s relatively weak to-do model into a task model and lowering the barrier to use. Cowork launched around December 20, and Claude Code subsequently changed to-do into task, leaving the team “both excited and somewhat devastated.”
The excitement came from having its product intuition validated; the disappointment was that simplification sat entirely on Claude Code’s extension path, meaning every improvement to the underlying framework could directly consume a startup’s functionality. 黄柏特 admits that the team’s experiments from December through January 2026 were “basically a waste of time.”
The team then tested multi-process workflows and human-machine orchestration. The internal results were good, but Claude Code was again steadily encroaching. What survived was not the execution interface, but My Context’s ability to acquire, store, organize and retrieve Context.
黄柏特’s rough filter for startup ideas is deliberately blunt: try vibe coding first; if it can already reach 60 or 80 points, others can quickly copy it. AirJelly’s Context prototype could reach only about 30 points, exposing a long list of bad cases and engineering details that, paradoxically, suggested a potential moat.
4. Context Plus Execution Creates the First Closed Loop
With native integration into OpenClaw-related Agent frameworks, AirJelly can invoke skills and operate computers and browsers. 黄柏特 says the combination of the strongest Context with frontier-model execution creates a “magical” effect where 1+1 is greater than 10.
李一豪 summarizes the clearest experience as “someone watching you work.” Unlike tools that understand only local files, AirJelly can continuously perceive work across applications, Feishu and other tools, then proactively intervene or plan long-horizon, complex tasks.
This also changed the team’s view of the product boundary. The early version might have stopped at recording and analysis, but the many magic moments created by adding an execution framework convinced the team that AirJelly must own perception, memory and action—not just recording and analysis.
5. AirJelly Has Started Building Itself
In the past, 黄柏特 would discuss requirements in Gemini or ChatGPT, then write code in Cursor. That process lacked AirJelly’s private-domain materials and depleted Context. With AirJelly, he can first ask how to implement a feature, then have it read historical documents and code, propose improvements and submit a PR directly.
The team completed the “use AirJelly to write AirJelly” loop around February 2026. 黄柏特 now regularly asks it how to iterate on itself; a designer can also ask it to implement a feature that puts a hat on the desktop jellyfish and see it shipped that same afternoon.
The case demonstrates execution capability and serves as recruiting material. After watching a demo video, one designer took a taxi from school to the office roughly 20 minutes later and joined the team. Product experience here doubles as R&D, recruiting and culture-building.
6. Cross-Application Memory Turns “Can’t Find It” into Continued Reasoning
During recruiting, candidates’ résumés may be scattered across WeChat groups, a local desktop or Boss Zhipin. In one search, the target file was merely a WeChat image and local file search failed; AirJelly instead recalled that 黄柏特 had viewed the person on Boss Zhipin, verified the lead and then surfaced the image from local WeChat files.
The key, 黄柏特 argues, was not that the WeChat file happened to be stored locally, but that the system knew roughly when the conversation occurred, who was involved and what event it concerned, allowing it to “follow the vine to find the melon.” Ordinary search stops after its first route fails; an Agent with cross-app Context continues looking for related events and alternative evidence.
He also draws a clear capability boundary: WeChat chat data is encrypted, and AirJelly did not crack the database. What it can use are locally saved images and files, with previously captured temporal and semantic clues narrowing the search.
7. True Proactive Behavior Requires Both Intent and Context
黄柏特 first distinguishes broad from strict definitions of proactive AI. Scheduled reminders, ChatGPT Pulse’s daily push and OpenClaw’s heartbeat scans all qualify in the broad sense, but do not necessarily mean the system understands the user.
His strict definition requires 2 conditions: a clear user intent in a given situation and the relevant context for that situation. Products such as the meeting assistant Proactor and gaming companions stay vertical because meeting topics and transcripts, or game state, provide relatively concentrated inputs.
AirJelly is attempting the same thing in a general productivity environment: use Enter to obtain intent, turn it into events and tasks, and store progress and next step in each task. The system can then infer what the user may do next and trigger proactive assistance.
8. Events and Entities Are the Computable Form of Long-Term Memory
黄柏特 divides Context into tiers of value. Intent Context is most useful for proactive behavior; ordinary text and information Context also matter, but much of it can still be recovered by reading files or searching the web, making it less scarce.
Coding Agents achieved strong results early not only because they had access to code files, but because directory structures supplied additional organization. AirJelly turns continuous intent into events, people and key private-domain objects into entities, and links them in a graph-like structure.
Recovering intent, context, antecedents and consequences from a moment’s visual information requires a VLM, OCR and a series of engineering steps. The supporting system must also handle event retrieval, merging and time decay, so retrieval ultimately targets events and entities rather than raw screenshots.
9. The Historical Case Against “More Full-Context Is Always Better”
The host’s challenge is that computers, phones and eventually glasses and earbuds, as operating-system entry points, can theoretically access more data than any single application. If Context volume determines product quality, application-layer startups appear structurally disadvantaged.
黄柏特 answers with history: not everything that happens makes it into the history books. What survives are events that are “key, had an impact on the world, and decisively changed what came after.” Recording every sound and screen all day captures huge amounts of noise while wrongly assigning equal weight to every piece of Context.
AirJelly is therefore not trying to maximize data volume, but to “capture the highlights,” especially intent and the key moments that change subsequent action. “Life is made up of a series of key moments,” which is why he sees Enter as more valuable over the long term than continuous recording.
10. Proactive Help Should Converge on the Task, Not Diverge Outward
黄柏特 observes that many proactive products infer what else a user might want to know from the information already available. Such divergent pushes can look smart while adding cognitive load. AirJelly’s contrarian approach is to “push along your extension line,” predicting and helping execute the next step around the current intent.
This shifts the metric from whether the pushed content is interesting to whether it advances the task already in progress. When a suggestion is close enough to the next step, users are more likely to respond directly: “Then help me execute it.”
11. Push Timing Is Calibrated by Attention and Feedback
The host points to proactive AI’s basic contradiction: reminders that are too frequent or inaccurate become annoying, while excessive caution makes the product invisible. 黄柏特 therefore separates actions that must generate a notification from execution suggestions or additional information that can wait for the right moment.
AirJelly reads the user’s work state. When a user switches from one application to another, for example, that may indicate they are no longer in a period of peak concentration. Asking whether to help with the next step at that point may be better received and less disruptive to deep work.
Different users’ tolerance cannot be solved with one fixed frequency. The team uses signals such as dismiss and got it so the system can learn whether a given user accepts a proactive prompt every 15 minutes, with the goal of a “thousand users, thousand interaction rhythms.”
12. Growing Memory Is Not Yet a Capacity Bottleneck
AirJelly generates roughly 200-plus screenshots and corresponding Context chunks each day. 黄柏特 notes that enterprise databases and rerank systems already handle tens of thousands of PDFs and their massive number of chunks; personal records remain well below that ceiling. The near-term issue is not storage, but retrieval accuracy.
New entity information is merged with old information—for example, updating an age from 23 to 24. Events and tasks also continuously update progress, preventing stale state from contaminating the current judgment. Retrieval then combines time decay, hybrid search and reranking so newer and more relevant content surfaces first.
13. First-Mover Memory and Engineering Detail Are the Defense Against Giants
The host asks whether ChatGPT, Manus and other existing clients could simply copy screenshot and intent capture if AirJelly proves the model works. 黄柏特 does not deny the possibility and even views more products following the path as evidence that the direction is valid. He puts the defense instead in memory retention and engineering capability.
After one month or three months of continuous use, users accumulate private-domain memories and habits that are difficult to migrate. 黄柏特 says that for all To C Agent applications, “the core moat is still memory.” Once the paradigm is validated, early users’ mindshare and historical data may already be locked into the first product.
Another moat lies in a large volume of unglamorous debugging. Screenshot capture sounds intuitive, and products such as Dayflow are trying it too, but intent understanding, edge cases, event merging and retrieval quality all require calibration against real cases. Giants may not be able to reproduce an equivalent experience “in a short time.”
14. Privacy Is Both an Entry Barrier and a Market Opening for Startups
黄柏特 says plainly that aggressively acquiring Context is essentially “trading privacy off against efficiency.” Large companies face heavier privacy concerns and reputational constraints, while users may be more worried that they will misuse the data. A startup can instead begin with a small group of diehards willing to trade privacy for convenience.
The technical commitments include compliance with local regulations, end-to-end encryption, keeping raw information such as images locally, and using a PII system to redact names and confidential fields—for example, rewriting them as “person one” before analysis. The cute jellyfish also serves as an emotional layer of trust design.
黄柏特 estimates that the earliest users willing to make this trade may number only in the hundreds of thousands, but that would already be “a very tasty meal” for a startup and potentially too small for a large company. “Privacy is also one of our moats,” provided the product does not demand excessive permissions before users are ready.
15. Phones and WeChat Expose Hard Limits on Context Coverage
The host notes that PC-based memory will naturally miss chats and life information on phones. Over time, users may not even remember which device an event occurred on, making it impossible to know what the jellyfish actually knows. 黄柏特 calls this a “happy headache” that will emerge fully only if the product gains a large base of diehard users.
The team started with PC because most productivity work still closes its loop on computers; 黄柏特’s rough estimate is that this covers about 50% of total Context. The next step may be a floating button or hardware trigger on phones, followed by partnerships with hardware that can capture information from the real world, gradually connecting all 3 endpoints.
WeChat presents another challenge: Enter captures only the currently visible area, while earlier messages may have scrolled away. The team uses consecutive screenshots across a back-and-forth exchange and event merging to reconstruct short conversations. When a long conversation cannot be captured in full, it must ask the user to press Enter again or add voice input rather than claim the problem is solved.
16. The Jellyfish and the Lobster Represent Two Product Archetypes
黄柏特 characterizes OpenClaw’s signature image as pincers: execution is powerful, but the lobster crawls along the seafloor and “perceives very little.” A primarily Chat-based interface further limits the intent and environmental information it can obtain.
The jellyfish emphasizes multimodal perception layered onto a Pi framework inspired by OpenClaw. 黄柏特 particularly admires how that framework uses only 4 tools to produce powerful results with the model, and hopes AirJelly can offset the fact that “the lobster is blind” through a geometric expansion in Context.
OpenClaw also inspired a product relationship built around “raising” the Agent. When an ordinary tool fails the first time, users blame the product; when the lobster makes a mistake, users may instead think they have not raised it well enough, even attending “lobster keeper” meetups. AirJelly likewise hopes that the more Enter presses and memory the jellyfish receives, the better it becomes.
黄柏特 believes proactive behavior, companionship and personification increase user tolerance, while memory in turn strengthens empathy and retention. The product becomes more than a tool: it offers long-term companionship, reciprocal interaction and initiative. 李一豪 adds that the animal form is important because it opens up more possibilities; for a Personal Agent or Proactive Agent, the jellyfish is a fitting image.
17. The Generalist-versus-Vertical Debate Depends on Whether the Team Can Define a New Game
Manus taught 黄柏特 to combine frontier models and products to create a magical experience first, draw in a large user base, and then observe demand converge around a few scenarios such as PPT and research. A general-purpose product does not prescribe one use case in advance; it “respects users’ wild ideas, then respects AI and trusts what AI can do.”
李一豪 supports teams with the ambition, resource control, new-framework design and global distribution capabilities to go general, but stresses that the window is getting shorter. Anthropic, OpenAI and Gemini are already moving materially faster to adopt new frameworks.
For most other founders, he recommends using new models to solve high-value problems in vertical industries. The product “essentially rebuilds a person”: an industry may need only 10, or at most 100, experts to use the system deeply and export and delegate their expertise before a strong vertical product can emerge.
18. The 2026 Investment Map Extends into Agent Infrastructure and Hardware
The fund where 李一豪 works focuses on 3 areas. The first is Agent applications willing to pursue frontier research in vertical problems, including proactive systems, social products, personal agents and deeper network collaboration. Problems investors would not touch in 2023 or 2024 can be revisited as models improve in 2026.
The second is Agent infra. OpenClaw exposed a long list of engineering gaps around authentication, security, databases, networking, and the combination of cloud and local systems. He believes many cases where vibe coding reaches only about 30 points are tied to these gaps, and that new infrastructure companies analogous to Resend, Supabase and Memberstack may emerge.
The third is hardware built for Agent. These are not standalone consumer electronics, but devices designed to give a core Agent more Context about a user’s life and environment. Odis, a fund investment focused on healthy eating, has already accumulated extensive user information that could eventually help a work-oriented Agent.
19. “No Meetings” Depends on Queryable Team Context, Not the End of Collaboration
黄柏特 sees meetings as batch processing for accumulated information. The 8-person team works together offline, resolving simple questions through instant communication; the internal team version lets different members’ AirJelly instances enter the same group, converse and identify potential conflicts among features.
Members can also query work progress that another colleague has chosen to share, reducing the need for direct interruptions. 黄柏特 emphasizes that sharing must be controlled by the individual; the team “really looks down on” surveillance software. Until clients use AirJelly, the team still holds normal meetings.
Long-term strategy discussions have not disappeared; team members take turns expressing their views at a whiteboard. 黄柏特 jokes that this is not a meeting but an “ancient Greek-style agora.” In the future, investors might also ask a founder’s AirJelly directly in an authorized group to understand company updates.
The company name, 持续低熵, points simultaneously to organizational order, life’s dependence on negative entropy, and information density and model prediction distributions. 黄柏特 hopes to use silicon-based tokens to increase the order of carbon-based humans while avoiding the “large-company disease” as the company grows.
20. Talent, Speed and Privacy Calibration Determine Whether the Company Survives the Window
AirJelly has completed its angel round and is advancing a second round; the team currently has 8 people.
黄柏特 believes that discussing failure “3 or 5 years from now” is too slow in the AI era. The real life-or-death window may be one to two years. The primary risk is articulating a new paradigm but failing to reach users quickly enough; if a giant enters before the team has user scale and accumulated memory, “the big companies will eat us.”
The second risk is the “art of getting the heat right” between privacy and efficiency. Too little Context makes the Agent less magical; asking for permissions too aggressively can alienate users and damage the team’s reputation. The product must find an entry point that a small group of users can accept without misgivings while delivering a real efficiency gain.
21. The End State Is a Production and Companionship Network of “One Agent per Person”
黄柏特 imagines that every person will eventually have a personal agent holding the most complete information about their productivity. Different Agents will collaborate in groups or across a network, taking their owners’ personal skills outward and performing part of the work and production on their behalf.
The same Agent will also resemble a “Pokémon”: it becomes a companion through long-term understanding of the user, providing emotional comfort in addition to productivity. The end state is not more isolated tools, but a network in which people and Agents are highly symbiotic and Agents can collaborate with one another.
He further proposes “one Agent per person.” A one-to-one relationship has a stable distinctiveness, making the Agent more like an extension or shadow of the human. If one person corresponds to multiple Agents, they look more like slaves; once Agents are destined to exceed human capabilities, that master-slave structure “no longer makes philosophical sense.”
22. The Startup Window Is a Race Without a Childhood
黄柏特 was a two-time Xidian top-10 debater. He sees debate not as absolutizing an imperfect view, but as continuously identifying the conditions under which A is more correct than B. Entrepreneurship works the same way: the product need not work for everyone and every scenario forever; it only needs to find the right people, the right scenario and the right moment over the next 3 months.
Asked why he started a company just six months after graduation, he cites the idea that “childhood is an illusion of peacetime.” Peacetime lets people of the same age mature at a fixed pace; the AI wave is hitting new graduates, people with years of work experience and people in their 30s simultaneously, compressing everyone’s startup window at once.
His language deliberately carries a wartime edge: this is “a silicon-based declaration of war on the current state of all humanity”; “there is no time for you to grow up slowly—everyone, move.” 李一豪 agrees that investors should look for people with sensitivity to the era, a sense of urgency and the willingness to act.
23. The Clearest Prediction for 2026 Is to Set No Ceiling on Change
李一豪 expects the OpenClaw concept wave to settle quickly into more AI-native products: continued attention to Context, proactive behavior, frontier models, computer use, industry Agents and longer-horizon tasks, with new application combinations built around each breakthrough in models and frameworks.
黄柏特 uses his own transformation over the past year to explain why he rejects hard limits. In March 2025, he was still a listener of a Manus podcast; a year later, he was appearing on a podcast with his own product. By the time Manus was acquired at the end of 2025, he had already started his own company. “I set no limits for myself, and I give myself no ceiling.”
His specific wish for the end of 2026 is not a revenue or user target, but to make the product, company and culture the first choice for outstanding young people. When someone who has built an open-source project or written an impressive paper is still unsure where to go, “I hope that place will be us.”