130: The Mobile Agent Wave Begins! What AutoGLM 2.0 Says About How Foundation Models Will Reshape Phones | Agent #4
Summary
- AutoGLM 2.0’s real leap is not knowing a few more apps, but moving execution off the user’s screen and onto asynchronously running cloud phones and cloud computers. The product covers iOS, Android, and the web, requires no invite code, and is currently free; 刘潇 summarizes the core experience as “not taking over your screen,” allowing lifestyle services, Deep Research, PPT creation, website building, and content publishing to run in the background for extended periods. The number of people who use phones every day is almost equal to China’s total population, making the potential entry point more valuable than an Agent limited to computer users.
- The mobile Agent market is most likely to break out around “what should I eat?” rather than WeChat messages that take only a few seconds to write and send. 刘潇 focuses on restaurant discovery, changing tastes, and cross-platform filtering: an Agent could search Xiaohongshu first, then place an order on Meituan. The episode also discusses having it compare milk-tea prices and promotions across JD.com, Ele.me, and Meituan. Users order braised chicken with rice “not because I actually want braised chicken with rice, but because I’m too lazy to keep looking.” The real value comes when users’ hands are occupied, they are busy, or the search-and-execution chain is sufficiently long.
- GUI is currently the fastest path to deploying mobile Agents, while API looks more like a higher-efficiency long-term complement whose infrastructure has yet to close the loop. GUI can directly reuse existing apps’ account, payment, and risk-control systems, so vendors do not need to maintain hundreds or thousands of interfaces, but having an Agent interpret images and click through screens is slower and more token-intensive. APIs are faster and more accurate, yet account binding, ordering, and payment after search remain mostly disconnected. 刘潇 expects both approaches to coexist, eventually forming a network in which personal Agents and vertical Agents negotiate through Agent to Agent interactions.
- Payment is not a minor end-of-funnel feature but the key bottleneck determining whether mobile Agents can move from demos to transaction entry points. AutoGLM currently returns control of the device to users at sensitive steps such as login and payment; this is safer, but can make ordering takeout no faster than doing it yourself. Possible improvements include payment notifications, voice confirmation, automatic payment below RMB30, and Agent wallets with limits on amount, frequency, and product category. 刘潇 also suggested that Agent payments might need to be easier to reverse and refund, while repeatedly stressing that “the specific mechanism is still uncertain.”
- 智谱’s central bet on capability gains is end-to-end online reinforcement learning: give the model a real environment and reward only the final outcome, rather than requiring it to imitate humans step by step. Early models scored about 20 on WebArena, versus roughly 70–80 for humans; 智谱’s current model reaches around 48% on OSWorld, which 刘潇 says “should be the best among single models,” although humans score about 70% and the benchmark is “still relatively elementary.” His condensed view is: “What is mainly stopping AI from learning how to do something right now is the environment and the reward.”
- A per-task cost of about $0.2 already permits large-scale experimentation, but has not yet proved a commercial loop. 刘潇 estimates that one Google search costs about $0.02, while noting that the accounting bases may not be fully comparable; he believes Agents create more value, but it remains unclear whether that value will be captured through subscriptions, advertising, or transaction commissions. AutoGLM 2.0 currently has no explicit usage budget cap, and if demand becomes too high, server congestion may be the more realistic constraint. Scale, smaller models, and inference optimization should continue lowering costs, but where they ultimately land “can only be observed once the market is running.”
- The battle for the mobile Agent entry point will involve foundation-model companies, super apps, handset makers, and new hardware simultaneously, but 刘潇 believes the decisive variable will still be whose intelligence is strongest. A future market may contain a small number of personal Agents, service Agents embedded in individual apps, and device Agents; 智谱 is supplying stable versions to handset makers while using its To C product to pursue harder, longer tasks and iterate every 1–2 weeks. 刘潇’s practical lower-bound definition of AGI is an Agent that can run stably 24 hours a day and handle life and work at the level of an ordinary colleague or assistant: “AGI has already arrived; it just isn’t evenly distributed.”
Deep dive
1. AutoGLM 2.0 Has Shifted from a Local Remote Control to an Asynchronous Cloud Executor
程曼祺 first highlights the scale of the mobile market: in China, the number of people who use phones every day is almost equal to the total population, while daily computer users are far fewer. Mobile Agents could therefore reach a broader audience than web- and computer-based Agents such as Deep Research and Devin, making them a strategic battleground for foundation-model companies, internet platforms, and handset makers.
Last October’s AutoGLM mainly ran on the user’s own device, competing for control of the screen while the Agent was working—a particularly painful experience for heavy users. AutoGLM 2.0 now uses a two-sided architecture: it can control both a cloud phone and a cloud computer, allowing tasks to run asynchronously in the background for extended periods.
The launch has also moved from a limited beta to broad availability: iOS, Android, and the web are all supported, with no invite code required. 刘潇 says there are currently no plans to charge, because the priority is to let users experience “this kind of AI capability that is truly autonomous” for free.
2. One Cloud Phone Handles Lifestyle Services; One Cloud Computer Handles Knowledge Work
The phone side covers tasks that existing apps can perform: ordering takeout, finding restaurants, booking cleaning or plumbing services, and checking nearby information. It can also collect content from platforms such as Xiaohongshu, then filter and act on it according to the user’s requirements.
The computer side is aimed at longer production tasks: Deep Research, PPT creation, data analysis, building websites and designing cards with code, generating podcasts, videos, and images, and continuing to publish multimedia content to social platforms.
The system does not rewrite a separate tool for every service. It lets the model read interfaces and simulate human clicks and text input. 刘潇 stresses that its generality comes from “representing human intent in execution”: in theory, any app installed on the device can be operated.
3. The Security Architecture Preserves Login State First, Then Returns Sensitive Actions to the User
刘潇 compares the cloud device to a new phone: users must log in themselves, while the cloud-device partner preserves the login state. 智谱 does not record or know users’ account credentials.
Control of the device is fully returned to the user for critical steps such as login and payment. Users can also take over the phone at any time, log out, or check its status. AutoGLM does not attempt in the current version to unconditionally complete transaction confirmation on the user’s behalf.
The Agent reads the screen only while a task is running automatically. Once the user takes over or the task ends, it stops reading. 刘潇 says data is strictly de-identified before entering storage and is cleaned and encrypted, with the specific implementation governed by the privacy policy.
4. Payment Confirmation Protects Users but Exposes the Product’s Core Friction
程曼祺’s field-test objection is specific: if users still have to enter the payment page themselves at the end, AutoGLM may not order takeout any faster than manual operation. Once a task is sent to the background, users may also forget to confirm and discover half an hour later that the order was never placed.
刘潇 acknowledges that this is “a very important area for the next iteration.” A near-term solution would be a notification or pop-up at the critical step, allowing users to say “pay for it” by voice without re-entering the full process.
A more advanced design would let users set boundaries—for example, “don’t ask me again for orders below RMB30”—or give the Agent a separate wallet with limits on spending, frequency, and product categories. Agent payments could also be designed for easier reversal and one-click refunds, but the mechanism remains undecided.
5. Proactive Tasks Are Not Live Yet, While WeChat Is Feasible but Not a Priority
The feature 刘潇 wants most but cannot currently deliver is scheduled and proactive execution: if he has not woken up by 9 a.m., the Agent would order coffee automatically so he could pick it up outside the office at 9:30 without issuing another instruction.
WeChat operations are technically feasible, including checking Moments on explicit instruction, but the team has not promoted them heavily. WeChat contains more privacy-sensitive information that is harder to de-identify, and the intent and wording of each message vary; for many messages, typing them yourself still takes only a few seconds.
He draws a simple boundary around the use case: actions that take a few seconds or ten seconds are not worth forcing through an Agent. The high-value zone is when users are driving, doing housework, showering, or running and cannot use their hands, or when the task involves booking travel, researching information, or planning a family weekend.
6. “What Should I Eat?” Could Become the First High-Frequency Mobile-Agent Use Case
刘潇 believes the easiest way for ordinary users to remember an Agent is through the daily problem of deciding what to eat. That includes ordering takeout, but also finding restaurants, choosing where to meet friends, and locating something they have not tried that matches their current tastes from a large pool of recommendations.
He describes the intensity of the need through his own weekend routine: “You wake up on Saturday afternoon, lie in bed without knowing what to eat that night, scroll for 2 hours, and still reach no conclusion.” The task looks mundane, but doing it seriously requires cross-platform search, comparison, and memory.
An Agent could search Xiaohongshu for a particular type of food and then place an order on Meituan. With explicit authorization, it could also check what friends have recently eaten in Moments. 刘潇 summarizes the opportunity this way: existing recommendation systems do not sufficiently cover what users actually have in mind, and people often choose braised chicken with rice simply because they are “too lazy to keep looking.” When platform recommendations disappoint, users have few better options.
7. General-Purpose Agents Will Reallocate Attention, but May Not Reduce Platform Visits
程曼祺 raises the platforms’ most direct concern: when users access Meituan or Ele.me through a cloud phone, the person actually browsing information is no longer human. The app could lose the value of attracting attention, displaying ads, and managing traffic.
刘潇’s response is that Agents may even visit platforms more often. They would still obtain information from the platforms and faithfully relay recommendations or ads from the feed. He calls himself “a porter of information”; the difference is that future filtering will incorporate the user’s choices from yesterday and today.
程曼祺 also imagines having AutoGLM compare milk-tea prices and coupons across JD.com, Ele.me, and Meituan. 刘潇 replies that the need to compare already exists: even without an Agent, users may open all 3 platforms themselves. The Agent simply automates the behavior.
8. GUI’s Biggest Advantage Is Reusing the Entire Mobile Internet, Not a Single Interface
The GUI approach lets the model read and operate graphical interfaces like a human, receiving information broadly equivalent to what an ordinary user sees. Whatever service an app offers to humans can naturally be offered to an AI representing humans, without separate adaptation.
For platforms, this is a low-burden integration path: they do not need to build or maintain APIs, nor determine whether a visitor is a human or an Agent. For Agents, GUI is especially suitable for long-tail, complex, and varied tasks.
The cost shifts to the Agent developer: the model must understand screenshots, identify exact locations, and maintain virtual devices, while visual processing consumes more tokens. 刘潇 does not avoid the speed issue—GUI is “definitely still going to be a little slow” for now.
9. APIs Are Faster and More Accurate, but Accounts, Payments, and Risk Controls Remain Disconnected
The advantages of APIs are straightforward: information is provided in structured form by the official source, making the process faster and more accurate. Alongside GUI, 智谱 is also pursuing formal API and MCP partnerships with third parties.
The real difficulty comes after search. Even if a platform opens a product-search API, there is still no natural loop for an AI account to bind to the user’s account in the shopping app, place an order, and process payment. “You need to rebuild the user system and rebuild the payment chain”; this is not a hurdle one app can easily clear.
Platforms also worry about API abuse: map POIs, merchant data, and dish data could be scraped in bulk, while the existing risk controls based on user accounts and GUI behavior cannot be reused directly. Platforms want the traffic AI could bring, but remain concerned that API risk control and payment are unresolved.
Maintenance costs are also significant. Whenever a product attribute or function changes, the interface must be updated; a large app could eventually accumulate hundreds or thousands of APIs and must continually mark which ones are active or obsolete. 刘潇 believes this may not be simpler than building an entirely new app.
10. GUI and API Will Not Be an Either-Or Choice but Long-Term Complements
Platforms remain willing to discuss APIs because they can deliver information to AI faster and turn questions that previously ended in advice into service opportunities. If a user asks what to do about a broken toilet, an Agent could recommend a platform and directly call an on-site repair service rather than merely provide instructions.
Large apps have more resources to build interfaces, but they also face more complicated internal business considerations and may build their own Agents first rather than immediately open APIs to outside partners.
刘潇 uses autonomous driving as an analogy: “Autonomous driving cannot possibly eliminate all human drivers,” and GUI and API will be similar. GUI also provides an interpretable sense of security—the user can see how the Agent makes its choice, rather than watching an order appear instantly behind an opaque interface.
11. General-Purpose Agents Stand with Users; Vertical Agents Stand with Services
AutoGLM aims to understand user context across scenarios and manage the user’s life and work as a whole. Vertical Agents from Meituan, Amap, and similar platforms are better suited to users whose intent is already fixed: “I specifically want to use your app to do this.”
刘潇 believes vertical Agents can activate capabilities inside an app that are difficult to use and improve the quality of existing traffic, but they may not generate incremental traffic because the user still has to open the app first. A general-purpose Agent can discover a service externally and then send the demand to the platform.
His proposed interaction model is simple: the personal Agent is the user’s private assistant, while the platform Agent is the service provider’s front desk. The assistant tells the front desk, “My principal wants to do something,” the front desk completes the task and returns the result, and the general-purpose Agent delivers it to the user.
12. Agent to Agent Will Form a Service Network Hidden Behind the Social Graph
刘潇 imagines every person and every app having its own assistant. They would automatically negotiate, search for information, and prepare proposals, while humans would only need to say, “This one looks good,” or, “No, find me another option.”
Even arranging a podcast could be handled by both sides’ Agents reading their owners’ calendars, finding a common opening, and separately requesting confirmation. Much of the preparation would therefore move from repeated human-to-human communication into a service network among Agents.
程曼祺 notes that Google has already proposed an Agent to Agent protocol, while mobile still lacks a mature ecosystem comparable to MCP on the web. She also mentions that Apple is reportedly considering App Intents. At the same time, she warns that GUI is already a unified interface, so building another protocol would carry a cost and depend on whether platforms are willing to migrate.
13. The Most Valuable Position Is the Agent Closest to the User
程曼祺 believes every participant wants to become the entry point closest to the user because that Agent can compare services on the user’s behalf, distribute demand, and control long-term cross-app context. This is a different kind of power from optimizing the experience within a single app.
刘潇 insists that the two roles can coexist: users need an Agent that “always thinks about your personalization and your value,” while platforms also need their own service Agents. How the two cooperate will have to emerge through actual use.
Near the end of the episode, the competition expands to hardware. The web may accommodate many knowledge-work Agents, while mobile is more concentrated; if glasses build an Agent-oriented operating system from scratch, there may be fewer direct user-facing entry points on each device and the competition could become even more intense.
14. The First Skill Ordinary Users Need Is Not Prompting but Managing an Agent
刘潇 observes that even after Agents have developed for 6–12 months inside the AI community, they remain abstract to people outside it. He usually describes one as an assistant working 24 hours a day: “You be the boss, and it becomes your employee.” The most common response is: “Then show me one I can use.”
One reason for the abstraction is that most people have little experience directing other people’s work. He even believes that “how to order an Agent to work for you” may become a skill children start learning in primary school; without it, people will struggle to use AI fully.
Copilot partially amplifies the output of a particular hour. An Agent that can work in parallel could expand a day from 24 hours to 48 hours, with multiple Agents creating further multiplicative effects. The prerequisite is not that the Agent perform familiar actions faster than the user, but that the user can effectively divide, delegate, and review the work.
15. 智谱 Is Pursuing To B Partnerships with Handset Makers and To C AutoGLM in Parallel
In partnerships with handset makers such as Honor and Samsung, 智谱 mainly supplies the underlying capabilities while the manufacturers build their own intelligent assistants. These versions prioritize product-level stability; once the model version is fixed, it cannot be changed frequently.
Existing handset-maker solutions typically use a cloud model to make decisions while executing on the user’s own device, so they still occupy the screen and are suitable only for short tasks such as sending messages, hailing a car, or placing an order. Their form is closer to last October’s AutoGLM.
The To C product uses cloud phones and cloud computers to handle harder, longer, and more challenging tasks, with versions that can iterate quickly based on user feedback. 刘潇 acknowledges that this route also serves as an algorithm test bed: users “pull” the model forward by giving it more difficult tasks.
16. Cloud-Phone Capabilities Can Reconnect Glasses, Refrigerators, and Other Devices to the Mobile Internet
刘潇 imagines a refrigerator detecting that it has run out of soda and replenishing the supply automatically, or AI glasses seeing an attractive piece of clothing and sending a single instruction to Taobao to find and order the same item. These devices would not need to pack the compute and operating system of an entire phone.
AutoGLM’s developer program and future commercial partnerships aim to let wearables, smart-home devices, and new terminals call on cloud-based mobile capabilities. The Agent would no longer be confined to a dialogue box that requires opening a web page on a computer, but become “ubiquitous.”
程曼祺 asks why a foundation-model startup might accomplish this rather than an established hardware ecosystem such as Huawei or Xiaomi. 刘潇 does not claim that 智谱 has an exclusive opportunity; he says 智谱 is more focused on algorithms and AGI. A system that can connect everything requires a “smarter brain” with both breadth and an understanding of user preferences.
智谱 has discussed developing its own hardware internally, but it is not currently a priority. The company’s more realistic position remains algorithms, cloud capabilities, and partnership interfaces.
17. The First Competition in Mobile Agents Is a Long Market-Education Campaign
刘潇 uses the iPhone to explain why he is willing to release a product that still has substantial room for refinement. The period from iPhone 1 to iPhone 4 took about 3 years of iteration before the product was broadly accepted; technology does not wait until it is fully ready before the market adopts it overnight.
His rough estimate is that perhaps 10%–20% of people have encountered Chatbots, while fewer than 1%—possibly only one in 1,000—have actually used an Agent. For “a considerable number, even 99%, of people,” Agents remain an abstract concept with no tangible reality.
The market therefore cannot be educated by 智谱, OpenAI, or Anthropic alone. Hardware, software, algorithms, and model vendors must grow the category together. 刘潇 welcomes competition not as a gesture, but because the industry first needs to move from demos to sustainable use.
On specific competitors, he only says that OpenAI will do something similar and expects computer-side competition to remain intense; he does not know the progress of each company on mobile. 程曼祺 believes companies such as ByteDance will not stay out, while 刘潇 says their participation is “inevitable,” although figuring out how to capture the reward will be a painful exploration.
18. GPT-4 First Convinced 刘潇 That Language Models Would Eventually Break Out of the Chat Box
In March 2023, GPT-4 and ChatGLM launched almost back to back. 刘潇, who was working on pre-training, post-training, and alignment, found that GPT-4 could already search Reddit posts and filter specific products on Amazon—even when almost nobody was discussing Agents. Its success rate was only “about 50–50,” but the direction was clear enough.
The key development was not adding search to a language model, but turning a single step into multi-step interaction. Web content and UIs change dynamically, yet the model can still use each new observation to choose the next step and complete 4 or 5, then 5 or 6, consecutive actions.
This differs from traditional RPA, which relies on fixed positions. When browser dimensions change or elements are added or removed, a model can still “find that point of certainty amid the mess,” and may even recover after a wrong click. The early implementation used Playwright to capture HTML and screenshots, then asked the model to locate the elements.
19. AgentBench Turned Intuition into the First Comparable Capability Gap
刘潇 believes GPT-4’s ability was largely generalized from internet data. Step-by-step knowledge such as WikiHow represented only a small share of the corpus, yet the model’s parameters and accumulated knowledge somehow allowed it to learn how to search and operate.
From April to August 2023, he led a team in building AgentBench, which he calls the world’s first comprehensive benchmark for evaluating foundation-model Agents. It included 8 dynamic environments to quantify task-completion accuracy.
The results showed that well-trained models could reach roughly 30%–40% through prompting alone, with GPT-4 still the best at the time. Many open-source models claimed performance close to ChatGPT, but on AgentBench even the best open-source model could not beat Google’s worst-performing Text Bison model at the time; 刘潇 no longer remembers the specific version numbers.
The secondary-priority research then became his main direction. The team grew from a handful of people to dozens, while demos, leaderboard milestones, market feedback, and commercial certainty successively attracted more algorithm, engineering, and product resources.
20. Agent Capabilities Must Be Prepared from Pre-Training, Not Assembled Only at the Product Layer
AgentBench convinced 智谱 that reaching AGI would require pre-training, post-training, algorithms, and applications to advance together. 刘潇 describes this as building foundation models “in the manner of heavy industry” to produce systems that generalize and work out of the box.
By GLM-4.5, Agents had become a company-wide consensus. The team began considering from the pre-training stage how to make the model good at Agent tasks rather than patching in the capability during post-training. 刘潇 attributes the performance to coordination across pre-training, post-training, and application teams.
The data path is model as agent: first use engineering Agents to generate trajectories, select the parts that completed tasks well, and transfer that experience back into model training. Agent data on the internet is scarce, so systems and models form a synthetic-data flywheel—a “chicken-and-egg” loop.
21. AutoWebGLM’s Cold Start Used Human Demonstrations to Bridge Planning and Execution
From August 2023 to early 2024, the team started with browser use through AutoWebGLM. Early GPT-4 had only a 30%–40% success rate and could not reliably generate its own data, so the model first had to imitate human experts browsing the web.
Materials such as WikiHow were “not very useful” because they described high-level planning—searching, entering a category—but did not tell the model where the search box was or which button expanded the category. Humans filled in those execution details through common sense; AI could not yet do so automatically.
Screen localization had to be extremely precise: clicking slightly outside the target area would fail the task, and the data also needed to be sufficiently diverse. The team combined semi-automation with automation, having AI generate candidate points for humans to select, while algorithm engineers personally built the earliest annotation system.
22. Offline Reinforcement Learning Improved Imitation but Did Not Teach Models to Take Responsibility for Outcomes
AutoWebGLM experimented with relatively simple, offline-oriented reinforcement-learning strategies, including DPO, and on public benchmarks such as Mind2Web and MiniWoB, as well as internal tests, it nearly surpassed GPT-4’s prompting performance on almost every web-browsing task at the time.
刘潇 explains that ordinary SFT gives every step the same weight, even though scrolling and key decisions clearly matter differently. Offline RL estimates the advantage of each step and dynamically adjusts learning weights, producing a better fit to human trajectories. He cites behavior cloning as a typical offline method.
The limitation follows directly: the data is static, so the model is only imitating humans, and may learn both correct and incorrect behaviors. A trajectory that resembles human operation does not mean the model can independently execute the task and achieve the same outcome.
23. WebArena Proved That Operating Like a Human Is Not the Same as Getting the Job Done
WebArena turns shopping and other websites into offline-deployable replicas. It does not compare intermediate actions, but checks the final state: if the task is to buy red shoes, the system only checks whether a pair of red shoes actually appears in the paid orders.
In this outcome-oriented evaluation, models scored only about 20 at the time, versus roughly 70–80 for humans. Other static benchmarks looked better because they mainly compared how closely AI predictions matched human trajectories.
程曼祺 summarizes the point sharply: being slightly wrong at every step creates a large error by the end. 刘潇 emphasizes that the same task may be completed through search or categorization; users do not care about the route, only delivery. This kind of evaluation is much closer to product intuition.
24. Online Reinforcement Learning Trains Recovery from Errors, Not the Ability to Never Make Them
刘潇 uses Apple’s troubleshooting guidance as an example: even professional instructions can be poorly written, forcing people to try repeatedly. In the real world, humans are not correct at every step; they detect mistakes, backtrack, change strategies, and continue until they find a workable path.
If the model is given only correct trajectories, it becomes highly confident and assumes every step it takes is right. Adding fixed counterexamples can simply confuse it. An online environment allows the model to make real mistakes and learn from subsequent feedback when to backtrack and when to retry.
His philosophical summary is that humans themselves are intelligent agents evolved in a complex planetary environment, with ancestors paying for failed reinforcement learning with their lives. Without practicing, an Agent will also struggle to develop the ability to handle a complex world.
25. O1 and DeepSeek Reinforced the Priority of Outcome Supervision
刘潇 believes the industry initially placed too much faith in process supervision, hoping models would produce concise, elegant, textbook-like proofs with every step correct, while rejecting hundreds of thousands of tokens of edits, attempts, and reversals.
Work on O1, DeepSeek, and similar systems quickly taught the industry that “we need to supervise the model through outcomes, rather than supervise the model through the process.” Closed domains such as mathematics and code first demonstrated that the route was viable.
Agents are harder. Theorems and compiler rules are relatively fixed, while real websites continue to change; the model must actually open the page to know whether a button is correct. Engineering environments, multi-step decisions, and sparse rewards are all substantially more difficult than single-step reasoning.
26. The Industrial Scale of Environments Determines How Much Practice an Agent Can Get
刘潇 compresses online reinforcement learning into 2 conditions: environment and reward. If a real environment and suitable incentives can be built, many tasks can be solved “to a certain extent” through RL, but neither has a standardized SOP.
Training cannot rely on a single computer. If one task takes 10 minutes, sampling 5,000 trajectories once requires 50,000 minutes. To compress iteration into an acceptable timeframe, roughly 1,000 virtual computers must explore in parallel.
This requires containerized simulation, virtual devices, scheduling systems, and algorithms to work together; nobody can simply bring in 1,000 physical computers. The team has internal algorithm and infrastructure staff and also relies on external cloud partners. On the computer side, it works with Alibaba Cloud’s Wuying, drawing on years of cloud-computer experience.
刘潇 uses this to revise the simple narrative that model progress alone can drive AGI. The model remains the core, but once Agents connect to everyday life and social services, engineering, product design, and cloud infrastructure will “matter” increasingly more.
27. The Hard Part of Reward Is Turning Expert Judgment into a Verifiable Signal
When training on real platforms such as Taobao, 智谱 cannot access the back-end database and can only infer task success from screen trajectories. Deep Research, industry research, and PPT creation are even harder: whether information is material, research is sufficiently deep, or a layout looks good all require professional expertise.
One construction method is to find a determinate answer first and then generate a complex question in reverse. For example, the team can find a basketball team’s score in a particular game, on a particular date and at a particular location, from a webpage or knowledge graph, then ask the Agent to navigate JavaScript interactions and database pages to retrieve the answer. A direct search is unlikely to hit it.
Humans remain indispensable in reward design because “although sometimes we don’t know how to do something, we can still tell whether you did it well.” Someone who cannot cook can judge the taste, but specialized areas such as classical Chinese still require experts in ancient Chinese to identify errors hidden beneath an apparently “period-appropriate” style.
刘潇 uses a humorous analogy for the model’s level: it may be “the best violin player in table tennis.” In an expert’s own field, it may not beat the expert, but in fields unfamiliar to the expert, it can easily outperform ordinary people. Annotation work is therefore becoming more specialized, not more mechanical.
28. AutoGLM 2.0 Turned 10 Months of Online Training into Visible Self-Correction
The internal version from late October 2024 already used an early online algorithm, but its concurrent environments were in the single digits and updates were slow. It could complete only short chains and single-app tasks; cross-app execution, complex intent, and multi-round initiation still lacked stable model-level methods.
The team then expanded concurrent training and simulation environments while completing the asynchronous engineering chain across cloud phones and cloud computers. The product no longer occupies the local screen and can handle longer tasks around the clock.
The clearest model improvement is not success along an ideal path, but recovery from anomalies. After clicking the wrong entry point, it backs out; when a page fails to load because of network problems, it tries again. Last year’s version might have “kept clicking the wrong thing without stopping.”
刘潇 is not anxious that Manus became widely known first in March this year. He believes it improved product communication and market education. 智谱 had discussed Agents earlier but had not clearly told users, “This thing is ready to use.” A shared industry breakout lowers the cost of explanation.
29. Capability, Infrastructure, and Ecosystem Standards Still Face an Entry-Level Barrier
刘潇 says 智谱’s current score on the OSWorld computer-operation benchmark is around 48%, which should be the best among single models; humans score about 70%. More importantly, OSWorld is still “relatively elementary,” and harder benchmarks will appear after it is solved.
Beyond model capability, bandwidth, carrier infrastructure, and the robustness of cloud virtual devices must improve. Platforms also need new standards: a device could proactively declare, “This is an Agent,” allowing an app to restrict unwanted behavior while permitting normal operations that generate orders and traffic.
On the API side, accounts and payments need to be designed jointly by Agent companies, applications, and payment institutions. 刘潇 says everyone is interested but “still has no idea where to start,” because there is no reference implementation showing how users want to authorize Agents or where they will actually use them.
30. At $0.2 per Task, Large-Scale Experimentation Is Possible, but the Business Model Is Still Unformed
AutoGLM currently costs about $0.2 to complete a single task. 刘潇 estimates that one Google search costs about $0.02, while warning that the accounting bases may not be fully comparable. He believes the value of an Agent interaction is clearly higher, but it remains unclear whether the commercial loop will come through subscriptions, advertising, or something else.
He does not believe Agents must fall to search-level pricing, although scale should continue to spread costs over more users. Smaller models, engineering optimization, and inference teams are already contributing. When discussing GLM-4.5’s input price, he only remembers that it was either $2 or $4 per million tokens: “I don’t remember.”
The free launch currently has no explicit budget ceiling, but resource volume will impose a natural limit. If too many users arrive, servers may become busy. He refuses to predict how far costs can fall: “That can only be observed once the market is running.”
The biggest pre-launch concern was engineering stability, because the number of companies able to run the full chain end to end “could probably be counted on one hand.” What he most wants to see is users discovering scenarios the team never anticipated and new hardware such as AI glasses connecting quickly.
31. 智谱 Is Feeding Post-Launch User Feedback Directly into Its AGI Roadmap
The new version removes the biggest barriers to use. The old version required both an invite code and complicated permissions on different Android devices, which “90% or 99% of users” would probably never configure. Now all major terminals can directly access cloud-based tasks.
The team plans to add new features every 1–2 weeks while seeking developers, applications, hardware makers, and payment partners. What 刘潇 most wants to know is not whether users like the preset demos, but how they actually use the product to reshape their lives and work.
His product philosophy is that users often do not know what they want before a product exists. The company therefore has to release a version that “indeed still has many shortcomings,” allowing real demand to shape the algorithms and engineering in return.
32. The Final Market May Contain Many Agents, but Users Will Directly Manage Only a Few
刘潇 expects 3 types of Agents to coexist: personal Agents standing with users, vertical Agents representing apps and service providers, and device Agents embedded in glasses, home appliances, and other hardware. An individual may even have 3 or 4 assistants managing different areas.
程曼祺 applies a management constraint: a boss should generally not directly manage more than 8 people, and the number of frequently used apps is also limited. Users may therefore interact frequently with only a small number of Agents, keeping the battle for entry points extremely intense.
刘潇 calls these commercial structures “a by-product of AGI.” If every system becomes sufficiently intelligent, the decisive variable may simply be whose AGI is stronger, because users will choose the one that is “smarter, more capable, and wiser.”
33. An Ordinary Assistant That Works Around the Clock Is 刘潇’s Lower-Bound Definition of AGI
刘潇 does not view AGI as a point crossed in a single instant, but as a range. A few people first feel that AGI has arrived, then everyone becomes convinced it has happened, while model capability and social adoption continue spreading in between.
His lower-bound standard is concrete: if an Agent can run autonomously and stably for 24 hours and handle users’ life and work at the level of an ordinary colleague or assistant, “I think it has already touched the lower bound of AGI.”
In AutoGLM’s dark interface, orange-yellow light gradually spreads. 程曼祺 compares it to dawn reaching more and more people. 刘潇 picks up the image: “AGI has already arrived; it just isn’t evenly distributed.” 智谱’s task is to help more people see and use it first, then push its capabilities toward the upper bound together.