B2B to A2A: Agent Infrastructure for Global One-Person Enterprises
Summary
Zhang Kuo’s core judgment is that roughly $30T in global B2B trade could eventually move from digitalization to A2A; the central proof point for Agents is not chat, but whether they can turn tokens into transaction efficiency and customer ROI. Alibaba.com currently handles about $70B in transaction volume, but penetration remains low; Accio had reached 10M MAU by March and was still growing rapidly month on month. It has cut the communication cycle from inquiry to a reliable commercial relationship from one week to one day—“to one-fifth of the original.”
The moat for a vertical Agent is not another layer on top of a foundation model, but a closed loop combining accurate data, reinforcement learning, safety and reliability, and ultra-long context. Tiny errors in actual product prices, logistics, tariffs, and landed cost can determine profit; Alibaba’s 26 years of transaction history and more than 1M daily conversations could bring results “90% or 99% close to the facts,” but Zhang Kuo explicitly does not promise 100%. Completed transactions, repeat purchases, and failed experiments then feed back as long-term reward signals.
OpenClaw shows the explosive potential of open Agents, while Cowork looks more like a next-generation workbench for knowledge workers; both still require step-by-step verification. OpenClaw is open and general-purpose, but comes with configuration barriers plus reinforcement-learning and security challenges; Cowork has a clearer target and tighter fit with Anthropic’s models. Zhang Kuo rejects the idea that “one-shot completion” is the definition of AI-native: if an 18-step workflow runs at 90% accuracy per step, the outcome is “90% to the 18th power” and nearly unusable.
AI will not kill SaaS across the board, but the billing basis may shift from seats to usage, while products move from standardized workflows toward low-cost personalization. Zhang Kuo cited an hourly-worker recruiting Agent that can explain dress code, check-in, parking, sick leave, and social-security rules; compared with SMS interactions, it lifts the show-up rate. That suggests AI can also strengthen vertical SaaS. The real dividing line is whether tokens can become ROI that customers can feel, not whether a product carries the SaaS label.
Accio Work targets the full back office of a “one-person enterprise”: research through sourcing may account for only 10%–20% of daily work, while operations, replenishment, customer service, inventory, and multi-platform publishing make up the rest. Zhang Kuo says at least 30%–40% of small and midsize businesses may be solo entrepreneurs; the platform will combine its own research and sourcing capabilities, browser use/computer use on platforms such as Shopify, and third-party sub-agents for HR, payroll, finance, and tax into an out-of-the-box multi-Agent system. The long-term form is a master agent coordinating multiple sub-agents, with businesses both inside and outside the enterprise potentially shifting to Agent to Agent.
Generative search will compress ad inventory, but not necessarily ad revenue; the ads that survive will be explainable, tightly matched performance ads. Hong Jun noted that shrinking results from 1,000 to 5 necessarily reduces traditional impression inventory, and Zhang Kuo agreed that “display advertising clearly won’t mean much.” His counterargument is that richer intent and multimodal inputs will expand query volume, while candidate ads that genuinely meet requirements for price, certification, capacity, and so on could see higher click-through and conversion.
Revenue will still develop along three tracks: high-value tokens, precise supply-demand matching, and supply-chain services such as payments and logistics. Accio may use usage-based billing with a free allowance or monthly cap; advertising will remain primarily performance-based; guaranteed-transaction take rate is only about 1–2 percentage points, with payment terms and logistics added on top. The winner will not be whoever offers the cheapest tokens, but whoever delivers the highest “intelligence density per token” and economic value.
The worst metrics for measuring AI transformation are tokens or lines of code; the most sensitive signal is whether the organization feels “excited, anxious, or nothing” when a new model appears. Alibaba.com generates about 300 ideas per quarter, launches 150, and ends up with about 50 that work; AI’s goal is to increase the number that create real business impact. Zhang Kuo also said EBIT should reach roughly 18% this year and that growth remains rapid versus last year. Excitement means new capabilities can solve old problems; anxiety suggests it may be just a wrapper, while “the worst outcome is that you feel nothing.”
Deep dive
1. Silicon Valley’s Advantage Is Not the Absence of Anxiety, but a More Finely Divided Ecosystem
Zhang Kuo observes that the U.S. AI startup ecosystem is not only concentrated in application companies; model access, inference cost reduction, and building tools have also formed multilayered markets. Companies such as Together AI and Fireworks provide model infrastructure to upper-layer applications, while vertical teams are “running dozens of meters ahead of the bulldozer,” surviving by maintaining that lead.
Speech-to-text is one such independent niche. Alibaba.com and Accio are also adopting products such as Wispr Flow, because multimodal query volume could grow by roughly 100%: users are increasingly willing to send voice, photos, or video instead of entering only keywords to find products.
SaaS has not become inherently obsolete because of AI. Zhang Kuo cited an hourly-worker recruiting product serving supermarket chains and restaurant businesses that uses an Agent to explain dress code, check-in, parking, sick leave, and social-security rules to blue-collar workers; compared with the former SMS interaction, it delivers a higher show-up rate. AI directly improves the existing product here.
A Silicon Valley pitch night can draw 500 or even 1,000 people, and anxiety is no less intense than in China; but once outside Silicon Valley, customers do not care about abstract token economics. Small and midsize businesses ask first, “What value does this actually create?” The industry still has to prove how underlying tokens become real ROI.
2. Whether OpenClaw Outlasts the Hype Depends on How Many Workflows Stick
Zhang Kuo summarizes the GTC narrative as token economics: compute and cost required per token are falling exponentially, spawning more applications; OpenClaw is merely one of the most conspicuous products of the recent wave.
OpenClaw is open source and general-purpose, and can connect to multiple models and tools, but it is not truly out of the box. Configuration still requires technical ability; its openness makes it hard to define “what it does well and how to make it remember that it does it well,” while reinforcement learning and security require long-term tuning.
Zhang Kuo sets a retention test for the hype: if users can institutionalize durable workflows, it has truly solved a problem; if “very few workflows stick,” the tide may recede after the hype. Open architecture itself is not a substitute for commercial value.
By contrast, Cowork, still in the research preview stage, already shows the outline of a next-generation workbench or Agent platform: its users are concentrated among programmers, analysts, and knowledge workers in research, finance, and law; its layered architecture is clear, and its fit with Anthropic models is stronger.
3. Enterprise Agents Do Not Eliminate Steps; They Make Every Step Verifiable
Hong Jun draws a sharp line: OpenClaw seems better suited to To C tasks in a chat box—scraping webpages, transcribing, and cleaning up audio—while an enterprise embedding AI into podcast production must define steps 1, 2, and 3 and guarantee quality at each step. Cowork therefore looks more like a To B work pattern.
Zhang Kuo’s correction is that one-shot completion versus step-by-step is not the core distinction. The most agentic products are “the tool of all tools” (所有工具的工具); different professionals will use them to create their own workflows, define success by their own expertise, correct intermediate steps, and then have the system remember what counts as acceptable.
His error example supplies the key constraint of the conversation: if podcast production has 18 steps and each step carries a 10% error rate, the end result is “90% to the 18th power” (百分之九十的十八次方). Core nodes therefore must be verified; only after learning can an Agent complete later reasoning and execution with fewer interactions.
Acceptable thresholds vary with stakes: collecting daily news only needs to clear a preference threshold; using it for quantitative trading demands more; if information is to guide production over the next year and affect overall ROI, the model, tools, and working methods all must be redesigned for the industry.
4. Alibaba.com and Accio Are Pursuing Two Parallel Rewrites of B2B
The first track is to rebuild the marketplace. Alibaba.com is embedding AI in search, recommendations, communication, transactions, logistics, and payment terms; buyers no longer screen huge volumes of information alone, but let AI drive the order process from expressing intent through online closing.
The second is moving from B2B to A2A. Accio is built around sourcing, had reached 10M MAU by March according to the interview, and was still growing rapidly month on month; it extends forward into research, ideation, and product design, and backward across supplier screening, communication, transactions, transportation, and after-sales.
For small and midsize businesses, choosing the next killer product often determines the direction of cash flow. A wrong choice means high trial-and-error costs; only after the right one is chosen can the company scale from one store and one product. Accio is therefore trying to automate not simple product search, but the full product-definition process.
From inquiry to establishing a reasonably reliable commercial relationship, Zhang Kuo says the time has been cut to one-fifth of the original: sourcing communication that used to take a week may now be completed in a day. AI can also turn rough ideas into a design pack containing images, text, and 3D content, reducing friction from language, time zones, and specialized expression.
5. Vertical Models Must Solve Facts Before Generating Smoother Answers
Product selection starts with indexing fragmented information across the web: comparable products; search and transaction performance on Amazon and other platforms; reviews, sales volume, price bands; and potential margin. Web 2.0 tools can provide data, but struggle to ensure sufficient breadth, timeliness, and accuracy after reading each page.
Zhang Kuo summarizes Accio’s extra investment in four areas: data accuracy, effective reinforcement learning, safety and reliability, and ultra-long context. Hong Jun notes that general-purpose models are already very strong at deep research, but these B2B constraints still require extensive engineering and tool combinations beyond the model.
Most critical are the actual product price and landed cost: the platform must identify hidden prices or additional fees from conversational context and contracts, then factor in logistics and tariffs. A slight difference in price can determine a merchant’s profit, so “the best solution in the market” cannot be built on hallucination.
Alibaba.com’s millions of daily conversations around design and technical details, together with 26 years of accumulated history and internet-indexing tools, provide signals from ideation to the design pack. Zhang Kuo carefully says 100% accuracy cannot be promised, but getting “90% or 99% close to the facts” is already clearly better than unconstrained generation.
6. In B2B, Reinforcement-Learning Rewards Come from the Transaction Loop—and Arrive Very Slowly
High-stakes tasks require getting each link in the chain as right as possible first. Users continuously report whether a design branch is good, whether the technology is feasible, and whether the margin works; the platform records the logic chain, creating more granular process supervision than a one-off like.
Zhang Kuo wants to reward “more value per token,” not aimless reasoning and search. If the same commercial task can be delivered through a shorter chain with a higher completion rate, unit-token economic value and system efficiency genuinely improve.
Alibaba.com’s distinctive signal comes at the outcome end: whether an idea ultimately converts, whether the customer continues purchasing, or whether several attempts show that it “doesn’t work.” The more industries and users there are, the more chances research, supplier matching, and execution workflows have to be corrected through closed-loop feedback.
Asked whether every industry needs a different reinforcement-learning method, Zhang Kuo did not pretend to have an answer; he only confirmed that To B and To C differ substantially. Logistics, tariffs, delivery, usage, and repeat-purchase signals run on long cycles, so the team is investing more in long-context reasoning; there is not yet enough information to draw conclusions at the industry level.
7. Security, Rollback, and Memory Are Not Side Features; They Are the Core of B2B Agents
B2B tasks are both serious and expensive. The platform must assess whether information is complete and whether each reasoning step is rigorous, while protecting data security in layers. When an Agent accesses internal enterprise systems, user data must be protected and sandbox isolation enforced.
When a long-chain reasoning process goes wrong, the system must roll back to an earlier state while preserving context continuity. Reliability here is not making the final answer look more plausible; it is locating the error node, restoring state, and rerunning execution.
Accio began as a browser-based Agent; the next version, Accio Work, runs on the desktop, expanding from design and sourcing to daily operations. The first phase may take about one month, while product sales, replenishment, and after-sales continue for six months or even a year, making the context nearly “infinite.”
Memory must be layered: the current conversation feeds immediate reasoning, multimodal design materials build an index, other information is accessed as needed, and customer feedback is distilled into the next round of product optimization. The system should not permanently stuff everything into the context window; it must know when to retrieve which kind of evidence.
8. Accio Work Aims to Become the Global Operating Back Office for the “One-Person Enterprise”
Product capability has three layers: proprietary Agents handle research through sourcing; general-purpose computer use and browser use help merchants open stores, list products, and operate them on Shopify and other platforms; HR, payroll, finance, and tax connect to U.S. vertical products as sub-agents within the platform.
Zhang Kuo says at least 30%–40% of small and midsize businesses may be solo entrepreneurs. Spending 3–4 days designing a product is acceptable, but the business then has to publish across multiple stores and social-media platforms, respond to customer feedback, and manage inventory and replenishment; these repetitive operating tasks are what keep consuming a one-person enterprise’s time.
His workload breakdown is that the front end from research to sourcing may account for only 10%–20% of daily work; the remaining roughly 80% is operations. Accio Work aims to have multiple Agents collaborate on most back-office tasks, with humans providing standards and limited interaction only at key points.
Customers are not limited to micro-merchants: Walmart and Amazon also source products such as packaging through the international site; the customer base stretches from individual operators, single stores, and chains to large enterprises. The core audience remains physical small and midsize businesses that start from real problems and grow gradually through global supply chains.
9. SMBs Want Out-of-the-Box Outcomes, Not Agent Components
General-purpose platforms expose skills, hooks, agents, connectors, and plugins. That means freedom to technical users, but a configuration burden for nail salons, restaurant chains, and small back offices. Hong Jun notes that many brick-and-mortar operators who have never touched AI can get stuck even at the step of setting up OpenClaw.
Agentic products change this by letting small businesses enter their own success criteria and have the model generate bespoke product-definition, launch, marketing, and social-media workflows. In the past, only large enterprises with scale or budgets in the millions or tens of millions of dollars could typically afford dedicated customization; now marginal cost could fall sharply.
Hence Zhang Kuo repeatedly narrows the product objective: “out of the box,” solve critical problems, and deliver ROI within a reasonable range. The openness of the underlying architecture is not itself a reason to buy; what matters is configuring one fewer tool and maintaining one fewer system.
10. Pricing Will Shift from Seats to Usage, while the Platform Keeps a Dual-Track Revenue Model
Accio’s first charging model is token-based: tasks run by its own tools consume tokens, and tax, HR, payroll, and finance partners connected as sub-agents consume the same allowance. Zhang Kuo gives only an uncertain example: it might set a cap of about $20/month, or offer a free allowance and let users increase their plan based on usage.
He sees this as a potentially stronger foundation for the next-generation business model than traditional per-seat pricing; price ultimately needs to correspond to actual usage and value created, not simply the number of accounts.
The second track extends the marketplace business model, including fees for advertising, payments, logistics, and other supply-demand services. Accio Work is therefore not pure subscription software; usage revenue runs alongside the existing transaction infrastructure.
11. Generative Search Reduces Candidate Count and Pushes Ads toward Tight Matching
AI lets users express parameters, constraints, and purpose in full context, and upload voice, photos, design packs, or PDFs. Matching no longer looks only at titles and keywords; it also reads the full product context, factory capabilities, certifications, use cases, and suppliers’ private databases.
Zhang Kuo believes better results will raise query and interaction frequency; Hong Jun worries that Accio can already match suppliers to requirements and even trigger an email-contact step. What used to display 1,000 results now needs only 5, so ad inventory will clearly shrink. Zhang Kuo did not evade it: “That understanding is definitely right.”
His boundary is clear: “display advertising clearly won’t mean much” in the future; performance will dominate. Any ad among the 5 results must itself be precise enough, clearly disclose that it is an ad, and explain which criteria it satisfies; irrelevant sellers cannot be forced onto buyers.
Zhang Kuo cites Google’s recent earnings as an observation, saying AI Overview or AI Mode have simultaneously increased time spent, query count, and ad revenue because more precise clicks reduce advertiser waste; he also says Alibaba.com’s search and advertising are two of its faster-growing tracks this year.
12. A $30T Market Is Large Enough to Support Three Forms of Long-Term Value Capture
The first is charging by usage for design, research, and global sourcing. The competitive standard is not absolute token price, but higher “intelligence density per token” and greater commercial value created by each token.
The second is supply-demand matching. There may be 10 or 100 suppliers capable of making similar products; sellers can compete for better-quality demand through price concessions or performance advertising. As long as the match is real, this model can carry into the Agent era.
The third is supply-chain services. Alibaba.com’s guaranteed-transaction take rate is about 1–2 percentage points, essentially like insurance: when a dispute occurs, the platform assesses and compensates first, then seeks recovery from the responsible party. Payment terms and logistics can follow similar pricing logic.
Zhang Kuo’s scale comparison is global B2B at about $30T, versus roughly $70B currently on Alibaba.com. The growth thesis is not to extract a little more from existing customers, but to bring more SMBs into global trade, expand platform scale, and create “incremental GDP and incremental value.”
13. The Data Flywheel and Master-Agent Positioning Form Accio’s Moat
Alibaba’s 26 years of platform accumulation generate millions to tens of millions of conversation, query, and transaction signals each day; on this base, the team conducts mid- and late-stage model training and injects some data into early-stage training. In China, the main open-source model is Qwen; for overseas users, the team tries to use the most suitable SOTA model for each scenario.
Agentic products also have a workflow flywheel: more users leave records of which tools perform better in which situations, which then improves task selection and execution in the next round. Zhang Kuo says this vertical feedback is harder to replicate than simply possessing a general-purpose model.
But target customers have smaller budgets and are more demanding about ROI, which is both a limitation and a source of discipline. The system must continually compress token usage so “intelligence per token” and cost-performance beat products serving high-budget enterprises.
The A2A vision is not one Agent monopolizing everything, but a high-frequency master agent coordinating multiple sub-agents. Zhang Kuo acknowledges that Amazon has a vertical advantage in optimizing its own complex back office; his confidence is limited to multi-platform operations, broad product design, and global supply chains—areas where Accio started earlier and has invested more know-how.
14. The Right Metrics Are Customer Value and Launch Hit Rate, Not Token or Code Volume
Accio first looks at retention: whether users keep coming back for more design and sourcing. Daily tools also need simultaneous tracking of breadth of use, completion for each vertical task, user adoption, and continued use, rather than only whether a feature is called.
Economic metrics are sustained increases in intelligence, value, and cost-effectiveness per token. Zhang Kuo believes total token volume likely will not be a benchmark, because that would encourage teams to “make you waste more tokens.”
Alibaba.com’s original vision was to “make cross-border B2B as simple as online shopping,” but a cross-border transaction can actually be broken into 28 steps involving countries, currencies, payment sequencing, fund security, and delivery. AI’s value is to keep compressing those steps, not hide the complexity inside a single unverifiable answer.
Similarly, code volume is not a measure of AI engineering success: “You can generate the most garbage code in the world without changing any business outcome.” The endpoint remains whether the product is adopted and whether business results change nonlinearly.
15. The Real Stress Test of an AI-Native Organization Is Its First Reaction When a New Model Appears
When hiring model and infra talent, degrees matter less than papers, research, or open-source contributions; interviews can discuss original work directly. Product managers must design products around model capabilities 3–6 months out. PRD, interaction, and engineering roles are also converging, with fewer people accountable for the end-to-end result.
Hong Jun recounts a roughly 3-hour Anthropic disruption: she observed Silicon Valley engineers stop coding and panic. A more mature structure is for a master agent to manage sub-agents handling code writing, document reading, code review, and check-ins, while sandboxing and guardrails control release risk.
The quantitative firm she and Zhang Kuo discussed takes a more aggressive route: researchers discuss model changes in Slack, then @ an Agent to read context, write a Markdown document, and generate code; sometimes a human looks once, sometimes it commits directly under a canary rollout. The key is not to copy the workflow, but to make the Agent adapt to the company’s engineering and risk environment.
Over the past 4 years, Alibaba.com generated roughly 300 ideas per quarter, of which about 150 could launch and about 50 ultimately worked; if AI raises effective outcomes from 50 to 100, that is 100% growth; at 300, it reaches the current human ceiling on ideas. Zhang Kuo also says EBIT should be around 18% this year and that growth remains rapid versus last year, while stressing that this reflects many links “adding up together.”
The final diagnostic is “excited, anxious, or nothing”: excitement means a new model unlocks old problems; anxiety means a product built on the model may just be a wrapper whose value will be squeezed out by the next iteration; “the worst outcome is that you feel nothing” (最差的结果就是你没感觉), like a high-speed train roaring past while the organization keeps driving an old car beside it.
Nano Banana makes the relationship between product-image dimensions and scenes easier to control; computer use, browser use, and long context finally make daily operations viable, so the team has been working overtime on Accio Work since Spring Festival. The difficulty is that Alibaba.com has 26 years of accumulated code—some of 吴泳铭’s early code may still be in the repository; making the full system AI-native will require incremental rewriting.