Interview: Huawei Cloud CTO 张宇昕 on the Internet as AI's Vanguard
Summary
- This AI cycle has turned enterprise technology upgrades from an option into a survival imperative, especially for To B industries that historically sat farther from the Internet. China’s cloud adoption once lagged the US by at least 5 years, but 张宇昕 says at least half of companies—and possibly more than half—are now actively embracing AI. The reason is straightforward: “Every To B company realizes that this time, the wolf is really here.”
- Huawei’s core call is not to copy the US general-purpose model playbook, but to turn China’s manufacturing base into a data advantage for industry models. Of the more than 600 industrial sectors worldwide, China has the broadest industrial coverage, with more than 220 sectors leading globally. Once the technology, experience and engineering capabilities accumulated in those sectors are captured as data, they become AI’s best training sets. That is why Pangu chose AI for Industry, while Huawei Cloud positions itself as the “black soil” in which customers grow.
- Large models are upgrading the cloud from utility-like infrastructure into a platform that brings together compute, data, use cases and ecosystems. Cloud providers are no longer selling only compute, storage and databases; they are helping companies digitize and unlock the value of data warehouses, data lakes and the experience stored in experts’ heads, so machines can assist with analysis, decisions and production. 张宇昕’s view is that for most enterprises and government and enterprise customers, “I can’t think of anywhere other than the cloud” that can organize these elements in a practical and convenient way.
- Huawei Cloud does not view domestic compute as its only moat; the real differentiation is software-hardware co-design plus a To B toolchain. 张宇昕 acknowledges compute as one advantage, but Huawei also supplies that compute to peers and industry customers, so the ecosystem must be expanded collectively. Huawei’s more distinctive capability is to work backward from customer use cases to design systems, hardware and chips. Rather than chase API traffic and general-purpose leaderboards, Huawei provides the foundation and training toolchain for customers to build proprietary models with their own industry know-how: “We only provide the black soil; the trees and flowers belong to the customers.”
- The long-term compute variable is inference rather than training, and 张宇昕 still believes scaling law will continue. “Exaggerating a little,” he estimates that perhaps around 5 institutions each in China and the US will need to keep training large models from scratch. Most customers, by contrast, face inference demand that is far larger than training demand. If AI is to replace basic work, its unit cost must approach or fall below the cost of human labor; on that basis, he believes compute could still grow by tens or hundreds of times over the next 5 or 10 years, but commercialization requires it to be “high-performing without being expensive.”
- Enterprise models will be won on modality, precision and unit cost—not total parameter count. NLP models start at tens of billions or hundreds of billions of parameters and may reach 1T; CV models may need only several billion, while inference, decision-making and scientific-computing models may work with hundreds of millions to several billion. These are not old-fashioned small models, but “small large models” that receive general-knowledge training before professional data is added. A weather model compresses the computing required from 3,000-5,000 high-end servers for about half a day into 10 seconds on one card; drug screening can fall from months to days or even hours.
- AI’s most verifiable value is already showing up in industrial output, trial-and-error costs and safety, but capturing it requires a full-stack rebuild. A mine in Shandong using the relevant models can produce about 2,000 additional tons of coal a year; a steelmaking model digitizing experience around furnace temperature and blend ratios has improved quality, with initial results potentially cutting costs by more than RMB1 per ton. Companies must also upgrade AI Native compute, cloud services, industry models and application systems in parallel, because “advantages cannot hold back the trend.”
- Under US-China competition, 张宇昕’s answer is industrial differentiation and collaboration—not every company retraining its own general-purpose model. He says China is unlikely to surpass the Western world in English-language models, which benefit from access to more public English knowledge, but can find another path through its leading industrial sectors. To C and To B companies should each play to their strengths, while industry supplies academia with problems and use cases, and academic results return to industry for validation, creating a cycle that is difficult to blockade or sever.
Deep dive
1. Huawei bet on AI before ChatGPT—but chose industry from day one
The interview, recorded at Huawei’s Chuangyuan Forum in Wanning, opens with 张宇昕’s technical background. He joined Huawei in 1999 as an entry-level software developer, then moved from communications into IT and cloud.
After Huawei established its 2012 Lab in 2011, 张宇昕 founded the Euler division and became its first head. He moved into the IT business in 2013 to lead storage R&D, and became Huawei Cloud CTO after the cloud BU was established in 2017.
Huawei unveiled its AI strategy at the 2018 Huawei Connect conference. It released the Pangu large model at the 2021 developer conference, upgraded it to 3.0 in 2023, and to 5.0 this year—roughly in step with the ChatGPT wave.
The real difference was the choice of route. 张宇昕 says Huawei “positioned itself from the very beginning as AI for Industry,” not because it failed to see the general-purpose model opportunity, but because it believed China’s most defensible advantage lay in industry.
2. China’s manufacturing base is the best data set of the AI era
张宇昕’s foundational argument starts with industrial structure: there are more than 600 industrial sectors globally, and China not only covers the broadest range but leads the world in more than 220 of them.
China’s edge in these sectors is not just output, but technology, experience and engineering capability. Once captured in data, “it becomes the best data set for artificial intelligence.” That is why China does not need to copy the US route.
卫诗婕 asked whether China could be the “black soil” for AI development. 张宇昕 agreed directly: “That is determined by our structure and by the characteristics of our industries.”
3. To B companies are treating the technology wave as a survival risk for the first time
张宇昕 divides customers into 2 groups. Traditional companies carry the burden of legacy systems, business models and hardware-software stacks, and prefer a gradual transition; innovative companies are both sensitive and anxious, worried that markets and value distribution will be rewritten while lacking the capital and accumulated technology for long-term R&D.
He estimates the 2 groups are “at least 50-50,” with proactive adopters already accounting for more than half. Companies do not need to rebuild an entire technology stack from scratch; they need low-friction access to compute, databases, big data, AI platforms, industry expertise and “everything as a service.”
China’s cloud adoption lagged the US by at least 5 years, and many companies are still wavering between public and private cloud. But after AI arrived, “not a single company says it won’t use artificial intelligence,” because companies that fail to change themselves risk being disrupted by others.
ChatGPT broke the old assumption that AI could solve only isolated problems. Massive compute and unsupervised, automated training revealed the possibility of general-purpose intelligence. AI was no longer a “local truth”—something that worked for one company but might not work for another—and instead became something “everyone is connected to.”
4. AI is not omnipotent, but waiting for 100% maturity means arriving too late
张宇昕 first tempered the enthusiasm: “If humans can’t do something, you can’t expect a large model to do it.” A model cannot solve the still-unsolved Goldbach conjecture; it remains bounded by the limits of existing human knowledge, experience and skills.
Nor did he erase the capability gap simply because Sora can generate video. Compared with real-world filming and production, it remains far behind, lacking intellectual and aesthetic judgment and failing to understand physical realities such as light and fluid dynamics. Hallucinations, controllability and authenticity all still need work.
But technological immaturity is not a reason to stay out: “By the time artificial intelligence is especially mature—by the time it is 100% ready for industrialization—you may already be behind the times if you start using it then.”
Technology and application do not have a fixed sequence. 张宇昕 compares them to a DNA double helix: new technology creates applications, while applications generate new requirements, provide use cases and serve as training grounds. Models and products must mature together in real business environments.
5. Large models are pushing the cloud from selling resources to digitizing enterprise knowledge
AI deployment requires 3 elements to arrive together: large-scale compute, massive data, and customer use cases plus upstream and downstream partners. For most enterprises and government and enterprise customers, 张宇昕 says, the cloud is the most practical and convenient platform for bringing those elements together and coordinating innovation.
The cloud’s first layer is infrastructure such as compute and storage. The second is technology and services including databases, big data, security and software development. AI will make that second layer materially thicker, because the cloud can help customers unlock the value in databases, data warehouses and data lakes rather than merely store the data.
张宇昕 uses financial statements to explain the shift: an ordinary person sees numbers, while a finance expert sees the relationships behind them. Training an enterprise model means digitizing and turning into software the knowledge and experience that previously lived only in experts’ heads.
One expert may understand 50 or 100 dimensions, while another understands a different set; machines can handle far more dimensions and assist with analysis. In 张宇昕’s view, the cloud can therefore do more than provide infrastructure and basic technology: it can pass on experience and unlock the value of “treasure data,” bringing its role closer to that of a business adviser.
6. Huawei’s compute advantage is not exclusive; software-hardware coordination is harder to replicate
卫诗婕 put the industry’s sharpest question on the table: with the US restricting China’s access to advanced chips, some peers also buying Huawei chips, and government and enterprise customers treating Huawei as the preferred domestic-compute option, is compute Huawei Cloud’s biggest advantage?
张宇昕 rejected the idea that it was the “only advantage.” Compute is indeed one advantage, but Huawei also supplies it to peers and the broader industry. The compute ecosystem therefore needs to be expanded collectively, making it “neither a unique advantage nor the only advantage.”
He sees the more distinctive capability in software-hardware coordination. The cloud team can start with a customer’s business and use case to design the entire system, then use system requirements to drive hardware and chip development, rather than optimizing isolated software and hardware metrics separately.
7. The inference market is far larger than training, and scaling law is nowhere near finished
“Exaggerating a little,” 张宇昕 estimates that only around 5 institutions each in China and the US will truly need to keep training large models from scratch. Training is expensive and arduous, while most customers simply use models, leaving far more room for inference compute than training compute.
He believes scaling law will continue because the unit cost of AI compute remains too high and efficiency too low. The human brain consumes limited energy, while a computing system with comparable cognitive ability is enormously expensive; if replacing basic work does not cost less than human labor, it fails the test of economics.
PCs became widely adopted only after escaping the cost structure of the mainframe era, and AI will be no different. 张宇昕 therefore expects compute scale to grow by tens or hundreds of times over the next 5 or 10 years. The fact that investment is becoming more rational only means that, at current costs, blindly expanding compute is unaffordable for some providers.
8. “High-performing without being expensive” is closer to commercialization than reaching for the stars
卫诗婕 cited 周鸿祎’s view that large models should be brought down from the altar: “Don’t reach for the stars; use one model to solve one specific problem first.” 张宇昕 said he “strongly agrees.”
The first meaning is “high-performing without being expensive”: technology that is both advanced and expensive can only float above the real economy. The second is that early efforts should not aim for omnipotence; reducing labor costs, raising automation or improving quality in a production system is already value creation.
The first customers to “eat the crab” and “drink the first bowl of soup” also provide feedback. Every problem that an existing model cannot solve points technology toward a new research target; industry practice is not the endpoint of mature technology, but part of the process by which technology matures.
9. To C and To B large models are diverging, and general-purpose leaderboards are not industrial metrics
张宇昕 expects the 2 paths to separate over time. To C providers serve the mass market, continually add general knowledge and expand their models, and aggregate users through conversational APIs; To B customers care more about whether AI actually creates value in a production system than about adding another chatbot.
That also explains why Pangu is not eager to compete on general NLP leaderboards. 张宇昕 acknowledges that leaderboard performance “may help marketing somewhat,” but a general benchmark cannot prove that a model will solve professional problems in coal mining, steelmaking, healthcare or other real-world settings.
Huawei provides the foundation model and a training toolchain; customers then apply their own industry know-how through incremental training, fine-tuning and business integration. “We are not the final tree or flower growing from the black soil. The trees and flowers belong to the customers.”
Huawei also debated whether to prioritize a public large-model API. It ultimately concluded that this would offer limited help for government and enterprise production applications, while data from government and healthcare customers involves confidentiality or privacy. Huawei therefore shifted toward being an enabler rather than a traffic platform.
10. The L0-to-L2 industry stack is the practical route for smaller companies to use AI
卫诗婕’s challenge was that nearly all the examples on stage were megacap companies. If only large companies can afford AI, does the claim of lowering the barrier still hold?
张宇昕’s architecture starts with an L0 general-knowledge model as the foundation, then moves upward to L1 industry models and L2 models for specific use cases. Industry leaders can use those models themselves while also providing general industry capabilities to suppliers, peers and smaller companies.
Cloud providers can supply the algorithms and compute; data is the hardest gap to fill. Smaller companies have accumulated less of it, while industry leaders hold more sector knowledge. As long as the data does not contain private or confidential information, general or anonymized data can potentially become open data sets or industry models.
Asked why an industry leader would do something that appears philanthropic, 张宇昕 answered with coopetition. Large companies have high costs and cannot cover small projects or narrow use cases, where smaller vendors may be more specialized. Sharing capabilities can fill out the ecosystem—“peers helping peers,” as in open-source software—while allowing each side to complement the other.
11. Fewer parameters do not make an old-style small model; enterprises need “small large models”
Traditional small models worked only within a single company or on a narrow data set and lacked generalization. Today’s industry models first undergo general-knowledge training and then add professional data; the training methods and engineering stack are fundamentally different.
Parameter scale also depends on modality. 张宇昕 says NLP models start at tens of billions or hundreds of billions of parameters, and even that is still viewed as insufficient, with some potentially reaching 1T. CV models may work with several billion parameters, while inference, decision-making, drug-molecule and other scientific-computing models may need only hundreds of millions to several billion.
Asked whether the tension between specialization and generalization can be resolved, 卫诗婕 received no false certainty from 张宇昕. He compares it to students: an elementary-school student may forget when learning several things at once, while a university student can absorb them simultaneously. More compute, new model architectures or new data-processing methods may eventually ease the trade-off.
Enterprises must still choose among size, precision and cost. Quantization and distillation can reduce compute requirements, and if the loss in precision is controllable, a smaller model may be better suited to production. 张宇昕 calls these “small large models.”
12. The value of scientific-computing models is compressing equation solving into rapid learning
A traditional weather forecast requires roughly 3,000-5,000 high-end servers running for half a day. A weather large model can produce a result on one card in 10 seconds. It is not a language model and does not pursue 100B- or 1T-scale parameters.
Drug-molecule models follow the same scientific-computing route. 张宇昕 says developing a drug could previously take 10 years, with drug screening alone measured in months; it can now be shortened to days or even hours.
The core of these models is a scientific-computing equation. Problems once solved by repeatedly working through equations and iterations can now be accelerated through learning. Smaller parameter counts do not prevent these models from generating significant value.
13. Coal mining and steelmaking show that industry models can flow directly into output and cost
In coal mining, the first issue is safety. Remote cameras and CV models monitor conditions at the coal face, while sound can help determine whether the material ahead is water, coal or rock, digitizing the experience once held by veteran workers who identified it by ear on site.
Extracted coal must also be separated from foreign material, while multiple blend ratios must be optimized during clean-coal production. Previously, workers adjusted the mix repeatedly after batching, combustion and slag analysis; a model can combine historical blending data and directly produce the optimal parameters. One mine in Shandong can consequently produce about 2,000 additional tons of coal a year.
Steelmaking turns workers’ experience with furnace temperature, raw materials and trace-element ratios into a model. One furnace holds about 60 tons of steel, so repeated trial and error is not an option. 张宇昕 says that if costs fall by RMB1 per ton, the initial result could be more than RMB1; applied to China’s output of hundreds of millions of tons, the scale effect is substantial.
14. To B products come from getting “both hands dirty,” not from office feature lists
Huawei now operates in more than 30 industries and serves hundreds to more than 1,000 customers, with each customer potentially bringing 3-5 requirements. Given limited staffing, teams must go to production sites to determine which use cases are most mature and valuable, rather than relying only on customers’ verbal descriptions.
Technical teams must also extract common patterns from customers’ individual 1-2-3-4 requirements—“discarding the rough and retaining the refined, separating truth from falsehood, then moving from the surface to the core and from one case to adjacent cases”—so that one solution covers related problems customers have not yet articulated.
Requirements management includes annual high-level planning, quarterly and monthly versions, and emergency decision processes. Once a product is ready, it must return to the customer site for validation and continued iteration. 张宇昕 calls this the To B mindset of getting “both hands dirty.”
The targets themselves also change. Sometimes the direction is wrong, as when Huawei abandoned the priority on a public API; sometimes only the sequencing across automotive, coal mining and energy needs to change; sometimes the target was set too low, requiring further gains in cost-performance, training-cluster stability, and interruption and recovery times.
15. AI Native requires a full-stack rebuild, while China must follow its own collaborative path
张宇昕 breaks AI Native into 4 layers: adopting new compute built for high compute density, high bandwidth and low latency; upgrading cloud services such as databases, big data, security and development with AI; building industry- or enterprise-specific models; and finally rebuilding application systems around those models.
Huawei has moved from data centers, chips and hardware down through the underlying software, and is upgrading cloud services under the “Reshape Cloud Service” initiative. Some services are adding Intelligent Assistant, while others are strengthening database operations, security situational awareness and vulnerability and attack detection. The process is unfinished and will continue to iterate for a long time.
Legacy code and accumulated team experience inevitably become baggage, potentially requiring years of work to be rewritten or discarded. 张宇昕’s internal principle is that “advantages cannot hold back the trend”: better to participate early and capture the upside than be overturned by the transformation.
卫诗婕 used the news that day of the “father of ChatGPT” leaving OpenAI and the flow of talent in the US to ask about US-China competition. 张宇昕 believes China will struggle to surpass the West in English-language general-purpose models and should instead “play to its strengths, avoid its weaknesses and find another path” through its industrial lead. To C and To B companies should each focus on what they do best, while industry and academia form a loop of “posing problems, researching technology and validating it in industry.”