The artificial intelligence industry is undergoing a monumental paradigm shift, transitioning rapidly from static generative models to autonomous, decision-making agents capable of executing complex multi-step workflows. Recognizing this fundamental evolution in enterprise computing demands, Huawei Cloud unveiled a comprehensive suite of cloud innovations during the second day of the annual HUAWEI CONNECT 2026 conference in Shanghai. At the center of this announcement was the launch of an enhanced computing architecture, specifically highlighted by the new Huawei AI Cluster Service for agentic AI workloads.
Delivered during a keynote presentation by Dr. Peter Zhou, Director of the Board at Huawei and Chief Executive Officer of Huawei Cloud, the release marks a significant milestone in silicon-to-cloud infrastructure development. As enterprises transition from simple chatbots to sophisticated digital agents that run continuously, existing cloud setups face unprecedented bottlenecks in memory capacity, token processing efficiency, system resilience, and unified workload scheduling. The introduction of the Huawei AI Cluster Service for agentic AI development directly addresses these challenges, laying a robust foundation for scalable enterprise intelligence.
Defining the Shift Toward the Agentic Infra Paradigm
In his keynote address, Dr. Peter Zhou emphasized that computing infrastructure can no longer serve merely as a passive provider of raw processing power. The era of autonomous agents requires a complete reimagining of the underlying architecture. Today's digital agents perform long-horizon tasks that demand continuous learning, rapid context retrieval, and seamless integration between general computing and specialized artificial intelligence processing.
To satisfy these requirements, Huawei Cloud formally established the Agentic Infra framework. This paradigm focuses on four core pillars: extreme token efficiency, petabyte-scale memory capabilities, unified scheduling across general and artificial intelligence compute workloads, and secure autonomous execution. By aligning hardware innovation with software optimization, the Agentic Infra framework provides an end-to-end environment capable of hosting thousands of active agents simultaneously without sacrificing stability or inflating operational costs.
Technical Innovations Driving the New Cluster Service
The primary technical foundation of this new infrastructure rollout is the upgraded cloud cluster management system. The deployment of the Huawei AI Cluster Service for agentic AI applications relies on an ultra-high bandwidth UnifiedBus network architecture. This interconnect technology allows cloud environments to scale up to clusters containing more than 100,000 compute cards, yielding a total aggregate computing capacity reaching up to 200 EFLOPS.
To maximize operational stability during large-scale model training and agent coordination, the platform incorporates a sophisticated five-level fast recovery mechanism combined with full-chain observability. In real-world cloud testing environments, the service demonstrated the ability to maintain uninterrupted, stable model training for over 40 days. In the event of hardware or network failures, the automated system executes complete fault recovery within approximately 10 minutes, drastically reducing downtime. Furthermore, through deep co-optimization of scheduling systems, caching mechanisms, and underlying algorithms, the platform achieves a 20 percent higher token throughput compared to previous-generation cloud computing offerings.
Solving Memory Bottlenecks with Context Memory Storage
A major bottleneck hindering the real-world deployment of autonomous agents is memory retention during extended multi-day operations. Conventional cloud setups often clear context state between sessions, forcing models to reprocess large volumes of historical data, which increases latency and overall cloud expenses. To overcome this hurdle alongside the cluster service, Huawei Cloud introduced Context Memory Storage.
By leveraging direct NPU passthrough technology to dedicated memory hardware, Context Memory Storage creates a petabyte-scale memory architecture specifically optimized for long-horizon agent tasks. This solution offers twice the storage capacity of comparable industry solutions while supporting high-speed memory access with a 50 percent increase in read performance. Supported by tiered Key-Value cache pooling, this technology enables continuous learning for digital agents and drastically cuts inference expenses across complex enterprise workflows.
Streamlining Model Access via Agentic Model as a Service
At the model interaction layer, Huawei Cloud expanded its software ecosystem to simplify how developers build and deploy autonomous systems. The newly launched Agentic Model as a Service platform consolidates a wide array of state-of-the-art foundation models from multiple leading developers into a single, unified interface.
Through this centralized model routing mechanism, developers can invoke diverse top-tier foundation models with a single API call without managing individual model deployments. The platform offers three distinct routing policies: experience-first, efficiency-first, and a balanced mode. By dynamically matching incoming requests to the optimal model based on task requirements, the system achieves a scheduling accuracy exceeding 95 percent while reducing model call expenses by an average of 20 percent. This seamless integration ensures that companies adopting the Huawei AI Cluster Service for agentic AI deployments can leverage the best available model for every specific step in a complex workflow.
Empowering Enterprises through AgentArts and Open Source Ecosystems
Beyond raw hardware and model routing, building functional enterprise agents requires structured developer tools and rich asset libraries. Huawei Cloud is advancing its enterprise-grade ecosystem through a dual commercial and open-source strategy centered on its AgentArts platform and the open-source openJiuwen framework.
These platforms provide enterprise developers with direct access to more than 5,000 general Model Context Protocol assets and over 1,000 industry-specific assets. By standardizing context protocols, organizations can rapidly build custom digital agents for specialized domain tasks ranging from financial analysis to municipal governance. To date, over 100 large enterprises and public entities, including Kingsoft Office and local government authorities, have implemented agent solutions built on this framework. By pairing these developer platforms with the Huawei AI Cluster Service for agentic AI workloads, enterprises receive both the software tools and high-density compute needed to scale operations efficiently.
Expanding the Industry AI Foundry Across Vertical Sectors
To further accelerate practical adoption, Huawei Cloud announced major expansions to its Industry AI Foundry platform. This initiative serves as a shared knowledge repository, allowing successful artificial intelligence deployment templates, domain-specific data structures, and pre-built agent workflows to be reused across different economic sectors.
During the conference, Huawei Cloud introduced two new dedicated zones to the Foundry: the Smart Government Zone and the AI Hardware Zone. In total, the Industry AI Foundry now encompasses over 1,000 verified industry assets and supports more than 1,000 active implementation projects globally. This modular approach helps enterprise clients drastically reduce the time required to move proof-of-concept AI agents into production environments.
Global Rollout Schedule and Market Outlook
The commercial rollout of these cloud technologies is structured across distinct phases to serve both domestic and international markets. Commercial availability for the new cluster service in mainland China is scheduled to begin on September 30, 2026. Following the initial domestic release, international enterprise clients will gain commercial access starting November 30, 2026. Meanwhile, the international commercial release of the AgentArts enterprise platform is set for December 30, 2026.
As global enterprises increasingly rely on autonomous digital systems to streamline operations, the demand for robust, high-density cloud infrastructure will continue to surge. By uniting cluster resilience, petabyte-scale memory systems, unified scheduling, and comprehensive developer assets under the Agentic Infra framework, Huawei Cloud is positioning itself as a primary enabler of the next generation of artificial intelligence innovation.
Read More

Friday, 25-09-26
