Select a Country or Region
- Australia - English
- Brazil - Português
- China - 简体中文
- Europe - English
- France - Français
- Germany - Deutsch
- Ireland - English
- Italy - Italiano
- Japan - 日本語
- Kazakhstan - Қазақ тілі
- Kazakhstan - Pусский
- Kenya - English
- Korea - 한국어
- Malaysia - English
- Mexico - Español
- Mongolia - Mонгол
- New Zealand - English
- Netherlands - Nederlands
- Poland - Polski
- Romania - Română
- Singapore - English
- South Africa - English
- Spain - Español
- Switzerland - Deutsch
- Switzerland - Français
- Switzerland - Italiano
- Switzerland - English
- Tanzania - English
- Thailand - ภาษาไทย
- Turkiye - Türkçe
- Ukraine - Українська
- United Kingdom - English
- Uzbekistan - Pусский
- Uzbekistan - O’zbek
- Vietnam - Tiếng Việt
- Global - English
This site uses cookies. By continuing to browse the site you are agreeing to our use of cookies. Read our privacy policy
AI is entering a new chapter: from training-intensive development to large-scale inference deployment. As key players in the construction of national digital infrastructure, carriers are uniquely positioned with their cloud-network-edge-device capabilities to seize new opportunities in this shift. In sectors like government, education, healthcare, industrial manufacturing, and the Internet, carriers have already begun piloting a new "token monetization" model — offering token packages, full-stack computing hosting services, and AI appliances for rapid delivery and scaled deployment.
In AI inference, efficiency, experience, compute resource scheduling, and security resilience are top concerns for enterprise customers. Huawei Data Storage is working with carriers to build AI data infrastructure around six core pillars: data lake, knowledge and memory platform, compute, models, agent framework, and data resilience. Together, they aim to deliver high-quality, secure token services for enterprise customers, accelerating carriers' transition to the "token monetization" era.
Compute has been widely discussed elsewhere. This article focuses on the other five pillars: data lake, knowledge and memory platform, models, agent framework, and data resilience.
Figure 1: AI data infrastructure
Data lake: High-quality data aggregation and supply
For enterprises, the first hurdle to AI adoption is massive data volumes and complex data management. In industrial quality inspection, for example, a single ultra-HD camera generates tens of thousands of high-precision images daily. One production line can add PB-scale data per month and tens of PB annually. Yet when a customer complaint triggers a quality traceback, the enterprise must search through tens of billions of records for a specific defect. Traditional methods take hours, severely slowing production and after-sales response.
Huawei's AI data lake solution addresses this with OceanStor Pacific all-flash distributed storage, offering 11 PB per 2U — industry-leading density at optimal TCO. The DME Omni-Dataverse unified data space enables real-time, multi-modal, cross-site data ingestion with global visibility and management. The solution also supports second-level retrieval across hundreds of billions of high-dimensional vectors, delivering high-quality data aggregation and supply.
Knowledge and memory platform: Better retrieval, faster inference, smarter models
To address bottlenecks in inference — frequent hallucinations, poor response experience, and lack of inference memory — Huawei has pioneered a "3+1" AI data platform that optimizes storage for knowledge, KV cache, and memory, with the Unified Cache Manager (UCM) technology enabling intelligent orchestration and management. This helps carriers enhance the inference experience for their B2B customers.
Knowledge base: High-precision, multi-modal knowledge for more accurate retrieval
In enterprise AI applications, retrieval accuracy directly determines system usability. Take intelligent query as an example: Customers handle massive amounts of information daily — product manuals, industry white papers, policy documents, and more. Traditional keyword-based retrieval only matches document titles or abstracts and cannot extract information from images or tables, often returning irrelevant results.
The knowledge base built on OceanStor AI storage transforms unstructured text, images, and video into fine-grained knowledge units through multi-modal lossless parsing and token-level encoding. Combined with multi-dimensional semantic retrieval and comparison, it accurately understands user intent. Field tests show retrieval accuracy exceeds 95%, delivering a highly precise, hallucination-free search experience backed by verifiable sources.
PB-scale KV cache: Accommodating massive historical data for more efficient inference
Inference tasks with long texts, long sequences, and multi-user concurrency are becoming increasingly common in applications like medical literature analysis and internet search and recommendation. As input sequences grow longer, the volume of generated KV cache data increases exponentially. System responses slow down if there is insufficient infrastructure for large-scale cache storage. This hinders inference and degrades the user experience.
Using the PB-scale KV cache expansion of OceanStor AI storage and OceanDisk intelligent disk enclosures, the system retains large-scale, multi-round cache data. When a new request arrives, the system directly reuses the historical cache to minimize redundant computation, significantly reduce inference latency, and boost inference throughput for a superior user experience. This provides key performance support for intelligent agent inference with long sequences and complex logic.
Memory bank: Context management that makes models smarter over time
Large models often struggle with "short-term forgetfulness" when handling complex business logic. In commercial data insights, for example, enterprises process massive volumes of sales data, campaign performance feedback, and inventory turnover monthly. This means inference cannot start from scratch every time — models need to retain insights from the past six months or even a year, and actively recall that historical memory when generating new strategies.
With the memory capability built on OceanStor AI storage, Huawei uses context state retention and context distillation to extract and refine historical data and experience, distilling it into recallable memory. The more a model remembers, the more accurate its reasoning becomes, enabling models to grow smarter over time and moving business insights from post-hoc review to intelligent decision-making.
UCM: Full-lifecycle management and scheduling of inference memory data
To further optimize token efficiency across business processes and improve inference performance and experience, Huawei has launched Unified Cache Manager (UCM). Its core logic is a three-layer collaboration between inference frameworks, compute, and storage, enabling tiered management and intelligent scheduling of knowledge bases, KV cache, and memory bank. In field tests, UCM has reduced time to first token (TTFT) by up to 90% and increased system throughput by as much as 22 times.
UCM consists of three main components. At the top is a connector that interfaces with mainstream inference frameworks such as vLLM, SGLang, and MindIE. The middle layer is an accelerator running on compute servers, responsible for tiered cache management of memory data. The bottom layer is a coordinator that works with dedicated shared storage to improve pass-through efficiency and reduce latency.
Model engineering and resource scheduling: Out-of-the-box models and fine-grained compute management
In carrier B2B operations, intelligent services often need to support multiple industry scenarios simultaneously, each with different requirements for model architecture, precision, and inference performance. Frequent model swapping is also required when shifting between business scenarios.
Huawei ModelEngine provides out-of-the-box model deployment and gateway capabilities, supporting one-click model deployment. Combined with fine-grained compute slicing and intelligent scheduling, it can partition a single compute card into multiple virtual compute units, achieving up to 1:10 xPU card slicing. This "one card, many virtual cards" approach significantly improves resource utilization.
Agent framework: Low-code development and autonomous evolution
AI agents are rapidly gaining traction in government services and healthcare, increasingly becoming permanent "digital employees". However, traditional agent development relies heavily on specialized programming skills. Once deployed, agents require ongoing manual tuning and prompt engineering to adapt to changing business rules and user interaction patterns, driving up maintenance costs.
Huawei ModelEngine Nexent enables agents to be generated directly through natural language interaction, dramatically lowering the development barrier and cutting deployment cycles by 80%. The platform also supports automatic optimization of skills, prompts, and memory, helping agents continuously evolve and grow smarter over time.
Data resilience: Trusted data spaces for end-to-end data circulation, security, and compliance
In industries like healthcare and manufacturing, data sharing often hits a "trust deadlock". Data providers are reluctant to share data due to concerns over loss of control, traceability, and potential misuse. Data consumers, on the other hand, hesitate to use data because they cannot verify its authenticity and security.
Huawei leverages its global data management technologies to build trusted data spaces. Addressing the two key pain points of data trust governance and secure, efficient transmission, Huawei offers over 30 usage control policies, covering access permissions, validity periods, view counts, and more. These ensure that data consumers strictly adhere to the policies and authentication set by data providers, with all data usage activities fully auditable and traceable.
Carrier intelligent computing platform in practice
Huawei and a Chinese carrier have jointly built an intelligent computing service platform that has been deployed at scale across the carrier's group as an AI capability backbone. It fully supports internal IT systems and multi-dimensional business innovation across consumer-facing, business-facing, and household-facing segments. The platform centers on KV cache storage, significantly improving storage resource efficiency while maintaining full compatibility with the carrier's self-developed models and open-source models like DeepSeek and Qwen.
For long-sequence input scenarios such as research report analysis, algorithm optimization has overcome engineering challenges such as inference stalls and low KV cache hit rates. Field data shows that overall throughput has increased by more than 10 times, inference costs reduced by approximately 50%, and response time shortened to under 1 second, delivering efficient, stable performance support for large-scale, multi-service concurrent inference.
AI has unlocked a new growth model for carriers: token monetization. Huawei is ready to work with global carriers to build AI data infrastructure, improve inference efficiency and experience, ensure data security, and drive new B2B growth — enabling carriers to make the leap from traffic monetization to token monetization.
- Tags: