As AI transitions from simple chatbots to complex, multi-turn agents, the demand for massive context windows has pushed traditional DRAM and on-chip memory to their physical limits. Huawei’s OceanStor M900 addresses this by creating a global, multi-tier KV cache that pools resources across compute, network, and storage layers. Utilizing a high-speed UnifiedBus interconnect, the system allows clusters to scale up to 64 PB of capacity, effectively shifting the bottleneck away from memory size for large-scale deployments.
The architecture integrates CPU, network, and NAND controllers to enable one-hop connections between NPUs and SSDs. This design slashes access latency to 60 microseconds—a 90% reduction compared to standard protocols—and pushes aggregate bandwidth to 40 TB/s. Beyond performance, the M900 utilizes KV-aware adaptive storage to intelligently manage data lifecycles. By extending SSD endurance 16-fold, the system lowers long-term operational costs, facilitating the transition toward more sustainable and efficient large-scale AI infrastructure.

Comments (0)
No comments yet. Be the first!