What performance gains does NPO unlock?
With that much-needed context out the way, let’s get back to the Atlas 960E SuperPoD. Huawei says that it features an advanced orthogonal architecture and a fully liquid-cooled design.
With the unified memory address that UnifiedBus technology enables, a single SuperPoD can have up to 4,096 neural processing units (NPUs), making it capable of delivering 8 EFLOPS at FP8 precision and 16 EFLOPS at FP4 precision.
The use of 5,500 Hi-ONE units eliminates the need for the 48,000 800G optical modules that would normally connect the NPUs, cutting power consumption by more than 550kW.
Huawei also says that this doubles the system’s fault-free running time and brings the system’s availability up to 99.8%, while greatly accelerating ‘innovation in training and inference for frontier models’.
With the UnifiedBus network or RDMA over Converged Ethernet (RoCE), multiple Atlas 960E SuperPoDs can be connected to build a larger SuperCluster. With a two-tier, four-plane Clos architecture, Huawei states that the SuperCluster can interconnect up to 512,000 NPUs, rising to up to one million when combined with a multi-rail topology.
By applying the SuperPod architecture to general-purpose computing systems, Huawei has created products like the ‘fully upgraded’ TaiShan 950 SuperPoD. This uses UnifiedBus all-optical network, supports up to 4,096 nodes and has a unified memory pool of up to 256TB.
According to the vendor, this allows higher-density agent sandboxes, faster sandbox startup and improved agent vector search performance.
