TL;DR
Huawei is shipping a 4,096-chip AI supercluster built around its new Ascend 960 accelerator, using optical interconnects and HBM to close the gap with Nvidia and AMD through sheer scale rather than per-chip performance.
What happened
- Huawei's Atlas 960E Superpod packs up to 4,096 Ascend 960 accelerators into a single system, reaching 16 Exaflops of FP4 compute and over 1 petabyte of HBM capacity.
- The Ascend 960DT (training variant) ships Q1 2027: 4 FP4 Petaflops per chip, 288 GB HBM across eight stacks, 9.6 TB/s memory bandwidth.
- The Ascend 960PR (inference variant) follows Q3 2027: 8 FP4 Petaflops per chip but drops to 192 GB at 2.4 TB/s, likely using cheaper LPDDR RAM.
- HBM sourcing is unconfirmed: Chinese manufacturer CXMT's HBM3e is a candidate, but only for limited production.
- Huawei uses on-chip optical engines on the sides of each accelerator to convert signals, enabling scaling beyond a single Superpod toward a stated goal of one million Ascend chips in one system.
Why it matters
- Per-chip performance gap is severe: one Nvidia Rubin GPU delivers 50 FP4 Petaflops and one AMD Instinct MI455X delivers 40, meaning Huawei needs at least 10 Ascend 960DTs to match either rival chip.
- Scale-out is the workaround: by clustering 4,096 chips with optical interconnects, Huawei sidesteps the node disadvantage and delivers a system-level number that is competitive on paper for large training runs.
- SMIC's 7 nm process is the hard ceiling: Nvidia is at 3 nm and AMD at 2 nm, so Huawei's per-chip efficiency deficit is structural and will persist as long as export controls block access to advanced foundries.
- Domestic HBM dependency is a new risk: if CXMT cannot supply HBM3e at volume, the 960DT's memory specs become aspirational rather than deliverable.
- A published annual roadmap through 2029 (Ascend 970 at 14 Petaflops in 2028, Ascend 980 doubling that with 38.4 TB/s in 2029) signals Huawei is committing to a sustained cadence, giving Chinese cloud and AI labs a credible domestic alternative to plan around.
What to watch next
- Q1 2027 960DT delivery volumes: whether Huawei can ship at scale will reveal whether CXMT HBM supply is real or a bottleneck.
- Power consumption disclosure: Huawei withholds wattage figures; data center operators need that number before committing to Superpod deployments, so any leak or official release will shift procurement decisions.
- Export control responses: if the U.S. tightens restrictions on SMIC or CXMT, the entire roadmap through Ascend 980 is at risk of slipping.
Originally published on Present of AI, a daily source-linked AI news timeline. Read the full timeline or browse the open dataset.