EN / 中文

Evolution Roadmap of Cloud-Side AI Chips

by xinpianneixieshier·August 26, 2026

In recent years, the deployment speed of self-developed AI (Artificial Intelligence) chips by cloud providers and large model developers has significantly accelerated. Overseas, OpenAI and Broadcom are collaborating to advance custom chip R&D, and self-developed chips from providers like Google and AWS have become primary computing infrastructure. Domestically, Baidu AI Cloud has lit up the Kunlun Xin P800 10,000-card cluster, Alibaba T-Head has released the Zhenwu M890, and ByteDance's self-developed chips are making rapid progress. Downstream manufacturers develop their own chips to collaboratively optimize chip architecture, memory, interconnects, software stacks, and cluster deployment based on workloads. This reflects a shift in the competitive logic of cloud-side AI chips: from merely improving single-chip computing power to system-level competition centered on overall performance, cost, and efficiency.

This article will review the product evolution roadmaps of overseas and domestic cloud-side AI chips, analyze the underlying trends, and further explore: as mainstream technical paths become increasingly consistent, how should domestic AI chips break free from homogeneous competition and find new breakthroughs?

I. Overseas: From Single-Card Performance Improvement to System-Level Computing

From Volta, Ampere, and Hopper to Blackwell, NVIDIA's four generations of products reflect the shift in AI computing bottlenecks: from insufficient computing power to memory, interconnects, and system efficiency.

First Generation: AI-Specific Computing. Volta introduced the Tensor Core for the first time, beginning the design of dedicated hardware for AI matrix computing. This marks the shift of AI computing from relying on general-purpose GPU capabilities to being handled by independent, dedicated computing units.

Second Generation: Comprehensive Performance Improvement. Compared to Volta, Ampere increased FP16 computing power by approximately 150%, maximum video memory capacity by 150%, and video memory bandwidth by about 127%, while doubling the NVLink bandwidth. This marks the transition of AI chips from merely improving computing power to the simultaneous upgrade of computing, memory, and interconnects.

Third Generation: Optimization for Large Models. Compared to Ampere, Hopper increased FP16 computing power by about 3.2 times and NVLink bandwidth by 50%, while introducing FP8 precision. The representative product, H200, increased video memory capacity by approximately 76% and video memory bandwidth by about 43% with basically unchanged computing power. This indicates that computing power is no longer the only bottleneck for AI chips; low precision, large video memory, and high bandwidth have begun to directly affect the actual computing efficiency of large models.

Fourth Generation: System-Level Computing. Compared to Hopper, Blackwell increased video memory capacity by about 28% and video memory bandwidth by approximately 67%, while doubling the NVLink bandwidth, with single GPU power consumption rising to the kilowatt level. The representative product, GB200 NVL72, further integrates 72 GPUs into a unified computing system. This marks the shift in the competitive focus of AI chips from single-chip performance to the performance and efficiency of the entire computing system.

In addition to NVIDIA, AMD, Broadcom, and Qualcomm have also deployed cloud-side AI chip businesses, with basically converging development directions: low precision, high-capacity and high-bandwidth memory, high-speed interconnects, and system-level optimization.

II. Domestic: Continuous Catch-Up Along the Technical Path Defined by NVIDIA

Domestic AI chips have a generational gap with NVIDIA, but the direction of technological evolution is basically consistent.

First Generation: AI-Specific Computing. Huawei's Ascend 910 and Cambricon's Siyuan 100 have completed the productization of cloud-side AI chips, established dedicated AI computing capabilities, and formed a product system of chips, boards, and servers. This marks the leap of domestic chips from technological R&D to product deployment.

Second Generation: Comprehensive Performance Improvement. Huawei's 910B continues to improve training and inference capabilities; Cambricon's Siyuan 270, compared to Siyuan 100, has increased the theoretical peak for non-sparse models by about 4 times. For training-oriented Siyuan 290, compared to Siyuan 270, peak computing power has increased by about 4 times, memory bandwidth by about 12 times, and inter-chip communication bandwidth by about 19 times. Domestic chips have begun to simultaneously catch up in computing power, memory, and interconnect capabilities.

Third Generation: Large Model Optimization. Huawei's 910C and Cambricon's Siyuan 370 and Siyuan 590 have begun to strengthen memory and multi-chip collaboration for large models. Compared to Siyuan 270, Siyuan 370 has doubled its maximum INT8 computing power, increased memory bandwidth to about 3 times, and introduced Chiplet and MLU-Link; Huawei's 910C has further entered the Atlas 900 A3 SuperPoD, with a single system capable of connecting hundreds of NPUs. The competitive focus of domestic chips has begun to shift from single-chip performance to large video memory, high bandwidth, and multi-chip collaboration.

Fourth Generation: System-Level Computing. Compared to 910C, Huawei's 950 has increased chip interconnect bandwidth by about 2.5 times and further supports low-precision computing such as FP8, MXFP8, and MXFP4; the planned NPU scale of Atlas 950 SuperPoD is expanded by more than 20 times compared to Atlas 900 A3. Huawei has begun to use system-level computing to bridge the gap in single NPU performance with NVIDIA.

In addition to Huawei and Cambricon, manufacturers such as Hygon, Kunlun Xin, MetaX, and Iluvatar CoreX are also evolving in the same direction, gradually moving from single-card computing power improvement to larger video memory capacity, higher video memory bandwidth, and interconnect bandwidth.

III. Beyond Catch-Up: Domestic AI Chips Must Find New Breakthroughs

If domestic AI chips always catch up along the existing path, facing the disadvantage of upstream supply resources, the cost will be higher power consumption, cost, and system complexity. Therefore, to narrow the gap, it is not enough to just chase parameters; they must also find their own solutions at the underlying architecture level. RISC-V provides new possibilities for this differentiation.

The value of RISC-V is first reflected in reducing computing costs. The open and extensible instruction set enables chips to add dedicated instructions for AI workloads and optimize computing, caching, memory, and interconnects around specific data flows. Compared to simply increasing computing units or improving memory bandwidth, this approach can reduce redundant computing and data movement at the architectural level, improve hardware resource utilization, and improve power consumption and cost under the same computing tasks.

Second is improving customization capabilities. AI computing workloads are developing towards diversification, with significant differences among various workloads in computing density, memory access requirements, and communication methods. RISC-V allows microarchitecture adjustments according to different workloads, enabling chip design to more directly correspond to specific application needs, providing greater design space for AI chips to develop from general-purpose to specialized.

Third is enhancing the independent control of the underlying technology system. RISC-V adopts an open instruction set and does not rely on closed processor architecture licensing. Chip manufacturers can independently design processor cores and further extend to key links such as computing, memory, and interconnects. For domestic AI chips, this means that technological competition can be extended from catching up at the product specification level to independent design at the underlying architecture level.

IV. Conclusion

Reviewing the development path of cloud-side AI chips, it is essentially about continuously solving new system bottlenecks: after insufficient computing power comes memory, followed by interconnects and system efficiency. As these technological directions gradually become industry consensus, domestic AI chip companies that only improve computing power, video memory, and interconnect specifications along the existing path will find their differentiated competitive advantages shrinking, facing the risk of falling into homogeneous competition.

RISC-V provides a new competitive path for the breakthrough of domestic AI chips. It further sinks product innovation to the underlying processor architecture, enabling enterprises to reconfigure computing resources and optimize data flows around specific workloads, forming stronger product customization and independent design capabilities. Domestic companies choosing the RISC-V path, such as Eswin Computing, are exactly practicing this idea: no longer limited to chasing parameters along the existing path, but forming sustainable technological differentiation advantages through underlying architecture innovation.