GLM inference stack sparks debate over China’s AI hardware push
Z.ai’s GLM team drew attention for building a production-grade inference service from scratch on a cluster of more than 100,000 Chinese-made AI accelerators. The system reportedly runs all production inference for GLM-5.3-Flash and, compared with an initial baseline on the same hardware, achieved a 3× improvement in end-to-end serving performance, with hardware efficiency and per-token cost described as comparable to mainstream NVIDIA GPUs.
The milestone fed a broader debate over whether US chip export restrictions have accelerated China’s domestic AI infrastructure. Several commenters argued that constrained access to NVIDIA hardware has pushed Chinese companies toward local accelerators, software optimization, and efficiency-focused model serving. Others said China was already pursuing indigenous chips and that restrictions may still limit access to the most advanced compute capacity.
Discussion also touched on Huawei Ascend processors, SMIC manufacturing, ASML lithography constraints, and the role of software optimization in closing hardware gaps. Some users questioned whether the service feels fast enough in practice, while others pointed to pricing, token limits, and open-weight availability as factors shaping GLM’s appeal against US model providers.