Z.ai Details GLM-5.3-Flash Service Built On 100,000 Chinese AI Accelerators
The account is a rare public description of a production inference stack running at that cluster scale on Chinese-made accelerators, and it credits an agent-driven feedback loop, not an infrastructure team, with most of the work.
Reporting from 1 source: GIGAZINE.
Z.ai explained in a blog post how it built the inference infrastructure behind the production service for GLM-5.3-Flash, running on a cluster of more than 100,000 Chinese-made AI accelerators. The company says most of the work was carried out by an infrastructure agent powered by GLM-5.3, and that optimizations brought end-to-end service performance to about 3x and cost per token in line with mainstream NVIDIA GPUs.
GLM-5.3-Flash also ran under the name Ox-Alpha on OpenCode and OpenRouter. Z.ai says that within one week of launch it was the most-used AI model on both platforms, processing more than 62 trillion tokens in six days.
The engineering account covers memory optimization that trades computation for bandwidth, intra-node tensor parallelism for linear attention and the LM head, ReplaySSM, W8A8 quantization, mixed-precision cache quantization using INT8, FP8 and BF16, layer splitting, and an encode-prefill-decode disaggregated architecture. From initial model adaptation to production operation took just under two weeks.
Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.