Anime, manga, and games, with a take · A Yukimedia publication

← all stories other 1 sources · 1h ago ·

Z.ai Details GLM-5.3-Flash Service Built On 100,000 Chinese AI Accelerators

The account is a rare public description of a production inference stack running at that cluster scale on Chinese-made accelerators, and it credits an agent-driven feedback loop, not an infrastructure team, with most of the work.

Reporting from 1 source: GIGAZINE.

Z.ai Details GLM-5.3-Flash Service Built On 100,000 Chinese AI Accelerators

Z.ai explained in a blog post how it built the inference infrastructure behind the production service for GLM-5.3-Flash, running on a cluster of more than 100,000 Chinese-made AI accelerators. The company says most of the work was carried out by an infrastructure agent powered by GLM-5.3, and that optimizations brought end-to-end service performance to about 3x and cost per token in line with mainstream NVIDIA GPUs.

GLM-5.3-Flash also ran under the name Ox-Alpha on OpenCode and OpenRouter. Z.ai says that within one week of launch it was the most-used AI model on both platforms, processing more than 62 trillion tokens in six days.

The engineering account covers memory optimization that trades computation for bandwidth, intra-node tensor parallelism for linear attention and the LM head, ReplaySSM, W8A8 quantization, mixed-precision cache quantization using INT8, FP8 and BF16, layer splitting, and an encode-prefill-decode disaggregated architecture. From initial model adaptation to production operation took just under two weeks.

Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.

Sources