AI Affairs, home

Tuesday 29 September 2026

Technology

Inspur launches 128-chip MetaBrain SD200 Ultra for Kimi K3 inference

Inspur claims a single machine can run Moonshot AI’s 2.8-trillion-parameter model with token-generation latency under 5.85 milliseconds per token.

Detailed image of a server rack with glowing lights in a modern data center
Photo: panumas nikhomkhai via Pexels

Inspur Information launched its MetaBrain SD200 Ultra supernode at the 2026 Artificial Intelligence Computing Conference, Pandaily reported on 23 September. The machine brings together 128 domestic AI chips. Inspur says it can run Moonshot AI’s 2.8-trillion-parameter Kimi K3 on one machine, generating tokens with latency under 5.85 milliseconds per token.

The chips have access to 8 TB of unified-address accelerator memory, alongside 64 TB of host memory. A shared address space lets an accelerator reach memory attached to another chip, rather than treating each chip’s memory as an isolated store. Inspur’s 3D Hyper Mesh connects those chips over short-reach copper links and, according to the company, carries memory-related communications with latency of 0.69 microseconds.

Writing a long answer could feel quicker if each new piece of text arrived at the claimed rate. That would depend on the system sustaining Inspur’s Kimi K3 result during the particular request, rather than on chip count alone.

Inspur also claims its network reduces AllReduce time by about 3.5× against earlier paths. AllReduce combines values calculated on separate chips and makes the result available across them — rather like passing around a running total until everyone has the same figure. For three stages of Kimi K3, Inspur says fused operations reduce the number of operators roughly tenfold and raise inference performance more than threefold.

Moonshot AI’s Kimi K3 launched on Amazon Bedrock and may trigger a $20m revenue licence, AI Affairs reported. Pandaily said the accounts it reviewed identified the SD200 Ultra’s chips only as domestic, without naming their supplier or model. Inspur also announced the MetaBrain HC2000 rack for capacity-focused inference, claiming tenfold token throughput per investment when service-level requirements are matched.

Topics: Chips, Data centres