AI Affairs, home

Tuesday 29 September 2026

Technology

China Telecom opens weights for Xing4.0-29B-A4B, trained on Ascend chips

The mixture-of-experts model activates about four billion of its 29 billion parameters per token. TechForge reported a software-repair benchmark score, while secondary coverage cited by Pandaily described a smaller local memory footprint.

China Telecom stall with circuit boards and monitors at a trade exhibition.
Photo: Dev Jadiya, CC BY-SA 4.0, via Wikimedia Commons (cropped)

China Telecom Artificial Intelligence Technology Co., Ltd. has released the weights and configuration files for Xing4.0-29B-A4B, a mixture-of-experts model trained on Ascend chips with MindSpore, Pandaily reported on 23 September. The files are available through Hugging Face and ModelScope under the XingChen-AGI organisation, giving developers a model they can download rather than access solely through a hosted service.

Key points

  • The model has 29 billion parameters in total and activates about four billion for each token.
  • China Telecom says Ascend 910C training optimisations raised throughput by roughly 96% against an out-of-the-box baseline.
  • A reported score of 75 out of 100 on SWE-bench Verified concerns software-repair tasks.
  • Secondary coverage cited by Pandaily says a 4-bit version can run locally using about 15 GB of graphics memory.

Four active experts in Xing4.0-29B-A4B

Xing4.0-29B-A4B contains 29 billion parameters, though about four billion are activated when it processes each token. Its mixture-of-experts design routes a token to four of 64 specialised groups of parameters and also uses one shared group. The rest remain part of the model without being used for that token — rather like keeping a library available while opening only the books needed for a particular question.

The model has 40 layers, uses MLA attention and accepts a native context of 256K tokens, extendable to 512K. Context is the material a model can work with in one request, so that capacity matters when a task involves code spread across files or instructions that must be considered together. China Telecom presents the architecture as suited to multi-step planning and calls to external tools.

Checking a faulty program against its other files could involve supplying more of the relevant code together under that context limit. A proposed repair would still have to be checked against the program, even if more of the surrounding material could be considered in the same request.

Only part of the model is active for any one token, but the full set of weights still needs somewhere to reside when the model is run. Secondary coverage cited by Pandaily says 4-bit quantisation, which stores the weights in a more compact form, can bring the memory required for a local run to about 15 GB. That figure concerns a compressed version, rather than the model’s parameter count.

Ascend 910C training and its 96% claim

China Telecom says Xing4.0-29B-A4B is the first model at this scale trained entirely on Ascend NPUs using MindSpore. Training took place on Ascend 910C clusters with MindSpore and the MindFormers toolkit. The chips performed the computation, while the software provided the training framework in which the model’s work could be arranged across them.

The training work included fused mHC operators and adjustments to communication between the model’s experts. China Telecom says those changes increased training throughput by roughly 96% compared with out-of-the-box baselines, Pandaily reported. This is the company’s comparison for its training setup, with the stated baseline, rather than a measurement of how quickly the released model answers a request.

For developers using Ascend hardware, the release combines downloadable weights with a model trained through the MindSpore software stack. It also lists serving routes through Transformers, vLLM, SGLang and KTransformers, and fine-tuning support through LLaMA-Factory and MindFormers. Those options give builders several stated ways to run or adapt the model without tying its use to the training framework.

The release uses what Pandaily describes as an Apache-style China Telecom open-source licence. Its weights and configuration files are distributed under the XingChen-AGI name on Hugging Face and ModelScope. Xing4.0 is part of the Xing series, previously called TeleChat.

SWE-bench Verified and China Telecom deployments

The model scored 75 out of 100 on SWE-bench Verified, TechForge reported. The evaluation concerns software-repair tasks involving repositories, bug identification and proposed patches against issue tickets. The score is a result on that test, whose tasks provide a narrower measure than all the work involved in running an automated software workflow.

China Telecom has also put Xing4.0-29B-A4B into its group-level customer service platform, where it routes enquiries, plans multi-step responses and calls external resolution systems, according to TechForge. The company reports improved first-contact resolution and handling speed in that setting. It has deployed the model for interactive home services that provide technical assistance on domestic hardware.

Those deployments concern Xing4.0-29B-A4B itself. Across the wider Xingchen model family, China Telecom maintains more than 500 operational enterprise workflows in 20 commercial sectors, TechForge reported.

Topics: Agents, Chips, Foundation models, Open source