iFLYTEK shipped Spark-ASR-2.0 on 23 September 2026, an updated automatic speech recognition large model built for noisy environments, regional dialects, low-volume speech and Chinese–English code-switching. The company says word-error rates fell across general recognition, complex acoustics, contextual recognition and text fluency compared with version 1.0, while inference cost rose only about 10 percent.
The architecture combines non-autoregressive recognition with an LLM-enhanced autoregressive correction branch. Speculative decoding decides when the LLM branch intervenes, and joint Chinese–English text and acoustic augmentation addresses scarce code-switch training data. Dynamic context injection draws on correction history, multi-turn memory and hotword libraries from large candidate pools.
The model handles regional speech without manual switching across 202 Chinese cities, according to the company, and shows improved grasp of medical, chemical and scientific terminology. iFLYTEK Input will absorb the model gradually from 24 September, and an open API will appear on the company’s open platform with a public trial gateway. Subsequent uses are planned for wearable displays, smart office notebooks and the Tingjian transcription service in both consumer and business ranges.
Company materials claim advantages over leading peers on dialects, high noise and low volume; those comparisons remain vendor-stated. Future work, according to the firm, targets filtering out unwanted speakers when several people talk at once, spoken commands for revising text, and adjusting layout to match tone.