Bilibili’s Index LLM team released Index-Translate on 30 September, Pandaily reported. Built on Alibaba’s Qwen3.5, the Apache-2.0 translation family covers 150 languages and includes text models alongside specialised versions for dubbing, syllable control and long documents.
Key points
- The text models come in 2B, 9B and 35B-A3B preview sizes.
- Index-Echo handles translated subtitles and speech, while Index-Homura targets a requested syllable count.
- On the team’s low-resource instruction test, the 9B model had the best instruction score among the models compared.
Index-Translate follows instructions beyond the words
The 2B and 9B text models and the 35B-A3B preview are available on Hugging Face and ModelScope. The release also includes a technical report, code on GitHub and an online demo. Beyond translating words, the models can follow requests about terminology and formatting, including instructions to leave particular content untouched.
In one project example, the 9B model translated a game-maintenance notice into Korean while retaining its JSON structure, star symbols and a specified hashtag, Pandaily reported. That kind of instruction matters when the text is also data: changing a label or removing a character can alter something other than the sentence a person reads.
Index-Echo and Index-Homura shape translated speech
Index-Echo produces translated subtitles or dubbed speech while preserving characteristics of the original speaker’s voice. Its packaged speech-to-speech release covers Chinese into English, Spanish and Japanese, and English into Chinese, Spanish and Japanese. Index-Homura instead steers a translation towards a requested number of syllables, a constraint relevant when words must fit a spoken passage or a subtitle.
Index-NativeLong, distributed under model IDs named Index-Nailong, works across whole documents to keep references consistent. In a team example involving about 32,000 tokens of fantasy writing, it retained a character’s name where the 9B model, translating separate chunks, used different renderings. The distinction is like keeping the same name on every page rather than deciding afresh each time a page is turned.
Translating a long document could therefore keep an ambiguous name consistent from beginning to end, rather than let its rendering change between passages. A translation aimed at a syllable count could also fit spoken material more closely, though that count is only one part of how a line sounds.
The 9B model leads the team’s instruction test
On the team’s benchmarks, the 35B-A3B preview scored 0.8794 on FLORES using COMET-22 and 76.76 on a WMT26 judge score. The 9B model scored 0.8789 and 75.35 respectively. The same comparison table gave DeepSeek-V4.1-Flash a higher WMT26 score of 83.55.
For low-resource instruction following, the team’s 9B model had the best instruction score among the models compared, at 0.7725, and the lowest off-target rate, at 3.47%. On the team’s SandGlass test, Index-Homura-9B came within 10% of the requested syllable count 81.92% of the time.
The repository offers a video-dubbing pipeline alongside a browser extension that can translate web pages using a model deployed locally.