AI Affairs, home

Saturday 3 October 2026

Technology

Huawei adds full-duplex voice and video to its WorkSwarm agent platform

The update separates live conversation from background tool tasks and provides a route for deploying multimodal models on Ascend NPUs through Huawei Cloud.

Huawei Beijing Research Institute Building Q14 behind a lawn and paved plaza
Photo: HoweyYuan, CC BY-SA 4.0, via Wikimedia Commons (cropped)

Huawei’s openJiuwen added full-duplex voice and video to WorkSwarm in an update announced on 1 October, Pandaily reported on 2 October. The change is intended to let a conversation continue, including interruptions, while a separate agent works through longer tasks.

Key points

  • WorkSwarm separates live camera and voice interaction from background tasks handled by a Core Agent.
  • Voice activity detection and a noise gate are designed to reduce interruptions caused by background sound.
  • Multimodal models can be deployed on Ascend NPUs through ModelArts using vLLM-Omni.

WorkSwarm separates conversation from tool tasks

Full duplex lets an assistant listen while it speaks, rather than making each side wait for the other to finish. In WorkSwarm, a valid new voice input stops the spoken answer and becomes the next instruction. Text already produced remains in the task record. Voice activity detection identifies speech, while a noise gate helps prevent background sound from triggering that interruption.

The system divides work between a real-time model, which handles the camera feed and conversation, and a Core Agent, which breaks down longer requests and calls tools. In the example Pandaily described, a camera question about a product can lead to a request for research and a document. WorkSwarm displays the background task’s status, tool calls, results and generated files while the conversation continues.

Putting together a document about a product could therefore continue after a spoken question interrupts the exchange. The research would remain a separate task, allowing the conversation to carry on while that work proceeds.

When a background task finishes, WorkSwarm posts a written summary before giving a shorter spoken account once the current exchange ends. Further requests enter a queue that can be reordered. Stopping a spoken reply does not stop a tool task, which must be cancelled separately, and work already passed to the Core Agent continues if the audio and video session closes. Its results return to the original conversation.

ModelArts provides an Ascend deployment route

WorkSwarm supports two model protocols. Qwen Omni uses a real-time WebSocket connection, either through an official endpoint or a local deployment. JoyAI instead links separate models for speech recognition, language and speech synthesis through official APIs or other platforms. Those are different ways to supply the conversation side of the system.

For Ascend NPUs, full-duplex multimodal models can be deployed through Huawei Cloud’s ModelArts platform using the vLLM-Omni inference framework, Pandaily reported. WorkSwarm also offers Windows, Mac and HarmonyOS installers and pip packages. Its GitHub repository carries an Apache-2.0 licence, giving developers access to the open-source workbench as well as the described deployment path.

Release v0.2.7 adds cross-agent sessions

OpenJiuwen was built by Huawei’s 2012 Laboratories, Huawei Cloud and its device and computing teams alongside universities and other developers. WorkSwarm is its workbench for office and coding tasks. The project’s latest tagged release, v0.2.7, added a Proactive Context Service and cross-agent sessions on 30 September.

Topics: Agents, Open source