AI Affairs, home

Sunday 4 October 2026

Technology

Suno launches Speech beta to generate spoken voice and music together

The feature is available on web and mobile, with an option to remove the soundtrack. Suno warns that accents and dramatic pauses can behave unpredictably.

Suno CEO Michael Shulman speaks into a microphone at TechCrunch Disrupt.
Photo: TechCrunch, CC BY 2.0, via Flickr (cropped)

Suno announced Speech on 1 October 2026, opening a public beta that generates spoken voice and background music together on its web and mobile platforms, Unite.AI reported. The company, best known for generating songs, says the new feature makes the voice and its soundtrack as one track. It can also produce speech without music.

Key points

  • Speech takes a written prompt or script and a description of the desired voice and music.
  • A toggle removes the background music; clips can run to about eight minutes.
  • Suno cautions that the beta can drift between accents and exaggerate pauses.
  • The launch extends Suno beyond songs while its music generator faces several lawsuits.

Suno puts voice and music in one track

Speech begins with text: an idea, a poem or something the user has written. The user then describes the voice and musical style they want, t2ONLINE reported. Suno says its model generates the spoken words and accompanying music together. That differs from a workflow in which a voice recording and a separate piece of music must be combined afterwards.

Reading a bedtime story aloud could instead start with its written words and a request for a voice over soft piano, one of Suno’s examples. Switching off the music could leave the story as speech alone. For a longer reading, the roughly eight-minute clip length reported for Speech would bound each generated passage.

The feature sits under Create, where Simple mode accepts a short description and Advanced mode allows a custom script. Advanced mode also offers controls for the AI voice’s gender, speech style and the amount of variety in a generation, NewsBytes reported. The Verge reported that a toggle removes the soundtrack. Those controls give a user two ways to begin: describe what should be said, or supply the words directly and adjust how they sound.

Jack Brody, Suno’s chief product officer, described the move as an extension of the company’s ambitions beyond songs. “Music will always be at the heart of Suno and what we build. At the same time, our vision has always extended to other forms of human expression,” he said in the announcement, according to The Verge. Suno calls Speech the first audio model to generate voice and music together as one cohesive track. That ranking is the company’s claim.

Suno’s beta can change an accent

Suno tested Speech with a small group for about a month before opening the beta to all users on web and mobile, t2ONLINE reported. The company says its team tried it on friends’ text messages, voice notes, meditations, poems, pep talks and bedtime stories. Those are examples of use during development, rather than measurements of how reliably it handles each kind of material.

Brody cautioned that a requested British accent can wander towards an Australian one and back. He also warned that “dramatic pauses may be very dramatic.” In speech, timing and accent carry part of the message: a pause can give a sentence emphasis, while an unexpected change of accent can distract from its words. Suno says it will continue developing Speech in response to user feedback.

Generating spoken audio is an established line of work. DeepMind has experimented with deep-learning speech synthesis for about a decade, Adobe offers text-to-speech, and ElevenLabs launched in 2023, t2ONLINE reported. Speech brings Suno’s music generation into that field rather than introducing synthetic speech itself. AI Affairs has also reported on ElevenLabs’ $300 million tender offer at a $22 billion valuation.

Speech follows Suno’s 9 September v6 release

The launch follows Suno’s introduction of v6 music models on 9 September 2026. Suno said it developed that generation with Warner Music Group, BMG and Believe. The company described tools for changing part of a song through a written instruction, combining multiple sources in a request and altering a lyric without rebuilding the whole song. Speech moves from editing or generating songs to producing spoken material, while retaining music as an option.

Suno’s music generator has been the subject of several lawsuits, t2ONLINE reported. The company says it has added checks for unauthorised use of audio and song text submitted to the platform.

Topics: Copyright, Foundation models