AI Affairs, home

Thursday 24 September 2026

Technology

Google releases Gemini 3.8 Flash TTS and Flash-Lite TTS with support for 100+ languages

The models introduce generative voice design, replication with consent checks, and benchmark leads across multiple languages, rolling out immediately in the Gemini API and AI Studio.

Google sign on Charleston Road at Google headquarters
Photo: Dietmar Rabich, CC BY-SA 4.0, via Wikimedia Commons (cropped)

Google introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on 23 September 2026, calling them its most expressive audio generation models yet. Both models began rolling out the same day in the Gemini API and Google AI Studio.

Key points

  • Two new TTS models launched 23 September 2026, rolling out same day in Gemini API and Google AI Studio
  • Flash TTS targets creative character design; Flash-Lite targets high-volume dubbing and voice agents
  • Over 100 languages supported with 2,000+ production-ready voices including regional varieties
  • Voice replication from 30-second samples includes consent verification, SynthID watermarking, C2PA credentials
  • Google reports top benchmark positions on Hume AI and Voice Arena across multiple languages

Two models for different workloads

Google positions the pair as complementary tools. Gemini 3.8 Flash TTS targets deep creative direction and character design, letting users craft original voices through natural language prompts for a range of creative applications. Gemini 3.8 Flash-Lite TTS is designed for high-volume, cost-efficient workloads such as dubbing and voice agents, offering detailed adjustment of vocal delivery. The announcement, written by Group Product Manager Leland Rechis and Director of Research Science Alan Cowen on behalf of the Gemini Audio Team, presents the models as the latest additions to a Gemini Audio family that already includes 3.5 Live Translate, 3.5 Transcribe, 3.8 Live, and 3.8 Live Extended Thinking.

According to the Gemini 3.8 Audio model card, the TTS variants are based on Gemini 3 Pro, accept text input up to 8K tokens, return audio output up to 64K tokens, and carry a knowledge cutoff of January 2025.

Generative voice design and replication

The release shifts voice generation from a preset library to a generative workspace. Gemini 3.8 Flash TTS can create bespoke voices from scratch using natural language prompts that specify role, accent, and vocal characteristics across more than 100 languages and dialects. Google says the original set of 30 voices has expanded into a library of more than 2,000 production-ready voices, including regional varieties such as Mexican Spanish, Quebec French, and Scots English. The system lets users store and manage their designed voices for consistent output, while a future update will allow modification of pitch, pace, accent and tonal quality on library voices.

Voice replication builds a consistent vocal profile from a 30-second sample of a person’s own voice or a voice they are authorised to employ. The feature includes built-in consent verification requiring a matching consent sample from the voice holder, along with imperceptible SynthID watermarking and C2PA credentials embedded in every generated clip so AI speech stays detectable. Every audio clip the Gemini Audio models generate carries the SynthID watermark.

Reported benchmarks and safety assessment

Google reports that Gemini 3.8 Flash TTS scored 71.4 on Hume AI’s Voice Design Benchmark, taking the overall #1 position, and 60.8 on accent modeling, also #1. Flash TTS and Flash-Lite TTS hold the #1 and #2 spots on Hume AI’s Overall Quality Index. Google states the new models represent a significant advance over Gemini 3.1 Flash TTS, particularly for extended narration and two-voice script handling.

In blind human preference evaluations on Voice Arena, Google says the two models took top positions among competitors in Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi.

Google DeepMind’s frontier safety assessment found no meaningful new capabilities or material performance gains in the Gemini 3.8 Audio models versus Gemini 3.7 Flash, which did not reach any Tracked or Critical Capability Levels under the company’s Frontier Safety Framework. Accordingly, the assessment concludes the Gemini 3.8 Audio models are not expected to reach such levels.

Availability and known limitations

Gemini 3.8 Flash TTS also appears in Gemini Notebook, while Flash-Lite TTS is integrated into Google Vids, and API access for enterprise customers through Gemini Enterprise is slated for a future release. Google AI Studio now includes an audio workspace resembling a voice design studio, letting developers create new vocal profiles from prompts or clone a voice and then guide each line’s performance in a two-voice script editor.

The model card notes that the system can produce hallucinated output and sometimes runs slowly or times out. The release notes that voice replication in AI Studio is restricted in Illinois, Texas, the European Economic Area, the United Kingdom, and Switzerland.

Topics: Foundation models, Inference