Gemini 3.8 Flash TTS (gemini-3.8-flash-tts) is Google's flagship creative
text-to-speech model, engineered for studio-grade voice fidelity, expressive
acting, authentic regional accents, and rock-solid long-form multi-turn
stability.
Overview and capabilities
Gemini 3.8 Flash TTS sets a new benchmark for expressive audio generation:
- High acoustic fidelity and acting nuance: Delivers rich emotional
range, natural cadence, and precise adherence to turn-level
styledirections and inline vocal events (<laugh>,<sigh>,<short pause>). - Long-form multi-turn stability: Maintains consistent voice identity, timbre, volume, and acoustic room tone across extended dialogues and multi-minute narrations without voice drift.
- Authentic regional accents and pronunciation: Supports regional accents, minority dialects, and inline International Phonetic Alphabet (IPA) overrides.
- Full voice ecosystem support: Works seamlessly with prebuilt voices,
the Extended Voice Library (
GET /v1beta/voices), custom Voice design personas, and Voice replication (persistent stored voices by default, plus optional stateless keys). You can also design and replicate voices interactively in Google AI Studio.
Visit the Text-to-speech guide for full coverage of features, prompting best practices, and code examples.
When to use which TTS model
Both Gemini 3.8 TTS models share the same API schema and prompting structure. Choose the model that fits your workload:
| Feature / workload | Gemini 3.8 Flash TTS (gemini-3.8-flash-tts) |
Gemini 3.8 Flash-Lite TTS (gemini-3.8-flash-lite-tts) |
|---|---|---|
| Primary strength | Maximum voice fidelity, acting nuance, and dialect coverage | High throughput, low latency, and cost efficiency |
| Best use cases | Audiobooks, studio narration, complex multi-speaker dialogue, heavy vocal-burst acting, difficult pronunciation, regional dialects | High-volume production, real-time voice agent cascades, read-aloud features, voice replication, everyday single-speaker generation |
| Supported languages | 130 languages | 101 languages |
| Recommended replacement for | New flagship creative tier | gemini-3.1-flash-tts-preview |
Migration guide
If you are migrating from gemini-3.1-flash-tts-preview or earlier Gemini TTS
models, update your requests for the Gemini 3.8 TTS schema:
- Move turn-level directions into
speech_metadata: Gemini 3.8 TTS treats input text strictly as a verbatim transcript. Inline text directions like"Say cheerfully: Hello!"or"Speaker 1: Hello!"may be spoken aloud. Move sustained delivery instructions (style) and speaker labels (speaker) into structured metadata:- Interactions API: Attach an annotation with
"type": "speech_metadata","speaker", and"style"to each text content block. - GenerateContent API: Attach
"speech_metadata": {"speaker": "...", "style": "..."}to eachpart.
- Interactions API: Attach an annotation with
- Use angle-bracket inline tags only for point-in-time vocal events: Keep
momentary non-speech vocalizations and pauses inline in the transcript using
angle brackets (such as
<laugh>,<sigh>,<cough>,<breath>, or<short pause>). Put delivery styles like whispering inspeech_metadata.style. - Specify
speakeron every turn in multi-speaker requests: Every turn in a multi-speaker request must explicitly includespeakerinsidespeech_metadatamatching one of the configured speakers. - Design personas upfront with Voice design: Replace long legacy Audio
Profile / Director's Notes blocks with a custom voice created in
Voice design, then carry that
voice_...ID through your TTS requests with minimal or emptystylestrings. - Account for default WAV (
audio/wav) output on unary requests: Unlikegemini-3.1-flash-tts-previewand earlier TTS models (which returned headerless raw PCMaudio/l16by default), Gemini 3.8 TTS returns WAV audio (audio/wav/AUDIO_WAV) with a standard RIFF header by default for unary requests.- If your code previously wrapped raw PCM bytes in a WAV header (for
example, using Python's
wavemodule orffmpeg), remove the manual header wrapper and write the returned bytes directly to a.wavfile. - If your pipeline requires headerless raw PCM, mu-law, or A-law audio,
explicitly set
response_formatto"audio/l16"("AUDIO_L16"),"audio/mulaw"("AUDIO_MULAW"), or"audio/alaw"("AUDIO_ALAW"). See Audio output formats.
- If your code previously wrapped raw PCM bytes in a WAV header (for
example, using Python's
gemini-3.8-flash-tts
| Property | Description |
|---|---|
| Model code | gemini-3.8-flash-tts |
| Supported data types |
Inputs Text Output Audio |
| Token limits[*] |
Input token limit 8,192 Output token limit 16,384 |
| Capabilities | Supported Supported Not supported Not supported Not supported Not supported Not supported Not supported Not supported Not supported Not supported Not supported |
| Consumption options |
Supported Supported Supported |
| Versions |
|
| Latest update | July 2026 |
Supported languages
gemini-3.8-flash-tts detects the input language automatically and supports
130 languages:
| Language | Language | Language |
|---|---|---|
| Acehnese (Arab script) | Greek | Nepali (individual language) |
| Afrikaans | Guarani | Nigerian Fulfulde |
| Akan | Gujarati | North Azerbaijani |
| Amharic | Haitian Creole | Northern Sotho |
| Armenian | Halh Mongolian | Northern Uzbek |
| Assamese | Hausa | Norwegian Bokmål |
| Awadhi | Hebrew | Norwegian Nynorsk |
| Balinese | Hindi | Nyanja |
| Bangla | Hungarian | Occitan |
| Banjar (Arab script) | Icelandic | Odia (individual language) |
| Banjar (Latn script) | Igbo | Pangasinan |
| Bashkir | Iloko | Persian (Afghanistan) |
| Basque | Indonesian | Polish |
| Belarusian | Iranian Persian | Portuguese |
| Bemba | Italian | Punjabi |
| Bhojpuri | Japanese | Romanian |
| Bosnian | Javanese | Russian |
| Buginese | Kabyle | Santali |
| Bulgarian | Kamba | Serbian |
| Burmese | Kannada | Sindhi |
| Cantonese | Kashmiri (Arab script) | Sinhala |
| Catalan | Kashmiri (Deva script) | Slovak |
| Cebuano | Kazakh | Slovenian |
| Central Kurdish | Khmer | Somali |
| Chhattisgarhi | Kikuyu | South Azerbaijani |
| Chinese (Hans script) | Kinyarwanda | Southern Pashto |
| Chinese (Hant script) | Kongo | Southern Sotho |
| Crimean Tatar | Korean | Spanish |
| Croatian | Kyrgyz | Standard Arabic (Arab script) |
| Czech | Lao | Standard Arabic (Latn script) |
| Danish | Latgalian | Standard Latvian |
| Dutch | Lingala | Standard Malay |
| Dyula | Lithuanian | Swahili (individual language) |
| Dzongkha | Luxembourgish | Swati |
| Egyptian Arabic | Macedonian | Swedish |
| English | Magahi | Tajik |
| Estonian | Maithili | Tamil |
| Filipino | Malayalam | Telugu |
| Finnish | Maltese | Thai |
| French | Manipuri | Tigrinya |
| Galician | Marathi | Tosk Albanian |
| Ganda | Minangkabau (Arab script) | Uyghur |
| Georgian | Minangkabau (Latn script) | — |
| German | Mizo | — |