Gemini 3.8 Flash TTS

Type a script, shape each line's emotion, and let Gemini 3.8 Flash TTS deliver natural speech in 130 languages — free to try online.

Gemini 3.8 Flash TTS
Craft expressive speech on the flagship tier, or scale affordably with Flash-Lite TTS
AI Video Prompt Generator

Support

Pro AI Tools

Explore elite tools including the Suno AI Music Generator.

placeholder hero

Why Creators Are Switching to Gemini 3.8 Flash TTS

Released on September 23, 2026, Google's Gemini TTS family pairs a creative flagship with a budget-friendly engine built for bulk audio.

  • Two Models, Two Different Workloads
    Gemini 3.8 Flash TTS handles nuanced, character-rich reads, while Flash-Lite TTS keeps cost and latency low for mass production.
  • Direction Instead of Presets
    Per-turn style notes, structured speech metadata and inline vocal events control tone, tempo, feeling and accent.
  • Design a Voice or Clone With Consent
    Describe the voice you want in plain language, or copy a real speaker using a reference clip plus a matching consent recording.

How to Prompt Gemini 3.8 Flash TTS for Clean Audio

Four habits that keep your transcript readable and let performance metadata do the acting.

What Gemini 3.8 Flash TTS Can Do

From line-by-line performance control to 130-language output, these are the flagship Gemini TTS model's real strengths.

Performance Direction That Feels Like Acting

Per-turn styling plus inline laughs, sighs, coughs, breaths and pauses — you direct a performance rather than pick a preset voice.

Design Voices With Plain Language

Prompts can specify age range, personality, accent, vocal texture and role, supported by 2,000+ production voices through the Voices endpoint.

Replication Locked Behind Consent

A clean reference recording plus a matching consent recording from the same adult speaker, protected by SynthID watermarking and C2PA credentials.

Dialogue Built for Two Speakers

Scripts carry the conversation for podcasts, teaching dialogues, product demos and game scenes without stitching lines by hand.

Steady Voice Across Long Recordings

Google documents consistent identity, timbre, volume and room tone across multi-minute narration and extended back-and-forth dialogue.

130 Languages With Regional Accents

Flash TTS supports 130 languages versus Flash-Lite's 101, including regional accents, minority dialects and IPA pronunciation overrides.

FAQ

Gemini 3.8 Flash TTS: Questions Answered

Quick answers about Gemini 3.8 Flash TTS pricing, choosing a model, benchmark results and safety rules.

1

What does Gemini 3.8 Flash TTS cost per audio minute?

Roughly 1.35 cents per minute, based on launch pricing of $0.50 per million input tokens and $9 per million output tokens.

2

Which tier should I pick, Flash TTS or Flash-Lite TTS?

Choose Flash TTS when acting nuance and long-form audio matter; pick Flash-Lite TTS for bulk jobs and low latency.

3

How does it score against rival voice models?

Google cites 71.4 on Hume's Voice Design Benchmark, and Voice Arena ranks it second with 1,260 Elo.

4

What changed from Gemini 3.1 Flash TTS Preview?

Flash-Lite TTS takes over from the 3.1 preview and drops audio output pricing from $20 to $6 per million tokens.

5

Can I clone a voice, and what safeguards exist?

Cloning requires a reference clip plus a matching consent recording from the same adult speaker.

6

Why is the model reading my stage directions out loud?

Because the input is treated as a verbatim transcript — move lasting directions into speech metadata instead.

Hear Gemini 3.8 Flash TTS Read Your Own Script

Try both tiers in the Gemini API or Google AI Studio — moving between Gemini 3.8 Flash TTS and Flash-Lite TTS is just a model ID swap. Compare batch and priority inference before you lock in a production budget.