Stronger delivery
Better prosody and pacing than the Flash tier, closer to a natural read.
Use Pro TTS when the voiceover is going to ship — stronger delivery and the same picker-driven voice selection as Flash.
Use Pro TTS when the line needs to carry emotion, pacing, and a premium finish.
Better prosody and pacing than the Flash tier, closer to a natural read.
Pass two speakers for short dialogues with distinct voices in one request.
Same 30 Gemini voices with filters for character, language hints, and previews.
Write the line like a performance cue, pick the voice, and re-run until the read matches the scene.
Use Flash TTS to settle the line and voice choice before paying for a Pro render.
Switch to Pro for the final delivery when pacing and quality matter.
Small punctuation and phrasing changes will shift the read more than voice swaps.
Use these notes as production guidance inside Creativ Studio. Provider availability and exact limits can vary by account, region, and model rollout.
Final voice delivery and multi-speaker dialogue
Text and a voice name
8,000 characters per request
wav
Use these as starting shapes, then swap in your own subject, references, and constraints.
This is how teams move from idea to asset. Not next week. Not tomorrow. Today.
Speaker A: We can actually ship this tonight. Speaker B: Then let's stop talking and start recording.
Use Pro when the voice output will ship as final audio. Use Flash for fast drafts and iteration.
Not yet. Pro TTS uses the 30 Gemini prebuilt voices. Custom voice cloning is on the roadmap.
Yes. Provide a speakers array with two distinct voices and labeled lines in the text.