Google text-to-speech

Gemini 2.5 Flash TTS

Prompt: Broadcast-style microphone with soft lighting for voice generation

Use Flash TTS for quick voiceovers with prebuilt voices and optional multi-speaker dialogue in a single request.

Inputs
  • Text prompt
  • Voice selection30 Gemini prebuilt voices
  • Multi-speakerUp to 2 speakers in one request
  • Voice cloning
Settings
  • Voice
  • Multi-speaker labels
  • Output formatAlways wav
  • Voice tuning
Outputs
  • Audio outputwav
  • Voice metadata
  • Multi-speaker track
Strengths

Fast, free-to-start Google TTS

Use Flash TTS when you want a fast Google voice that plugs into the same generation surface as your other models.

30 prebuilt voices

Filter by character — bright, firm, breezy, mature, warm — and preview before committing.

Multi-speaker dialogue

Pass a speaker list with up to two voices to synthesize a short dialogue in one pass.

Low latency

Flash tier trades a small amount of quality for faster speech synthesis.

Workflow

A quick voice pass

Flash is the right pick when you want Google voice quality without waiting on a heavier model.

  1. 1

    Write the line

    Keep it short and natural so the prebuilt voice reads it cleanly.

  2. 2

    Pick a voice

    Open the voice picker and filter down to the character you want.

  3. 3

    Switch to Pro if needed

    Move to 2.5 Pro TTS for higher-quality reads when the scene deserves it.

Details

Technical notes

Use these notes as production guidance inside Creativ Studio. Provider availability and exact limits can vary by account, region, and model rollout.

Best for

Fast Google voice synthesis and dialogue prototypes

Inputs

Text and a voice name

Prompt limit

8,000 characters per request

Output

wav

Prompt recipes

Use these as starting shapes, then swap in your own subject, references, and constraints.

Intro line

Welcome to the Creativ Studio demo. Today we'll cover everything you need to ship your first generation.

Two-speaker dialogue

Speaker A: Ready to review the shot list? Speaker B: Yes, let's start from the top.

Questions

Which voices can I use?

Gemini Flash TTS uses a set of 30 prebuilt voices. Open the voice picker to filter by character.

Can I synthesize dialogue?

Yes. Provide a speakers array with up to two voices and speaker labels in the text.

What output format do I get?

Gemini TTS always returns wav audio. Convert to mp3 downstream if delivery requires it.