All posts

#gemini#artificial intelligence#text-to-speech

Gemini 3.8 Text-to-Speech: What Do the New Models Offer?

Google's new Gemini 3.8 text-to-speech models cut prices while boosting quality. We examined the differences between Flash and Flash-Lite.

Gemini 3.8 Text-to-Speech: What Do the New Models Offer?

Gemini 3.8 text-to-speech is the voice generation system Google introduced on September 23, 2026. It comes in two versions: Flash TTS and Flash-Lite TTS. For businesses, it lowers costs and raises quality across voice assistant, voiceover, and customer experience projects.

In short:

  • Two models: Flash TTS (high quality) and Flash-Lite TTS (high volume, low cost)
  • Support for 100+ languages, 2000+ prebuilt voices, and custom voice design
  • Price dropped to $9 per million tokens, roughly 55 percent cheaper than the previous model
  • Accessible via Google AI Studio and the Gemini API, includes SynthID watermarking

What is Gemini 3.8 text-to-speech?

This system is Google DeepMind's newest generation of voice generation models. It turns text into speech that sounds natural and close to a human voice. The model family consists of two products: Flash TTS and Flash-Lite TTS (DeepMind).

These models make it possible to design voices using natural language. Line-by-line performance direction lets you control the tone, pace, and emotion of the speech. They also support more than 100 languages.

For Turkish-speaking businesses, this makes it easier to build multilingual voice products. A single model can serve different markets at once.

What's the difference between Flash TTS and Flash-Lite TTS?

The Flash version is designed for custom voice design and high-quality output. It can create an entirely new voice from scratch, or clone an existing voice from just a 30-second audio sample (DeepMind).

Flash-Lite TTS serves a different purpose. It's focused on high-volume, cost-efficient production scenarios. It's suited for voice agents, automated read-aloud features, and bulk content generation.

In short, Flash is built for studio-quality work, while Flash-Lite is built for operations at scale. Businesses can choose between the two based on their needs.

How has the pricing changed?

This model family has become noticeably cheaper compared to the previous version. The new price sits at $9 per million tokens, effective as of late 2026 (OrcaRouter).

The previous Gemini 3.1 Flash TTS Preview model cost $20 per million audio tokens. That amounts to roughly a 55 percent price cut.

TTS Fiyatı (milyon token, $)
Gemini 3.8 Flash TTS9
Gemini 3.1 Flash TTS Preview20

Kaynak: OrcaRouter

Audio is calculated at roughly 25 tokens per second. That works out to a cost of around 0.02 cents per second, which is quite low. For businesses producing high volumes of voice content, this difference offers a meaningful budget advantage.

How does it compare to competitors on quality?

These models performed well in an independent quality test. Flash TTS ranked first overall in the Hume AI Voice Design Benchmark, with a score of 71.4 (DeepMind).

Hume AI Voice Design Benchmark
Gemini 3.8 Flash TTS* (genel)71.4
Gemini 3.8 Flash TTS* (aksan)60.8

Kaynak: DeepMind

It also leads in the accent modeling category, with a score of 60.8. These results show that the model isn't just fast — it also produces high-quality audio. This is a significant advantage for projects that require precise accent and intonation.

What are the standout features?

This system allows for adding realistic conversational textures. Thanks to a back-channeling feature, you can insert natural vocal responses like "|mhm|" or "|yeah|" (DeepMind).

Controllability is another standout aspect. Style, accent, pace, and tone can be directed through structured speech_metadata descriptions and inline speech tags (Google AI Developers).

Below are the key technical limits and features:

Feature Detail
Language support 100+ languages
Prebuilt voice options 2000+
Custom voice sample length 30 seconds
Multi-speaker support Exactly 2 speakers
Maximum output duration Roughly 655 seconds

Multi-speaker configuration currently supports only two speakers (Firebase). The maximum output duration is roughly 655 seconds; text beyond that limit gets cut off (Google Cloud).

How is safety and verification handled?

This model family includes additional safeguards around voice identity verification. The models come with built-in permission verification, which makes unauthorized voice cloning harder to pull off.

It also uses C2PA credentials and SynthID watermarking (DeepMind). These two technologies help verify that generated audio was created by AI.

For businesses, this matters both for legal compliance and for user trust. It builds a layer of defense against fake voice content.

Where can you access it?

These models are available through Google AI Studio and the Gemini API. Developers can use these tools to add voice generation to their existing applications.

For teams building voice assistants, e-learning platforms, or customer service bots, integration is relatively quick. Similarly, anyone curious about the price-performance balance across AI models can check out our Claude Opus 5.5 analysis.

Frequently asked questions

How much cheaper is Gemini 3.8 TTS compared to previous models?

The new price is $9 per million tokens, compared to $20 for the previous Gemini 3.1 Flash TTS Preview. That's roughly a 55 percent reduction.

Should I choose Flash or Flash-Lite?

Flash is suited for studio work that requires custom voice design and high quality. Flash-Lite is designed for high-volume, cost-efficient production scenarios.

How do you create a custom voice?

You can create an entirely new voice from scratch using natural language instructions. Alternatively, you can clone an existing voice using just a 30-second audio sample.

Which languages does it support?

The system supports more than 100 languages, including regional accents. It also offers more than 2000 prebuilt voice options.

With its lower price and advanced control options, this model family offers a practical choice for voice products. If you'd like to design the right voice experience for your business, you can discuss these kinds of integrations with the EngerekTech team.

Sources

Source: orcarouter.ai

ShareLinkedInXWhatsApp
Need help with this?

If you would like to apply what this post covers to your own project, let’s look at it together.

Write to us
YE
Yunus Emre Şenyiğit

From the EngerekTech team. We build web, mobile and enterprise software for businesses and share what we learn here.