Drafted with AI tools and edited by a human. The figures and conclusions were checked by a GPUniq editor.
TL;DR
Gemini 3.8 Flash TTS costs $0.50 per 1M input tokens and $9.00 per 1M audio output tokens through December 31, 2026, then doubles to $1.00 and $18.00 from January 1, 2027. According to the Hume AI leaderboard, it scores 0.920 on the Real-World VoiceEQ Frontier Expressivity-Reliability benchmark, topping 34 competing models. No live streaming. No live streaming.
Pricing structure
Two tiers, one cliff.
| Tier | Input, $/1M tokens | Audio output, $/1M tokens | Effective dates |
|---|---|---|---|
| 2026 standard | $0.50 | $9.00 | through 2026-12-31 |
| 2027 standard | $1.00 | $18.00 | from 2027-01-01 |
Both line items double on January 1, 2027. No graduated tier between them.
The context window is 8,192 tokens and max output is 16,384 tokens per request. General availability launched September 22, 2026.
The math is simple. A job consuming 10M input tokens and 5M audio output tokens costs $5.00 + $45.00 = $50.00 in 2026. From January 2027 that same job costs $10.00 + $90.00 = $100.00. Front-load as much 2026 processing as your pipeline allows.
Expressivity vs. rivals
According to the Hume AI leaderboard (September 2026), Gemini 3.8 Flash TTS sits at the top of a 34-model field.
| Benchmark | Score | Field size |
|---|---|---|
| Hume Real-World VoiceEQ Frontier Expressivity-Reliability | 0.920 | 34 models |
| Hume Voice Design | 71.4 | multiple models |
| Hume Accent Modeling | 60.8 | multiple models |
The Voice Design and Accent Modeling scores come from BenchLM Voice Benchmarks. Note that 0.920 and 60.8 are on different scales; comparing them directly is meaningless. What matters: expressivity is the model's strongest suit, accent coverage is decent but not dominant.
The competitive set in the Hume benchmark includes 34 models.
What changed from 3.1 Preview
The predecessor is gemini-3.1-flash-tts-preview. Four concrete areas improved.
Fidelity and prosody. Long-form content shows better naturalness.
Accent authenticity. The Hume Accent Modeling score of 60.8 reflects the model's performance on regional accent coverage, as measured by BenchLM's launch evaluation.
Language coverage. Support expanded to 130+ languages with automatic per-sentence input detection. A single request can mix languages and the model detects each segment without explicit tagging.
API compatibility. No breaking changes. Preview users migrate to gemini-3.8-flash-tts with minimal refactoring. The same API structure is shared with Gemini 3.8 Flash-Lite TTS.
Real-time streaming
No. Batch only. There's no live streaming or real-time API call path.
That's fine for pre-recorded content, podcasts, and audiobooks. It rules out interactive voice response systems needing sub-second response and any real-time conversational agent where latency matters.
If you need lower latency and can accept a fidelity trade-off, Gemini 3.8 Flash-Lite TTS shares the same API structure and is the documented lower-cost alternative.
Voice control and accent features
The model ships with pre-built speaker variants covering a range of profiles, available through the full voice ecosystem.
Fine-grained control comes through inline vocal events: <laugh>, <sigh>, <short pause>, and similar tags. Two-speaker dialogue scenes with overlapping backchannels are supported natively.
One hard limit: speaker similarity in voice replication lags compared to leading models despite strong naturalness. You can use the Extended Voice Library and custom voice design tools in Google AI Studio, including voice replication with consent verification. Speaker similarity in voice replication lags compared to leading models in that capability, despite strong naturalness.
Language and voice coverage
130+ languages, with automatic input language detection. Automatic input language detection is built in, so a request mixing English and Japanese gets each segment handled correctly without manual tagging.
Default audio output is WAV (RIFF format), with alternatives including audio/l16, mulaw, and alaw.
Known limitations
Key limitations to know before deploying:
- No live streaming or real-time API calls. Batch only.
- No function calling, grounding via search, or structured output reasoning inside the TTS pipeline.
- Voice replication requires consent verification; speaker similarity in replication lags leading models. Speaker similarity in replication is a documented weak point.
- Embedded instructions in text may be spoken verbatim if not handled correctly. Multi-speaker dialogue requires explicit speaker and style metadata.
- Fine control over volume and young-adult voice accuracy is less reliable than some competitors.
- Max output is 16,384 tokens per request. Jobs longer than that need chunking.
- Accent modeling (60.8) trails expressivity (0.920) by a meaningful margin on their respective scales. Niche dialects will disappoint.
- Pricing doubles January 1, 2027.
Should you migrate from 3.1 Preview?
If long-form audio is your core product, yes. The fidelity gains in multi-paragraph synthesis and the top-of-field expressivity score (0.920 across 34 models) are real. The 130+ language coverage with auto-detection removes a class of preprocessing work.
If you need real-time streaming, don't migrate yet.
The 2027 pricing cliff is the biggest planning item. Budget a 2x cost increase from January 1. The 2026 rates of $0.50/1M input and $9.00/1M audio output are gone on January 1, 2027. The preview deprecation date has not been published. I'd migrate workloads before that's announced rather than after.
Related on the GPUniq blog
Methodology
Facts are compiled from the vendor's own announcement and documentation, linked below, at the time of writing; where the vendor hasn't published a number it is marked as such. GPUniq prices, when the model is already available here, come from our live catalog.
Sources
- 1.Gemini 3.8 Flash TTS | Gemini API Model Card — Google AI for Developers (accessed )
- 2.Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are our most expressive audio generation models yet — Google Blog (accessed )
- 3.Newly-released Google’s Gemini 3.8 Flash TTS tops Hume’s Real-World VoiceEQ leaderboard — Hume AI (accessed )
- 4.Gemini 3.8 Flash TTS Launch Evaluation — BenchLM.ai (accessed )
Want cheap GPUs for your next project?
Browse live GPU prices and rent the right card in seconds — H100, A100, RTX 4090, and 50+ more models.
