Google did not make a lot of noise about it. On September 23, in the same week that attention was locked on chatbots and coding agents, the company quietly rolled out two new voice models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, into the Gemini API and Google AI Studio. No keynote, no splashy demo reel. Just a pricing page, a developer blog post, and more than 2,000 voices sitting behind it.
That understatement is a little misleading, because the numbers underneath are not small. A voice library Google puts at over 2,000 prebuilt profiles. Coverage across 130 languages with automatic detection. A “voice design” tool that turns a plain-English description of a character into a reusable synthetic voice. And a price that works out to well under a dollar an hour of generated audio, for now. Google has already published the date that changes.
The short version
- Gemini 3.8 Flash TTS and Flash-Lite TTS launched September 23, 2026 in the Gemini API and Google AI Studio, Google’s first dedicated text-to-speech release since last year’s Gemini 2.5 TTS models
- The library spans more than 2,000 prebuilt voices across 130 languages, plus a “voice design” feature that builds a new one from a text description
- Introductory pricing runs $0.50 per million text tokens, plus $9 per million audio tokens for Flash TTS or $6 per million for Flash-Lite, which works out to roughly 81 cents and 54 cents per hour of audio
- That pricing is only guaranteed through December 31, 2026. Every rate doubles on January 1, 2027, according to Google’s own published terms
- Flash TTS is aimed at creative work like games, audiobooks and podcasts; Flash-Lite is the cheaper tier built for high-volume dubbing and voice agents
- It is the same playbook Google ran with Gemini 3.6 Flash’s pricing cut in August: an aggressive introductory rate with an expiration date already printed on the label
Two tiers, one library of voices
Google split the release into two models with different jobs. Gemini 3.8 Flash TTS is pitched at anything where the voice itself is part of the product: narration for games, audiobook chapters, podcast segments, a character who needs to sound like a specific kind of person. Gemini 3.8 Flash-Lite TTS drops some of that fidelity in exchange for a lower price, built for the unglamorous, high-volume end of the business: dubbing a back catalog of videos into a dozen languages, or running the voice layer of a customer service agent that fields thousands of calls a day.
Both models draw from the same underlying library, which Google describes as more than 2,000 prebuilt vocal profiles, with a smaller curated set of 30 voices exposed directly for selection in the API and AI Studio at launch. Flash TTS carries automatic language detection across 130 languages, which matters more than it sounds like it should. A dubbing pipeline that has to manually flag which language each clip is in already loses most of the speed advantage that made an AI model worth using in the first place.
The feature Google leaned on hardest in its own materials is voice design: describe a character in plain language, something like an older narrator with a slight rasp and a slow, deliberate pace, and the model produces a synthetic voice matching that description, which can then be reused across a project. It is not a new idea in the voice AI space, ElevenLabs and ReadSpeaker have both shipped versions of it, but Google folding it directly into Gemini’s existing developer tooling, rather than a separate product, is the part worth watching. It turns voice generation into one more thing a Gemini-based application can do inline, instead of a separate integration a developer has to go build.
The pricing, and the date already stamped on it
Here is where the numbers get specific enough to actually plan around. Text input for both models runs $0.50 per million tokens. Audio output is where the two tiers separate: Flash TTS costs $9 per million audio tokens, Flash-Lite costs $6. Converted into something closer to how a producer or developer actually thinks about cost, that comes out to roughly $0.81 per hour of generated audio on Flash TTS and about $0.54 per hour on Flash-Lite.
Those are introductory rates, and Google has been unusually specific about their shelf life. The pricing holds through December 31, 2026. On January 1, 2027, every rate on both models doubles. There is no ambiguity being managed here the way there sometimes is with “promotional pricing” that quietly becomes permanent. Google put a date on the page.
For context on how that compares with the rest of the industry: OpenAI’s standard TTS API runs $15 per million characters, roughly a dollar or two per hour of finished audio depending on how dense the script is, with a higher-fidelity HD tier at $30 per million. ElevenLabs prices per character rather than per token, generally landing between $50 and $100 per million characters across its main voice models, which puts it well above both Gemini tiers on raw cost, though ElevenLabs’ voice cloning and acting controls remain more mature than what Google shipped this week.
| Model | Approx. cost / hour | Notes |
|---|---|---|
| Gemini 3.8 Flash-Lite TTS | ~$0.54 | Introductory rate through Dec 31, 2026. Built for dubbing and voice agents at volume. |
| Gemini 3.8 Flash TTS | ~$0.81 | Introductory rate through Dec 31, 2026. Built for games, audiobooks, podcasts. |
| OpenAI TTS (standard) OpenAI | ~$1 to $2* | $15 per million characters. HD tier runs $30 per million. |
| ElevenLabs (API) ElevenLabs | ~$3 to $6* | $0.05 to $0.10 per 1,000 characters depending on model tier. |
*Per-hour figures for character-billed models are rough estimates based on typical narration pacing and will vary by script density. Prices as published or reported, September 2026.
Why the expiration date is the actual story
A pricing table with a built-in deadline is becoming the default shape of an AI product launch, not an exception to it. Google cut Gemini 3.6 Flash’s price by 50 percent in August, an introductory rate that was likewise set to expire, in that case on January 1, 2027 as well. Anthropic ran a version of the same move in September, leaving Claude’s standard token pricing alone while cutting cached-token costs 75 percent on its newest model, a discount aimed specifically at the repeated-context workloads that make agents expensive to run. The shape is consistent across every major lab: cut the number that scales with volume, put a clock on it, and let developers build fast while the price is low.
For voice specifically, the volume math is unusually direct. A dubbing operation converting a video library into a dozen languages is not making one API call, it is making thousands, one per clip per language. A customer service agent handling calls all day is generating audio continuously, not in the occasional burst a chat interface produces. Cutting the per-hour cost in half turns a workload that was marginal into one that is comfortably worth automating, at least until the rate resets. Google is betting that six-plus months at aggressive pricing is enough time to get developers building products that assume Flash-Lite is the default cost of a synthetic voice, so that January’s doubled rate becomes a cost of doing business rather than a reason to switch providers.
It is also worth noticing what Google did not ship alongside this: no mention of real-time streaming latency figures, no update to how Gemini Live handles voice, no bundling with the Pixel line the way Gemini has been folded into Pixel Watch’s offline assistant features this month. This is a developer-facing API release, aimed squarely at the businesses building dubbing pipelines and voice agents on top of Gemini, not at end users noticing a new voice in a Google app. The audience is whoever is running the invoice, not whoever is listening to the output.
Signals to watch
- What actually happens on January 1, 2027. Gemini 3.6 Flash’s own standard rate is due back the same week; whether Google lets both reversions stand or quietly extends the discount again says a lot about how price-sensitive this market still is
- Whether ElevenLabs or OpenAI answer with their own cuts. Every major price move in AI this year has drawn a competitive response within weeks, not months
- Real-world quality outside Google’s own demos. Google’s promotional material is, unsurprisingly, Google’s best-case audio. Independent comparisons of the 2,000-voice claim against actual usable variety are still thin
- How the “voice design” feature is licensed for commercial reuse. A custom voice built from a text prompt raises the same ownership and consent questions that AI-generated voices have been raising all year, and Google’s terms here have not gotten much outside scrutiny yet
None of this makes Gemini 3.8 Flash TTS a dramatic leap in what synthetic voices can do. Voice cloning, character design and multilingual dubbing have all existed in some form for a few years now, spread across ElevenLabs, Amazon Polly, Microsoft’s Azure voices and Google’s own earlier TTS models. What changed this week is the price of doing all of it inside the same developer account already running Gemini, with a library large enough that most projects will not need to look elsewhere, at least while the introductory rate holds. The real test is not the demo. It is what a dubbing studio’s or a call center’s actual bill looks like in February, once the discount Google is currently offering has expired and Flash-Lite is charging what Google always intended to charge for it.

