← 深度专栏/产品观察
产品观察

The 3-Cent Radio Play: Inside Gemini 3.8's Text-to-Speech

What happens when you combine three different AI models to create a mini radio play? You get a glimpse into the hyper-efficient, frictionless future of content...

潜
作者
潜龙编辑部
关注 AI 与社会议题
发布于
2026/10/4
READ
长读
The 3-Cent Radio Play: Inside Gemini 3.8's Text-to-Speech
illustration · QianLong editorial

What happens when you combine three different AI models to create a mini radio play? You get a glimpse into the hyper-efficient, frictionless future of content creation.

Google recently rolled out its Gemini 3.8 text-to-speech (TTS) models, specifically the Flash-TTS and Flash-Lite-TTS versions. While AI-generated voices are not new, the sheer scale and accessibility of this release are turning heads. The models boast a staggering library of over 2,000 preset voices. More impressively, they allow users to create custom voice clones using just a 30-second audio sample—provided, as Google stipulates, that the user has the legal rights to use that voice.

Beyond the vast library, the real magic lies in how the API handles complex audio. Historically, producing a scene with multiple AI voices meant generating separate audio files and manually stitching them together in an editing suite. The Gemini 3.8 API is designed to effortlessly process full, multi-character conversations in a single go, allowing creators to assign distinct voices and stylistic instructions to different speakers simultaneously.

Developer Simon Willison recently demonstrated this capability by generating a one-minute-and-eighteen-second audio clip of two pelicans debating a move to a new pier. The generation took a mere 20 seconds and cost just 2.74 cents using the higher-end Flash model. When producing a minute of professional-sounding dialogue costs less than three cents, independent game developers, solo podcasters, and small educational platforms gain production capabilities that once required a rented studio and a roster of voice actors.

Perhaps the most fascinating aspect of Willison's experiment is the workflow itself. He didn't write the code for his testing interface; he used GPT-6 Astra to "vibe code" it. He didn't write the script, either; he handed that task to Claude 4.5 Opus. Gemini 3.8 simply stepped in as the final voice actor. This seamless handoff between different AI ecosystems highlights a shift in how digital tools are used. We are moving past single-model interactions into an era where AI models act as specialized collaborators on an automated assembly line.

As high-quality voice synthesis becomes cheaper than a piece of candy and faster than brewing a cup of coffee, the barriers to rich audio production are effectively gone. However, this accessibility also brings ethical guardrails into sharp focus. As the technology becomes entirely frictionless, enforcing vocal copyright and preventing deepfakes will be the next major hurdle for the industry. The challenge for tomorrow's creators won't be how to produce multi-character audio, but rather what stories are actually worth telling.

Key Points

  • Google's new Gemini 3.8 TTS models offer 2,000+ voices and 30-second custom voice cloning.
  • The API is specifically designed to handle full multi-character conversations with distinct styles in one request.
  • Audio generation is hyper-efficient, costing just 2.74 cents for over a minute of high-quality dialogue.
  • The developer ecosystem increasingly relies on chaining different AI models (like GPT-6, Claude 4.5, and Gemini) for seamless workflows.

Why It Matters

By driving the cost and time of audio production down to near zero, this technology democratizes multimedia creation while amplifying the urgent need for robust voice copyright enforcement.


Sources:

潛
本文完
潜龙编辑部 · 2026/10/4
潜龙 QianLong · 中文 AI 内容与工具平台