Gemini 3.8 text-to-speech says hello

Hacker News by 7 min read 66x views
Gemini 3.8 text-to-speech says hello

Share Post

Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are our most expressive audio generation models yet. Generate tradition character voices and straightforward environment conversation throughout Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids.


Leland Rechis

Group Product Manager

Alan Cowen

Director, Research Science, on Behalf of the Gemini Audio Team


a content cardstock depiction study "Introducing Gemini 3.8 Flash TTS and 3.8 Flash-Lite TTS"

Today, we’re introducing two new text-to-speech models to the Gemini family, transforming sound generation from fixed presets into a energetic imaginative studio. These models allow creators, developers, and enterprises to create richer, additional expressive audio experiences, during enabling improved person experiences in products akin Gemini Notebook and Google Vids.

  • Gemini 3.8 Flash TTS: Built for profound imaginative direction and character design. Create entirely new voices from scratch using natural tongue prompts to bring characters to existence throughout gaming, immersive audiobooks, podcasts, and interactive media. Direct all achievement row by row alongside granular authority complete acting cues, pacing, dialect shifts, and backchanneling.
  • Gemini 3.8 Flash-Lite TTS: Built for high-volume, cost-efficient scale. Optimized for high-volume dubbing, audio satisfied creation, and expressive sound agents alongside fine-grained authority complete tone, pacing, and expressive nuance.

These models complement our fast-growing Gemini Audio family, following 3.5 Live Translate, 3.5 Transcribe, 3.8 Live, and 3.8 Live Extended Thinking.

Create and customize your own voices

Scale up from 30 first voices to an infinite library. Whether you need an entirely first character sound or a accordant brand ambassador, our 3.8 Flash TTS example powers a complete vocal studio. This enables you to create and use expressive, natural-sounding voices for all moment, during empowering developers and enterprises to effortlessly build tradition audio experiences.

  • Generative sound design: With Gemini 3.8 Flash TTS, create bespoke voices from scratch by customizing role, accent and sound characteristics throughout additional than 100 languages and dialects using natural tongue prompting — whether you're bringing a dramatic, fire-breathing dragon to existence or crafting a influential narrator alongside a distinct local cadence.
  • Expansive sound library: Access 2,000+ production-ready voices alongside broad tongue safety — including local varieties akin Mexican Spanish, Quebec French, and Scots English.
  • Voice replication: Recreate accordant vocal profiles from fair a 30-second audio example of your sound or a sound you have the entitlements to use, backed by built-in consent verification, SynthID watermarking, and C2PA credentials to defend the two developers and their vocal talent.
  • Save and scale: Save and oversee the tradition voices you designed to justify accordant achievement and minimal drift throughout ongoing projects.
  • Voice remixing: Coming soon, choice a sound from our sound archive and fine-tune timbre, pitch, pace, and accent. Use prompts to dial in characteristics (e.g. “add subtle Southern US accent” or “soften the delivery”).

Direct the performance, row by line

Once you've selected your voices, the two TTS models provision you exact authority complete how all row is delivered.

  • Direct achievement row by line: Write your own phase directions or let Gemini steer shipment alongside natural manuscript cues — from a calm client assistance delegate to a whispered suspense scene.
  • Long-form generation: Maintain elevated sound quality, natural pacing, and character timbre throughout hours of uninterrupted audio alongside minimal speaker drift — ideal for podcasts and audiobooks.
  • Native two-speaker environment staging: Direct multi-turn conversations seamlessly from a sole manuscript —whether for a podcast or theatrical storytelling—while keeping the two voices distinctly divided alongside natural conversational turn-taking.
  • Scripted vocal bursts & backchanneling: Add realistic conversational texture using non verbal cues (like <laughs>, <sigh>, <gasp> and active-listening interjections (like |mhm| or|yeah|) for exact comedic timing and reply beats.

Get expressive high-quality address generation built for earth scale

Gemini 3.8 Flash TTS delivers foremost sound customization capabilities, securing the #1 general place on Hume AI’s Voice Design Benchmark (71.4) and additionally foremost in accent modeling (60.8).

Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS allow really expressive performances without sacrificing reliability, additionally securing the #1 and #2 spots respectively on Hume AI’s Overall Quality Index. The example shows important improvements on a broad range of use cases specified as long-form satisfied and dual-speaker screenplay authority compared to Gemini 3.1 Flash TTS.

In blind individual penchant evaluations on Voice Arena, Gemini 3.8 Flash and Flash-Lite TTS safe top positions amongst competitors in key earth languages, including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic (MSA), Mexican Spanish and Hindi. With assistance for complete 100 languages, these models empower creators, developers, and enterprises to build high-quality, multilingual sound experiences worldwide.

An evaluation showing text-to-speech norm benchmark Hume AI

an evaluation diagram showing content to address sound scheme leaderboard Hume AI

an evaluation diagram showing content to address leaderboard for Voice Arena

We built our sound innovation and replication capabilities alongside strict safeguards to assistance defend sound talent, regard identity, and justify satisfied transparency. For sound replication our scheme leverages consent verification: users must provision a verbal consent record from the sound owner that matches the citation speaker before a sound can be created.

More broadly, all audio clip generated by our Gemini Audio models is watermarked with SynthID. This imperceptible watermark is woven immediately into the audio output, ensuring AI-generated address remains detectable to assistance forestall misinformation. For additional particulars on our method to safety and responsibility, assessment the model card.

Try our new Google AI Studio audio playground

Starting today, developers can cognition these new address generation capabilities in Google AI Studio. Built akin a sound scheme workspace, you can immediate entirely new vocal identities from scratch or replicate your own voice 1 , afterward bring them immediately into a dual-speaker screenplay publishing company to straightforward line-by-line delivery.

Try sound replication in Google AI Studio.

Deploy high-performance sound interfaces alongside ease

By using the Gemini API, developer platforms specified as Agora, LiveKit, Pipecat, Vercel allow developers to build and deploy high-performance address generation experiences alongside ease.

We’re partnering alongside companies akin Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang, who are integrating our latest TTS models to assistance accelerate earth dubbing, localize media alongside nuanced local accents, and power conversational sound agents at scale.

a citation from Darius Cheung, CEO and Co-Founder of 99 Group

a citation from Mason Adams, Developer Evangelist, Agora

a citation from Jonathan Gur-Zeev, Director of Product, Figma Weave

a citation from Bin Liu, VP of Engineering for Hygen

a citation cardstock from Luke Pane, Developer, katsuyo

quote from Ritwik Baranwal, Associate Director AI/ML of kuku.

a citation from Oded Shafran, Co-Founder & CTO of linguana

Aziz Ulak, CTO & Co-founder, Ollang

a citation from Phil Marshall, Founder and CEO of Spoken

quote from Zina Rahman, Co-Founder and CEO, Transforms.AI

quote from Mei Ki Yiu, CTO of Wondercraft

Start using our latest Gemini Audio models:

Gemini 3.8 Flash TTS is rolling out starting today:

Gemini 3.8 Flash-Lite TTS is rolling out starting today:

Get the latest news from Google in your inbox

Sign up for our newsletters alongside merchandise updates, event information, particular offers, and more.

Your data volition be used in accordance alongside Google's privacy policy. You may opt out at any time.

Other Article Hacker News
Close Right Ads
Close Left Ads