Meet the new Soniox voice library

September 9, 2026 by Soniox Team
Soniox voice library

Soniox Text-to-Speech now includes more than 200 built-in voices, covering a wide range of accents, ages, speaking styles, and use cases.

Every voice works across all 60+ supported languages, so you can choose the voice that best fits your product and use that same voice globally.

A voice you select for an English-speaking application can also speak Spanish, French, Japanese, Hindi, Arabic, German, Italian, Korean, Portuguese, and dozens of other languages while preserving its voice identity and character.

This makes it possible to build products with a consistent voice across languages and markets, without selecting a different voice for every language.


Find the right voice faster

With more than 200 voices, discovery matters.

We added detailed metadata to each voice so you can quickly narrow the library based on the characteristics you need:

  • Gender: perceived speaker gender
  • Age: youthful, middle-aged, or mature
  • Accent: regional or origin characteristics of the voice
  • Use case: applications where the voice works particularly well
  • Style: timbre, energy, personality, and emotional character

Use-case tags include conversational, narration, educational, social media, and entertainment. Style tags include characteristics such as warm, confident, friendly, smooth, soft, calm, bright, energetic, deep, expressive, formal, dramatic, playful, raspy, and breathy.

A single voice can have multiple use-case and style tags, so you can search for combinations such as:

  • Conversational + warm + middle-aged
  • Narration + deep + confident
  • Social media + young + energetic
  • Educational + calm + friendly

Accent stays part of the voice across languages

Accent is an important part of a voice's identity, especially for multilingual applications.

You might choose a British, Indian, Australian, Japanese, Spanish, or another accented voice not because your application only speaks that language, but because that is how you want your product to sound.

The selected voice can then speak across Soniox's 60+ supported languages while retaining the recognizable character and regional coloring of that voice.

For global applications, this means you can maintain a consistent product identity while speaking to users in their own language.


Voices selected for different applications

Different applications place different demands on a voice.

A voice agent typically benefits from a conversational voice that sounds natural during short, back-and-forth interactions.

For customer support, you may prefer a voice that is calm, friendly, clear, and reassuring.

For narration or audiobooks, smoothness, expressiveness, pacing, and longer-form consistency become more important.

For education and training, clarity and measured delivery may matter most.

For social media, advertising, and entertainment, brighter, more energetic, dramatic, playful, or expressive voices may fit better.

We manually selected, listened to, curated, and tagged the voices with these applications in mind. The goal is for the metadata to reflect how the voice actually sounds, so developers can narrow the library quickly instead of auditioning hundreds of voices individually.


Choose by age and gender

Age and gender provide another way to shape the personality of a product.

The Soniox library includes young, middle-aged, and older voices, allowing you to choose between more youthful, settled, or mature characteristics depending on the application.

Combined with accent, style, and use case, these attributes make it much easier to find a small set of voices that fit what you are building.


Or create your own voice with voice cloning

If none of the built-in voices matches the identity you want, Soniox also supports high-quality instant voice cloning.

You can create a custom voice from up to 2 minutes of reference audio.

The longer reference sample gives the model more information about the speaker, including phonemes, intonation, pauses, rhythm, pacing, emphasis, emotion, and speaking style.

This results in:

  • Higher voice similarity
  • Greater consistency, especially across longer sentences
  • Better pronunciation coverage
  • More natural prosody, including rhythm, emphasis, pacing, and emotion
  • Better preservation of the speaker's original style

Cloned voices get the same multilingual capabilities as Soniox TTS, making it possible to create a custom voice and use it across languages rather than limiting it to the language of the original recording.

This gives developers two ways to choose the voice for their product: select from more than 200 curated built-in voices, or create a custom voice of their own.


A new Playground for exploring voices

We also redesigned the Soniox TTS Playground to make voice discovery and experimentation easier.

You can browse and filter the expanded library, test different voices with your own text, switch between languages, and quickly compare how different voice characteristics fit your application.

Because every built-in voice works across all 60+ supported languages, you can test the same voice across the languages your product needs to support.


Built for real applications

The expanded voice library and voice cloning work with the rest of Soniox Text-to-Speech, including natural and expressive speech generation, multilingual speech across 60+ languages, low-latency streaming for real-time applications, and precise control over generated speech.

Whether you start with a built-in voice or clone your own, the goal is the same: find or create the voice that fits your application, then use it across languages and markets.

Explore Soniox Text-to-Speech