Meet Soniox Text-to-Speech v2

August 11, 2026 by Soniox Team

Soniox TTS v2 is our most powerful text-to-speech model yet.

It brings extraordinary voice quality, expressive control, exceptional precision, high-quality voice cloning, support for more than 60 languages, and low-latency streaming together in one model.

Soniox TTS v2 is available globally today for $0.70 per generated hour.

Try Soniox TTS v2


Direct every performance

Natural speech is no longer enough. Voice experiences need emotion, personality, and delivery that adapt to the moment.

Soniox TTS v2 introduces powerful audio tags that let you direct how every part of the text should be performed. A voice can whisper, laugh, hesitate, become tense, sound relieved, or shift naturally between emotions within the same passage.

The gift
0:000:00
The magic trick
0:000:00
The breakup
0:000:00
The fridge
0:000:00

Audio tags make voice performance programmable while keeping the delivery natural and coherent.

Try audio tags


Natural speech in 60+ languages

Without audio tags, Soniox TTS v2 automatically adapts rhythm, pacing, emphasis, and expression to the meaning of the text.

The same quality extends across more than 60 languages. Soniox was built as a multilingual model, not an English model adapted to other languages after the fact.

Hear the same voice across languages

English
Spanish
Japanese
French
German
Italian
Portuguese
Chinese
Korean
Arabic

The voice remains recognizable across languages, with consistent quality, identity, pronunciation, and expressive range.

Mix languages naturally

Real-world text often includes foreign names, international addresses, technical terminology, and phrases from multiple languages. Soniox follows these language changes naturally within one continuous utterance.

English and Spanish
0:000:00
English and Korean
0:000:00
English and Hindi
0:000:00

No separate model, voice, or request is required when the language changes.

Explore supported languages


Voice cloning that captures the person

A convincing voice clone requires more than matching how someone sounds. It must preserve the characteristics that make the voice recognizable.

Soniox TTS v2 creates a high-quality voice from seconds of audio, preserving the speaker’s identity, accent, rhythm, pacing, personality, and expressive range.

It also works with everyday recordings. Soniox removes background noise, echo, and recording artifacts to create a clean, faithful voice clone, even when the original audio was captured on a phone or outside a studio.

Original phone recording
A casual recording with background noise.
Clean voice clone
The same voice, cloned clean and studio-ready.
Original recording
A short recording of the original speaker.
Voice clone
New speech generated with the cloned voice.

The cloned voice supports the full capabilities of Soniox TTS v2, including audio tags, multilingual speech, language mixing, and real-time streaming.

Explore Soniox Voice Cloning


Exact speech for production

A voice can sound extraordinary and still fail in production if it mispronounces a name, changes a number, or scrambles a verification code.

Soniox TTS v2 combines exceptional voice quality with the precision required for real applications.

Built for every domain

From medicine and science to law, finance, and engineering, Soniox accurately pronounces specialized terminology across fields of human knowledge.

Medicine
0:000:00
Science
0:000:00
Engineering
0:000:00

Alphanumerics spoken correctly

Phone numbers, email addresses, verification codes, prices, dates, addresses, and identifiers often contain the most important information in a message. They are also among the easiest things for TTS systems to get wrong.

Phone number
0:000:00
Email address
0:000:00
Verification code
0:000:00
Account identifier
0:000:00
Price
0:000:00

Names pronounced naturally

The pronunciation of people, places, brands, and organizations depends on language, origin, and context. Soniox uses that context to pronounce them naturally and accurately.

People
0:000:00
Places
0:000:00
Brands and technology
0:000:00

Test your own text


Built for real-time conversation

Voice agents need to respond quickly, stop cleanly when interrupted, and know exactly what has already been spoken.

Soniox TTS v2 provides low-latency streaming with precise control over audio playback.

Speech starts before the sentence ends
Start playback while text is still streaming, without waiting for the full message.

Precise timing for every character
Character-level timestamps let applications synchronize text and audio, stop cleanly during interruptions, and continue from the correct point.

Faster, more responsive speech
Voice agents should not leave users waiting through unnecessary pauses. Reduce silence between sentences and punctuation while preserving fluent, natural delivery.

Standard pacing
Speaks with natural, unhurried pauses.
Reduced silence
Trims the pauses without rushing the voice.

Together, low-latency streaming, precise timestamps, and reduced silence enable voice applications that respond quickly, handle interruptions cleanly, and feel naturally conversational.


Frontier TTS at $0.70 per hour

Production voice applications can generate millions of hours of speech. At that scale, efficiency becomes part of the product architecture.

Soniox TTS v2 combines a new model architecture, a new audio codec, and an extremely efficient inference engine. Every layer of the system was developed in-house from scratch and optimized together for maximum quality, speed, and compute efficiency.

The result is frontier text-to-speech at $0.70 per generated hour, a fraction of the cost of many leading models, without compromising voice quality, precision, language coverage, cloning, or real-time performance.

View pricing


Available today

Soniox TTS v2 is available today in the United States, Europe, and Japan under the model name tts-rt-v2.

The same model, voices, capabilities, and API are available in every region, with low-latency streaming and regional data processing.

The API is fully backward compatible with tts-rt-v1. To upgrade an existing integration, simply change the model name to tts-rt-v2.

Read the documentation


A new foundation for voice experiences

Soniox TTS v2 brings together a combination of capabilities no other text-to-speech model delivers.

It combines extraordinary voice quality, direct control over emotion and performance, exceptional precision, high-quality voice cloning, consistent quality across more than 60 languages, natural language mixing, low-latency streaming, global deployment, and production pricing of $0.70 per generated hour.

This creates one complete foundation for voice agents, global products, customer experiences, interactive characters, assistive tools, media, games, and entirely new voice applications.

Soniox TTS v2 is available today.

Try Soniox TTS v2