Meet Soniox Text-to-Speech v2

August 11, 2026 by Soniox Team

Soniox TTS v2 is our most powerful text-to-speech model yet.

It brings extraordinary voice quality, expressive control, exceptional precision, high-quality voice cloning, support for more than 60 languages, and low-latency streaming together in one model.

Soniox TTS v2 is available globally today for $0.70 per generated hour.

Try Soniox TTS v2


Direct every performance

Natural speech is no longer enough. Voice experiences need emotion, personality, and delivery that adapt to the moment.

Soniox TTS v2 introduces powerful audio tags that let you direct how every part of the text should be performed. A voice can whisper, laugh, hesitate, become tense, sound relieved, or shift naturally between emotions within the same passage.

The gift

[excited]Okay okay open it, open it! [gasp][awe]No way is this the [voice cracking][happy]oh my gosh, I can't believe you remembered! [sobs][laughing]I'm crying, I'm actually crying.

0:000:00
The magic trick

[confused]Uhhhh, hi? Wait a second. How did you just DO that?

0:000:00
The breakup

[shaky breath]I... I don't know how to say this. [sad]We've grown apart, haven't we? [trembling voice]I never wanted to hurt you... [sobs][sniffles]please don't make this harder. [quietly]I'm sorry.

0:000:00
The fridge

[disgusted]Ugh, what is that smell. [sniffs][gagging tone]Eww, no [pfft]Pfft, that has been in there for weeks. [sucks teeth]I told you to throw it out.

0:000:00

Audio tags make voice performance programmable while keeping the delivery natural and coherent.

Try audio tags


Natural speech in 60+ languages

Without audio tags, Soniox TTS v2 automatically adapts rhythm, pacing, emphasis, and expression to the meaning of the text.

The same quality extends across more than 60 languages. Soniox was built as a multilingual model, not an English model adapted to other languages after the fact.

Hear the same voice across languages

The voice remains recognizable across languages, with consistent quality, identity, pronunciation, and expressive range.

Mix languages naturally

Real-world text often includes foreign names, international addresses, technical terminology, and phrases from multiple languages. Soniox follows these language changes naturally within one continuous utterance.

English and Spanish

Your reservation is confirmed. Cuando llegues al hotel, muestra este código en recepción: ES-4928.

0:000:00
English and Korean

To enable the new feature, open Settings and select 음성 복제, then tap Continue.

0:000:00
English and Hindi

Your appointment is confirmed for Monday at 11 a.m. कृपया दस मिनट पहले पहुँचें.

0:000:00

No separate model, voice, or request is required when the language changes.

Explore supported languages


Voice cloning that captures the person

A convincing voice clone requires more than matching how someone sounds. It must preserve the characteristics that make the voice recognizable.

Soniox TTS v2 creates a high-quality voice from seconds of audio, preserving the speaker’s identity, accent, rhythm, pacing, personality, and expressive range.

It also works with everyday recordings. Soniox removes background noise, echo, and recording artifacts to create a clean, faithful voice clone, even when the original audio was captured on a phone or outside a studio.

Original phone recording
A casual recording with background noise.
Clean voice clone
The same voice, cloned clean and studio-ready.
Original recording
A short recording of the original speaker.
Voice clone
New speech generated with the cloned voice.

The cloned voice supports the full capabilities of Soniox TTS v2, including audio tags, multilingual speech, language mixing, and real-time streaming.

Explore Soniox Voice Cloning


Exact speech for production

A voice can sound extraordinary and still fail in production if it mispronounces a name, changes a number, or scrambles a verification code.

Soniox TTS v2 combines exceptional voice quality with the precision required for real applications.

Built for every domain

From medicine and science to law, finance, and engineering, Soniox accurately pronounces specialized terminology across fields of human knowledge.

Medicine

The differential diagnosis includes pheochromocytoma, hyperthyroidism, and supraventricular tachycardia.

0:000:00
Science

Deoxyribonucleic acid is transcribed into messenger RNA before translation occurs at the ribosome.

0:000:00
Engineering

The system uses asynchronous replication, Byzantine fault tolerance, and hardware-accelerated cryptography.

0:000:00

Alphanumerics spoken correctly

Phone numbers, email addresses, verification codes, prices, dates, addresses, and identifiers often contain the most important information in a message. They are also among the easiest things for TTS systems to get wrong.

Phone number

You can reach our support team at +1 415 682 9074.

0:000:00
Email address

Send the completed form to alex.chen+support@example.com.

0:000:00
Verification code

Your verification code is 7Q4M9B. I repeat: 7Q4M9B.

0:000:00
Account identifier

Your case number is CX-8047-A19, and the affected device is model XR-12 Pro.

0:000:00
Price

Your total is $1,284.37, including tax and delivery.

0:000:00

Names pronounced naturally

The pronunciation of people, places, brands, and organizations depends on language, origin, and context. Soniox uses that context to pronounce them naturally and accurately.

People

Nguyễn Minh Anh will be joining the meeting shortly.

0:000:00
Places

Your hotel is located near the historic center of Aix-en-Provence.

0:000:00
Brands and technology

The workload runs on NVIDIA H100 GPUs and is orchestrated through Kubernetes.

0:000:00

Test your own text


Built for real-time conversation

Voice agents need to respond quickly, stop cleanly when interrupted, and know exactly what has already been spoken.

Soniox TTS v2 provides low-latency streaming with precise control over audio playback.

Speech starts before the sentence ends
Start playback while text is still streaming, without waiting for the full message.

Precise timing for every character
Character-level timestamps let applications synchronize text and audio, stop cleanly during interruptions, and continue from the correct point.

Faster, more responsive speech
Voice agents should not leave users waiting through unnecessary pauses. Reduce silence between sentences and punctuation while preserving fluent, natural delivery.

Standard pacing
Speaks with natural, unhurried pauses.
Reduced silence
Trims the pauses without rushing the voice.

Together, low-latency streaming, precise timestamps, and reduced silence enable voice applications that respond quickly, handle interruptions cleanly, and feel naturally conversational.


Frontier TTS at $0.70 per hour

Production voice applications can generate millions of hours of speech. At that scale, efficiency becomes part of the product architecture.

Soniox TTS v2 combines a new model architecture, a new audio codec, and an extremely efficient inference engine. Every layer of the system was developed in-house from scratch and optimized together for maximum quality, speed, and compute efficiency.

The result is frontier text-to-speech at $0.70 per generated hour, a fraction of the cost of many leading models, without compromising voice quality, precision, language coverage, cloning, or real-time performance.

View pricing


Available today

Soniox TTS v2 is available today in the United States, Europe, and Japan under the model name tts-rt-v2.

The same model, voices, capabilities, and API are available in every region, with low-latency streaming and regional data processing.

The API is fully backward compatible with tts-rt-v1. To upgrade an existing integration, simply change the model name to tts-rt-v2.

Read the documentation


A new foundation for voice experiences

Soniox TTS v2 brings together a combination of capabilities no other text-to-speech model delivers.

It combines extraordinary voice quality, direct control over emotion and performance, exceptional precision, high-quality voice cloning, consistent quality across more than 60 languages, natural language mixing, low-latency streaming, global deployment, and production pricing of $0.70 per generated hour.

This creates one complete foundation for voice agents, global products, customer experiences, interactive characters, assistive tools, media, games, and entirely new voice applications.

Soniox TTS v2 is available today.

Try Soniox TTS v2