Soniox TTS v2 is our most powerful text-to-speech model yet.
It brings extraordinary voice quality, expressive control, exceptional precision, high-quality voice cloning, support for more than 60 languages, and low-latency streaming together in one model.
Soniox TTS v2 is available globally today for $0.70 per generated hour.
Direct every performance
Natural speech is no longer enough. Voice experiences need emotion, personality, and delivery that adapt to the moment.
Soniox TTS v2 introduces powerful audio tags that let you direct how every part of the text should be performed. A voice can whisper, laugh, hesitate, become tense, sound relieved, or shift naturally between emotions within the same passage.
[excited]Okay okay open it, open it! [gasp][awe]No way — is this the — [voice cracking][happy]oh my gosh, I can't believe you remembered! [sobs][laughing]I'm crying, I'm actually crying.
[confused]Uhhhh, hi? Wait a second. How did you just DO that?
[shaky breath]I... I don't know how to say this. [sad]We've grown apart, haven't we? [trembling voice]I never wanted to hurt you... [sobs][sniffles]please don't make this harder. [quietly]I'm sorry.
[disgusted]Ugh, what is that smell. [sniffs][gagging tone]Eww, no — [pfft]Pfft, that has been in there for weeks. [sucks teeth]I told you to throw it out.
Audio tags make voice performance programmable while keeping the delivery natural and coherent.
Natural speech in 60+ languages
Without audio tags, Soniox TTS v2 automatically adapts rhythm, pacing, emphasis, and expression to the meaning of the text.
The same quality extends across more than 60 languages. Soniox was built as a multilingual model, not an English model adapted to other languages after the fact.
Hear the same voice across languages
The voice remains recognizable across languages, with consistent quality, identity, pronunciation, and expressive range.
Mix languages naturally
Real-world text often includes foreign names, international addresses, technical terminology, and phrases from multiple languages. Soniox follows these language changes naturally within one continuous utterance.
Your reservation is confirmed. Cuando llegues al hotel, muestra este código en recepción: ES-4928.
To enable the new feature, open Settings and select 음성 복제, then tap Continue.
Your appointment is confirmed for Monday at 11 a.m. कृपया दस मिनट पहले पहुँचें.
No separate model, voice, or request is required when the language changes.
Voice cloning that captures the person
A convincing voice clone requires more than matching how someone sounds. It must preserve the characteristics that make the voice recognizable.
Soniox TTS v2 creates a high-quality voice from seconds of audio, preserving the speaker’s identity, accent, rhythm, pacing, personality, and expressive range.
It also works with everyday recordings. Soniox removes background noise, echo, and recording artifacts to create a clean, faithful voice clone, even when the original audio was captured on a phone or outside a studio.
The cloned voice supports the full capabilities of Soniox TTS v2, including audio tags, multilingual speech, language mixing, and real-time streaming.
Exact speech for production
A voice can sound extraordinary and still fail in production if it mispronounces a name, changes a number, or scrambles a verification code.
Soniox TTS v2 combines exceptional voice quality with the precision required for real applications.
Built for every domain
From medicine and science to law, finance, and engineering, Soniox accurately pronounces specialized terminology across fields of human knowledge.
The differential diagnosis includes pheochromocytoma, hyperthyroidism, and supraventricular tachycardia.
Deoxyribonucleic acid is transcribed into messenger RNA before translation occurs at the ribosome.
The system uses asynchronous replication, Byzantine fault tolerance, and hardware-accelerated cryptography.
Alphanumerics spoken correctly
Phone numbers, email addresses, verification codes, prices, dates, addresses, and identifiers often contain the most important information in a message. They are also among the easiest things for TTS systems to get wrong.
You can reach our support team at +1 415 682 9074.
Send the completed form to alex.chen+support@example.com.
Your verification code is 7Q4M9B. I repeat: 7Q4M9B.
Your case number is CX-8047-A19, and the affected device is model XR-12 Pro.
Your total is $1,284.37, including tax and delivery.
Names pronounced naturally
The pronunciation of people, places, brands, and organizations depends on language, origin, and context. Soniox uses that context to pronounce them naturally and accurately.
Nguyễn Minh Anh will be joining the meeting shortly.
Your hotel is located near the historic center of Aix-en-Provence.
The workload runs on NVIDIA H100 GPUs and is orchestrated through Kubernetes.
Built for real-time conversation
Voice agents need to respond quickly, stop cleanly when interrupted, and know exactly what has already been spoken.
Soniox TTS v2 provides low-latency streaming with precise control over audio playback.
Speech starts before the sentence ends
Start playback while text is still streaming, without waiting for the full message.
Precise timing for every character
Character-level timestamps let applications synchronize text and audio, stop cleanly during interruptions, and continue from the correct point.
Faster, more responsive speech
Voice agents should not leave users waiting through unnecessary pauses. Reduce silence between sentences and punctuation while preserving fluent, natural delivery.
Together, low-latency streaming, precise timestamps, and reduced silence enable voice applications that respond quickly, handle interruptions cleanly, and feel naturally conversational.
Frontier TTS at $0.70 per hour
Production voice applications can generate millions of hours of speech. At that scale, efficiency becomes part of the product architecture.
Soniox TTS v2 combines a new model architecture, a new audio codec, and an extremely efficient inference engine. Every layer of the system was developed in-house from scratch and optimized together for maximum quality, speed, and compute efficiency.
The result is frontier text-to-speech at $0.70 per generated hour, a fraction of the cost of many leading models, without compromising voice quality, precision, language coverage, cloning, or real-time performance.
Available today
Soniox TTS v2 is available today in the United States, Europe, and Japan under the model name tts-rt-v2.
The same model, voices, capabilities, and API are available in every region, with low-latency streaming and regional data processing.
The API is fully backward compatible with tts-rt-v1. To upgrade an existing integration, simply change the model name to tts-rt-v2.
A new foundation for voice experiences
Soniox TTS v2 brings together a combination of capabilities no other text-to-speech model delivers.
It combines extraordinary voice quality, direct control over emotion and performance, exceptional precision, high-quality voice cloning, consistent quality across more than 60 languages, natural language mixing, low-latency streaming, global deployment, and production pricing of $0.70 per generated hour.
This creates one complete foundation for voice agents, global products, customer experiences, interactive characters, assistive tools, media, games, and entirely new voice applications.
Soniox TTS v2 is available today.
Try Soniox TTS v2