Text-to-speech API for language mixing and multilingual speech

Generate mixed-language speech in one request. Switch languages mid-sentence, pronounce foreign names and phrases naturally, and keep the same voice across 60+ languages without splitting text or stitching audio.

$0.70 per generated hour.

Trusted by teams building global voice products

Livekit
Krisp
Pipecat
Summary AI
Perplexity
Samsung
Wispr Flow
LG
Agora
Retell AI
Fireflies.ai
Skit.ai
Kindroid
Deliver Health
Truecaller
Journalia
Mobius
TranscribeMe
Vapi
Zomato
SLNG
Japan AI
Boost.ai
Convin
Genspark
HappyRobot
Uniscribe
Jamie.ai
InteractCX
The Plato
MobilApp
Onvego
Wonderful.ai
Manifone
Tana
Transync AI
SotaTek

Mixed-language TTS without splitting or stitching

Real-world text rarely stays inside one language. Customer names, addresses, brands, technical terms, quotes, and entire phrases often come from somewhere else. Soniox Text-to-Speech API speaks mixed-language input as one continuous utterance, so your application does not need to split text by language and combine separate audio clips.

Switch languages naturally mid-sentence

Send text containing multiple languages in the same request. Soniox follows the language changes while preserving one continuous voice and delivery.

1 / 2
English and Spanish

Your reservation is confirmed. Cuando llegues al hotel, muestra este código en recepción: ES-4928.

English and French

The meeting has been moved to Tuesday. Merci de confirmer votre disponibilité avant la fin de la journée.

English and Korean

To enable the new feature, open Settings and select 음성 복제, then tap Continue.

English and Japanese

Your order is ready for pickup. 店頭で注文番号 A-7392 をお見せください.

Handle code-switching in one utterance

Build speech for bilingual and multilingual content where languages change naturally inside a sentence or response. Keep the full utterance together instead of routing each language segment through a separate TTS request.

Send mixed text

Send the complete multilingual sentence, message, or response to Soniox Text-to-Speech API as one input.

Keep one request

Set the overall delivery language without splitting every foreign name, phrase, or language change into another request.

Generate one voice

Receive continuous speech that preserves the same speaker identity as the utterance moves between languages.

Pronounce foreign names, places, and terms naturally

Keep names, addresses, brands, product terms, and borrowed words inside the surrounding text. Soniox pronounces them naturally without requiring you to change the language setting for every embedded term.

Multilingual names and places

Your meeting with Nguyễn Minh Anh is scheduled for 3 p.m. at our office near Place de la République in Paris.

International product terms

Open the Soniox Developer Console, select Text-to-Speech, and set the output to 24-kilohertz PCM before deploying on Kubernetes.

Person name

Your appointment with Siobhan O’Connor is confirmed.

Person name

Nguyễn Minh Anh will be joining the meeting shortly.

Place

Our Asian engineering office is located in Guangzhou.

Place

Your hotel is located near the historic center of Aix-en-Provence.

Set the delivery language, not every language switch

The language field controls the overall delivery and applies a light accent bias to the voice. Embedded words and phrases from other languages can still be pronounced naturally inside the same utterance.

English delivery

English is the primary delivery language, so the utterance follows English prosody. The Spanish street name and title remain part of the same request and are pronounced naturally.

{
  "model": "tts-rt-v2",
  "language": "en",
  "voice": "Adrian",
  "text": "Please deliver the package to Calle de Alcalá 45 in Madrid, attention Señora García."
}

Spanish delivery

The text stays exactly the same, but Spanish now controls the overall delivery. The same voice speaks the complete mixed-language utterance with Spanish driving its prosody.

{
  "model": "tts-rt-v2",
  "language": "es",
  "voice": "Adrian",
  "text": "Please deliver the package to Calle de Alcalá 45 in Madrid, attention Señora García."
}

Use the main language of the text

Choose this when most of the utterance is in one language and names, addresses, technical terms, or shorter phrases come from other languages.

Or choose the desired delivery language

Choose this when you want the overall voice delivery to follow a particular language community even when the input contains substantial text from another language.

Keep the same voice across languages

Language mixing should not sound like several different speech systems stitched together. Every Soniox voice works across all 60+ supported languages, preserving one speaker identity as the language changes.

One voice across 60+ languages

Use the same built-in voice for multilingual applications, localized products, and mixed-language speech while maintaining consistent voice identity and quality.

English
Spanish
Japanese
French
German
Italian
Portuguese
Chinese
Korean
Arabic

Clone once, speak every supported language

Create a cloned voice from up to 20 seconds of audio and use that same voice across all 60+ supported languages, including mixed-language utterances.

Emma
Conversational voice agent
Original voice
Cloned voice

Mixed-language speech still needs every detail right

Multilingual applications often combine language changes with names, numbers, contact details, addresses, and other structured information. Soniox keeps those details clear inside the same spoken response.

Speak numbers, addresses, and identifiers clearly

Phone numbers, email addresses, verification codes, prices, and postal addresses remain clear and precise alongside multilingual text.

Phone number

You can reach our support team at +1 415 682 9074.

Email address

Send the completed form to alex.chen+support@example.com.

Verification code

Your verification code is 7Q4M9B. I repeat: 7Q4M9B.

Postal address

Your delivery is going to 1427 North St. Andrews Place, Apartment 6B, Los Angeles, California 90028.

Price

Your total is $1,284.37, including tax and delivery.

Language mixing for real-time voice applications

Use mixed-language TTS in both single requests and streaming sessions. Generate multilingual speech progressively while keeping one continuous voice and precise control over what has already been spoken.

Stream mixed-language speech as text arrives

Send text progressively and begin playback before the complete message is available, including responses that contain multiple languages, foreign names, and embedded phrases.

Incoming textStreaming

Generated speechSpeaking

Keep timing aligned through language changes

Character-level timestamps let your application synchronize generated speech with text, stop cleanly during interruptions, and track exactly what has already been spoken.

Agent
User
Actually, can I change the dates?
Agent
Character timestamps

Estimate your multilingual text-to-speech cost

Set your expected generated speech volume to estimate the cost of using Soniox Text-to-Speech API for multilingual and mixed-language applications.

Pricing calculator

Stop overpaying for speech AI

Sonioxvs

1,000 hours of speech / month

1025501002505001k2.5k5k10k100k

Pricing assumptions

Based on public pay-as-you-go pricing. Enterprise discounts and committed-use contracts may differ. Some providers charge separately for certain features. The calculator uses the public price for the provider configuration that most closely matches Soniox.

Speech infrastructure for massive scale

Soniox Text-to-Speech API performance and reliability

Build on one API and deploy in your region

Use the same models and API everywhere, with in-region processing to meet latency, data residency, and regulatory requirements.

Available: US, EU, Japan, India
Coming soon: Korea, Australia, Canada, Saudi Arabia, UK, Brazil

View data residency docs
Soniox Text-to-Speech API performance and reliability

Run mission-critical systems with confidence

  • 99.9% uptime
    Production-hardened infrastructure with monitoring and redundancy.
  • low-latency streaming
    Process speech in real time with low latency for responsive voice applications.
  • Priority support
    Severity-based incident response with direct access to the Soniox team.
Onvego uses Soniox Text-to-Speech API for multilingual voice experiences

"Before Soniox, our international users always had a noticeably different experience. Now accuracy and responsiveness match across all regions…it feels like one system instead of five."

Alon Yair CTO of Onvego

Privacy and compliance, built right in

Never stored, never saved.

Audio stays in memory, everything is processed in real-time.

Built for privacy-critical use cases.

Adhering to leading global security, privacy, and compliance standards.

Trusted where privacy matters most.

Used in industries where speech is sensitive, from healthcare to enterprise.

Soniox is Soc 2 Type 2 compliant
Soniox is ISO 27001:2022 compliant
Soniox is HIPAA compliant
Soniox is GDPR compliant
SOC 2 Type 2 · ISO/IEC 27001:2022 · HIPAA · GDPR

Frequently asked questions about language-mixing text-to-speech

What is language mixing in text-to-speech?
Language mixing lets one text-to-speech request contain text from more than one language. Soniox follows the language changes inside the utterance while maintaining one continuous voice and delivery.
Can Soniox speak multiple languages in the same sentence?
Yes. Send mixed-language text in a single request or streaming session. Soniox follows language changes within the continuous utterance, so you do not need to split each language into separate requests and stitch the resulting audio together.
Does Soniox Text-to-Speech support code-switching?
Yes. Soniox can generate speech from text that switches between languages inside the same utterance. This is useful for bilingual conversations, language-learning content, multilingual voice agents, and other applications where speakers naturally move between languages.
Do I need to split mixed-language text into separate TTS requests?
No. Mixed-language text can be sent as one request or stream session. Soniox generates one continuous spoken utterance without requiring separate synthesis and audio stitching for each language segment.
Which language should I set in the request?
The language field is required. Set it to the main language of the text, or to the language you want to drive the overall delivery. It controls the overall prosody and applies a light accent bias to the utterance.
Do I need to change the language setting for foreign names and phrases?
No. Foreign names, addresses, terms, and embedded phrases can remain inside the surrounding text. The language field controls the overall delivery rather than requiring you to switch settings for every foreign word.
Can Soniox pronounce foreign names and places naturally?
Soniox handles person names, place names, brand names, addresses, and borrowed words from languages other than the surrounding text, so multilingual content can remain in one continuous utterance.
Does the same text-to-speech voice work across languages?
Yes. Every Soniox voice works across all 60+ supported languages, so the same speaker identity can continue as the language changes or across localized versions of your product.
Do cloned voices support multilingual language mixing?
Yes. A cloned voice works like a built-in voice and can be used across all 60+ supported languages, including mixed-language text.
Can language mixing be used with streaming text-to-speech?
Yes. Language mixing works in streaming sessions as well as single requests, so applications can begin generating and playing multilingual speech while text is still arriving.
Which languages does Soniox Text-to-Speech support?
Soniox Text-to-Speech supports 60+ languages. The same voice can be used across supported languages and in mixed-language utterances.
How much does Soniox Text-to-Speech cost?
Pricing is token-based: $4.00 per 1M input text tokens and $21.50 per 1M output audio tokens, which comes to about $0.70 per hour of generated speech.
How do I start building with language-mixing TTS?

Send your mixed-language text directly to Soniox Text-to-Speech API and set the language field according to the primary or desired delivery language.

Read the language mixing documentation for request examples and implementation details.

Speak every language change naturally

Build multilingual speech with one Text-to-Speech API. Send mixed-language text in one request, keep one consistent voice across 60+ languages, and generate natural speech without splitting text or stitching audio.

Ready to get started?

Create an account instantly, or contact us to design a custom package for your business.

Build with API

Documentation

Get up and running in minutes and spend your time building, not wrestling with the API.

Explore docs

See what you’ll pay

Pay only for what you use with our flexible pricing. Built to scale with you.

Pricing details