Text-to-speech API with emotion and tone control

Generate expressive speech you can direct line by line. Add emotion, tone, pitch, pauses, and human sounds with audio tags written directly into the text you send.

$0.70 per generated hour.

Trusted by teams building global voice products

Livekit
Krisp
Pipecat
Summary AI
Perplexity
Samsung
Wispr Flow
LG
Agora
Retell AI
Fireflies.ai
Skit.ai
Kindroid
Deliver Health
Truecaller
Journalia
Mobius
TranscribeMe
Vapi
Zomato
SLNG
Japan AI
Boost.ai
Convin
Genspark
HappyRobot
Uniscribe
Jamie.ai
InteractCX
The Plato
MobilApp
Onvego
Wonderful.ai
Manifone
Tana
Transync AI
SotaTek

Expressive text-to-speech you can direct

Natural speech is only the starting point. Soniox Text-to-Speech API lets your application control how each line is performed, from excitement and reassurance to whispers, hesitation, laughter, tension, and dramatic pauses.

Control emotion line by line

Place audio tags directly before the text you want to affect. Change emotion and delivery throughout the same utterance instead of choosing one fixed speaking style for the entire response.

The emergency
Panic

The magic trick
Confusion

The breakup
Sadness

Bedtime
Coziness

The fridge
Disgust

Surprise party
Excitement

The countdown
Tension

The gift
Joy

Nervous confession
Hesitation

The reunion
Surprise

Sports commentary
Euphoria

Natural expression even without tags

Soniox already adapts rhythm, pacing, emphasis, and expression to the meaning of the text. Audio tags give your application additional control whenever a line needs a specific performance.

Natural and expressive

I thought I knew exactly how the evening would unfold. Then the phone rang. For a moment, I considered letting it go to voicemail—but something told me to answer.

Conversational

Absolutely. I found three flights that arrive before noon. The first is the least expensive, but the second gives you a much shorter connection. Would you like me to compare them?

Storytelling

By the time we reached the top of the hill, the sun was already beginning to set. We stopped for a moment, looked back at the road behind us, and realized the entire valley had turned gold.

Add emotion directly to your TTS input

Emotion and tone control lives inside the text your application already sends. Add bracketed audio tags wherever you want the delivery to change, with no separate performance timeline or editing workflow.

Control speech with audio tags

Write tags directly into the text field. The tag changes how the following words are delivered.

{
  "model": "tts-rt-v2",
  "language": "en",
  "voice": "Adrian",
  "audio_format": "wav",
  "text": "[excited] INCREDIBLE news! [calm] Let me explain what happened."
}

Shape delivery with punctuation

Formatting also influences delivery. Use uppercase for strong emphasis, *stress* for individual words, ... for hesitation, and ?! for incredulity.

You did WHAT?!
Oh, *you* noticed the haircut?
Wait... N-n-no. That can't be right.

Control emotion, tone, pace, pitch, volume, and more

Use audio tags to direct different dimensions of the performance. Tags can change throughout the utterance, giving your application line-by-line control over how speech sounds.

Emotion

[happy][sad][angry][excited][nervous][surprised][relieved][curious][calm]

Tone and manner

[warm][stern][serious][playful][sarcastic][deadpan][reassuringly][dramatically]

Human sounds

[laughs][chuckles][sighs][gasps][sobs][sniffles][clears throat][yawns]

Volume

[whispering][softly][loudly][shouting][getting louder][trailing off][muttering]

Pace, pitch, and voice quality

[slowly][quickly][hesitantly][low voice][monotone][breathy][trembling voice]

Pauses

[pause][long pause]

Emotional text-to-speech across 60+ languages

Use the same emotion and tone controls across every supported language. Audio tags stay in English while the spoken text can be generated in any of 60+ languages.

One set of audio tags for every language

Use the same English tags regardless of the language being spoken. For example, [excited] works with Spanish text just as it does with English text.

{
  "model": "tts-rt-v2",
  "language": "es",
  "voice": "Adrian",
  "text": "[excited] ¡Tengo noticias INCREÍBLES! [calm] Déjame explicarte lo que pasó."
}

The same expressive voice in 60+ languages

Keep the same voice identity while changing languages, and use the same emotion, tone, and delivery controls throughout your multilingual product.

English
Spanish
Japanese
French
German
Italian
Portuguese
Chinese
Korean
Arabic

Give cloned voices the same expressive control

Create a high-fidelity cloned voice and direct it with the same emotion, tone, pace, pitch, volume, pauses, and human-sound tags available to built-in voices.

Emma
Conversational voice agent
Original voice
Cloned voice

Expressive TTS also works in real time

Stream expressive speech while text is still arriving, so voice agents, interactive characters, tutors, and other dynamic applications can generate directed delivery at runtime.

Stream expressive speech as text arrives

Send text progressively and begin playback before the complete message is available, including responses that contain audio tags and directed delivery.

Incoming textStreaming

Generated speechSpeaking

Control pacing without losing natural delivery

Adjust pauses and overall speaking rate while preserving fluent, expressive speech for the pace your product and users need.

Standard pacing
Speaks with natural, unhurried pauses.
Reduced silence
Trims the pauses without rushing the voice.

Estimate your expressive text-to-speech cost

Set your expected generated speech volume to estimate the cost of using Soniox Text-to-Speech API for expressive and emotional speech.

Pricing calculator

Stop overpaying for speech AI

Sonioxvs

1,000 hours of speech / month

1025501002505001k2.5k5k10k100k

Pricing assumptions

Based on public pay-as-you-go pricing. Enterprise discounts and committed-use contracts may differ. Some providers charge separately for certain features. The calculator uses the public price for the provider configuration that most closely matches Soniox.

Speech infrastructure for massive scale

Soniox Text-to-Speech API performance and reliability

Build on one API and deploy in your region

Use the same models and API everywhere, with in-region processing to meet latency, data residency, and regulatory requirements.

Available: US, EU, Japan, India
Coming soon: Korea, Australia, Canada, Saudi Arabia, UK, Brazil

View data residency docs
Soniox Text-to-Speech API performance and reliability

Run mission-critical systems with confidence

  • 99.9% uptime
    Production-hardened infrastructure with monitoring and redundancy.
  • low-latency streaming
    Process speech in real time with low latency for responsive voice applications.
  • Priority support
    Severity-based incident response with direct access to the Soniox team.
Onvego uses Soniox Text-to-Speech API for multilingual voice experiences

"Before Soniox, our international users always had a noticeably different experience. Now accuracy and responsiveness match across all regions…it feels like one system instead of five."

Alon Yair CTO of Onvego

Privacy and compliance, built right in

Never stored, never saved.

Audio stays in memory, everything is processed in real-time.

Built for privacy-critical use cases.

Adhering to leading global security, privacy, and compliance standards.

Trusted where privacy matters most.

Used in industries where speech is sensitive, from healthcare to enterprise.

Soniox is Soc 2 Type 2 compliant
Soniox is ISO 27001:2022 compliant
Soniox is HIPAA compliant
Soniox is GDPR compliant
SOC 2 Type 2 · ISO/IEC 27001:2022 · HIPAA · GDPR

Frequently asked questions about emotional and expressive text-to-speech

Can text-to-speech generate speech with emotion?
Yes. Soniox Text-to-Speech API supports audio tags that let your application direct emotions such as [happy], [sad], [angry], [excited], [nervous], and [calm] directly inside the input text.
How do I control emotion in Soniox Text-to-Speech?
Write audio tags like [excited], [whispering], or [laughs] directly into the text you send. Place the tag before the words you want to affect, and those words inherit that delivery.
Can I control tone as well as emotion?
Yes. Audio tags can direct tone and manner such as [warm], [stern], [serious], [playful], [sarcastic], and [reassuringly], in addition to emotional states.
Can I control pace, pitch, volume, and voice quality?
Yes. Soniox supports tags for local pacing, volume, pitch, and voice quality, including directions such as [slowly], [quickly], [whispering], [shouting], [low voice], [breathy], and [trembling voice].
Can text-to-speech generate laughter, sighs, and other human sounds?
Yes. Audio tags can direct vocal reactions and human sounds such as [laughs], [chuckles], [sighs], [gasps], [sobs], and [yawns].
Can I shape delivery without using audio tags?
Yes. Formatting and punctuation also influence delivery. Use UPPERCASE for stronger emphasis, *stress* to stress a word, elongated spellings, and punctuation such as ... or ?!. These techniques can also be combined with audio tags.
Do emotion and tone controls work in other languages?
Yes. Write audio tags in English while generating speech in any supported language. The same set of tags works across 60+ languages.
Can cloned voices use emotion and tone control?
Yes. Cloned voices can be directed with the same audio tags used with built-in voices, including emotion, tone, volume, pace, pitch, pauses, and vocal reactions.
How many audio tags can I combine?
Two tags before a clause, such as [warm] [softly], normally combine well. Keep each tag short and single-purpose. Large stacks of conflicting directions are less reliable.
What happens if Soniox does not recognize an audio tag?
An unrecognized tag may be spoken aloud as text. Use the documented tags for production workloads and test other tags before relying on them.
Can expressive TTS be generated in real time?
Yes. Soniox supports streaming text-to-speech, so applications can begin receiving and playing expressive speech while text is still arriving.
How do I change the overall speaking rate?
Use the speed parameter for overall speaking rate. Audio tags and punctuation control expressive delivery and local pacing rather than replacing the global speed setting.
How much does Soniox Text-to-Speech cost?
Pricing is token-based: $4.00 per 1M input text tokens and $21.50 per 1M output audio tokens, which comes to about $0.70 per hour of generated speech.
How do I start building expressive text-to-speech?

Add audio tags directly to the text you send to Soniox Text-to-Speech API wherever you want the delivery to change.

Read the emotion and tone documentation for supported tags, examples, and implementation details.

Make every line sound the way you want

Build expressive text-to-speech with programmatic control over emotion, tone, pace, pitch, volume, pauses, and human sounds. Direct every line with Soniox Text-to-Speech API across 60+ languages.

Ready to get started?

Create an account instantly, or contact us to design a custom package for your business.

Build with API

Documentation

Get up and running in minutes and spend your time building, not wrestling with the API.

Explore docs

See what you’ll pay

Pay only for what you use with our flexible pricing. Built to scale with you.

Pricing details