Text-to-speech API with character-level timestamps

Get start and end times for every spoken character, streamed alongside the audio. Derive word-level timestamps for live text highlighting, captions, read-along experiences, and voice-agent interruptions.

$0.70 per generated hour.

Trusted by teams building global voice products

Livekit
Krisp
Pipecat
Summary AI
Perplexity
Samsung
Wispr Flow
LG
Agora
Retell AI
Fireflies.ai
Skit.ai
Kindroid
Deliver Health
Truecaller
Journalia
Mobius
TranscribeMe
Vapi
Zomato
SLNG
Japan AI
Boost.ai
Convin
Genspark
HappyRobot
Uniscribe
Jamie.ai
InteractCX
The Plato
MobilApp
Onvego
Wonderful.ai
Manifone
Tana
Transync AI
SotaTek

TTS timestamps streamed with the audio

Soniox Text-to-Speech API returns character-level timing data while speech is being generated. Keep text and audio synchronized in real time without running a separate alignment step after synthesis.

Highlight words as they are spoken

Use character timestamps directly for character-by-character highlighting, or group them into words for karaoke-style captions, transcripts, reading interfaces, and other synchronized text experiences.

Natural and expressive

I thought I knew exactly how the evening would unfold. Then the phone rang. For a moment, I considered letting it go to voicemail—but something told me to answer.

Conversational

Absolutely. I found three flights that arrive before noon. The first is the least expensive, but the second gives you a much shorter connection. Would you like me to compare them?

Storytelling

By the time we reached the top of the hill, the sun was already beginning to set. We stopped for a moment, looked back at the road behind us, and realized the entire valley had turned gold.

Know exactly what a voice agent finished saying

When a user interrupts an agent mid-response, character timestamps show how far playback reached. Use the spoken-so-far text to keep your LLM or dialogue system aligned with what the user actually heard.

Agent
User
Actually, can I change the dates?
Agent
Character timestamps

Synchronize every script and language

Character-level alignment works naturally for multilingual text, language mixing, and writing systems where word boundaries are not always represented by spaces.

English and Japanese

Your order is ready for pickup. 店頭で注文番号 A-7392 をお見せください.

English and Korean

To enable the new feature, open Settings and select 음성 복제, then tap Continue.

English and Hindi

Your appointment is confirmed for Monday at 11 a.m. कृपया दस मिनट पहले पहुँचें.

English and German

The deployment completed successfully. Bitte prüfen Sie jetzt die Produktionsumgebung und bestätigen Sie, dass alles funktioniert.

English with multilingual names and places

Your meeting with Nguyễn Minh Anh is scheduled for 3 p.m. at our office near Place de la République in Paris.

Need word-level TTS timestamps? Start with character timing

Soniox returns the start and end time of every spoken character. That gives you the information needed to create word-level timestamps while preserving finer-grained timing whenever your application needs it.

Group character timestamps into words

Concatenate the character arrays from each audio frame, then group characters between whitespace. A word begins at the start of its first character and ends at the end of its last.

function toWords(characters, starts, ends) {
  const words = [];
  let word = null;

  characters.forEach((char, i) => {
    if (/\s/.test(char)) {
      word = null;
      return;
    }

    if (!word) {
      word = {
        text: '',
        start: starts[i],
        end: ends[i],
      };
      words.push(word);
    }

    word.text += char;
    word.end = ends[i];
  });

  return words;
}

Word-level timestamp output

The character timing can be reduced to the word timing most caption, transcript, and read-along interfaces need.

[
  {
    "text": "Hello",
    "start": 0.0,
    "end": 0.5
  }
]

Why Soniox returns character-level timestamps

Character timing preserves more information. You can derive word-level timing from it, highlight individual characters, handle scripts without whitespace boundaries, or determine where playback stopped inside a word.

Turn on character timestamps with one field

Set return_timestamps to true when you start a Soniox Text-to-Speech WebSocket stream. Audio responses can then include character-to-audio alignment data alongside the generated audio.

Request timestamps

{
  "api_key": "<SONIOX_API_KEY>",
  "model": "tts-rt-v2",
  "language": "en",
  "voice": "Adrian",
  "audio_format": "pcm_s16le",
  "sample_rate": 24000,
  "stream_id": "stream-001",
  "return_timestamps": true
}

Receive audio and alignment

{
  "stream_id": "stream-001",
  "audio": "<base64-encoded-audio-chunk>",
  "timestamps": {
    "characters": ["H", "e", "l", "l", "o"],
    "character_start_times_seconds": [0.0, 0.1, 0.2, 0.3, 0.4],
    "character_end_times_seconds": [0.1, 0.2, 0.3, 0.4, 0.5]
  }
}

Real-time text-to-audio alignment

Timestamps are designed for streaming applications. Alignment data arrives incrementally with generated speech, so your interface can react while the audio is still playing.

Streamed with the audio

Character timestamps arrive incrementally alongside generated audio, so your application can synchronize text while speech is still playing.

Aligned to spoken output

Each timestamp maps generated speech back to the normalized text Soniox actually speaks, giving your application precise text-to-audio alignment.

One continuous timeline

Times are non-decreasing across the stream and remain aligned as audio arrives across multiple chunks.

One timing entry per character

Receive a start and end time for every spoken character, giving you finer-grained alignment than word-level timing alone.

Word highlighting and captions in real time

Because timestamp data is returned during synthesis, applications can update captions, transcripts, and read-along interfaces progressively instead of waiting for the complete generated clip.

Start highlighting while speech is still generating

Send text progressively, begin playback before the complete response is available, and receive timing data alongside the generated audio.

Incoming textStreaming

Generated speechSpeaking

TTS timing for more than captions

Character-level speech timing gives applications a precise connection between generated text and generated audio.

Live text highlighting

Follow generated speech word by word or character by character in captions, transcripts, karaoke-style interfaces, and read-along experiences.

Interruption tracking

Determine exactly how much of a generated response was actually spoken before playback stopped or a user interrupted.

Text-audio synchronization

Keep text interfaces, subtitles, learning content, and other application state synchronized with generated speech.

Estimate your text-to-speech cost

Set your expected generated speech volume to estimate the cost of using Soniox Text-to-Speech API with character-level timing.

Pricing calculator

Stop overpaying for speech AI

Sonioxvs

1,000 hours of speech / month

1025501002505001k2.5k5k10k100k

Pricing assumptions

Based on public pay-as-you-go pricing. Enterprise discounts and committed-use contracts may differ. Some providers charge separately for certain features. The calculator uses the public price for the provider configuration that most closely matches Soniox.

Speech infrastructure for massive scale

Soniox Text-to-Speech API performance and reliability

Build on one API and deploy in your region

Use the same models and API everywhere, with in-region processing to meet latency, data residency, and regulatory requirements.

Available: US, EU, Japan, India
Coming soon: Korea, Australia, Canada, Saudi Arabia, UK, Brazil

View data residency docs
Soniox Text-to-Speech API performance and reliability

Run mission-critical systems with confidence

  • 99.9% uptime
    Production-hardened infrastructure with monitoring and redundancy.
  • low-latency streaming
    Process speech in real time with low latency for responsive voice applications.
  • Priority support
    Severity-based incident response with direct access to the Soniox team.
Onvego uses Soniox Text-to-Speech API for multilingual voice experiences

"Before Soniox, our international users always had a noticeably different experience. Now accuracy and responsiveness match across all regions…it feels like one system instead of five."

Alon Yair CTO of Onvego

Privacy and compliance, built right in

Never stored, never saved.

Audio stays in memory, everything is processed in real-time.

Built for privacy-critical use cases.

Adhering to leading global security, privacy, and compliance standards.

Trusted where privacy matters most.

Used in industries where speech is sensitive, from healthcare to enterprise.

Soniox is Soc 2 Type 2 compliant
Soniox is ISO 27001:2022 compliant
Soniox is HIPAA compliant
Soniox is GDPR compliant
SOC 2 Type 2 · ISO/IEC 27001:2022 · HIPAA · GDPR

Frequently asked questions about TTS timestamps

Does Soniox Text-to-Speech return character-level timestamps?
Yes. The Soniox Text-to-Speech WebSocket API returns the start and end time of every spoken character alongside generated audio when timestamps are enabled.
Does Soniox Text-to-Speech provide word-level timestamps?

Soniox natively returns character-level timestamps rather than a separate word-level array.

Word-level timestamps are straightforward to derive: group the characters belonging to each word, use the first character's start time as the word start, and the last character's end time as the word end.

What are character-level TTS timestamps?
Character-level timestamps are text-to-audio alignment data that identify when each spoken character begins and ends in generated speech. They provide finer-grained timing than word-level timestamps alone.
What are word-level timestamps in text-to-speech?
Word-level timestamps identify when each spoken word starts and ends in generated audio. With Soniox, you derive them from the character-level timestamps returned by the WebSocket API.
Are TTS timestamps the same as speech marks?
The terms describe closely related timing metadata. Some speech APIs use speech marks for mappings between synthesized text and audio. Soniox calls its output character-level timestamps and returns start and end times for each spoken character.
Can I use Soniox timestamps for word-by-word text highlighting?
Yes. Group character timestamps into words and use their start and end times to highlight captions, transcripts, or read-along text as each word is spoken.
Can I use timestamps for character-by-character highlighting?
Yes. Because Soniox returns timing for every spoken character, your interface can highlight at character granularity without first grouping the alignment into words.
Do Soniox TTS timestamps stream in real time?
Yes. Timestamp data arrives incrementally with generated audio frames, so applications can synchronize text while speech is still being generated and played.
How do I enable timestamps in Soniox Text-to-Speech API?
Set return_timestamps: true in the configuration message when starting a Text-to-Speech WebSocket stream. Timestamp output is disabled by default.
Are TTS timestamps available through the REST API?
No. Character-level timestamps are available through the Soniox Text-to-Speech WebSocket API. The REST endpoint returns raw audio and does not return timestamp metadata.
How do timestamps help with voice-agent interruptions?
When playback stops because a user interrupts, timestamps show how far the generated speech actually progressed. Your application can use the spoken portion of the response to keep its LLM or dialogue state consistent with what the user heard.
Can I use timestamps for text-to-audio alignment?
Yes. Character timestamps map generated speech to the normalized text Soniox actually speaks, allowing your application to synchronize text interfaces with the audio timeline.
Do timestamps work with multilingual and mixed-language TTS?
Yes. Character-level timing can be returned for generated multilingual speech and mixed-language utterances, including writing systems where whitespace is not a reliable word boundary.
Why can the timestamp text differ from my original input?
Soniox timestamps align to the preprocessed spoken text, not necessarily the raw input string. Whitespace may be normalized and characters the model cannot pronounce, such as emojis, can be removed before synthesis.
How much does Soniox Text-to-Speech cost?
Pricing is token-based: $4.00 per 1M input text tokens and $21.50 per 1M output audio tokens, which comes to about $0.70 per hour of generated speech.
How do I start building with TTS timestamps?

Start a Soniox Text-to-Speech WebSocket stream with return_timestamps: true and process the character, start-time, and end-time arrays returned with generated audio.

Read the TTS timestamps documentation for the complete request and response format.

Keep every word in sync with the voice

Get character-level timestamps directly from Soniox Text-to-Speech API and derive the word-level timing your product needs for highlighting, captions, read-along experiences, and real-time voice interactions.

Ready to get started?

Create an account instantly, or contact us to design a custom package for your business.

Build with API

Documentation

Get up and running in minutes and spend your time building, not wrestling with the API.

Explore docs

See what you’ll pay

Pay only for what you use with our flexible pricing. Built to scale with you.

Pricing details