Speaker diarization API for real-time and recorded audio

Know who said what in every transcript. Soniox separates up to 15 speakers in live streams and recordings across 60+ languages, with a speaker label on every token.

$0.10/hour async · $0.12/hour real-time · diarization included

Start the demo and talk with someone next to you. Speaker labels appear as you speak.

Trusted by teams building global voice products

Livekit
Krisp
Pipecat
Summary AI
Perplexity
Samsung
Wispr Flow
LG
Agora
Retell AI
Fireflies.ai
Skit.ai
Kindroid
Deliver Health
Truecaller
Journalia
Mobius
TranscribeMe
Vapi
Zomato
SLNG
Japan AI
Boost.ai
Convin
Genspark
HappyRobot
Uniscribe
Jamie.ai
InteractCX
The Plato
MobilApp
Onvego
Wonderful.ai
Manifone
Tana
Transync AI
SotaTek

Know who said what, word by word

Soniox detects speaker changes and attaches a speaker label to every token in the transcript. Group tokens by speaker to build readable turns, or show the labels directly in your UI.

“Soniox knows who’s speaking and when each thought ends. The real-time transcripts read like true dialogue, not data dumps.”
Mobius MDAdam Strom, Co-Founder & President at Mobius MD

Up to 15 speakers

Separate up to 15 different speakers in one transcription session, from two-person calls to panel discussions.

Speakers and text in one response

Each token includes its speaker label alongside text, timestamps, and language. There is no second diarization pass to run or align with the transcript.

Enable with one parameter

No manual labeling or extra metadata. Set enable_speaker_diarization to true and speaker labels appear in the output.

Real-time or async diarization

Use real-time diarization when you need speaker labels while people talk. Use async diarization for the highest speaker accuracy on recorded audio.

Real-time speaker diarization

Speaker labels stream with each token for live meetings, support calls, captions, and conversational AI. Labels can briefly switch and then settle as more audio arrives.

Async speaker diarization

The model sees the full recording before assigning speakers, which gives significantly higher diarization accuracy for recorded calls, meetings, interviews, and podcasts.

Add speaker diarization with one parameter

Turn on diarization in the request and read the speaker field on each token. The same flag works in the real-time and async APIs.

Request

{
  "model": "stt-rt-v5",
  "enable_speaker_diarization": true
}

Output tokens

{"text": "How",    "speaker": "1"}
{"text": " are",   "speaker": "1"}
{"text": " you",   "speaker": "1"}
{"text": "?",      "speaker": "1"}
{"text": "I",      "speaker": "2"}
{"text": " am",    "speaker": "2"}
{"text": " fan",   "speaker": "2"}
{"text": "tastic", "speaker": "2"}

Speaker diarization in Python

from soniox import SonioxClient
from soniox.types import RealtimeSTTConfig
from soniox.utils import start_audio_thread, throttle_audio

client = SonioxClient()
config = RealtimeSTTConfig(
    model="stt-rt-v5",
    audio_format="mp3",
    enable_speaker_diarization=True,
)

audio = throttle_audio("meeting.mp3", delay_seconds=0.1)

with client.realtime.stt.connect(config=config) as session:
    start_audio_thread(session, audio)

    current_speaker = None
    for event in session.receive_events():
        for token in event.tokens:
            if not token.is_final:
                continue
            if token.speaker != current_speaker:
                current_speaker = token.speaker
                print(f"\nSpeaker {current_speaker}:", end="")
            print(token.text, end="", flush=True)

Speaker diarization in 60+ languages

Diarization works in every language Soniox supports, with one model and the same parameter. No per-language setup.

Spanish

Japanese

Arabic

Compare real-time speaker diarization

Compare diarization and related transcript metadata across each provider’s real-time model.

FeatureSoniox
stt-rt-v5
OpenAI
gpt-4o-transcribe
Google
gemini-3.5-transcribe-live
Azure
en-US-Conversation
Speechmatics
enhanced
Deepgram
nova-3
AssemblyAI
Universal-3.5 Pro
Cartesia
ink-2
ElevenLabs
Scribe v2 Realtime
Meta
muse-voice-transcribe-1.0
Smallest AI
Pulse
xAI
grok-voice-transcribe-2.0
Inworld
inworld-stt-1
Speaker diarization
Timestamps
Language identification
Confidence scores
SupportedPartialNot supported

Diarization while streaming

The OpenAI, Google, and ElevenLabs models shown above do not return speaker labels in their streaming output. Soniox labels speakers in the live stream.

No add-on fee

Soniox includes diarization in its hourly transcription rate. Some providers charge separately for speaker diarization.

One API instead of separate models

Self-hosted pipelines run a diarization model next to a transcription model, then match speaker turns to word timestamps. Soniox returns both from one API call, with no models to host.

Speech infrastructure for massive scale

Soniox Text-to-Speech API performance and reliability

Build on one API and deploy in your region

Use the same models and API everywhere, with in-region processing to meet latency, data residency, and regulatory requirements.

Available: US, EU, Japan, India
Coming soon: Korea, Australia, Canada, Saudi Arabia, UK, Brazil

View data residency docs
Soniox Text-to-Speech API performance and reliability

Run mission-critical systems with confidence

  • 99.9% uptime
    Production-hardened infrastructure with monitoring and redundancy.
  • low-latency streaming
    Process speech in real time with low latency for responsive voice applications.
  • Priority support
    Severity-based incident response with direct access to the Soniox team.
Onvego uses Soniox Text-to-Speech API for multilingual voice experiences

"Before Soniox, our international users always had a noticeably different experience. Now accuracy and responsiveness match across all regions…it feels like one system instead of five."

Alon Yair CTO of Onvego

Diarization included in the hourly rate

Choose real-time or async transcription and set your monthly audio volume. Speaker diarization has no add-on fee.

Pricing calculator

Stop overpaying for speech AI

Sonioxvs

1,000 hours of audio / month

1025501002505001k2.5k5k10k100k

Pricing assumptions

Based on public pay-as-you-go pricing. Enterprise discounts and committed-use contracts may differ. Some providers charge separately for certain features. The calculator uses the public price for the provider configuration that most closely matches Soniox.

Privacy and compliance, built right in

Never stored, never saved.

Audio stays in memory, everything is processed in real-time.

Built for privacy-critical use cases.

Adhering to leading global security, privacy, and compliance standards.

Trusted where privacy matters most.

Used in industries where speech is sensitive, from healthcare to enterprise.

Soniox is Soc 2 Type 2 compliant
Soniox is ISO 27001:2022 compliant
Soniox is HIPAA compliant
Soniox is GDPR compliant
SOC 2 Type 2 · ISO/IEC 27001:2022 · HIPAA · GDPR

Frequently asked questions about speaker diarization

What is speaker diarization?
Speaker diarization answers the question who spoke when. Soniox detects speaker changes in the audio and assigns each spoken segment a label such as Speaker 1 or Speaker 2, so transcripts split into clear, speaker-attributed sections. Read more in what is speaker diarization.
Does Soniox support real-time speaker diarization?
Yes. Diarization works on streaming audio through the real-time API, and speaker labels arrive with each token. Real-time diarization is harder than async because of low-latency constraints, so you may see a temporary speaker switch that settles as more context arrives.
Is real-time or async diarization more accurate?
Async. The model has the full audio context, which gives significantly higher diarization accuracy. Use real-time diarization when you need speaker attribution immediately, and async for recorded audio where accuracy matters most.
How many speakers can Soniox separate?
Up to 15 different speakers per transcription session. Accuracy may decrease when many speakers have similar voice characteristics.
How do I enable speaker diarization?
Set enable_speaker_diarization to true in your request. In the Python SDK, pass enable_speaker_diarization=True to RealtimeSTTConfig. See the speaker diarization docs for full examples.
What does the diarization output look like?
Every token includes a speaker field, for example {"text": "How", "speaker": "1"}. Group consecutive tokens by speaker to build readable turns, or render the labels directly in your UI.
Does speaker diarization cost extra?
No. Speaker diarization is included in the hourly rate: $0.10/hour for async and $0.12/hour for real-time transcription. See pricing.
Which languages support speaker diarization?
Speaker diarization is available in all 60+ supported languages.
Is diarization the same as speaker identification?
No. Diarization separates voices and gives them consistent labels like Speaker 1 and Speaker 2 within a session. It does not tell you a speaker’s name. If your application knows who is in the conversation, it can map each label to a person. See diarization vs speaker identification.
Does endpoint detection affect diarization accuracy?
Yes. Endpoint detection and manual finalization force tokens to finalize early, which reduces diarization accuracy. For the highest accuracy, do not use endpoint detection. If you use manual finalization, provide enough audio context before finalizing.
Can Soniox replace pyannote, WhisperX, or NeMo?
Soniox provides a hosted alternative to running separate transcription and diarization models. It returns text and speaker labels from one API call, in real time or async, so there is no diarization model to host or alignment step between speaker turns and words.
How do I get started?
Create an API key in the Soniox Console, then follow the speaker diarization guide.

Ready to get started?

Create an account instantly, or contact us to design a custom package for your business.

Build with API

Documentation

Get up and running in minutes and spend your time building, not wrestling with the API.

Explore docs

See what you’ll pay

Pay only for what you use with our flexible pricing. Built to scale with you.

Pricing details