Speaker diarization API for real-time and recorded audio
Know who said what in every transcript. Soniox separates up to 15 speakers in live streams and recordings across 60+ languages, with a speaker label on every token.
$0.10/hour async · $0.12/hour real-time · diarization included
Start the demo and talk with someone next to you. Speaker labels appear as you speak.
Trusted by teams building global voice products
Know who said what, word by word
Soniox detects speaker changes and attaches a speaker label to every token in the transcript. Group tokens by speaker to build readable turns, or show the labels directly in your UI.
“Soniox knows who’s speaking and when each thought ends. The real-time transcripts read like true dialogue, not data dumps.”
Up to 15 speakers
Separate up to 15 different speakers in one transcription session, from two-person calls to panel discussions.
Speakers and text in one response
Each token includes its speaker label alongside text, timestamps, and language. There is no second diarization pass to run or align with the transcript.
Enable with one parameter
No manual labeling or extra metadata. Set enable_speaker_diarization to true and speaker labels appear in the output.
Real-time or async diarization
Use real-time diarization when you need speaker labels while people talk. Use async diarization for the highest speaker accuracy on recorded audio.
Real-time speaker diarization
Speaker labels stream with each token for live meetings, support calls, captions, and conversational AI. Labels can briefly switch and then settle as more audio arrives.

Async speaker diarization
The model sees the full recording before assigning speakers, which gives significantly higher diarization accuracy for recorded calls, meetings, interviews, and podcasts.
Add speaker diarization with one parameter
Turn on diarization in the request and read the speaker field on each token. The same flag works in the real-time and async APIs.
Request
{
"model": "stt-rt-v5",
"enable_speaker_diarization": true
}Output tokens
{"text": "How", "speaker": "1"}
{"text": " are", "speaker": "1"}
{"text": " you", "speaker": "1"}
{"text": "?", "speaker": "1"}
{"text": "I", "speaker": "2"}
{"text": " am", "speaker": "2"}
{"text": " fan", "speaker": "2"}
{"text": "tastic", "speaker": "2"}Speaker diarization in Python
from soniox import SonioxClient
from soniox.types import RealtimeSTTConfig
from soniox.utils import start_audio_thread, throttle_audio
client = SonioxClient()
config = RealtimeSTTConfig(
model="stt-rt-v5",
audio_format="mp3",
enable_speaker_diarization=True,
)
audio = throttle_audio("meeting.mp3", delay_seconds=0.1)
with client.realtime.stt.connect(config=config) as session:
start_audio_thread(session, audio)
current_speaker = None
for event in session.receive_events():
for token in event.tokens:
if not token.is_final:
continue
if token.speaker != current_speaker:
current_speaker = token.speaker
print(f"\nSpeaker {current_speaker}:", end="")
print(token.text, end="", flush=True)Speaker diarization in 60+ languages
Diarization works in every language Soniox supports, with one model and the same parameter. No per-language setup.
Spanish
Japanese
Arabic
Compare real-time speaker diarization
Compare diarization and related transcript metadata across each provider’s real-time model.
| Feature | Soniox stt-rt-v5 | OpenAI gpt-4o-transcribe | Google gemini-3.5-transcribe-live | Azure en-US-Conversation | Speechmatics enhanced | Deepgram nova-3 | AssemblyAI Universal-3.5 Pro | Cartesia ink-2 | ElevenLabs Scribe v2 Realtime | Meta muse-voice-transcribe-1.0 | Smallest AI Pulse | xAI grok-voice-transcribe-2.0 | Inworld inworld-stt-1 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Speaker diarization | |||||||||||||
| Timestamps | |||||||||||||
| Language identification | |||||||||||||
| Confidence scores |
Diarization while streaming
The OpenAI, Google, and ElevenLabs models shown above do not return speaker labels in their streaming output. Soniox labels speakers in the live stream.
No add-on fee
Soniox includes diarization in its hourly transcription rate. Some providers charge separately for speaker diarization.
One API instead of separate models
Self-hosted pipelines run a diarization model next to a transcription model, then match speaker turns to word timestamps. Soniox returns both from one API call, with no models to host.
Built for multi-speaker audio
Use speaker-labeled transcripts in any conversation with more than one speaker.
Meeting notes
Attribute every statement, decision, and action item to the right participant in speaker-labeled meeting transcripts.
Call centers
Separate agent and customer speech in live and recorded calls for agent assist, QA, and compliance review.
Medical transcription
Keep clinician and patient speech apart so clinical notes reflect who reported what.
Media and podcasts
Produce speaker-labeled transcripts and captions for interviews, podcasts, panels, and broadcasts.
Speech analytics
Analyze talk time, turn-taking, and what each speaker said across large volumes of conversations.
Wearables
Label the voices around the wearer in real time for live captions, translation, and conversation notes.
Speech infrastructure for massive scale

Build on one API and deploy in your region
Use the same models and API everywhere, with in-region processing to meet latency, data residency, and regulatory requirements.
Available: US, EU, Japan, India
Coming soon: Korea, Australia, Canada, Saudi Arabia, UK, Brazil

Run mission-critical systems with confidence
- 99.9% uptime
Production-hardened infrastructure with monitoring and redundancy. - low-latency streaming
Process speech in real time with low latency for responsive voice applications. - Priority support
Severity-based incident response with direct access to the Soniox team.
"Before Soniox, our international users always had a noticeably different experience. Now accuracy and responsiveness match across all regions…it feels like one system instead of five."
Alon Yair CTO of Onvego
Diarization included in the hourly rate
Choose real-time or async transcription and set your monthly audio volume. Speaker diarization has no add-on fee.
Pricing calculator
Stop overpaying for speech AI
1,000 hours of audio / month
Pricing assumptions
Based on public pay-as-you-go pricing. Enterprise discounts and committed-use contracts may differ. Some providers charge separately for certain features. The calculator uses the public price for the provider configuration that most closely matches Soniox.
Privacy and compliance, built right in
Never stored, never saved.
Audio stays in memory, everything is processed in real-time.
Built for privacy-critical use cases.
Adhering to leading global security, privacy, and compliance standards.
Trusted where privacy matters most.
Used in industries where speech is sensitive, from healthcare to enterprise.




Frequently asked questions about speaker diarization
What is speaker diarization?
Does Soniox support real-time speaker diarization?
Is real-time or async diarization more accurate?
How many speakers can Soniox separate?
How do I enable speaker diarization?
enable_speaker_diarization to true in your request. In the Python SDK, pass enable_speaker_diarization=True to RealtimeSTTConfig. See the speaker diarization docs for full examples.What does the diarization output look like?
speaker field, for example {"text": "How", "speaker": "1"}. Group consecutive tokens by speaker to build readable turns, or render the labels directly in your UI.Does speaker diarization cost extra?
Which languages support speaker diarization?
Is diarization the same as speaker identification?
Does endpoint detection affect diarization accuracy?
Can Soniox replace pyannote, WhisperX, or NeMo?
How do I get started?
Ready to get started?
Create an account instantly, or contact us to design a custom package for your business.
Build with APIDocumentation
Get up and running in minutes and spend your time building, not wrestling with the API.
Explore docsSee what you’ll pay
Pay only for what you use with our flexible pricing. Built to scale with you.
Pricing details