A cheaper alternative to OpenAI Whisper
For developers building voice agents, live captions, meeting and call transcription, and other real-time speech applications that need accurate multilingual transcription.
Quick verdict
$0.12 per hour of audio
Real-time transcription with speaker diarization, language identification, and timestamps included in the rate.
$1.02 per hour of audio
GPT-Realtime-Whisper is listed at $0.017 per minute. Speaker diarization and timestamps are not available in real-time transcription sessions.
Turn detection in the stream
Endpoint detection and manual finalization tell your app when a speaker has finished.
Manual turn commits
Turn detection must be set to null for GPT-Realtime-Whisper, so apps send input_audio_buffer.commit to finish each audio turn.
Zero retention by default
Soniox does not store real-time audio or transcripts, and never trains its models on your data.
Up to 30 days of abuse monitoring logs
Abuse monitoring logs are retained for up to 30 days by default. Zero Data Retention requires prior approval from OpenAI and acceptance of additional requirements.
Soniox vs OpenAI Whisper
Languages, speakers, customization, timestamps, and real-time controls in each speech-to-text API.
| Single multilingual model | ||
|---|---|---|
| Language hints | ||
| Language identification | ||
| Speaker diarization | ||
| Customization | ||
| Timestamps | ||
| Confidence scores | ||
| Real-time latency config | ||
| Endpoint detection | ||
| Manual finalization |
Cost at volume
Speech-to-text costs become material when your application transcribes thousands of hours of audio.
$120 per month
1,000 hours of real-time audio at $0.12 per hour.
$1,020 per month
1,000 hours of real-time audio at $1.02 per hour on GPT-Realtime-Whisper.
Soniox
stt-rt-v5
OpenAI
gpt-realtime-whisper
At the published rates, the difference is $0.90 per hour of audio, or $900 a month at 1,000 hours. For exact costs on your workload, use the pricing calculator.
Calculate your costWhat the price includes
What each hourly rate covers, and what needs another model or service.
Features included
Speaker diarization, language identification, timestamps, confidence scores, and custom context are included in the hourly rate. Translation runs in the same real-time call at no extra cost.
Real-time transcription rate
The listed rate covers real-time transcription. Speaker labeling and timestamps are not supported in Realtime transcription sessions, and detected languages are returned only by gpt-transcribe, not this model.
Concurrency
Built for high concurrency
Soniox runs real-time speech-to-text on production infrastructure built for high-concurrency workloads, with 99.9% uptime and regional deployment.
60-minute sessions
Realtime sessions have a maximum duration of 60 minutes, with rate limits measured in minutes of audio per minute by usage tier.
Speakers and languages
Every token labeled by speaker and language
Speaker diarization and language identification run in the real-time stream, so each token carries its speaker and its language.
No real-time speaker or language labels
Speaker labeling is not supported in Realtime transcription sessions, and the detected languages field is returned by gpt-transcribe, not GPT-Realtime-Whisper.
Can Soniox transcribe speakers who switch languages?
Yes. One multilingual model covers every supported language, and language identification labels each token, so a speaker can switch languages mid-sentence.
Explore speaker diarizationReal-time transcription
Provisional and final tokens
Provisional tokens appear as audio arrives, and final tokens never change. Endpoint detection and manual finalization tell your app when a speaker has finished.
Partial and final events
Partial text arrives in conversation.item.input_audio_transcription.delta events and the final transcript in conversation.item.input_audio_transcription.completed events. Turn detection is disabled and apps use input_audio_buffer.commit to finish turns, with sessions limited to 60 minutes.
Switch from OpenAI Whisper in minutes
Soniox provides a real-time WebSocket API and an async REST API, plus official SDKs for Python, Node.js, Web, React, and React Native.
{
"api_key": "<SONIOX_API_KEY>",
"model": "stt-rt-v5",
"audio_format": "auto",
"enable_speaker_diarization": true,
"enable_language_identification": true,
"enable_endpoint_detection": true
}{
"type": "session.update",
"session": {
"type": "transcription",
"audio": {
"input": {
"format": {
"type": "audio/pcm",
"rate": 24000
},
"transcription": {
"model": "gpt-realtime-whisper",
"language": "en"
},
"turn_detection": null
}
}
}
}Switching from OpenAI Whisper means replacing its session config with one Soniox config message, sent first on the WebSocket. Speaker diarization, language identification, and endpoint detection are switched on in that same message.
Start your migrationPrivacy and data handling
Zero retention by default
Soniox does not store real-time audio or transcripts, and never trains its models on your audio or transcripts. Audio is processed within your selected region, with regional deployment for data residency requirements.
Up to 30 days of abuse monitoring logs
OpenAI says API data is not used to train or improve its models unless explicitly opted in. Abuse monitoring logs are retained for up to 30 days, Zero Data Retention requires prior approval, and real-time transcription can be processed in the United States or Europe under the stated residency requirements.
OpenAI Whisper alternative FAQ
How much does Soniox speech-to-text cost compared with OpenAI Whisper?
Does Soniox translate in real time?
Does Soniox support speaker diarization in real time?
Does Soniox identify the spoken language?
How many languages does Soniox speech-to-text support?
Does Soniox have an API and SDKs?
How does Soniox pricing work?
Ready to get started?
Create an account instantly, or contact us to design a custom package for your business.
Build with APIDocumentation
Get up and running in minutes and spend your time building, not wrestling with the API.
Explore docsSee what you’ll pay
Pay only for what you use with our flexible pricing. Built to scale with you.
Pricing details