Back to Soniox

Compare speech-to-text APIs live, on your own audio

See how speech-to-text models actually perform on the speech that matters to you. Stream the same audio through multiple providers at the same time and compare their transcripts as they arrive.

Who supports what feature

Which streaming features each provider gives you: language hints, context, speaker diarization, language identification, and endpoint detection.

FeatureSoniox
stt-rt-v5
OpenAI
gpt-4o-transcribe
Google
gemini-3.5-transcribe-live
Azure
en-US-Conversation
Speechmatics
enhanced
Deepgram
nova-3
AssemblyAI
Universal-3.5 Pro
Cartesia
ink-2
ElevenLabs
Scribe v2 Realtime
Meta
muse-voice-transcribe-1.0
Smallest AI
Pulse
xAI
stt-v1
Inworld
inworld-stt-1
Languages60+?7781566018190+25212530
Single multilingual model
Language hints
Language identification
Speaker diarization
Customization
Timestamps
Confidence scores
Real-time latency config
Endpoint detection
Manual finalization
SupportedPartialNot supported

Who supports what language

The 130 languages spoken by more than a million people, and which providers transcribe them. Pin the ones your product needs.

LanguageSoniox
stt-rt-v5
OpenAI
gpt-4o-transcribe
Google
gemini-3.5-transcribe-live
Azure
en-US-Conversation
Speechmatics
enhanced
Deepgram
nova-3
AssemblyAI
Universal-3.5 Pro
Cartesia
ink-2
ElevenLabs
Scribe v2 Realtime
Meta
muse-voice-transcribe-1.0
Smallest AI
Pulse
xAI
stt-v1
Inworld
inworld-stt-1
en
zh
hi
es
ar
fr
bn
pt
id
ur
ru
de
SupportedNot supported

Compare is open source

The complete source code is on GitHub under the MIT License.

Run the same comparison yourself

Soniox Compare is open source under the MIT License. Inspect every provider integration, see exactly which models and settings are used, and see how audio and results are handled.

Clone the repository, add your own API keys, and run the comparison locally. The same audio goes to every selected provider using its official streaming API.

View on GitHub
git clone https://github.com/soniox/soniox-compare
cd soniox-compare
for app in stt tts translate; do
  cp "$app/.env.example" "$app/.env"
done
./dev.sh

Contribute

Add a provider, improve an integration, add support for a model, or help make comparisons more useful for real speech applications.

Start contributing
The soniox-compare repository on GitHub

Why benchmarks mislead

Benchmarks don’t test your real-world audio.

A benchmark can be accurate and still be the wrong test

Speech-to-text benchmarks usually reduce performance to a small number of metrics, often Word Error Rate. WER is useful, but it measures how closely a transcript matches one reference, not whether the transcript is useful for your application.

Payment details

Alex Morgan

Enter a valid card number

MM / YY
123

The dataset decides the result

Clean read speech, accents, noisy calls, meetings, spontaneous conversation, and voice-agent speech expose different weaknesses. A model that performs well on one dataset may behave differently on the speech you actually receive.

Meeting transcript

Semantic word error rate

Not every textual difference is a meaningful transcription error. Numbers, names, IDs, phone numbers, punctuation, formatting, spelling, and spoken-versus-written forms can produce different strings with the same meaning. Modern benchmarks increasingly normalize these cases because literal string matching can penalize equivalent output.

But normalization has limits too. In some applications, the exact form matters: an ID, phone number, URL, code, address, or product name may need to be reproduced exactly.

Support for order #A7-4491

Closed benchmarks are hard to verify

Not every benchmark makes its audio, transcripts, normalization, prompts, or evaluation setup available. When the underlying dataset and methodology are closed, you can’t independently inspect what was tested or reproduce the result yourself. You are ultimately trusting the benchmark’s setup as much as its score.

Even an open benchmark only tells you how models performed on that particular dataset.

Benchmark dataset

7f3a…
d41c…
9be2…
26f8…
e0a5…

English isn’t enough

Multilingual speech is not just English speech with more languages added. Language switching, accents, and differences in names, numbers, and pronunciation can expose weaknesses that an English-focused benchmark never shows.

A benchmark can include many languages and still tell you very little about how a model handles the language mix your users actually speak.

World map with countries where English is an official language highlighted
Countries where English is an official language
World map with countries where English is not an official language highlighted
Everywhere else

Methodology

Same audio, every provider

Your browser streams audio as 16 kHz PCM to the Compare backend. The backend opens one session per selected provider, using that provider’s official streaming API and the model listed in the feature table, and sends every provider the same audio at the same time.

Same settings wherever supported

Each provider receives the same settings wherever supported: language hints, context, speaker diarization, language identification, and endpoint detection. Each integration declares which options it supports, and unsupported options are dropped for that provider only.

Outputs shown as returned

Outputs are shown as returned. Compare does not rewrite, correct, normalize, or clean up transcripts.

No hidden dataset

There is no hidden benchmark dataset and no aggregate score. You provide the audio, so you can test the speech that actually matters to your application.

Live, as you speak

Transcripts arrive live, so you can compare not only what each provider transcribes, but how the transcription develops while someone is still speaking.

No universal winner

Compare does not claim that one provider is universally better. It gives you the raw comparison so you can evaluate the tradeoffs for your own use case.

Frequently asked questions

Which speech-to-text API is most accurate?
It depends on your audio. Accents, background noise, language switching, names, and numbers all change the ranking, and a provider that wins on one recording can lose on the next. That is why this tool streams your audio to every provider at once, so you judge on the data that matters to you.
Why don’t you publish a WER leaderboard?
A leaderboard reduces everything to one number against a fixed reference. The reference is often not what a product needs, so a model can score well and still produce transcripts you would have to clean up. Showing outputs side by side lets you see the difference instead of trusting a score.
Is the comparison fair to every provider?
Every provider receives the same audio and the same settings through its official streaming API, using the model listed in the feature table. Each provider only gets the options it supports, and outputs are shown as returned. The code is open, so you can check every integration and run it yourself.
Can I add my own provider or model?
Yes. Each provider is a single Python module. Write the module, register it in the provider map, and add its API key to the environment file. The frontend picks it up from the backend without any change. Pull requests are welcome.
Which languages are supported?
The language table covers 130 languages with more than a million speakers, and marks which providers support each one. The source languages you can pick in the tool come from the Soniox speech-to-text model, which covers 60+ languages.
How is latency measured?
It is not scored. Transcripts stream in live, so you see how quickly each provider responds as it happens. Network distance, provider regions, and load all affect what you see, so the tool does not publish latency numbers.