Back to Soniox

Compare speech-to-text translation APIs live, on your own audio

See how different APIs handle accents, multilingual speech, language switching, names, numbers, and domain-specific terms, in real time.

Who supports what feature

Compare providers across the capabilities that matter for real-time speech translation. See which models support the languages, language pairs, streaming, language identification, translation, and other capabilities you need.

FeatureSoniox
stt-rt-v5
OpenAI
gpt-realtime-translate
Google
gemini-3.5-live-translate-preview
Speechmatics
enhanced
Azure
azure-speech-translation
Target languages60+137530+86
Source transcript
Single multilingual model
Language hints
Language identification
Speaker diarization
Timestamps
Endpoint detection
SupportedPartialNot supported

Who supports what language

See which providers support the languages and language pairs you need, including multilingual speech, language switching, and streaming translation.

LanguageSoniox
stt-rt-v5
OpenAI
gpt-realtime-translate
Google
gemini-3.5-live-translate-preview
Speechmatics
enhanced
Azure
azure-speech-translation
en
zh
hi
es
ar
fr
bn
pt
id
ur
ru
de
SupportedNot supported

Compare is open source

The complete source code is on GitHub under the MIT License.

Run the same comparison yourself

Inspect every provider integration, see exactly which models and settings are used, and check how audio is sent and translation results are returned.

Clone the repository, add your own API keys, and run it locally. Every provider receives the same audio through its official API. A missing key only disables that provider.

Clone the repo on GitHub
git clone https://github.com/soniox/soniox-compare
cd soniox-compare
for app in stt tts translate; do
  cp "$app/.env.example" "$app/.env"
done
./dev.sh

Contribute

Add a provider, add a model, improve an integration, or help make speech translation comparisons more useful for real applications.

Start contributing
The soniox-compare repository on GitHub

Methodology

Same audio, every provider

Your browser sends the same live audio to the Compare backend. The backend opens one session per selected provider, using that provider’s official speech-to-text translation API and the model listed in the feature table.

Same settings wherever supported

Every provider receives the same audio at the same time and the same requested settings wherever supported. Provider-specific options are used according to each integration, and unsupported options are dropped for that provider only.

Translations shown as returned

Translation results are returned as generated by each provider. Compare does not rewrite, normalize, correct, or otherwise alter the translated output before displaying it.

No hidden dataset

There is no hidden benchmark dataset and no aggregate quality score. You provide the audio, so you can test the speech, languages, and conversations your application actually needs to translate.

Live, as the tokens arrive

Results are shown as they arrive, so you can compare both the translated output and the experience of receiving it in real time.

No universal winner

Compare does not claim that one speech translation provider is universally better. It lets you evaluate the models against the requirements of your own product.

Frequently asked questions

Which speech-to-text API is most accurate?
It depends on your audio. Accents, background noise, language switching, names, and numbers all change the ranking, and a provider that wins on one recording can lose on the next. That is why this tool streams your audio to every provider at once, so you judge on the data that matters to you.
Why don’t you publish a WER leaderboard?
A leaderboard reduces everything to one number against a fixed reference. The reference is often not what a product needs, so a model can score well and still produce transcripts you would have to clean up. Showing outputs side by side lets you see the difference instead of trusting a score.
Is the comparison fair to every provider?
Every provider receives the same audio and the same settings through its official streaming API, using the model listed in the feature table. Each provider only gets the options it supports, and outputs are shown as returned. The code is open, so you can check every integration and run it yourself.
Can I add my own provider or model?
Yes. Each provider is a single Python module. Write the module, register it in the provider map, and add its API key to the environment file. The frontend picks it up from the backend without any change. Pull requests are welcome.
Which languages are supported?
The language table covers 130 languages with more than a million speakers, and marks which providers support each one. The source languages you can pick in the tool come from the Soniox speech-to-text model, which covers 60+ languages.
How is latency measured?
It is not scored. Transcripts stream in live, so you see how quickly each provider responds as it happens. Network distance, provider regions, and load all affect what you see, so the tool does not publish latency numbers.