Voice dictation API for fast, accurate, and multilingual voice typing

Build voice typing and voice dictation experiences that feel fast, accurate, and natural. Soniox Speech-to-Text API streams text as users speak, accurately captures names, numbers, dates, and delivers native-speaker accuracy across 60+ languages, with custom vocabulary support.

Trusted by teams building global voice products

Livekit
Krisp
Pipecat
Summary AI
Perplexity
Samsung
Wispr Flow
LG
Agora
Retell AI
Fireflies.ai
Skit.ai
Kindroid
Deliver Health
Truecaller
Journalia
Mobius
TranscribeMe
Vapi
Zomato
SLNG
Japan AI
Boost.ai
Convin
Genspark
HappyRobot
Uniscribe
Jamie.ai
InteractCX
The Plato
MobilApp
Onvego
Wonderful.ai
Manifone
Tana
Transync AI
SotaTek

Dictation has to feel like writing, not transcription

Users expect to press a microphone button, speak naturally, and see accurate text appear without breaking their train of thought. Delay, unstable words, misspelled names, or badly recognized numbers turn voice input into more work than typing.Soniox Speech-to-Text API is built for fast interaction: stream speech as it happens, finalize text when the user is done, and preserve the details that make dictated text useful.

Text appears while users speak

Stream provisional transcription immediately and replace it with confirmed final text as the model gains enough context.

Capture the details users dictate

Reliably recognize names, numbers, dates, times, email addresses, addresses, IDs, codes, and other structured information.

Understand specialized vocabulary

Provide names, product terms, professional vocabulary, project context, and domain knowledge to improve recognition.

Native-speaker accuracy across 60+ languages

Build high-quality voice typing across languages, accents, dialects, and multilingual speech with one unified model.

From microphone to text field

Soniox Speech-to-Text API handles the speech recognition layer while your application controls the dictation experience around it.

Start

Begin dictation from a microphone button, keyboard shortcut, push-to-talk control, or other voice input.

Stream

Send microphone audio to Soniox Speech-to-Text API and receive transcription continuously.

Finalize

Finalize the current speech segment when the user releases the button or finishes dictating.

Insert

Place confirmed text into your editor, message, document, form, prompt, or downstream workflow.

Speech recognition built for voice dictation

Dictation needs more than generic speech-to-text. Your product needs responsive streaming, predictable finalization, accurate vocabulary, and transcript output that can move directly into the user’s writing workflow.

Phone number

My number is +49 172 684 9317.

Reference ID

The case number is AX-4927-B.

Mixed alphanumeric

The device serial number is XR8Q-19F2.

International name

Nguyễn Minh Anh will join the meeting at three.

“As Germany’s leading voicebot provider for automotive dealerships, Soniox has transformed our recognition of customer IDs and alphanumerics, driving much higher voicebot acceptance rates.”
mobilAppDr. Steven Zielke, Founder & CEO of mobilApp

Low-latency streaming transcription

Soniox Speech-to-Text API returns provisional text while speech is still arriving, followed by confirmed tokens that remain stable once finalized.

Explore real-time transcription

Push-to-talk finalization

Tell Soniox Speech-to-Text API exactly when the current dictation segment is complete, making manual finalization a natural fit for push-to-talk and keyboard-shortcut workflows.

Explore manual finalization

Context-aware vocabulary

Give Soniox Speech-to-Text API the vocabulary behind each dictation session, including names, organizations, products, technical terms, projects, and domain-specific language.

Explore context customization

Structured speech details

Preserve numbers, dates, times, email addresses, names, addresses, IDs, and codes in forms that are easier for users and downstream software to work with.

Token-level confidence

Soniox Speech-to-Text API returns confidence information for recognized tokens so your application can identify uncertain words and design review or correction experiences around them.

Explore confidence scores

Automatic language identification

Detect the language being dictated automatically, including speech where users move between languages within the same session.

Explore language identification

Speech infrastructure for massive scale

Soniox Text-to-Speech API performance and reliability

Build on one API and deploy in your region

Use the same models and API everywhere, with in-region processing to meet latency, data residency, and regulatory requirements.

Available: US, EU, Japan, India
Coming soon: Korea, Australia, Canada, Saudi Arabia, UK, Brazil

View data residency docs
Soniox Text-to-Speech API performance and reliability

Run mission-critical systems with confidence

  • 99.9% uptime
    Production-hardened infrastructure with monitoring and redundancy.
  • low-latency streaming
    Process speech in real time with low latency for responsive voice applications.
  • Priority support
    Severity-based incident response with direct access to the Soniox team.
Onvego uses Soniox Text-to-Speech API for multilingual voice experiences

"Before Soniox, our international users always had a noticeably different experience. Now accuracy and responsiveness match across all regions…it feels like one system instead of five."

Alon Yair CTO of Onvego

Build voice typing into any kind of product

Use Soniox Speech-to-Text API anywhere speech can be faster or more convenient than a keyboard.

Universal voice typing

Add fast voice input to text fields, editors, forms, search boxes, messages, and other places where users normally type.

Mobile dictation

Turn microphone input into responsive text for mobile apps where typing is slow, inconvenient, or hands-busy.

Desktop productivity

Build keyboard-shortcut and push-to-talk dictation for documents, email, productivity tools, and desktop applications.

Professional documentation

Capture long-form speech, specialized terminology, names, numbers, and structured details for documentation workflows.

Field and hands-free input

Let users capture notes, reports, observations, and structured information while working away from a keyboard.

AI prompts and workflows

Turn spoken thoughts into accurate text for prompts, commands, forms, agents, and other AI-powered experiences.

Native-speaker dictation across 60+ languages

Voice typing products should not work beautifully in English and degrade as soon as users switch languages. Soniox Speech-to-Text API is built for native-speaker accuracy across more than 60 supported languages with one unified speech model.

English

“It just gets the words right, any language, any accent, any context. That’s what accuracy is supposed to look like.”
AgoraTony Wang, Cofounder & Chief Revenue Officer at Agora

High accuracy across supported languages

Build the same high-quality dictation experience across English, Spanish, German, French, Japanese, Korean, Arabic, Slovenian, and dozens of other languages.

Dictate naturally across languages

Handle accents, dialects, foreign names, and language switching without forcing users to manually change speech models every time their language changes.

Build for quick voice input or long-form dictation

Dictation experiences range from a two-second reply to paragraphs of continuous speech. Soniox Speech-to-Text API gives your product the streaming primitives to support both.

Short-form voice typing

Capture messages, searches, form fields, prompts, comments, and quick notes with fast streaming and explicit finalization when the user stops.

Long-form dictation

Let users speak naturally through longer documents, reports, notes, and professional writing while maintaining context across the session.

Privacy and compliance, built right in

Never stored, never saved.

Audio stays in memory, everything is processed in real-time.

Built for privacy-critical use cases.

Adhering to leading global security, privacy, and compliance standards.

Trusted where privacy matters most.

Used in industries where speech is sensitive, from healthcare to enterprise.

Soniox is Soc 2 Type 2 compliant
Soniox is ISO 27001:2022 compliant
Soniox is HIPAA compliant
Soniox is GDPR compliant
SOC 2 Type 2 · ISO/IEC 27001:2022 · HIPAA · GDPR

Frequently asked questions about voice dictation APIs

What is a voice dictation API?

A voice dictation API converts microphone audio into text that an application can insert into an editor, document, form, message, prompt, or other text input.

A good dictation experience also needs responsive streaming, reliable finalization, accurate vocabulary, and predictable handling of spoken details.

Is Soniox Speech-to-Text API suitable for voice dictation?

Yes. Soniox Speech-to-Text API supports real-time transcription with low-latency provisional results and confirmed final tokens.

It also supports context customization, manual finalization, language identification, confidence scores, and speech recognition across more than 60 languages.

Can I build push-to-talk dictation?

Yes. Soniox Speech-to-Text API supports manual finalization, which is specifically useful for push-to-talk, keyboard-shortcut, and client-side voice activity detection workflows.

When the user releases the dictation control, your application can request finalization and wait for the remaining transcript tokens to become final.

Can text appear while the user is still speaking?

Yes. Soniox Speech-to-Text API streams provisional transcription tokens while audio is still arriving.

Those tokens may update as additional speech context arrives, then become confirmed final tokens that no longer change.

Can I control when dictated text becomes final?

Yes. Soniox Speech-to-Text API can finalize text automatically or let your application explicitly finalize the current audio segment.

Manual finalization gives dictation products precise control over when a push-to-talk utterance or writing segment is complete.

Can Soniox Speech-to-Text API recognize custom vocabulary?

Yes. Soniox Speech-to-Text API lets your application provide context with each transcription session.

Context can include names, organizations, products, projects, technical terminology, professional vocabulary, topics, and other domain-specific information that can improve recognition.

How does Soniox Speech-to-Text API handle names, numbers, and structured details?

Soniox Speech-to-Text API is designed to recognize and format structured speech such as numbers, dates, times, email addresses, names, addresses, IDs, and codes.

This is especially important for dictation products where users expect spoken information to arrive as useful text, not just approximate words.

Does Soniox Speech-to-Text API provide confidence scores?

Yes. Soniox Speech-to-Text API includes confidence scores for recognized tokens.

Dictation applications can use them to identify uncertain words, highlight possible errors, or trigger their own review and correction experiences.

Does Soniox Speech-to-Text API support multilingual dictation?

Yes. Soniox Speech-to-Text API supports more than 60 languages with native-speaker accuracy through one unified speech model.

It can also handle accents, dialects, foreign terms, and sessions where users switch languages naturally.

Does the user need to select a language before dictating?

Not necessarily. Soniox Speech-to-Text API can identify spoken languages automatically.

Your application can also provide language hints when you know which languages are expected.

Can I use Soniox Speech-to-Text API for mobile, web, and desktop dictation?

Yes. Soniox Speech-to-Text API can receive live microphone audio over its real-time WebSocket interface, so it can sit behind mobile, web, desktop, and embedded voice-input experiences.

Your application controls microphone capture, the user interface, text insertion, shortcuts, editing, and any additional post-processing.

Can Soniox Speech-to-Text API transcribe recorded dictation?

Yes. In addition to real-time streaming, Soniox Speech-to-Text API supports asynchronous transcription for recorded audio files.

This can be useful for saved voice notes, uploaded dictation, background processing, and workflows that do not require text to appear immediately.

Does Soniox Speech-to-Text API rewrite or summarize dictated text?

Soniox Speech-to-Text API provides the speech recognition layer and returns the transcript and associated speech metadata.

If your product needs rewriting, summarization, tone changes, commands, or other transformations, you can pass the transcript into your own LLM or post-processing pipeline.

Does Soniox Speech-to-Text API store real-time dictation audio?

Real-time audio is processed transiently and is not stored by Soniox.

Recorded-file workflows use different storage behavior because asynchronous transcription requires the audio to be available while the job is processed.

Build dictation that feels faster than typing

Add responsive voice typing, context-aware recognition, structured speech handling, and native-speaker accuracy across 60+ languages with Soniox Speech-to-Text API.