Medical speech-to-text API for clinical transcription and dictation
Build medical transcription, clinical dictation, and healthcare voice products with Soniox Speech-to-Text API. Recognize medical terminology, separate clinicians and patients, transcribe live or recorded audio, and support 60+ languages with one secure speech recognition API.
Trusted by teams building global voice products
Medical speech recognition depends on getting the clinical details right
Clinical speech is dense with information. Medication names, specialist terminology, abbreviations, dosages, dates, measurements, patient names, and rapidly spoken details all need to survive transcription accurately. Soniox Speech-to-Text API gives healthcare developers a medical speech recognition foundation for clinical conversations, dictation, and documentation workflows.
Understand medical terminology
Recognize clinical language and provide medical terms, medication names, specialties, procedures, and organization-specific vocabulary as context.

Separate clinicians and patients
Detect speaker changes automatically so consultations, interviews, and patient encounters remain clear and properly attributed.

Handle dictation and conversation
Transcribe rapid clinical dictation, natural patient encounters, telehealth sessions, and recorded medical conversations with the same API.

Support multilingual healthcare
Transcribe speech across 60+ languages with one model and handle language switching without maintaining separate recognition pipelines.

From clinical speech to structured documentation
Soniox Speech-to-Text API turns live or recorded medical speech into structured transcript data your healthcare application can build on.
Capture
Stream live clinical audio to Soniox Speech-to-Text API or submit recorded encounters and dictation for asynchronous transcription.
Recognize
Transcribe clinical speech using medical context, speaker separation, and automatic language identification.
Structure
Receive speaker-aware transcripts with timestamps, language data, and useful speech metadata.
Integrate
Send structured transcript data into documentation, EHR, scribing, analytics, or other healthcare workflows.
Clinical speech recognition API built around healthcare workflows
Medical transcription requires more than generic speech recognition. Healthcare applications need control over clinical terminology, clear speaker attribution, reliable multilingual speech handling, and infrastructure designed for sensitive healthcare data.
Good morning, Mrs.Valetti. SinceCalypta didn't reduce your migraine days, we'll startVieppti infusions every 12 weeks, and review how you respond at your next visit.
Good morning, Mrs.Faletti. SinceQulipta didn't reduce your migraine days, we'll startVyepti infusions every 12 weeks, and review how you respond at your next visit.
“Soniox captures complex medical terminology with high accuracy, helping physicians finalize notes faster and focus on patient care.”
Max Malyk, Vice President at DeliverHealthMedical vocabulary and context
Give Soniox Speech-to-Text API the clinical context behind each session. Provide specialty terms, medication names, procedures, clinician names, organizations, and other domain-specific vocabulary to improve recognition.
Explore context customizationClinician-patient speaker separation
Automatically detect speaker changes and preserve who said what across consultations, interviews, telehealth sessions, and other multi-speaker clinical conversations.
Explore speaker diarizationAutomatic language identification
Identify spoken languages automatically, including encounters where clinicians and patients switch languages during the same conversation.
Explore language identificationRegional data residency
Keep audio and transcript content within supported regions for processing and storage, helping healthcare organizations meet data residency requirements.
Explore data residencySpeech infrastructure for massive scale

Build on one API and deploy in your region
Use the same models and API everywhere, with in-region processing to meet latency, data residency, and regulatory requirements.
Available: US, EU, Japan, India
Coming soon: Korea, Australia, Canada, Saudi Arabia, UK, Brazil

Run mission-critical systems with confidence
- 99.9% uptime
Production-hardened infrastructure with monitoring and redundancy. - low-latency streaming
Process speech in real time with low latency for responsive voice applications. - Priority support
Severity-based incident response with direct access to the Soniox team.
"Before Soniox, our international users always had a noticeably different experience. Now accuracy and responsiveness match across all regions…it feels like one system instead of five."
Alon Yair CTO of Onvego
Build medical transcription and dictation into healthcare products
Use Soniox Speech-to-Text API across ambient documentation, medical dictation, telehealth, patient intake, research, and clinical workflow products.
Ambient clinical documentation
Capture clinician-patient conversations with speaker attribution and timestamps, ready for downstream notes, summaries, and charting workflows.
Medical dictation
Transcribe fast, free-form clinical dictation while preserving medical terminology, names, numbers, dates, and other important details.
Telehealth transcription
Add real-time medical transcription to virtual visits with speaker separation, multilingual support, and structured transcript output.
Patient intake
Capture symptoms, medical history, medications, and patient-reported information as structured speech data for downstream workflows.
Clinical research
Create searchable, speaker-aware transcripts for research interviews, clinical studies, qualitative analysis, and recorded sessions.
Care coordination
Turn handoffs, consultations, and care-team conversations into accurate transcripts that can flow into documentation and workflow systems.
Real-time and asynchronous medical speech-to-text API
Healthcare products need both live and recorded transcription. Soniox Speech-to-Text API supports both processing modes so your application can choose the right workflow for each clinical experience.
Real-time medical transcription
Stream text as clinicians and patients speak. Use real-time medical speech recognition for ambient documentation, telehealth, live dictation, clinical assistants, and in-session workflows.

Asynchronous medical transcription
Process completed recordings for clinical encounters, dictated notes, research interviews, consultations, and other healthcare workflows where immediate transcription is not required.
Privacy and compliance, built right in
Never stored, never saved.
Audio stays in memory, everything is processed in real-time.
Built for privacy-critical use cases.
Adhering to leading global security, privacy, and compliance standards.
Trusted where privacy matters most.
Used in industries where speech is sensitive, from healthcare to enterprise.




Frequently asked questions about medical speech recognition APIs
What is a medical speech recognition API?
A medical speech recognition API converts clinical speech into text that healthcare applications can use programmatically.
Soniox Speech-to-Text API can process live or recorded clinical audio and return transcript data for medical transcription, dictation, ambient documentation, telehealth, research, and other healthcare workflows.
Is Soniox Speech-to-Text API suitable for medical transcription?
Yes. Soniox Speech-to-Text API supports real-time and asynchronous transcription for clinical conversations, medical dictation, telehealth, ambient documentation, and other healthcare speech workflows.
It includes capabilities such as context customization, speaker diarization, timestamps, language identification, and multilingual transcription.
Can Soniox Speech-to-Text API recognize medical terminology and vocabulary?
Yes. Soniox Speech-to-Text API lets your application provide domain-specific context with each transcription request.
You can supply medical terminology, medication names, procedures, specialties, clinician names, organizations, patient context, and other relevant vocabulary to help guide recognition.
Can I customize clinical speech recognition for different medical specialties?
Yes. Context can be supplied dynamically for each session, allowing your application to provide terminology and background information relevant to the specialty, provider, patient, or encounter.
This lets you adapt clinical speech recognition without maintaining a separately fine-tuned model for every healthcare workflow.
Does Soniox Speech-to-Text API support speaker diarization for clinician-patient conversations?
Yes. Speaker diarization detects speaker changes and assigns transcript content to different speaker labels.
Healthcare applications can use those labels to structure clinician-patient encounters, consultations, interviews, telehealth sessions, and other multi-speaker medical conversations.
Can I use Soniox Speech-to-Text API as a medical dictation API?
Yes. Soniox Speech-to-Text API can transcribe fast, free-form medical dictation in real time or process completed dictation recordings asynchronously.
Context customization is especially useful for medical dictation applications where clinicians regularly use specialty terms, medication names, abbreviations, people, and organization-specific vocabulary.
Can I use Soniox Speech-to-Text API as a clinical dictation API?
Yes. Clinical dictation applications can stream microphone audio to Soniox Speech-to-Text API and receive transcript results while the clinician is speaking.
Your application can then use the transcript inside documentation, EHR, editing, summarization, or other clinical workflows.
Can Soniox Speech-to-Text API transcribe clinical encounters in real time?
Yes. Soniox Speech-to-Text API streams transcript results while audio is arriving.
This makes it suitable for ambient clinical documentation, telehealth, live medical dictation, clinical assistants, and healthcare applications that need transcript data during the encounter.
Can Soniox Speech-to-Text API process recorded medical audio?
Yes. Soniox Speech-to-Text API provides asynchronous transcription for completed audio files.
This is useful for recorded consultations, dictated notes, clinical research interviews, uploaded audio, and other workflows where real-time results are not required.
Does Soniox Speech-to-Text API support multilingual medical transcription?
Yes. Soniox Speech-to-Text API supports speech recognition across more than 60 languages using one unified model.
Spoken languages can be identified automatically, including clinical conversations where speakers switch languages during the same encounter.
Can Soniox Speech-to-Text API translate clinician-patient conversations?
Yes. Soniox Speech-to-Text API supports speech translation alongside transcription for both real-time and asynchronous workflows.
Healthcare applications can use this capability to build multilingual transcript and communication experiences while preserving the original speech transcription.
Is Soniox HIPAA compliant?
Soniox lists HIPAA among its supported compliance frameworks and maintains SOC 2 Type 2 and ISO/IEC 27001:2022 certifications.
Healthcare deployments involving protected health information may require additional contractual terms, such as a Business Associate Agreement. Review the Soniox Security & Compliance documentation and your organization's requirements before production use.
Does Soniox Speech-to-Text API store patient audio?
Real-time API requests are processed transiently and are not stored by Soniox. Soniox also states that customer audio and transcripts are not used to train its models.
Asynchronous services require storage for processing and retrieval, and stored content follows the applicable retention and deletion behavior documented by Soniox.
Can healthcare data be processed in a specific region?
Yes. Soniox supports regional data residency for eligible projects. When a region is selected, audio and transcript content remains within that region for processing and storage.
Regional deployments are currently documented for the United States, European Union, Japan, and India.
Can I integrate Soniox Speech-to-Text API into an EHR or clinical documentation platform?
Yes. Soniox Speech-to-Text API returns transcript content and metadata your application can pass into its own clinical documentation, EHR, scribing, analytics, or workflow layer.
Your application remains responsible for mapping, validation, formatting, and integration with the target healthcare system.
Should medical transcripts be reviewed before clinical use?
Yes. Speech recognition systems can produce errors, and healthcare applications should apply the review, validation, safeguards, and compliance processes appropriate for their use case.
Soniox Speech-to-Text API provides the speech recognition layer; your application determines how transcript output is reviewed and used within the clinical workflow.
Build healthcare products on accurate medical speech recognition
Add medical transcription, clinical dictation, speaker separation, multilingual speech recognition, and secure speech infrastructure to the healthcare applications you build with Soniox Speech-to-Text API.