Meeting transcription API for AI note-taking

Build meeting notes, summaries, action items, and searchable conversation history on accurate speaker-labeled transcripts. The Soniox Speech-to-Text API combines meeting transcription and speaker diarization with native-speaker accuracy across 60+ languages, precise timestamps, and context-aware recognition.

Trusted by teams building global voice products

Livekit
Krisp
Pipecat
Summary AI
Perplexity
Samsung
Wispr Flow
LG
Agora
Retell AI
Fireflies.ai
Skit.ai
Kindroid
Deliver Health
Truecaller
Journalia
Mobius
TranscribeMe
Vapi
Zomato
SLNG
Japan AI
Boost.ai
Convin
Genspark
HappyRobot
Uniscribe
Jamie.ai
InteractCX
The Plato
MobilApp
Onvego
Wonderful.ai
Manifone
Tana
Transync AI
SotaTek

Better AI notes start with better meeting transcription

Summaries, action items, decisions, and searchable meeting history are only as reliable as the transcript underneath them. Miss a name and the summary changes. Attribute a statement to the wrong speaker and the meeting record becomes misleading. Lose a number, date, or product term and important context disappears.The Soniox Speech-to-Text API gives note-taking products accurate, structured meeting transcripts before the rest of your AI pipeline takes over.

Accurate meeting transcription

Capture names, numbers, dates, product terminology, accents, and other imporant data with the accuracy your downstream notes depend on.

Built-in speaker diarization

Automatically separate speakers so your application can preserve who said what throughout meetings and multi-person conversations.

Native-speaker accuracy across 60+ languages

Get high-quality meeting transcriptions across all supported languages, including accents, dialects, and mixed-language speech.

Real-time or post-meeting

Stream transcripts while the meeting happens or process completed recordings asynchronously after the conversation ends.

From meeting audio to AI-generated notes

The Soniox Speech-to-Text API handles the speech layer between your meeting audio and the AI features your product builds on top.

Capture

Stream live meeting audio or submit a completed recording for asynchronous transcription.

Transcribe

Convert meeting speech into accurate text across speakers, accents, terminology, and languages.

Structure

Preserve speakers, timestamps, language information, and useful transcript metadata.

Build on it

Generate notes, summaries, action items, decisions, search, and other meeting intelligence.

Speaker diarization API for meeting transcripts

A meeting transcript becomes much more useful when your application knows who said what. Speaker diarization in the Soniox Speech-to-Text API automatically detects speaker changes and adds speaker information directly to the transcript output.

“Soniox knows who’s speaking and when each thought ends. The real-time transcripts read like true dialogue, not data dumps.”
Mobius MDAdam Strom, Co-Founder & President at Mobius MD

Speaker-labeled transcripts

Keep each participant's speech separate so downstream AI can distinguish questions, answers, commitments, decisions, and contributions throughout the meeting.

Diarization in real time or async

Use speaker diarization during live transcription or process the complete recording asynchronously for post-meeting workflows.

Native-speaker accuracy across 60+ languages

Global meeting products should not deliver great transcription in English and noticeably weaker results everywhere else. The Soniox Speech-to-Text API is built for native-speaker accuracy across 60+ supported languages with one unified speech model.

English

“It just gets the words right, any language, any accent, any context. That’s what accuracy is supposed to look like.”
AgoraTony Wang, Cofounder & Chief Revenue Officer at Agora

High accuracy in every supported language

Build the same high-quality meeting experience across English, Spanish, German, Japanese, Korean, Hindi, Arabic and dozens of other supported languages.

Natural language switching

Transcribe meetings where participants switch languages, pronounce foreign names, or mix terminology without restarting the stream or routing to another model.

Transcript data your note-taking API can build on

Speaker labels are only part of a useful meeting transcript. The Soniox Speech-to-Text API also gives your note-taking application timing, language, and contextual information needed to build richer meeting experiences.

Precise timestamps

Connect words and transcript content back to the original meeting audio for playback, highlighting, search, clips, and citations.

Explore timestamps

Context-aware recognition

Provide participant names, companies, projects, products, and terminology so important meeting vocabulary is recognized correctly.

Explore context customization

Automatic language identification

Identify spoken languages throughout the meeting, including conversations where participants switch languages naturally.

Explore language identification

Structured transcript output

Receive speech data in a form your application can pass directly into LLMs, search systems, knowledge bases, and other downstream workflows.

Speech infrastructure for massive scale

Soniox Text-to-Speech API performance and reliability

Build on one API and deploy in your region

Use the same models and API everywhere, with in-region processing to meet latency, data residency, and regulatory requirements.

Available: US, EU, Japan, India
Coming soon: Korea, Australia, Canada, Saudi Arabia, UK, Brazil

View data residency docs
Soniox Text-to-Speech API performance and reliability

Run mission-critical systems with confidence

  • 99.9% uptime
    Production-hardened infrastructure with monitoring and redundancy.
  • low-latency streaming
    Process speech in real time with low latency for responsive voice applications.
  • Priority support
    Severity-based incident response with direct access to the Soniox team.
Onvego uses Soniox Text-to-Speech API for multilingual voice experiences

"Before Soniox, our international users always had a noticeably different experience. Now accuracy and responsiveness match across all regions…it feels like one system instead of five."

Alon Yair CTO of Onvego

Build AI note-taking for every kind of meeting

Use the same meeting transcription API across collaboration, sales, customer success, interviews, research, and learning products.

AI meeting assistants

Turn meetings into speaker-aware transcripts for summaries, action items, decisions, and searchable meeting history.

Sales meeting notes

Capture customer conversations accurately and feed meeting notes, follow-ups, decisions, and CRM workflows.

Customer success meetings

Create reliable meeting records for customer needs, commitments, feedback, next steps, and account history.

Recruiting interviews

Transcribe interviews with clear speaker attribution and timestamps for notes, review, search, and hiring workflows.

User research

Capture research interviews and multi-speaker conversations as structured transcripts ready for analysis and synthesis.

Lectures and learning

Turn lectures, workshops, and discussions into searchable transcripts for notes, summaries, and learning experiences.

Real-time meeting transcription or post-meeting processing

Choose the transcription workflow that matches your note-taking product. Stream text while the meeting happens or process the full recording after it ends.

Real-time meeting transcription API

Receive transcription continuously while participants speak for live notes, captions, collaborative interfaces, and AI features that operate during the meeting.

Asynchronous meeting transcription API

Process completed meeting recordings when immediate output is not required and your product is building post-meeting notes, search, summaries, or other AI workflows.

Privacy and compliance, built right in

Never stored, never saved.

Audio stays in memory, everything is processed in real-time.

Built for privacy-critical use cases.

Adhering to leading global security, privacy, and compliance standards.

Trusted where privacy matters most.

Used in industries where speech is sensitive, from healthcare to enterprise.

Soniox is Soc 2 Type 2 compliant
Soniox is ISO 27001:2022 compliant
Soniox is HIPAA compliant
Soniox is GDPR compliant
SOC 2 Type 2 · ISO/IEC 27001:2022 · HIPAA · GDPR

Frequently asked questions about meeting transcription and speaker diarization APIs

What is a meeting transcription API?

A meeting transcription API converts meeting audio into text that your application can use programmatically.

The Soniox Speech-to-Text API accepts live audio streams or completed recordings and returns transcript data that can include speaker labels, timestamps, language information, and other metadata for downstream note-taking and AI workflows.

Is Soniox a meeting transcription API?

Yes. The Soniox Speech-to-Text API provides real-time and asynchronous transcription for meeting audio.

Your application supplies the live audio stream or completed recording, and the Soniox Speech-to-Text API returns the transcript and structured speech data your product can build on.

Does Soniox join Zoom, Google Meet, or Microsoft Teams meetings?

The Soniox Speech-to-Text API provides the transcription layer rather than the meeting-bot or conferencing capture layer.

If your application, recording system, or meeting bot can access the meeting audio, you can stream that audio to the Soniox Speech-to-Text API in real time or send the completed recording for asynchronous transcription.

What is a speaker diarization API?

A speaker diarization API separates different speakers in an audio stream and associates transcript content with speaker labels.

In meetings, this answers the question of who said what and gives note-taking applications the structure needed for speaker-aware notes, summaries, action items, and decisions.

Does Soniox include speaker diarization in its API?

Yes. The Soniox Speech-to-Text API supports speaker diarization in both real-time and asynchronous transcription.

When enabled, transcript tokens include speaker information that your application can group into speaker-attributed sections.

Can Soniox identify meeting participants by name?

Speaker diarization distinguishes voices and assigns speaker labels such as Speaker 1 and Speaker 2. It does not by itself determine the real-world identity of each participant.

Your note-taking application can map those labels to known meeting participants when participant identity is available from your meeting or capture layer.

Can I use Soniox as a note-taking API?

The Soniox Speech-to-Text API provides the speech transcription layer underneath a note-taking API or AI note-taking product.

It returns accurate, structured meeting transcripts that your own AI pipeline can turn into notes, summaries, action items, decisions, chapters, search, or other meeting intelligence.

Does Soniox generate meeting notes and summaries?

The Soniox Speech-to-Text API provides the transcript and structured speech metadata rather than a finished meeting note.

Your application can pass that output into its own LLM or AI pipeline to generate summaries, notes, action items, decisions, and other downstream features.

How accurate is Soniox for multilingual meeting transcription?

The Soniox Speech-to-Text API is built for native-speaker accuracy across more than 60 supported languages rather than treating multilingual transcription as an English-first add-on.

The same unified model can handle different accents, dialects, foreign names, and meetings where participants switch languages during the same conversation.

Can I transcribe meetings in real time?

Yes. The Soniox Speech-to-Text API streams transcription while meeting audio is arriving, making it suitable for live notes, captions, collaborative interfaces, and AI functionality that operates while the meeting is still happening.

Can I transcribe recorded meetings?

Yes. The Soniox Speech-to-Text API supports asynchronous transcription for completed meeting recordings and uploaded audio.

This is useful when your note-taking application generates its final transcript and meeting intelligence after the call has ended.

Does Soniox provide timestamps?

Yes. The Soniox Speech-to-Text API provides precise timestamps for recognized transcript tokens by default.

Note-taking products can use timestamps for playback, highlighting, transcript navigation, clips, citations, and links from generated notes back to the source meeting.

Can I improve recognition of participant names and terminology?

Yes. Context in the Soniox Speech-to-Text API lets your application provide participant names, company names, products, projects, terminology, topics, and other relevant meeting information.

Context can improve recognition of uncommon vocabulary and provide useful information about the expected meeting content.

Can Soniox transcribe multilingual meetings?

Yes. The Soniox Speech-to-Text API supports transcription across more than 60 languages with one unified model.

It can automatically identify spoken languages and handle meetings where participants switch languages within the same conversation.

Build your AI note-taking product on better meeting transcripts

Add accurate meeting transcription, speaker diarization, native-speaker accuracy across 60+ languages, timestamps, and context-aware speech data to the note-taking experience you are building.