# Soniox docs Documentation from the [soniox.com/docs](https://soniox.com/docs) website. ## Links to docs content pages: - [AI engineering](https://soniox.com/docs/ai-engineering) - [Community and support](https://soniox.com/docs/community-and-support) - [Data residency](https://soniox.com/docs/data-residency) - [FAQ](https://soniox.com/docs/faq) - [Overview](https://soniox.com/docs/) - [Security and privacy](https://soniox.com/docs/security-and-privacy) - [Errors](https://soniox.com/docs/api-reference/errors) - [Soniox API reference](https://soniox.com/docs/api-reference) - [Soniox Live](https://soniox.com/docs/demo-apps/soniox-live) - [Speech-to-speech translation](https://soniox.com/docs/demo-apps/soniox-speech-to-speech-translation) - [Soniox Voice Agent](https://soniox.com/docs/demo-apps/soniox-voice-agent) - [Concurrency limits](https://soniox.com/docs/guides/concurrency-limits) - [Direct stream](https://soniox.com/docs/guides/direct-stream) - [Guides](https://soniox.com/docs/guides) - [Proxy stream](https://soniox.com/docs/guides/proxy-stream) - [Temporary API keys](https://soniox.com/docs/guides/temporary-api-keys) - [Usage logs](https://soniox.com/docs/guides/usage-logs) - [Usage summary](https://soniox.com/docs/guides/usage-summary) - [Community integrations](https://soniox.com/docs/integrations/community-integrations) - [Integrations](https://soniox.com/docs/integrations) - [n8n](https://soniox.com/docs/integrations/n8n) - [TanStack AI SDK](https://soniox.com/docs/integrations/tanstack-ai-sdk) - [Twilio](https://soniox.com/docs/integrations/twilio) - [Vercel AI SDK](https://soniox.com/docs/integrations/vercel-ai-sdk) - [Get started](https://soniox.com/docs/stt/get-started) - [Models](https://soniox.com/docs/stt/models) - [Get started](https://soniox.com/docs/translation/get-started) - [Real-time speech-to-speech translation](https://soniox.com/docs/translation/sts-translation) - [Supported languages](https://soniox.com/docs/translation/supported-languages) - [Get started](https://soniox.com/docs/tts/get-started) - [Models](https://soniox.com/docs/tts/models) - [SDKs](https://soniox.com/docs/sdk) - [React Native SDK](https://soniox.com/docs/sdk/react-native-SDK) - [Get concurrency limits](https://soniox.com/docs/api-reference/other/get_concurrency_limits) - [Get concurrent streams history](https://soniox.com/docs/api-reference/other/get_concurrent_streams_history) - [Get usage logs](https://soniox.com/docs/api-reference/other/get_usage_logs) - [Get usage summary](https://soniox.com/docs/api-reference/other/get_usage_summary) - [Get models](https://soniox.com/docs/api-reference/stt/get_models) - [WebSocket API](https://soniox.com/docs/api-reference/stt/websocket-api) - [Generate speech](https://soniox.com/docs/api-reference/tts/generate_tts) - [Get TTS models](https://soniox.com/docs/api-reference/tts/get_tts_models) - [WebSocket API](https://soniox.com/docs/api-reference/tts/websocket-api) - [LangChain.js (JavaScript)](https://soniox.com/docs/integrations/langchain/langchain-js) - [LangChain (Python)](https://soniox.com/docs/integrations/langchain/langchain) - [LiveKit](https://soniox.com/docs/integrations/livekit) - [Switch your LiveKit STT or TTS provider to Soniox](https://soniox.com/docs/integrations/livekit/migrate) - [Speech-to-Text](https://soniox.com/docs/integrations/livekit/stt) - [Text-to-Speech](https://soniox.com/docs/integrations/livekit/tts) - [Build a voice agent with LiveKit and Soniox](https://soniox.com/docs/integrations/livekit/voice-agent) - [Pipecat](https://soniox.com/docs/integrations/pipecat) - [Switch your Pipecat STT or TTS provider to Soniox](https://soniox.com/docs/integrations/pipecat/migrate) - [Speech-to-Text](https://soniox.com/docs/integrations/pipecat/stt) - [Text-to-Speech](https://soniox.com/docs/integrations/pipecat/tts) - [Build a voice agent with Pipecat and Soniox](https://soniox.com/docs/integrations/pipecat/voice-agent) - [Async transcription](https://soniox.com/docs/stt/async/async-transcription) - [Error handling](https://soniox.com/docs/stt/async/error-handling) - [Limits & quotas](https://soniox.com/docs/stt/async/limits-and-quotas) - [Webhooks](https://soniox.com/docs/stt/async/webhooks) - [Confidence scores](https://soniox.com/docs/stt/concepts/confidence-scores) - [Context](https://soniox.com/docs/stt/concepts/context) - [Language hints](https://soniox.com/docs/stt/concepts/language-hints) - [Language identification](https://soniox.com/docs/stt/concepts/language-identification) - [Language restrictions](https://soniox.com/docs/stt/concepts/language-restrictions) - [Speaker diarization](https://soniox.com/docs/stt/concepts/speaker-diarization) - [Supported languages](https://soniox.com/docs/stt/concepts/supported-languages) - [Timestamps](https://soniox.com/docs/stt/concepts/timestamps) - [Connection keepalive](https://soniox.com/docs/stt/rt/connection-keepalive) - [Endpoint detection](https://soniox.com/docs/stt/rt/endpoint-detection) - [Error handling](https://soniox.com/docs/stt/rt/error-handling) - [Limits & quotas](https://soniox.com/docs/stt/rt/limits-and-quotas) - [Manual finalization](https://soniox.com/docs/stt/rt/manual-finalization) - [Real-time transcription](https://soniox.com/docs/stt/rt/real-time-transcription) - [Async speech-to-text translation](https://soniox.com/docs/translation/stt-translation/async-translation) - [Speech-to-text translation](https://soniox.com/docs/translation/stt-translation) - [Real-time speech-to-text translation](https://soniox.com/docs/translation/stt-translation/rt-translation) - [Audio formats](https://soniox.com/docs/tts/concepts/audio-formats) - [Emotion & tone](https://soniox.com/docs/tts/concepts/emotion-and-tone) - [Language mixing](https://soniox.com/docs/tts/concepts/language-mixing) - [Reduce silence](https://soniox.com/docs/tts/concepts/reduce-silence) - [Speech speed](https://soniox.com/docs/tts/concepts/speech-speed) - [Supported languages](https://soniox.com/docs/tts/concepts/supported-languages) - [Voice cloning](https://soniox.com/docs/tts/concepts/voice-cloning) - [Voices](https://soniox.com/docs/tts/concepts/voices) - [Generate speech](https://soniox.com/docs/tts/rest-api/generate-speech) - [Limits & quotas](https://soniox.com/docs/tts/rest-api/limits-and-quotas) - [Node SDK](https://soniox.com/docs/sdk/node-SDK) - [Python SDK](https://soniox.com/docs/sdk/python-SDK) - [Sync vs async clients](https://soniox.com/docs/sdk/python-SDK/sync-vs-async-clients) - [React SDK](https://soniox.com/docs/sdk/react-SDK) - [Web SDK](https://soniox.com/docs/sdk/web-SDK) - [Delete file](https://soniox.com/docs/api-reference/stt/files/delete_file) - [Get file](https://soniox.com/docs/api-reference/stt/files/get_file) - [Get files](https://soniox.com/docs/api-reference/stt/files/get_files) - [Get files count](https://soniox.com/docs/api-reference/stt/files/get_files_count) - [Upload file](https://soniox.com/docs/api-reference/stt/files/upload_file) - [Create transcription](https://soniox.com/docs/api-reference/stt/transcriptions/create_transcription) - [Delete transcription](https://soniox.com/docs/api-reference/stt/transcriptions/delete_transcription) - [Get transcription](https://soniox.com/docs/api-reference/stt/transcriptions/get_transcription) - [Get transcription transcript](https://soniox.com/docs/api-reference/stt/transcriptions/get_transcription_transcript) - [Get transcriptions](https://soniox.com/docs/api-reference/stt/transcriptions/get_transcriptions) - [Get transcriptions count](https://soniox.com/docs/api-reference/stt/transcriptions/get_transcriptions_count) - [Create voice](https://soniox.com/docs/api-reference/tts/voices/create_voice) - [Delete voice](https://soniox.com/docs/api-reference/tts/voices/delete_voice) - [Get voice](https://soniox.com/docs/api-reference/tts/voices/get_voice) - [Get voices](https://soniox.com/docs/api-reference/tts/voices/get_voices) - [Get voices count](https://soniox.com/docs/api-reference/tts/voices/get_voices_count) - [Recompute voice](https://soniox.com/docs/api-reference/tts/voices/recompute_voice) - [Create temporary API key](https://soniox.com/docs/api-reference/auth/create_temporary_api_key) - [Connection keepalive](https://soniox.com/docs/tts/rt/connection-keepalive) - [Limits & quotas](https://soniox.com/docs/tts/rt/limits-and-quotas) - [Real-time generation](https://soniox.com/docs/tts/rt/real-time-generation) - [Streams](https://soniox.com/docs/tts/rt/streams) - [Stream termination](https://soniox.com/docs/tts/rt/termination) - [Timestamps](https://soniox.com/docs/tts/rt/timestamps) - [Classes](https://soniox.com/docs/sdk/node-SDK/reference/classes) - [Full Node SDK reference](https://soniox.com/docs/sdk/node-SDK/reference) - [Types](https://soniox.com/docs/sdk/node-SDK/reference/types) - [Async transcription with Node SDK](https://soniox.com/docs/sdk/node-SDK/stt/async-transcription) - [Async translation with Node SDK](https://soniox.com/docs/sdk/node-SDK/stt/async-translation) - [Handling files with Node SDK](https://soniox.com/docs/sdk/node-SDK/stt/files) - [Real-time transcription with Node SDK](https://soniox.com/docs/sdk/node-SDK/stt/realtime-transcription) - [Handling webhooks with Node SDK](https://soniox.com/docs/sdk/node-SDK/stt/webhooks) - [Real-time speech generation with Node SDK](https://soniox.com/docs/sdk/node-SDK/tts/realtime-speech-generation) - [REST speech generation with Node SDK](https://soniox.com/docs/sdk/node-SDK/tts/rest-speech-generation) - [Async Client](https://soniox.com/docs/sdk/python-SDK/Full-SDK-reference/async_client) - [Realtime Client](https://soniox.com/docs/sdk/python-SDK/Full-SDK-reference/realtime_client) - [Types](https://soniox.com/docs/sdk/python-SDK/Full-SDK-reference/types) - [Helpers](https://soniox.com/docs/sdk/python-SDK/Full-SDK-reference/utils) - [Async transcription with Python SDK](https://soniox.com/docs/sdk/python-SDK/stt/async-transcription) - [Handling files with Python SDK](https://soniox.com/docs/sdk/python-SDK/stt/files) - [Real-time transcription with Python SDK](https://soniox.com/docs/sdk/python-SDK/stt/realtime-transcription) - [Handling webhooks with Python SDK](https://soniox.com/docs/sdk/python-SDK/stt/webhooks) - [Real-time speech generation with Python SDK](https://soniox.com/docs/sdk/python-SDK/tts/realtime-speech-generation) - [REST speech generation with Python SDK](https://soniox.com/docs/sdk/python-SDK/tts/rest-speech-generation) - [Full React SDK reference](https://soniox.com/docs/sdk/react-SDK/reference) - [Types](https://soniox.com/docs/sdk/react-SDK/reference/types) - [Real-time transcription with React SDK](https://soniox.com/docs/sdk/react-SDK/stt/realtime-transcription) - [Real-time speech generation with React SDK](https://soniox.com/docs/sdk/react-SDK/tts/realtime-speech-generation) - [Classes](https://soniox.com/docs/sdk/web-SDK/reference/classes) - [Full Web SDK reference](https://soniox.com/docs/sdk/web-SDK/reference) - [Types](https://soniox.com/docs/sdk/web-SDK/reference/types) - [Real-time transcription with Web SDK](https://soniox.com/docs/sdk/web-SDK/stt/realtime-transcription) - [Real-time speech generation with Web SDK](https://soniox.com/docs/sdk/web-SDK/tts/realtime-speech-generation) - [REST speech generation with Web SDK](https://soniox.com/docs/sdk/web-SDK/tts/rest-speech-generation) # AI engineering URL: /ai-engineering Using MCP, AI assistant, and LLMs with Soniox for AI-powered development import Image from "next/image"; Soniox provides easy-to-use AI tools that help you explore documentation, generate code, and get guidance, even if you're new to programming. These tools work directly with your coding environment, so you can focus on building instead of searching for answers. With Soniox AI engineering, you can: * Browse documentation via the **MCP server** without leaving your coding tools * Ask the **AI assistant** for explanations, examples, or code help * Use **LLM context files** so AI models understand Soniox APIs and examples * Copy page content or open it directly in your preferred AI tool These features reduce friction, help you learn faster, and make working with Soniox APIs simple and efficient. *** ## MCP server The **MCP server** lets you access Soniox documentation right from tools like Cursor, Windsurf, or Claude Code. You can search guides, view examples, and explore APIs without switching windows. ### Available tools The Soniox MCP server exposes the following tools to your AI assistant: Search Soniox documentation by natural-language query. Use this first to find focused pages and headings before reading a page or section. **Inputs** * `query` (string, required) – natural-language search query. * `limit` (integer, optional) – maximum number of results to return (1–10). Read a specific Soniox documentation page by URL. Accepts relative URLs (e.g. `/stt/get-started`) and absolute URLs (e.g. `https://soniox.com/docs/stt/get-started`). Use `soniox_search_docs` first unless the URL is already known. **Inputs** * `url` (string, required) – page URL to read. Read one section from a Soniox documentation page by URL plus heading or anchor. Use `soniox_search_docs` first when the URL or heading is unknown. **Inputs** * `url` (string, required) – page URL to read from. * `heading` (string, optional) – heading text within the page. * `anchor` (string, optional) – anchor identifier within the page. Get the table of contents for broad Soniox documentation navigation. Use `soniox_search_docs` for natural-language search instead. Get a concise overview of Soniox documentation (`llms.txt`) for background on features, APIs, and concepts. Use `soniox_search_docs` when looking for a specific answer. Last resort: get the full content of all Soniox documentation pages (\~100k tokens). This is very large — prefer `soniox_search_docs`, then `soniox_read_section` or `soniox_read_page`.
### How to set it up Pick your AI tool below and follow the steps to connect the Soniox MCP server. Open **Cursor settings → Tools & MCPs** and click **New MCP Server**, or edit `~/.cursor/mcp.json` directly and add the following entry: ```json title="~/.cursor/mcp.json" { "mcpServers": { "soniox-docs": { "url": "https://soniox.com/docs/api/mcp/mcp" } } } ``` Prefer a one-click install? Use the button below:
[![Install MCP Server](https://cursor.com/deeplink/mcp-install-light.svg)](https://cursor.com/en/install-mcp?name=soniox-docs\&config=eyJjb21tYW5kIjoibnB4IC15IG1jcC1yZW1vdGUgaHR0cHM6Ly9zb25pb3guY29tL2RvY3MvYXBpL21jcC9tY3AifQ%3D%3D)
[![Install MCP Server](https://cursor.com/deeplink/mcp-install-dark.svg)](https://cursor.com/en/install-mcp?name=soniox-docs\&config=eyJjb21tYW5kIjoibnB4IC15IG1jcC1yZW1vdGUgaHR0cHM6Ly9zb25pb3guY29tL2RvY3MvYXBpL21jcC9tY3AifQ%3D%3D)
Open Claude Desktop and click **Customize** in the left sidebar. In the Customize panel, select **Connectors**. Click the **+** button and choose **Add custom connector**. Give it a name (for example, `Soniox docs`) and paste the server URL below: ```text title="Server URL" https://soniox.com/docs/api/mcp/mcp ``` Open [ChatGPT](https://chatgpt.com/) and go to **Settings**. Navigate to **Connectors** → **Developer mode**. Click **Add MCP server**. Paste the server URL below and click **Add**: ```text title="Server URL" https://soniox.com/docs/api/mcp/mcp ``` Codex accepts a remote MCP server through a built-in form — no config file to edit. Go to **Settings → MCP Servers**. Click **Add server**. Select **Streamable HTTP**. Paste the URL below into the **URL** field and save: ```text title="Server URL" https://soniox.com/docs/api/mcp/mcp ``` In Windsurf, open **Settings → Cascade → MCP Servers**. Click the gear icon — `mcp_config.json` opens automatically. Paste the snippet below into the file and save: ```json title="mcp_config.json" { "mcpServers": { "soniox-docs": { "serverUrl": "https://soniox.com/docs/api/mcp/mcp" } } } ``` Open Zed's agent panel. Click the **…** button at the top-right of the panel, then **Add MCP Server**. In the new window, click **Remote**. Paste the snippet below and save: ```json title="settings.json" { "soniox-docs": { "url": "https://soniox.com/docs/api/mcp/mcp" } } ``` In Antigravity, open the MCP store via the **…** dropdown at the top of the editor's agent panel. Click **Manage MCP Servers → View raw config**. Paste the snippet below into `mcp_config.json` and save: ```json title="mcp_config.json" { "mcpServers": { "soniox-docs": { "serverUrl": "https://soniox.com/docs/api/mcp/mcp" } } } ``` Open [Claude.ai settings](https://claude.ai/settings). Navigate to **Integrations** and click **Add Integration**. Paste the server URL below and click **Connect**: ```text title="Server URL" https://soniox.com/docs/api/mcp/mcp ```
*** ## AI assistant The **Soniox AI assistant** is available directly from the docs. It can: * Answer questions about Soniox APIs * Explain example code or suggest modifications * Provide guidance in context, so you don't need to guess Even if you're new to programming, the AI assistant can help you understand code and API workflows quickly. *** ## LLM context files Soniox provides two files that give AI models context about our APIs and examples: * [llms.txt](/llms.txt) – core context for general tasks * [llms-full.txt](/llms-full.txt) – extended context for advanced workflows Adding these files to your AI tool ensures the model can provide accurate, context-aware help. *** ## Copy and open buttons Copy button At the top of each documentation page, the **Copy page** button makes it easy to bring content into your workflow: * **Copy Markdown** – copy the full page content instantly * **Open in ChatGPT or Claude** – send the page context for live AI interaction These features help you experiment and learn by bringing examples and documentation directly into your coding environment. *** For more information about Soniox products, pricing, or general resources, visit our [website](https://soniox.com/). # Community and support URL: /community-and-support Engage with our community to explore new updates, participate in discussions, contribute to our projects, and report any issues you encounter. import { LinkCard } from "@/components/link-card"; import { GitHubIcon } from "@/components/github-icon"; ## Support We offer three levels of support depending on your plan: **Free**
Get help from the developer community through our [Discord server](https://discord.gg/rWfnk9uM5j). **Business and Enterprise support**
For production deployments, Soniox offers dedicated support channels, response-time SLAs, and escalation paths based on your support plan. Contact [support@soniox.com](mailto:support@soniox.com) to discuss the right support option for your team. *** ## Github We use GitHub to track issues related to official Soniox SDKs and integrations. Check our [Soniox GitHub](https://github.com/soniox) profile for all available code. *** ## Website For more information about our products, pricing or Soniox in general, visit out [website](https://soniox.com/). # Data residency URL: /data-residency Learn about data residency. Soniox keeps your data yours. Any content you send to the Soniox API (audio, transcripts, or metadata) is **never used to train or improve our models.** For more information, see our [Security and privacy](/security-and-privacy). *** ## What is data residency Data residency lets you choose **where** Soniox processes and stores your content. When you select a region for a project, **all audio and transcript data for that project stays in that region**, for both processing and storage. To get access to regional deployments send your inquiry to [support@soniox.com](mailto:support@soniox.com). *** ## How data residency works When data residency is enabled for your account: * You choose a **region** when creating a new project. * Any API requests made using that project's API key are handled **fully within the selected region.** * All **content data** (audio + transcripts) remains within that region for processing and storage. ### System data Data residency **does not apply to system data** such as account and project metadata, usage statistics, and billing data. This system data may be processed outside the selected region. **Your content (audio + transcripts) never leaves the region you choose.** *** ## Using data residency Data residency is set **per project** within your Soniox organization. ### 1. Create a project with a region When creating a new project: * Select the region from the **region** dropdown. * Each project receives region-specific API keys. ### 2. Use the region-specific API domain To ensure processing stays in the region, use: * The **API key** from the regional project. * The **correct API domain** for that region (see below). *** ## Regional endpoints | Region | Regional storage | Regional processing | Capabilities | API domain | | ------------------ | ---------------- | ------------------- | ------------------ | ------------------------------------------------------------------------------- | | **United States** | ✅ Yes | ✅ Yes | Full API supported | `api.soniox.com`
`stt-rt.soniox.com`
`tts-rt.soniox.com` | | **European Union** | ✅ Yes | ✅ Yes | Full API supported | `api.eu.soniox.com`
`stt-rt.eu.soniox.com`
`tts-rt.eu.soniox.com` | | **Japan** | ✅ Yes | ✅ Yes | Full API supported | `api.jp.soniox.com`
`stt-rt.jp.soniox.com`
`tts-rt.jp.soniox.com` | | **India** | ✅ Yes | ✅ Yes | Full API supported | `api.in.soniox.com`
`stt-rt.in.soniox.com`
`tts-rt.in.soniox.com` | If you need help enabling data residency reach out to [support@soniox.com](mailto:support@soniox.com). # FAQ URL: /faq Common troubleshooting guidance and answers for integrating with the Soniox API. This page answers common questions related to integrating with the Soniox API. High WebSocket connection startup time is usually caused by one or more of the following: * **Network latency:** High round-trip time between your client and Soniox increases the duration of the TLS handshake and WebSocket upgrade. * **Region selection:** Using an endpoint located far from your compute environment adds unnecessary cross-region latency during connection establishment. See [Data residency](/data-residency) for more info. * **Large initial context payload:** Sending a large [context](/stt/concepts/context) during initialization delays readiness because the server must fully receive and process the payload before the session becomes active. To minimize perceived startup delay, you should always **buffer audio locally before the WebSocket connection is established** and immediately stream all buffered audio chunks after sending the initial configuration message. You can request a limit increase from the [Soniox Console organization limits](https://console.soniox.com/org/limits) page. Requests are reviewed within 1-3 business days. Yes. Soniox provides standard legal and compliance documentation for companies integrating the Soniox API into their products or services. This may include an MSA, DPA, and security or compliance documentation required for procurement or security review processes. Documentation is available for Business and Enterprise customers. Please contact [support@soniox.com](mailto:support@soniox.com) to request access or begin the review process. # Overview URL: / Soniox provides powerful, production-ready APIs for transcribing, translating, generating, and understanding audio content. import { LinkCards, SpeechToTextIcon, TextToSpeechIcon, TranslationIcon, } from "@/components/link-card"; import { Step, Steps } from "fumadocs-ui/components/steps"; ## Get started with Soniox APIs Welcome to Soniox, the voice AI platform for speech-to-text, text-to-speech, and translation, with unmatched accuracy in 60+ languages. Build production-ready voice products with fast, accurate APIs for transcribing audio, generating natural speech, translating across languages, and extracting structure and meaning from conversations. Whether you are creating real-time voice interfaces, processing large audio volumes, or powering multilingual experiences, Soniox is designed to help you move quickly and scale with confidence. Integrate Soniox through simple REST and WebSocket APIs, with SDKs and real-time streaming support for modern applications. ## Products , description: "Transcribe speech in real time across 60+ languages, with native-speaker accuracy for multilingual, language-switching, and multi-speaker conversations.", arrowInSeparateLine: true, titleSize: "text-xl", cta: "Get started", }, { title: "Text-to-Speech", href: "/tts/get-started", headIcon: , description: "Generate natural, high-fidelity speech in 60+ languages, with precise handling of alphanumerics, names, borrowed words, and language switching.", arrowInSeparateLine: true, titleSize: "text-xl", cta: "Get started", }, { title: "Speech Translation", href: "/translation/get-started", headIcon: , description: "Translate speech in real time across 3,600 language pairs, with low-latency output before sentences finish and high-quality multilingual results.", arrowInSeparateLine: true, titleSize: "text-xl", cta: "Get started", }, ]} /> ## Before you begin To start using Soniox, create a [Soniox account](https://console.soniox.com/signup/). Visit the [Soniox Console](https://console.soniox.com/) to generate and manage API keys, view usage, logs, and billing. Soniox Console is your self-service control center for everything Soniox. # Security and privacy URL: /security-and-privacy Learn about security and privacy policies. At Soniox, we take security and privacy seriously. Our platform is designed to keep your data protected while reducing compliance burdens for your business. This page outlines how Soniox handles data, meets compliance requirements, and ensures secure communication. *** ## Compliance Soniox meets industry-leading certification standards: * **SOC 2 Type 2** – auditing standard that evaluates an organization’s controls for security, availability, processing integrity, confidentiality, and privacy over an extended period of time. * **ISO/IEC 27001:2022** – internationally recognized standard for Information Security Management Systems (ISMS). * **GDPR** – European Union regulation that governs the collection, processing, and protection of personal data and privacy rights. * **HIPAA** – U.S. regulatory framework that establishes requirements for protecting sensitive healthcare data, including Protected Health Information (PHI). All compliance documentation can be obtained through Soniox Console [Security & compliance section](https://console.soniox.com/org/security-and-compliance/). *** ## Data handling * **No model training** – your audio and transcripts are never used to improve Soniox models or services. * **No retention** – Soniox does not store your audio or transcript data unless explicitly requested through a service that supports storage, i.e. Async API. * **Storage** – when you choose to store data, it is securely isolated within your Soniox Account. * **Data deletion** – you can delete all stored audio and transcripts at any time via the Soniox Console or API. Audio and transcriptions stored via the Async API are automatically deleted after 30 days. *** ## Logging * Minimal logging is performed for service reliability, debugging, and billing. * Logs **never** contain raw audio or transcript content. * Diagnostic metadata (such as request IDs or error traces) may be retained temporarily for operational purposes. * Per-request usage metadata (model, audio duration, tokens, optional `client_reference_id`) is recorded as [usage logs](/guides/usage-logs) for debugging. *** ## Encryption & access control * **Encryption** – All data is encrypted in transit using TLS 1.2+ and at rest using industry-standard encryption. * **Access control** – Stored data is logically isolated within your Soniox account and accessible only using your authorized API keys. # Errors URL: /api-reference/errors Reference for every error type returned by Soniox APIs (REST, async, STT WebSocket, TTS WebSocket, TTS REST), with the cause and how to resolve each one. import { Callout } from "@/components/callout"; import { Badge } from "@openapi/ui/components/method-label"; ## Overview This page documents errors from every Soniox API surface: * **REST API** at `api.soniox.com`: synchronous endpoints for files, transcriptions, speakers, usage logs, temporary API keys, and async transcription submission. * **Async transcription failure**: a job created via `POST /v1/transcriptions` returned `200`, but the transcription itself ends in `status: "failed"` after the audio is downloaded or transcribed. * **STT WebSocket** at `stt-rt.soniox.com`: real-time speech-to-text. * **TTS WebSocket** at `tts-rt.soniox.com/tts-websocket`: real-time text-to-speech. * **TTS REST** at `tts-rt.soniox.com/tts`: single-request text-to-speech. All surfaces share one `error_type` taxonomy. Branch your client code on `error_type`, not on the human-readable message; the wording may change, the slug will not. Every error response carries: * `error_type`: stable, machine-readable identifier. See [API error types](#api-error-types) for the full list. * `request_id`: unique per request. Include it when contacting [support@soniox.com](mailto:support@soniox.com); server logs are keyed on it. * `more_info`: a URL pointing at the section on this page describing the `error_type`. The pattern is `https://soniox.com/docs/api-reference/errors#` with underscores replaced by hyphens. The HTTP status code mapped to each `error_type` is the same on every surface. *** ## Response structure ### REST API response REST API errors return a JSON body with the HTTP status code: ```json { "status_code": 401, "error_type": "unauthenticated", "message": "Incorrect API key provided. You can get an API key at https://console.soniox.com", "validation_errors": [], "request_id": "3d37a3bd-5078-47ee-a369-b204e3bbedda", "more_info": "https://soniox.com/docs/api-reference/errors#unauthenticated" } ``` HTTP status code, mirrored in the response body for convenience. Stable, machine-readable identifier of the error. Branch on this, not on `message`. See the [error reference](#api-error-types) below for the full list. Human-readable description of the error. Safe to show to operators or log; do not parse or pattern-match against it. Populated when `error_type` is `invalid_request` and the cause is a per-field validation failure. Each entry has `error_type`, `location` (dotted path of the offending field), and `message`. Unique identifier of this request. Server logs are keyed on it, so include it when contacting [support@soniox.com](mailto:support@soniox.com). Link to the section on this page describing the `error_type`. The URL pattern is `https://soniox.com/docs/api-reference/errors#` with underscores replaced by hyphens. ### Async transcription failure When you create a transcription with `POST /v1/transcriptions`, the call returns `200` and a transcription object as soon as the job is accepted. The work itself happens asynchronously: the audio is downloaded (if a URL was supplied), then transcribed. If the **job** fails, the HTTP layer doesn't see the failure. Instead, `GET /v1/transcriptions/{id}` returns `status: "failed"` with `error_type` and `error_message` fields populated. Always check `status` and read these fields when polling a transcription. ```json { "id": "84c32fc6-4fb5-4e7a-b656-b5ec70493753", "status": "failed", "error_type": "file_download_failed", "error_message": "Failed to download audio from the provided URL.", "...": "..." } ``` Failed async jobs are **not auto-retried**. Fix the cause and submit a new transcription. The `error_type` values that can appear on a failed job are flagged below with **Async** on their **Surfaces** line. ### Voice processing failure When you create a voice with `POST /v1/voices`, the call returns `201` as soon as the reference clip is accepted. The voice is then processed for each Text-to-Speech model asynchronously, so a per-model failure is not visible to the HTTP layer. Instead, `GET /v1/voices/{id}` reports a per-model status in the `models` array. A model the voice failed to process for has `status: "failed"` with `error_type` and `error_message` populated: ```json { "id": "21b9c8e2-1c3a-4d5e-9f8a-123456789abc", "name": "narrator", "models": [ { "model": "tts-rt-v2", "status": "failed", "error_type": "voice_audio_too_long", "error_message": "The reference audio clip is too long for this model. Upload a shorter clip." } ] } ``` A failure is **terminal and not auto-retried**. Fix the reference clip and create a new voice. The `error_type` values that can appear here are flagged below with **Voice** on their **Surfaces** line. Using a voice for a model whose status is not `ready` returns an HTTP error at stream start (see [`voice_not_ready`](#voice-not-ready), [`voice_not_prepared`](#voice-not-prepared), and [`voice_failed`](#voice-failed)). ### STT WebSocket error frame On error, the STT WebSocket sends a single JSON text frame and then closes the connection with WebSocket close code `1000` (Normal Closure). Branch on `error_type` inside the frame, not on the close code. ```json { "tokens": [], "error_code": 401, "error_type": "unauthenticated", "error_message": "Incorrect API key provided. You can get an API key at https://console.soniox.com", "more_info": "https://soniox.com/docs/api-reference/errors#unauthenticated", "request_id": "3d37a3bd-5078-47ee-a369-b204e3bbedda" } ``` The error frame reuses the shape of a normal response frame (hence the empty `tokens` array), with the error fields populated. `error_code` is the HTTP status as an integer, and the human-readable text lives in `error_message` rather than `message`. The frame does not carry a `validation_errors` array; a single failure is described in `error_message`. ### TTS WebSocket error frame The TTS WebSocket multiplexes multiple streams over a single connection. Every error frame carries a `stream_id`: ```json { "stream_id": "stream-1", "error_code": 400, "error_type": "invalid_stream_state", "error_message": "Stream stream-1 is already active on this connection. Choose a different stream_id, or cancel the existing stream first.", "more_info": "https://soniox.com/docs/api-reference/errors#invalid-stream-state", "request_id": "3d37a3bd-5078-47ee-a369-b204e3bbedda" } ``` **Per-stream errors do not close the connection.** Other streams on the same connection continue running and the client may keep using the WebSocket. A `{"terminated": true, "stream_id": "..."}` frame typically follows for the affected stream. Connection-level errors (malformed top-level message, server shutting down) carry `stream_id: ""` and close the connection. ### TTS REST response TTS REST has two error paths depending on whether response headers have been sent. **Headers not yet sent.** Validation and authentication errors, as well as errors raised before audio streaming starts, return a JSON body with the HTTP status code: ```http HTTP/1.1 401 Unauthorized Content-Type: application/json { "error_code": 401, "error_type": "unauthenticated", "error_message": "Missing API key. Provide it as an Authorization header (e.g. 'Authorization: Bearer '). You can get an API key at https://console.soniox.com.", "more_info": "https://soniox.com/docs/api-reference/errors#unauthenticated", "request_id": "3d37a3bd-5078-47ee-a369-b204e3bbedda" } ``` **Headers already sent (mid-stream failure).** The initial response was `200 OK` and audio bytes have already been flushed to the client. The connection closes abruptly with no in-band error payload. Clients should treat a truncated audio stream as a server-side failure. The exception is [`max_audio_duration_reached`](#max-audio-duration-reached), where the response ends cleanly with truncated audio. *** ## API error types Each section below documents one `error_type`. The badge shows the HTTP status code (the same on every surface), and the **Surfaces** line lists every API where this `error_type` can appear. ### `invalid_request` HTTP 400 **Surfaces:** [REST](#rest-api-response) · [STT WebSocket](#stt-websocket-error-frame) · [TTS WebSocket](#tts-websocket-error-frame) · [TTS REST](#tts-rest-response) **Cause.** The request body or query parameters failed validation. Common triggers: * **REST:** a required field is missing, a value is out of range, an enum value is unknown, a multipart upload is malformed or has unknown size, an uploaded file exceeds the per-file byte limit, or a request-level invariant is violated (for example, `end_time` not strictly after `start_time` on `GET /v1/usage-logs`). * **WebSocket / TTS REST:** a required start-request or per-stream field is missing or too long, a JSON message is malformed, an audio frame is not valid base64, a translation parameter combination is invalid, `speed` is outside the supported range (between `0.7` and `1.3`), or `reduce_silence` is enabled on a model that does not support it. When the cause on REST is a per-field validation failure, the response includes `validation_errors` with the exact `location` and reason for each field: ```json { "status_code": 400, "error_type": "invalid_request", "message": "Invalid request.", "validation_errors": [ { "error_type": "missing", "location": "file.file", "message": "Field required" } ], "request_id": "..." } ``` WebSocket and TTS REST responses do not carry `validation_errors`; the single failure is described in `error_message`. **Solution.** Read `validation_errors` (REST) or `error_message` (streaming) and resend the request with corrected values. ### `invalid_cursor` HTTP 400 **Surfaces:** [REST](#rest-api-response) **Cause.** The `cursor` query parameter on a paginated endpoint (`/v1/files`, `/v1/transcriptions`, `/v1/usage-logs`) could not be decoded. The cursor is opaque and only valid when copied verbatim from the previous page's `next_page_cursor`; any modification, truncation, or stale cursor triggers this error. **Solution.** Omit `cursor` to start from the first page, or pass the exact `next_page_cursor` from the previous response without modification. Keep all other query parameters constant across pages of the same listing. ### `model_not_available` HTTP 400 **Surfaces:** [REST](#rest-api-response) · [Async](#async-transcription-failure) · [STT WebSocket](#stt-websocket-error-frame) · [TTS WebSocket](#tts-websocket-error-frame) · [TTS REST](#tts-rest-response) **Cause.** The requested `model` is unknown, has been retired, or is not enabled for the calling project or region. **Solution.** Use a model from the [STT models catalog](/stt/models) or [TTS models catalog](/tts/models). If you've been pinning a specific model, switch to a current one before its retirement date. ### `max_concurrent_streams_reached` HTTP 400 **Surfaces:** [TTS WebSocket](#tts-websocket-error-frame) **Cause.** The TTS WebSocket connection has reached its server-imposed cap on the number of simultaneous streams it can run. This is a per-connection limit and is **not** the per-organization or per-project concurrency cap (which surfaces as [`limit_exceeded`](#limit-exceeded)). **Solution.** Send a `cancel` message for one of the active streams on this connection to free a slot, or open a new WebSocket connection. ### `invalid_stream_state` HTTP 400 **Surfaces:** [TTS WebSocket](#tts-websocket-error-frame) **Cause.** A request was issued against a stream that is in the wrong state for that operation. Concrete cases: * `start` with a `stream_id` that is already active on the connection. * `text` or `cancel` for a `stream_id` that does not exist. * `text` or `text_end` after the stream already received `text_end`. * `cancel` after the stream was already cancelled. **Solution.** Read `error_message` for the specific state issue. Most cases indicate a client-side bug; verify the stream lifecycle follows `start` → `text…` → `text_end` (or `cancel` from any state). ### `invalid_audio_file` HTTP 400 **Surfaces:** [REST](#rest-api-response) (`POST /v1/files` upload) · [Async](#async-transcription-failure) **Cause.** The audio could not be decoded, contains no audio stream, or exceeds the model's maximum audio duration. For async jobs this is set during the download / convert step or during the transcribe step's duration recheck. **Solution.** Try transcribing the file directly in the [Soniox Console](https://console.soniox.com) to confirm whether the file itself is the problem. If it is, re-encode to a supported container/codec or split long audio into shorter segments. See the [supported audio formats](/stt/async/async-transcription#audio-formats). ### `unauthenticated` HTTP 401 **Surfaces:** [REST](#rest-api-response) · [STT WebSocket](#stt-websocket-error-frame) · [TTS WebSocket](#tts-websocket-error-frame) · [TTS REST](#tts-rest-response) **Cause.** The request did not authenticate. The `Authorization` header was missing, or the API key was malformed, revoked, expired, scoped to a different environment, or longer than 256 characters. Temporary API keys can also fail here when they are expired, single-use and already consumed, or used for an action they were not issued for. **Solution.** Pass a valid key as `Authorization: Bearer ` on REST and TTS REST, or in the `api_key` field of the start message on the WebSocket APIs. Get or rotate API keys in the [Soniox Console](https://console.soniox.com). For client apps, use a [temporary API key](/guides/temporary-api-keys). ### `organization_balance_exhausted` HTTP 402 **Surfaces:** [REST](#rest-api-response) · [Async](#async-transcription-failure) · [STT WebSocket](#stt-websocket-error-frame) · [TTS WebSocket](#tts-websocket-error-frame) · [TTS REST](#tts-rest-response) **Cause.** The available balance has dropped to zero. Raised before the request runs, and also evaluated when an async job is about to start downloading or transcribing. In the async case the job moves to `failed` with this `error_type`. **Solution.** Top up at [Billing overview](https://console.soniox.com/org/billing/overview/) or enable autopay. Failed async jobs are not auto-retried; resubmit after funding. ### `organization_monthly_budget_exhausted` HTTP 402 **Surfaces:** [REST](#rest-api-response) · [Async](#async-transcription-failure) · [STT WebSocket](#stt-websocket-error-frame) · [TTS WebSocket](#tts-websocket-error-frame) · [TTS REST](#tts-rest-response) **Cause.** The organization has hit its configured monthly budget cap (`monthly_budget_usd`). Resets at the start of the next calendar month. **Solution.** Raise the cap on the [organization limits](https://console.soniox.com/org/limits/) page, or wait for the month to roll over. ### `project_monthly_budget_exhausted` HTTP 402 **Surfaces:** [REST](#rest-api-response) · [Async](#async-transcription-failure) · [STT WebSocket](#stt-websocket-error-frame) · [TTS WebSocket](#tts-websocket-error-frame) · [TTS REST](#tts-rest-response) **Cause.** The project has hit its configured monthly budget cap. **Solution.** Raise the cap on the [project limits](https://console.soniox.com/org/projects/limits/) page, or wait for the month to roll over. ### `file_download_blocked` **Surfaces:** [Async](#async-transcription-failure) (failed transcription, no HTTP error) **Cause.** The supplied `audio_url` could not be fetched for safety reasons: the URL resolved to a private or otherwise blocked IP (SSRF guard). **Solution.** Use a publicly reachable URL whose host resolves to a routable IP. If the project or organization was deleted, recreate it (or use a different one) and resubmit. ### `file_download_failed` **Surfaces:** [Async](#async-transcription-failure) (failed transcription, no HTTP error) **Cause.** Fetching `audio_url` failed: HTTP error (4xx / 5xx), DNS failure, TLS failure, connection timeout, or the server closed the connection mid-stream. 4xx responses are not retried; other transient failures are retried internally up to a bound before the job is marked as failed. **Solution.** Make the URL reachable and stable. For files you control, prefer uploading via `POST /v1/files` rather than relying on a remote URL. ### `transcription_output_too_long` **Surfaces:** [Async](#async-transcription-failure) (failed transcription, no HTTP error) **Cause.** The transcription completed but produced more output than the model is allowed to return in a single job. Not expected to occur in normal use. If you see this `error_type`, please contact [support@soniox.com](mailto:support@soniox.com) with the transcription ID. ### `file_not_found` HTTP 404 **Surfaces:** [REST](#rest-api-response) · [Async](#async-transcription-failure) **Cause.** No file with the given ID exists in the project the API key belongs to. Returned for IDs that never existed, IDs that were deleted, and IDs that exist but belong to a different project. Soniox does not leak cross-project existence. On async jobs it means the file referenced by `file_id` was deleted after the transcription was created but before it started processing. **Solution.** Confirm the ID by listing files via `GET /v1/files`, and make sure the API key is scoped to the project that owns the file. Keep an uploaded file until the transcription that uses it reaches a terminal state (`completed` or `error`), then delete it. ### `transcription_not_found` HTTP 404 **Surfaces:** [REST](#rest-api-response) **Cause.** No transcription with the given ID exists in the caller's project. Same semantics as [`file_not_found`](#file-not-found): deleted, never existed, or belongs to a different project. **Solution.** Confirm the ID with `GET /v1/transcriptions`, and check that the API key is scoped to the project that owns the transcription. ### `voice_not_found` HTTP 404 **Surfaces:** [REST](#rest-api-response) (`GET`/`DELETE /v1/voices/{id}`, `POST /v1/voices/{id}/recompute`) · [TTS WebSocket](#tts-websocket-error-frame) · [TTS REST](#tts-rest-response) **Cause.** No voice with the given ID exists in the calling project. Returned for IDs that never existed, were deleted, or belong to a different project; Soniox does not leak cross-project existence. On the TTS APIs this is raised at stream start when the `voice` field is a UUID that does not resolve to a voice in the project. **Solution.** Confirm the ID with `GET /v1/voices`, and make sure the API key is scoped to the project that owns the voice. To use a built-in voice instead, pass its name (for example `Adrian`); see [voices](/tts/concepts/voices). ### `voice_name_conflict` HTTP 409 **Surfaces:** [REST](#rest-api-response) (`POST /v1/voices`) **Cause.** A voice with the same `name` already exists in the project. Voice names are unique per project. **Solution.** Choose a different name, or delete the existing voice with `DELETE /v1/voices/{id}` first. ### `voice_not_prepared` HTTP 409 **Surfaces:** [TTS WebSocket](#tts-websocket-error-frame) · [TTS REST](#tts-rest-response) **Cause.** The voice exists but has not been prepared for the requested `model`, so it cannot be used with it yet. This typically happens when a model is released after the voice was created: the voice reports `status: "not_computed"` for that model. **Solution.** Call `POST /v1/voices/{id}/recompute` (optionally with a `model`) to prepare the voice for the model, poll `GET /v1/voices/{id}` until that model's status is `ready`, then retry the request. ### `voice_not_ready` HTTP 503 **Surfaces:** [TTS WebSocket](#tts-websocket-error-frame) · [TTS REST](#tts-rest-response) **Cause.** The voice is still being processed for the requested model (per-model status `processing`). This is transient: processing starts right after the voice is created or recomputed and usually finishes within seconds. **Solution.** Retry after a short delay, or poll `GET /v1/voices/{id}` until the model's status is `ready` before starting the stream. ### `voice_failed` HTTP 503 **Surfaces:** [TTS WebSocket](#tts-websocket-error-frame) · [TTS REST](#tts-rest-response) **Cause.** Processing the voice failed permanently for the requested model, so the voice cannot be used with it. Despite the 503 status this is **terminal, not retryable**. Inspect the voice's `models[].error_type` and `error_message` via `GET /v1/voices/{id}` for the underlying reason (for example [`voice_audio_too_long`](#voice-audio-too-long) or [`voice_invalid_audio`](#voice-invalid-audio)). **Solution.** Create a new voice from a valid reference clip. Recompute does not recover a voice that has terminally failed. ### `voice_audio_too_long` **Surfaces:** [Voice](#voice-processing-failure) (failed voice processing, no HTTP error) **Cause.** The reference audio clip is longer than the model allows (currently a maximum of 2 minutes). Clip duration is not checked at upload time, so this surfaces as a per-model processing failure rather than an upload error. **Solution.** Create a new voice from a shorter clip (at most 2 minutes). ### `voice_invalid_audio` **Surfaces:** [Voice](#voice-processing-failure) (failed voice processing, no HTTP error) **Cause.** The reference audio clip could not be processed. It may be in an unsupported format, corrupted, or contain no audio. **Solution.** Re-encode the clip to a supported audio format and create a new voice. ### `temp_api_key_session_expired` HTTP 403 **Surfaces:** [STT WebSocket](#stt-websocket-error-frame) · [TTS WebSocket](#tts-websocket-error-frame) · [TTS REST](#tts-rest-response) **Cause.** The temporary API key in use was created with a `max_session_duration_seconds` cap, and that duration has elapsed for the current session. The temporary key itself may still be valid for new sessions (subject to its other limits); only the current session has been cut off. **Solution.** Create a new temporary API key (or use the same one if it has remaining uses) and start a new session. Long-running clients should refresh before the per-session cap rather than after. ### `request_timeout` HTTP 408 **Surfaces:** [STT WebSocket](#stt-websocket-error-frame) · [TTS WebSocket](#tts-websocket-error-frame) · [TTS REST](#tts-rest-response) **Cause.** A deadline was exceeded before the request could complete. Most common triggers: * **STT WebSocket: client too slow on audio.** The client failed to send audio chunks fast enough. For example, the initial chunk never arrived, or the decoder starved mid-stream. * **TTS WebSocket: client too slow on text.** After `start`, the client did not send any text within the allowed window, or stopped sending intermediate text chunks and never sent `text_end` to close the stream. The server cannot keep an idle stream open indefinitely. **Solution.** Retry the request. For STT, ensure audio is streamed in real time or faster and that the first chunk is sent promptly after opening the connection. For TTS, send text chunks as soon as you have them and send `text_end` when the input is complete; if you have no more text to send but want to keep the stream alive, cancel it and start a new one when ready. If the timeout persists across retries, contact [support@soniox.com](mailto:support@soniox.com) with the `request_id`. ### `transcription_invalid_state` HTTP 409 **Surfaces:** [REST](#rest-api-response) The transcription exists but is in a state that doesn't allow the requested operation. Read `message` for the specific reason. #### Not completed yet `GET /v1/transcriptions/{id}/transcript` while `status` is `queued`, `downloading`, or `transcribing`. **Solution.** Poll `GET /v1/transcriptions/{id}` until `status` is `completed`, or configure a webhook on the transcription to be notified. #### Failed The transcript was requested for a transcription that ended in `status: "failed"`. **Solution.** Read `error_type` and `error_message` from `GET /v1/transcriptions/{id}` to see why, and submit a new transcription after fixing the cause. #### No longer available Transcript data was retained for a limited window after completion and has since been purged. **Solution.** Re-submit the original audio as a new transcription. #### Cannot delete while processing `DELETE /v1/transcriptions/{id}` while `status` is `transcribing`. **Solution.** Wait for `status` to reach `completed` or `failed`, then retry the delete. ### `max_duration_reached` HTTP 413 **Surfaces:** [STT WebSocket](#stt-websocket-error-frame) **Cause.** The real-time connection reached the maximum allowed session duration (a fixed server-side cap) and was closed. This is a hard ceiling on how long a single WebSocket session may run, independent of any per-temporary-API-key cap (which surfaces as [`temp_api_key_session_expired`](#temp-api-key-session-expired)). Reconnecting on the same connection cannot extend it. **Solution.** Open a new WebSocket connection and resume streaming. For long-running capture, roll over to a fresh connection before the cap is reached rather than waiting for the server to close the session. ### `max_audio_duration_reached` HTTP 413 **Surfaces:** [TTS WebSocket](#tts-websocket-error-frame) **Cause.** Generated audio is capped at 2 minutes per request. When the cap is reached, generation stops and the audio output is truncated. This limit is fixed and cannot be increased. It is distinct from [`max_duration_reached`](#max-duration-reached), which caps the session duration of an STT WebSocket connection. On the **TTS WebSocket**, the final audio chunk is delivered without `audio_end`, followed by an error response and a `terminated` message for the stream: ```json { "error_code": 413, "error_type": "max_audio_duration_reached", "error_message": "The generated audio reached the maximum audio duration per request and the output was truncated. Synthesize the remaining text with a new request." } ``` The WebSocket connection stays open; other streams continue, and you can generate the remaining text on a new stream. On **TTS REST**, audio is streamed, so by the time the cap is reached the `200 OK` response is already committed and no error can be reported. The response simply ends with the truncated audio. **Solution.** Split long text into multiple requests, each below the [audio duration cap](/tts/rest-api/limits-and-quotas). On the WebSocket, start a new stream for the remaining text. ### `limit_exceeded` HTTP 429 **Surfaces:** [REST](#rest-api-response) · [STT WebSocket](#stt-websocket-error-frame) · [TTS WebSocket](#tts-websocket-error-frame) · [TTS REST](#tts-rest-response) A single `error_type` covers many distinct limits. Read `message` (REST) or `error_message` (streaming) to find out which one was hit. The sub-causes below are grouped by what kind of limit triggered the response. All current limits and your usage are visible in the [Soniox Console](https://console.soniox.com). Per-product reference: [Async STT](/stt/async/limits-and-quotas), [Real-time STT](/stt/rt/limits-and-quotas), [REST TTS](/tts/rest-api/limits-and-quotas), [Real-time TTS](/tts/rt/limits-and-quotas). #### Requests per minute (RPM) Too many calls of the same kind in a 60-second window, evaluated per organization or per project. RPM limits exist for: * WebSocket transcription * Async transcription * File management * Text-to-Speech * Temporary API key creation * Usage logs **Solution.** Slow down, batch where possible, or raise the limit in the console: [organization limits](https://console.soniox.com/org/limits/) or [project limits](https://console.soniox.com/org/projects/limits/). #### Concurrent requests Too many simultaneous in-flight requests on the organization or project. Applies to streaming workloads such as STT WebSocket, TTS WebSocket, and TTS REST. **Solution.** Reduce parallelism in the client, or request a higher concurrent-requests cap in the console. The TTS WebSocket also enforces a per-connection cap on simultaneous streams, which surfaces as the separate [`max_concurrent_streams_reached`](#max-concurrent-streams-reached) `error_type`. #### Total file count / total file size (GB) Adding this file would put the organization or project over its retained-storage cap (`files_total_count` or `files_total_size_gb`). **Solution.** Delete unused files via `DELETE /v1/files/{id}`, or raise the cap in the console. #### Pending file count Too many uploaded files are awaiting transcription (`transcribe_async_pending_num_files`). **Solution.** Wait for in-flight transcriptions to complete, then retry. #### Total transcription count / pending transcription count Analogous caps on the number of stored transcriptions (`transcribe_async_total_num_files`, `transcribe_async_pending_num_files`). **Solution.** Delete completed transcriptions with `DELETE /v1/transcriptions/{id}`, or raise the cap. ### `internal_error` HTTP 500 **Surfaces:** [REST](#rest-api-response) · [Async](#async-transcription-failure) · [STT WebSocket](#stt-websocket-error-frame) · [TTS WebSocket](#tts-websocket-error-frame) · [TTS REST](#tts-rest-response) **Cause.** Unhandled server-side exception, or an unrecoverable internal inconsistency. **Solution.** Retry once. If the error persists, contact [support@soniox.com](mailto:support@soniox.com) and include the `request_id` from the response. ### `service_unavailable` HTTP 503 **Surfaces:** [STT WebSocket](#stt-websocket-error-frame) · [TTS WebSocket](#tts-websocket-error-frame) · [TTS REST](#tts-rest-response) **Cause.** The service cannot accept the request right now. Concrete sub-causes include a backend model service being overloaded, encoder or decoder cache exhausted, the server gracefully shutting down, or no eligible model backend being available. The numeric `(code N)` in the message identifies the specific sub-cause for support triage. **Solution.** Retry with exponential backoff. The condition is transient; the retry will typically be routed to a healthy backend. If retries consistently fail, contact [support@soniox.com](mailto:support@soniox.com) with the `request_id` and the `(code N)` from the error message. *** ## Getting help If you can't determine the cause from `error_type` and `message`, or if an `internal_error` keeps recurring, contact [support@soniox.com](mailto:support@soniox.com) and include: * The `request_id` from the error response. * The `error_type` and `message` (or `error_message`). * The endpoint and HTTP method, or for async failures the transcription ID. * A short description of what you were trying to do. # Soniox API reference URL: /api-reference Soniox API reference for Speech-to-Text and Text-to-Speech. ## Base URLs API base URLs depend on two factors: the specific service you are calling (Speech-to-Text, Text-to-Speech, Authentication) and your project's region. For a complete list of regional domains (e.g., US, EU, Japan, India) and their corresponding REST and WebSocket endpoints, please refer to our [Data residency](/data-residency) guide. *** ## REST API **OpenAPI schema**: [https://soniox.com/docs/openapi.yaml](https://soniox.com/docs/openapi.yaml) ### Authentication Create and manage API keys for authentication. * **[Create temporary API key](/api-reference/auth/create_temporary_api_key)**: Create short-lived API keys. ### Speech-to-Text REST API for Speech-to-Text is available at `https://api.soniox.com/v1` and includes: * **[Files](/api-reference/stt/files/get_files)**: Manage audio files by uploading, listing, retrieving, and deleting them. * **[Transcriptions](/api-reference/stt/transcriptions/get_transcriptions)**: Create and manage transcriptions for audio files. * **[STT models](/api-reference/stt/get_models)**: List available STT models. See [Get started](/stt/get-started) for an introduction. ### Text-to-Speech REST API for Text-to-Speech is available at `https://tts-rt.soniox.com/tts`. * **[Generate TTS](/api-reference/tts/generate_tts)**: Generate speech from text. * **[TTS models](/api-reference/tts/get_tts_models)**: List available TTS models. Voice management is available at `https://api.soniox.com/v1`. * **[Voices](/api-reference/tts/voices/create_voice)**: Create custom voices from a reference clip and list, retrieve, recompute, and delete them. See [Get started](/tts/get-started) for an introduction and [Voice cloning](/tts/concepts/voice-cloning) for a guide. ### Other Additional REST endpoints are available at `https://api.soniox.com/v1`. * **[Concurrency limits](/api-reference/other/get_concurrency_limits)**: Retrieve current concurrent counts and configured limits for the project and organization. * **[Concurrent streams history](/api-reference/other/get_concurrent_streams_history)**: Retrieve historical concurrent stream counts for the project, aggregated per minute, hour, or day. * **[Usage logs](/api-reference/other/get_usage_logs)**: Retrieve per-request usage log entries for the project. * **[Usage summary](/api-reference/other/get_usage_summary)**: Retrieve daily cost and activity for the project, broken down per model. *** ## Errors Every Soniox API surface (REST, async transcription, STT WebSocket, TTS WebSocket, TTS REST) returns a stable `error_type` you can branch on in code, alongside a human-readable message and a `request_id` for support. See the [Errors reference](/api-reference/errors) for the full list of error types, their causes, and how to resolve them. *** ## WebSocket API ### Realtime Speech-to-Text Use WebSocket API to transcribe and translate live audio streams in real-time. * **Endpoint**: `wss://stt-rt.soniox.com` * **Reference**: [WebSocket API](/api-reference/stt/websocket-api) ### Realtime Text-to-Speech Use WebSocket API for low-latency, streaming speech generation. * **Endpoint**: `wss://tts-rt.soniox.com/tts-websocket` * **Reference**: [WebSocket API](/api-reference/tts/websocket-api) # Soniox Live URL: /demo-apps/soniox-live Demo apps showing how to add Soniox Speech-to-Text to your product import Image from "next/image"; ## Overview Soniox Live is a **demo app** that shows how to stream audio from your microphone directly to the Soniox [Real-time API](/api-reference/stt/websocket-api) for instant transcription and translation. This is not the [Soniox App](https://soniox.com/soniox-app) (our end-user product). Instead, it is a **reference implementation** for developers who want to learn how to embed Soniox into their own web or mobile applications. Web and mobile demo apps screenshot ## Features * Stream audio from your mic to Soniox in real time * Low-latency, high-accuracy transcription in 60+ languages * Low-latency speech translation to 60+ languages * Runs in the browser (React) and mobile (React Native) * Lightweight server issues temporary client keys for secure access ## Usage flow 1. Tap **Start** to begin streaming from your mic 2. **Live captions** appear word by word, then finalize 3. Toggle **Translation** and choose a target language for live translated captions 4. Tap **Stop** to end the session ## Architecture * **Server (Python):** Stores your secret Soniox API key and issues **temporary API keys** to clients * **Frontend (React & React Native):** Requests a temporary API key from your server, then streams microphone audio directly to Soniox servers for real-time transcription and translation We provide all the implementations with links to GitHub: * [Python server](https://github.com/soniox/soniox_examples/tree/master/apps/soniox-live-demo/server) * [React frontend](https://github.com/soniox/soniox_examples/tree/master/apps/soniox-live-demo/react) (web) * [React Native frontend](https://github.com/soniox/soniox_examples/tree/master/apps/soniox-live-demo/react-native) (mobile) # Speech-to-speech translation URL: /demo-apps/soniox-speech-to-speech-translation Demo app showing how to add real-time speech-to-speech translation with Soniox. ## Overview [Soniox Speech-to-speech Translation](https://github.com/soniox/soniox_examples/tree/master/apps/soniox-speech-to-speech-translation-demo) demo app shows how to combine the Soniox real-time speech-to-text with translation and real-time text-to-speech WebSocket APIs into a complete speech-to-speech translation pipeline - voice input in one language, hear the translation in another, in real time. This is a **reference implementation** for developers who want to learn how to wire the two real-time APIs together at the protocol level, without an SDK. The backend is a small FastAPI service; the frontend is a vanilla HTML/JS page. ## Features * Stream audio from an audio file or your mic to Soniox in real time * Live transcription in 60+ languages, with automatic source-language detection * Mid-sentence speech translation to 60+ languages - translation tokens stream as you talk * Live spoken translation through one of the [Soniox voices](/tts/concepts/voices), played back to you * Optional speaker diarization and language identification * Toggle for text-only translation mode that skips TTS entirely ## Usage flow 1. Pick a **target language** and a **voice** in the sidebar 2. Tap **Start talking** to begin streaming from your mic 3. The **original transcript** appears on the left; the **translation** appears on the right, word by word 4. The translated speech plays through your speakers in the chosen voice 5. Tap **Stop** to end the session ## Architecture * **Server (Python / FastAPI):** Holds your Soniox API key, accepts a WebSocket from the browser, and proxies audio and tokens between the browser and Soniox. Manages the per-utterance TTS stream lifecycle, pre-warming, and connection keepalive. * **Frontend (vanilla HTML / JS):** Captures audio file or microphone audio with `MediaRecorder`, streams the bytes to the backend over WebSocket, renders incoming token JSON into the transcript columns, and plays incoming PCM audio through the Web Audio API. Source code available in our [GitHub examples repo](https://github.com/soniox/soniox_examples/tree/master/apps/soniox-speech-to-speech-translation-demo). # Soniox Voice Agent URL: /demo-apps/soniox-voice-agent Demo app showing how to implement a voice-to-voice AI agent with Soniox voice solutions import Image from "next/image"; ## Overview Soniox Voice Agent is a **demo app** that shows how to build a complete voice-to-voice conversational AI assistant. It demonstrates how to integrate streaming speech-to-text, a large language model (LLM), and streaming text-to-speech (TTS) into a seamless, low-latency application. The demo bot is pre-configured as an appointment booking assistant for a fictional car repair shop, "Soniox AutoWorks." It can book appointments for services such as oil changes and car repairs, collect customer names and vehicle information, provide available appointment slots, and interactively guide users through the booking process. The entire voice bot codebase is designed for easy customization and extension to other domains. You can quickly adapt the bot to different business needs, integrate new tools, or change its persona, making it a flexible starting point for any conversational AI application. {/* Voice bot web interface screenshot */} ## Features * **End-to-end real-time**: Fully streaming architecture (voice-in, voice-out) for natural, low-latency conversations * **Multilingual**: Understands and responds to users in multiple languages * **Customizable AI**: The bot's persona and business logic are defined in a single, easy-to-edit file * **Extensible tools**: Connect the LLM to your own APIs and databases to perform real-world actions * **Multiple ways to interact**: Web frontend, Twilio phone call, or any other WebSocket-based connection ## Usage flow 1. Connect via **web browser** or **phone call** 2. Speak naturally in any language to the AI agent 3. The bot transcribes your speech, understands intent, and generates a response 4. Listen to the AI's spoken response in real-time 5. The conversation continues with full context awareness ## Architecture * **Server (Python):** Orchestrates the conversation with modular processors (VAD, STT, LLM, TTS) * **Frontend (React):** Captures microphone audio, streams it to the backend, and plays back responses * **Twilio proxy (Python):** Optional bridge to connect phone calls to the voice bot backend We provide all the implementations with links to GitHub: * [Python server](https://github.com/soniox/soniox_examples/tree/master/apps/soniox-voice-bot-demo/server) * [React frontend](https://github.com/soniox/soniox_examples/tree/master/apps/soniox-voice-bot-demo/frontend) (web) * [Twilio proxy](https://github.com/soniox/soniox_examples/tree/master/apps/soniox-voice-bot-demo/twilio) (phone integration) ## How it works The system is built on a modular, asynchronous architecture. When a user connects, a session is created to orchestrate the entire conversation, managing the flow of data between four core processors: #### Voice Activity Detection (VAD) Processor Uses Silero VAD to detect speech boundaries in incoming audio. Emits events to interrupt TTS when the user starts speaking. #### Speech-to-Text (STT) Processor The user's voice is captured by a client (web app or phone call) and streamed to the backend. The STT Processor uses the Soniox API to transcribe the audio into text in real-time. #### Language Model (LLM) Processor The transcribed text is sent to the LLM Processor. It maintains the conversation history, determines the user's intent, and decides whether to generate a direct response or use a predefined tool (like checking available slots). #### Text-to-Speech (TTS) Processor The LLM's final text response is sent to the TTS Processor, which uses the Soniox API to convert it back into audio and streams it to the user, completing the conversational turn. ## Next steps Use this [project](https://github.com/soniox/soniox_examples/tree/master/apps/soniox-voice-bot-demo) as a starting point to build your own voice assistant: * Customize the bot's persona and instructions in `server/tools.py` * Implement your own tools to connect to external APIs and databases * Adapt the frontend or Twilio integration to your needs # Concurrency limits URL: /guides/concurrency-limits Live counts of active real-time requests alongside the configured concurrency limits for the project and the organization that owns it, plus historical per-period concurrency. ## Overview Soniox applies concurrency limits to real-time APIs to ensure stability and fair use. Each project and organization has a limit on the number of [Speech-to-Text WebSocket](/api-reference/stt/websocket-api) sessions and [Text-to-Speech WebSocket](/api-reference/tts/websocket-api) streams it can have open at the same time. Active [`POST /tts`](/api-reference/tts/generate_tts) REST requests count toward the same `tts_concurrent` limit. To monitor live concurrency, open the project's **Usage > Activity** page in the [Soniox Console](https://console.soniox.com/) - the dashboard shows real-time charts of concurrent requests against the configured limit for both the project and the organization. The organization limit is checked first, then the project limit. A request is rejected as soon as one is at its cap, and the `429` response identifies which tier rejected it. See [rate and usage limits](/api-reference/errors#limit_exceeded) for the response shape. *** ## What gets returned The response has two top-level scopes, `project` and `organization`. Each scope contains: * `current` - live counts of active requests right now. * `limits` - the configured cap for that scope, or `null` when no cap is configured. When a value under `project.limits` is `null`, the project has no cap of its own and only the organization limit applies. Two services are tracked under each scope: * `transcribe_concurrent` - open Speech-to-Text WebSocket sessions. * `tts_concurrent` - open Text-to-Speech WebSocket streams and active Text-to-Speech REST requests. ```json { "project": { "current": { "transcribe_concurrent": 2, "tts_concurrent": 0 }, "limits": { "transcribe_concurrent": 4, "tts_concurrent": 1 } }, "organization": { "current": { "transcribe_concurrent": 5, "tts_concurrent": 1 }, "limits": { "transcribe_concurrent": 10, "tts_concurrent": 2 } } } ``` See [`GET /v1/concurrency-limits`](/api-reference/other/get_concurrency_limits) for the full schema and field types. *** ## Historical concurrency `GET /v1/concurrency-limits` reports only the counts right now. To see how concurrency moved over time - for example to size a limit increase against actual peaks - use [`GET /v1/concurrent-streams-history`](/api-reference/other/get_concurrent_streams_history). It returns per-period aggregates for the authenticated project, which is the same data behind the **Usage > Activity** charts in the Console. Pick one `kind` per request: * `stt` - Speech-to-Text WebSocket sessions. * `tts` - Text-to-Speech WebSocket streams and REST requests. Choose the granularity with `period_sec`. The longer the period, the longer the window allowed, and the longer the data is kept: | `period_sec` | Granularity | Maximum window | Retained for | | ------------ | ----------- | -------------- | ------------ | | `60` | Per minute | 7 days | 7 days | | `3600` | Hourly | 90 days | 90 days | | `86400` | Daily | 366 days | 2 years | `start_time` is inclusive and `end_time` is exclusive, both ISO 8601 UTC. ```http GET /v1/concurrent-streams-history?start_time=2026-04-28T09:00:00Z&end_time=2026-04-28T10:00:00Z&period_sec=60&kind=tts Authorization: Bearer ``` ```json { "kind": "tts", "entries": [ { "period_start": "2026-04-28T09:00:00Z", "period_sec": 60, "sample_min": 0, "sample_max": 4, "sample_sum": 21, "sample_count": 9, "total_count": 1 }, { "period_start": "2026-04-28T09:01:00Z", "period_sec": 60, "sample_min": 0, "sample_max": 0, "sample_sum": 0, "sample_count": 0, "total_count": 0 } ] } ``` Periods with no activity are still returned, with every field set to `0`. Use `sample_max` for the peak concurrency in a period and `sample_sum / sample_count` for the average. See [`GET /v1/concurrent-streams-history`](/api-reference/other/get_concurrent_streams_history) for the full schema and field types. *** You can request higher limits in the [Soniox Console](https://console.soniox.com/org/limits). # Direct stream URL: /guides/direct-stream Stream directly from microphone to Soniox Speech-to-Text WebSocket API to minimize latency. import Image from "next/image"; ## Overview Applies to [Text-to-Speech](/tts/get-started) and [Speech-to-Text](/stt/get-started). Examples use Speech-to-Text. This guide walks you through capturing and transcribing microphone audio in real time using the Soniox [WebSocket API](/api-reference/stt/websocket-api) — optimized for the lowest possible latency. The direct stream approach enables the browser to send audio directly to the Soniox WebSocket API over a WebSocket connection, eliminating the need for any intermediary server. This results in faster transcription and a simpler architecture. Soniox's [Web Library](/sdk/web-SDK) handles everything client-side — capturing microphone input, managing the WebSocket connection, and authenticating using temporary API keys. Use this setup when you want real-time speech-to-text performance directly in the browser **with minimal delay**. Soniox Speech-to-Text direct stream flowchart *** ## Temporary API keys [Temporary API keys](/guides/temporary-api-keys) (obtained from the [REST API](/api-reference/auth/create_temporary_api_key)) are required solely to establish the WebSocket connection. Once the connection is established, it will be kept alive as long it remains active. The `expires_in_seconds` configuration parameter should be set to a short duration. Following parameters are required to create a temporary API key: ```json { "usage_type": "transcribe_websocket", "expires_in_seconds": 60 } ``` To attribute browser-side traffic to an end user or session, bind a `client_reference_id` to the temporary API key — it is recorded in [usage logs](/guides/usage-logs) for every request authenticated with the key. Clients cannot override it. API request limits apply when creating temporary API keys. See **Limits** section in the [Soniox Console](https://console.soniox.com). *** ## Example This is an example of a browser-based transcription, but same principle applies to any other type of client - you minimize latency by connecting the client directly to the WebSocket API using a temporary API key. First we create a simple HTTP server that on request: 1. Renders the `index.html` template. 2. Exposes an endpoint to serve the temporary API key (`/temporary-api-key`). Python server using [FastAPI](https://fastapi.tiangolo.com/): ``` import os import requests import uvicorn from dotenv import load_dotenv from fastapi import FastAPI, Request from fastapi.responses import HTMLResponse, JSONResponse from fastapi.templating import Jinja2Templates load_dotenv() templates = Jinja2Templates(directory="templates") app = FastAPI() @app.get("/", response_class=HTMLResponse) async def get_index(request: Request): return templates.TemplateResponse( request=request, name="index.html", ) @app.get("/temporary-api-key", response_class=JSONResponse) async def get_temporary_api_key(): try: response = requests.post( "https://api.soniox.com/v1/auth/temporary-api-key", headers={ "Authorization": f"Bearer {os.getenv('SONIOX_API_KEY')}", "Content-Type": "application/json", }, json={ "usage_type": "transcribe_websocket", "expires_in_seconds": 60, }, ) if not response.ok: raise Exception(f"Error: {response.json()}") temporaryApiKeyData = response.json() return temporaryApiKeyData except Exception as error: print(error) return JSONResponse( status_code=500, content={"error": f"Server failed to obtain temporary api key: {error}"}, ) if __name__ == "__main__": port = int(os.getenv("PORT", 3001)) uvicorn.run(app, host="0.0.0.0", port=port) ``` [View example on GitHub](https://github.com/soniox/soniox_examples/tree/master/speech_to_text/python/real_time/browser_direct_stream) Node.js server using [Express](https://expressjs.com/): ``` require("dotenv").config(); const http = require("http"); const express = require("express"); const fetch = require("node-fetch"); const path = require("path"); const fs = require("fs").promises; const app = express(); app.use("/templates", express.static(path.join(__dirname, "templates"))); app.get("/", async (req, res) => { const index = await fs.readFile( path.join(__dirname, "templates/index.html"), "utf8" ); res.send(index); }); app.get("/temporary-api-key", async (req, res) => { try { const response = await fetch( "https://api.soniox.com/v1/auth/temporary-api-key", { method: "POST", headers: { Authorization: `Bearer ${process.env.SONIOX_API_KEY}`, "Content-Type": "application/json", }, body: JSON.stringify({ usage_type: "transcribe_websocket", expires_in_seconds: 60, }), } ); if (!response.ok) { throw await response.json(); } const temporaryApiKeyData = await response.json(); res.json(temporaryApiKeyData); } catch (error) { console.error(error); res.status(500).json({ error: `Server failed to obtain temporary api key: ${JSON.stringify(error)}`, }); } }); // Create HTTP server with Express const server = http.createServer(app); server.listen(process.env.PORT, () => { console.log( `HTTP server listening on http://0.0.0.0:${process.env.PORT}` ); }); ``` [View example on GitHub](https://github.com/soniox/soniox_examples/tree/master/speech_to_text/nodejs/real_time/browser_direct_stream/server.js) Our HTML client template contains a single "Start" button, that when clicked: 1. Requests microphone permissions. 2. Calls the `/temporary-api-key` endpoint to obtain a temporary API key. 3. Creates a new [`RecordTranscribe`](/sdk/web-SDK) class instance passing temporary api key as `apiKey` parameter. 4. Connects to the [WebSocket API](/api-reference/stt/websocket-api). 5. Starts transcribing from microphone input and renders transcribed text into a `div` in real-time. ```

Browser direct stream example


```
[View example on GitHub](https://github.com/soniox/soniox_examples/tree/master/speech_to_text/python/real_time/browser_direct_stream)
# Guides URL: /guides Practical Soniox guides for streaming architectures, temporary API keys, usage logs, and concurrency limits. import { LinkCards } from "@/components/link-card"; Soniox guides help you design production-ready Speech-to-Text and Text-to-Speech integrations around real-time streaming, secure client access, usage attribution, and operational limits. Use these guides when you need architectural guidance beyond the API reference: how to connect browsers to Soniox, when to introduce a proxy, how to issue short-lived credentials, and how to monitor usage and capacity. # Proxy stream URL: /guides/proxy-stream How to stream audio from a client app to Soniox Speech-to-Text WebSocket API through a proxy server. import Image from "next/image"; ## Overview Applies to [Text-to-Speech](/tts/get-started) and [Speech-to-Text](/stt/get-started). Examples use Speech-to-Text. This guide explains how to stream microphone audio from a client to the Soniox [WebSocket API](/api-reference/stt/websocket-api) through a proxy server. In this architecture, the client captures audio and sends it over WebSocket to a proxy server. The proxy server establishes a connection to the Soniox WebSocket API, authenticates the session, streams the audio for transcription, and relays the transcribed results back to the client in real time. This setup is useful when you want to **inspect, transform, or store audio and transcription data on the server side** before passing it to the client. If your goal is simply to transcribe audio and return results with the lowest possible latency, consider using the [direct stream](/guides/direct-stream) approach instead. Soniox STT stream with proxy flowchart ## Example In the following example, we create a proxy HTTP server that: 1. Listens for incoming WebSocket connections from the client. 2. Forwards audio data from the client to the [WebSocket API](/api-reference/stt/websocket-api). 3. Relays transcription results back to the client. Authentication with the [WebSocket API](/api-reference/stt/websocket-api) is handled by the proxy server using the `SONIOX_API_KEY`. Python server that will act as a proxy between our client and [WebSocket API](/api-reference/stt/websocket-api). ``` import os import json import asyncio from dotenv import load_dotenv import websockets load_dotenv() async def handle_client(websocket): print("Browser client connected") # create a message queue to store client messages received before # Soniox WebSocket API connection is ready, so we don't loose any message_queue = [] soniox_ws = None soniox_ws_ready = False async def init_soniox_connection(): nonlocal soniox_ws, soniox_ws_ready try: soniox_ws = await websockets.connect( "wss://stt-rt.soniox.com/transcribe-websocket" ) print("Connected to Soniox STT WebSocket API") # Send initial configuration message start_message = json.dumps( { "api_key": os.getenv("SONIOX_API_KEY"), "audio_format": "auto", "model": "stt-rt-v5", "language_hints": ["en"], } ) await soniox_ws.send(start_message) print("Sent start message to Soniox") # mark connection as ready soniox_ws_ready = True # process any queued messages while len(message_queue) > 0 and soniox_ws_ready: data = message_queue.pop(0) await forward_data(data) # receive messages from Soniox STT WebSocket API async for message in soniox_ws: try: await websocket.send(message) except Exception as e: print(f"Error forwarding Soniox response: {e}") break except Exception as e: print(f"Soniox WebSocket error: {e}") soniox_ws_ready = False finally: if soniox_ws: await soniox_ws.close() soniox_ws_ready = False print("Soniox WebSocket closed") async def forward_data(data): try: if soniox_ws: await soniox_ws.send(data) except Exception as e: print(f"Error forwarding data to Soniox: {e}") # initialize Soniox connection soniox_task = asyncio.create_task(init_soniox_connection()) try: # receive messages from browser client async for data in websocket: if soniox_ws_ready: # forward messages instantly await forward_data(data) else: # queue the message to be processed # as soon as connection to Soniox STT WebSocket API is ready message_queue.append(data) except Exception as e: print(f"Error with browser client: {e}") finally: print("Browser client disconnected") soniox_task.cancel() try: await soniox_task except asyncio.CancelledError: pass async def main(): port = int(os.getenv("PORT", 3001)) server = await websockets.serve(handle_client, "0.0.0.0", port) print(f"WebSocket proxy server listening on http://0.0.0.0:{port}") await server.wait_closed() if __name__ == "__main__": asyncio.run(main()) ``` [View example on GitHub](https://github.com/soniox/soniox_examples/tree/master/speech_to_text/python/real_time/browser_proxy_stream) Node.js server that will act as a proxy between our client and [WebSocket API](/api-reference/stt/websocket-api). ``` require("dotenv").config(); const WebSocket = require("ws"); const http = require("http"); const server = http.createServer(); const wss = new WebSocket.Server({ server }); wss.on("connection", (ws) => { console.log("Browser client connected"); // create a message queue to store client messages received before // Soniox WebSocket API connection is ready, so we don't loose any const messageQueue = []; let sonioxWs = null; let sonioxWsReady = false; function initSonioxConnection() { sonioxWs = new WebSocket("wss://stt-rt.soniox.com/transcribe-websocket"); sonioxWs.on("open", () => { console.log("Connected to Soniox STT WebSocket API"); // send initial configuration message const startMessage = JSON.stringify({ api_key: process.env.SONIOX_API_KEY, audio_format: "auto", model: "stt-rt-v5", language_hints: ["en"], }); sonioxWs.send(startMessage); console.log("Sent start message to Soniox"); // mark connection as ready sonioxWsReady = true; // process any queued messages while (messageQueue.length > 0 && sonioxWsReady) { const data = messageQueue.shift(); forwardData(data); } }); // receive messages from Soniox STT WebSocket API sonioxWs.on("message", (data) => { // note: // at this point we could manipulate and enhance the transcribed data try { ws.send(data.toString()); } catch (err) { console.error("Error forwarding Soniox response:", err); } }); sonioxWs.on("error", (error) => { console.log("Soniox WebSocket error:", error); sonioxWsReady = false; }); sonioxWs.on("close", (code, reason) => { console.log("Soniox WebSocket closed:", code, reason); sonioxWsReady = false; ws.close(); }); } // forward message data to Soniox STT WebSocket API function forwardData(data) { try { sonioxWs.send(data); } catch (err) { console.error("Error forwarding data to Soniox:", err); } } // initialize Soniox connection initSonioxConnection(); // receive messages from browser client ws.on("message", (data) => { if (sonioxWsReady) { // forward messages instantly forwardData(data); } else { // queue the message to be processed // as soon as connection to Soniox STT WebSocket API is ready messageQueue.push(data); } }); ws.on("close", () => { console.log("Browser client disconnected"); if (sonioxWs) { try { sonioxWs.close(); } catch (err) { console.error("Error closing Soniox connection:", err); } } }); }); server.listen(process.env.PORT, () => { console.log( `WebSocket proxy server listening on http://0.0.0.0:${process.env.PORT}`, ); }); ``` [View example on GitHub](https://github.com/soniox/soniox_examples/tree/master/speech_to_text/nodejs/real_time/browser_proxy_stream) Next, we create a basic HTML page as the client (same concept works for any other app framework). The HTML client: 1. Connects to the proxy server via WebSocket. 2. Captures audio stream from the microphone through the [`MediaRecorder`](https://developer.mozilla.org/en-US/docs/Web/API/MediaRecorder). 3. Streams audio data to the proxy server. 4. Receives messages from the proxy server and renders transcribed text into a `div`. ```

Browser proxy stream example


```
[View example on GitHub](https://github.com/soniox/soniox_examples/tree/master/speech_to_text/python/real_time/browser_proxy_stream)
# Temporary API keys URL: /guides/temporary-api-keys Short-lived credentials that let untrusted clients connect directly to Soniox without exposing your long-lived API key. ## Overview A long-lived Soniox API key authenticates against your account and bills usage to it. If it leaks (embedded in a client app, captured in transit, or shared by mistake), anyone can spend your credits until you rotate the key. Temporary API keys are short-lived credentials your backend mints from a long-lived key and hands to a client. They are the only safe way for an untrusted client to connect directly to Soniox. Usage is still billed to the issuing account, so each key carries restrictions that limit what a leaked or replayed key can do: * `usage_type`: which Soniox service the key is valid for (always required). * `expires_in_seconds`: how long the key can be used to open new streams (always required). * `single_use`: the key may open at most one stream. * `max_session_duration_seconds`: the maximum duration of any one stream opened with the key. * `client_reference_id`: tracking identifier bound to the key; appears in [usage logs](/guides/usage-logs) for every request authenticated with it. The flow is always the same: 1. Client requests a temporary key from an endpoint on your backend. 2. Backend uses its long-lived API key to call Soniox and returns the temporary key to the client. 3. Client opens a stream directly to Soniox using the temporary key. See the [direct stream guide](/guides/direct-stream) for a complete client/server example. *** ## Creating a temporary API key Run this code on a **backend server you control**, never in the client. The long-lived API key must never leave your server. The minimal request needs only `usage_type` and `expires_in_seconds`. See the [API reference](/api-reference/auth/create_temporary_api_key) for the full request shape. ```python # Server-side endpoint your client calls @app.post("/temporary-api-key") def create_temporary_api_key(): response = requests.post( "https://api.soniox.com/v1/auth/temporary-api-key", headers={ "Authorization": f"Bearer {SONIOX_API_KEY}", "Content-Type": "application/json", }, json={ "usage_type": "transcribe_websocket", "expires_in_seconds": 60, }, ) response.raise_for_status() return response.json() ``` Our SDKs provide helpers for the same call: * [Node SDK](/sdk/node-SDK/reference/classes#sonioxauthapi-createtemporarykey): `client.auth.createTemporaryKey(...)` * [Python SDK](/sdk/python-SDK/stt/realtime-transcription): `client.auth.create_temporary_api_key(...)` *** ## Usage type `usage_type` is required and locks the key to exactly one Soniox service: * `transcribe_websocket`: for the [Speech-to-Text WebSocket API](/api-reference/stt/websocket-api). * `tts_rt`: for [Text-to-Speech](/api-reference/tts/websocket-api), both the WebSocket API and the [REST endpoint](/api-reference/tts/generate_tts). A key issued for one service cannot be used against the other. Issue a separate key for each service the client needs. *** ## Expiration `expires_in_seconds` controls **how long the key can be used to open new streams**. It does not terminate streams that are already open. Use `max_session_duration_seconds` (below) to bound an individual session. *** ## Single use When `single_use` is `true`, the temporary API key may be used to open a stream exactly once. Any later attempt to open a stream with the same key is rejected, even if the key has not yet expired. Use this whenever the client is expected to perform a single transcription or speech generation session. It ensures that a key intercepted in transit cannot be reused to start additional billable sessions. ```json { "usage_type": "transcribe_websocket", "expires_in_seconds": 60, "single_use": true } ``` *** ## Max session duration `max_session_duration_seconds` sets the maximum duration of a single stream opened with the temporary API key. The timer starts when the stream is opened (not when the underlying connection is established). Omitting the field disables the limit. The limit applies **per stream**. On a [multi-stream connection](/tts/rt/streams), each stream is timed independently using the value carried by the temporary API key. ```json { "usage_type": "tts_rt", "expires_in_seconds": 300, "max_session_duration_seconds": 60 } ``` Set `max_session_duration_seconds` based on the longest legitimate session you expect. A value that is too low will cut off real users; a value that is too high reduces protection if the key is leaked. For text-to-speech, the limit caps how long the stream stays open, not the duration of generated audio. In practice, TTS streams audio at roughly real-time pace, so a stream-duration cap also bounds how much audio a single session can produce. When the limit is reached, the stream is terminated with HTTP status `403` and the message `Temporary API key session duration limit exceeded.` How that error is delivered depends on the endpoint. ### Speech-to-Text WebSocket [`/transcribe-websocket`](/api-reference/stt/websocket-api) carries one stream per connection. When the session duration limit is reached, the server sends a final JSON response and closes the WebSocket with `StatusNormalClosure`: ```json { "error_code": 403, "error_type": "temp_api_key_session_expired", "error_message": "Temporary API key session duration limit exceeded." } ``` ### Text-to-Speech WebSocket [`/tts-websocket`](/api-reference/tts/websocket-api) supports multiple streams over a single connection. The limit is enforced per stream. Each `start_stream` starts its own timer using the value from the temporary API key. When a stream's limit is reached, the server returns a stream-scoped error response followed by a `terminated` message for that `stream_id`. Other streams on the same connection are unaffected and the connection stays open: ```json { "stream_id": "...", "error_code": 403, "error_type": "temp_api_key_session_expired", "error_message": "Temporary API key session duration limit exceeded." } ``` ### Text-to-Speech REST [`/tts`](/api-reference/tts/generate_tts) serves a single stream per request. How the error is reported depends on whether audio bytes have already been sent: * **Before any audio has been sent**, a standard HTTP error response: ```json { "error_code": 403, "error_type": "temp_api_key_session_expired", "error_message": "Temporary API key session duration limit exceeded. Create a new temporary API key to start a new session." } ``` Once audio streaming has started, errors cannot be delivered to the client. *** ## Tracking Bind a `client_reference_id` to the temporary API key at creation time. Every request authenticated with the key is recorded in [usage logs](/guides/usage-logs) with that identifier — clients cannot override it. ```json { "usage_type": "transcribe_websocket", "expires_in_seconds": 60, "client_reference_id": "user_8f2c4b1a" } ``` # Usage logs URL: /guides/usage-logs Per-request record of every transcription or speech generation processed by Soniox - model, audio duration, tokens, cost, and an optional client_reference_id. ## Overview Usage logs are the per-request record of all Speech-to-Text and Text-to-Speech traffic in your Soniox project. Each entry captures the model, audio duration, tokens, cost, timing, and an optional [`client_reference_id`](#tracking-with-client_reference_id). Use them to: * **Monitor cost** - track total Soniox spend, or split it per end user, tenant, or product feature using `client_reference_id`. * **Debug a session** - trace a specific request (model, duration, `uuid`) when a user reports an issue. * **Catch anomalies** - spot leaked API keys, runaway loops, or unusual user behavior through sudden usage spikes. When a client can't be trusted to set `client_reference_id` (browser, mobile, third-party integrations), bind it to a [temporary API key](/guides/temporary-api-keys) at creation - the identifier is then set server-side and cannot be overridden. Logs are available via the [`GET /v1/usage-logs`](/api-reference/other/get_usage_logs) API. For daily cost totals per model rather than individual requests, see [usage summary](/guides/usage-summary). *** ## What gets logged Each entry captures the request's identity (`uuid`, `model`, `request_scope`, `client_reference_id`), its timing (`start_time`, `end_time` in UTC), token and audio-duration counts for input and output, and the resulting cost broken down by input/output and text/audio. See [`GET /v1/usage-logs`](/api-reference/other/get_usage_logs) for the full schema and field types. Only successfully completed requests are logged. Requests that fail before completion - for example, due to an invalid API key, insufficient funds, or a validation error - do not appear in usage logs. *** ## Tracking with `client_reference_id` `client_reference_id` is an optional string you attach to a request to tag its log entry. Use it when you need to know exact Soniox API usage per end-user - for billing, monitoring, or debugging. * **Type:** string * **Max length:** 256 characters Supported on all four APIs: ### Speech-to-Text WebSocket Set in the [configuration message](/api-reference/stt/websocket-api): ```json { "api_key": "", "model": "stt-rt-v5", "audio_format": "auto", "client_reference_id": "user_8f2c4b1a" } ``` ### Speech-to-Text Async Set in the body of [`POST /v1/transcriptions`](/api-reference/stt/transcriptions/create_transcription): ```json { "model": "stt-async-v5", "audio_url": "https://example.com/audio.mp3", "client_reference_id": "user_8f2c4b1a" } ``` ### Text-to-Speech WebSocket Set in the [stream configuration message](/api-reference/tts/websocket-api): ```json { "api_key": "", "model": "tts-rt-v2", "language": "en", "voice": "Adrian", "audio_format": "wav", "stream_id": "stream-001", "client_reference_id": "user_8f2c4b1a" } ``` ### Text-to-Speech REST Set in the body of [`POST /tts`](/api-reference/tts/generate_tts): ```json { "model": "tts-rt-v2", "language": "en", "voice": "Adrian", "audio_format": "wav", "text": "Hello from Soniox.", "client_reference_id": "user_8f2c4b1a" } ``` *** ## Binding `client_reference_id` to a temporary API key A `client_reference_id` can also be bound to a [temporary API key](/api-reference/auth/create_temporary_api_key) at creation time. Every request authenticated with that key is logged with that identifier - no per-request field required from the client. ```json { "usage_type": "transcribe_websocket", "expires_in_seconds": 60, "client_reference_id": "user_8f2c4b1a" } ``` For temporary API keys, `client_reference_id` must be set at creation time - it cannot be set or overridden later via the request. If no `client_reference_id` was bound when the key was created, log entries for requests using that key will have no `client_reference_id`, even if one is passed in the request. *** ## Accessing logs Fetch entries with [`GET /v1/usage-logs`](/api-reference/other/get_usage_logs). # Usage summary URL: /guides/usage-summary Daily cost and activity for a project, broken down per model and summed across all models. ## Overview The usage summary is a project's cost and activity rolled up by UTC day, per model and across all models. It is the same data behind the project's **Usage** page in the [Soniox Console](https://console.soniox.com/). Use it to: * **Report spend** - total cost for a billing period, or per model. * **Chart trends** - plot daily cost or request volume over weeks and months. * **Attribute cost** - see which models account for most of the spend. For a per-request record instead, see [usage logs](/guides/usage-logs). Summaries are available via the [`GET /v1/usage/summary`](/api-reference/other/get_usage_summary) API. *** ## Choosing a window `start_time` is inclusive and `end_time` is exclusive, both ISO 8601 UTC. A UTC day is included when the window covers any part of it, so an `end_time` at exactly midnight leaves out its own day: `2026-04-01T00:00:00Z` to `2026-04-03T00:00:00Z` returns April 1 and April 2. A window may cover at most 366 UTC days. *** ## What gets returned The response has a `total` entry covering all models, and a `models` array with one entry per model that recorded usage. The `total` entry is the one with a `model` of `null`, and a project with no usage gets an empty `models` array. Every entry carries a `days` array of UTC dates, and the per-day arrays align to it by index. Every day in the window appears, so days with no usage are zeros. ```http GET /v1/usage/summary?start_time=2026-04-01T00:00:00Z&end_time=2026-04-03T00:00:00Z Authorization: Bearer ``` ```json { "total": { "model": null, "days": ["2026-04-01", "2026-04-02"], "total_cost_usd": "0.2567250000", "total_input_cost_usd": "0.0415500000", "total_output_cost_usd": "0.2151750000", "total_duration_cost_usd": "0.0000000000", "cost_usd": ["0.1711500000", "0.0855750000"], "input_cost_usd": ["0.0277000000", "0.0138500000"], "output_cost_usd": ["0.1434500000", "0.0717250000"], "duration_cost_usd": ["0.0000000000", "0.0000000000"], "total_num_requests": 285, "total_input_text_tokens": 4800, "total_input_audio_tokens": 15000, "total_input_audio_duration_ms": 1800000, "total_output_text_tokens": 10800, "total_output_audio_tokens": 8250, "total_output_audio_duration_ms": 990000, "total_duration_ms": 0, "num_requests": [190, 95], "input_text_tokens": [3200, 1600], "input_audio_tokens": [10000, 5000], "input_audio_duration_ms": [1200000, 600000], "output_text_tokens": [7200, 3600], "output_audio_tokens": [5500, 2750], "output_audio_duration_ms": [660000, 330000], "duration_ms": [0, 0] }, "models": [ { "model": "stt-async-v5", "days": ["2026-04-01", "2026-04-02"], "total_cost_usd": "0.0613500000", "total_input_cost_usd": "0.0235500000", "total_output_cost_usd": "0.0378000000", "total_duration_cost_usd": "0.0000000000", "cost_usd": ["0.0409000000", "0.0204500000"], "input_cost_usd": ["0.0157000000", "0.0078500000"], "output_cost_usd": ["0.0252000000", "0.0126000000"], "duration_cost_usd": ["0.0000000000", "0.0000000000"], "total_num_requests": 60, "total_input_text_tokens": 300, "total_input_audio_tokens": 15000, "total_input_audio_duration_ms": 1800000, "total_output_text_tokens": 10800, "total_output_audio_tokens": 0, "total_output_audio_duration_ms": 0, "total_duration_ms": 0, "num_requests": [40, 20], "input_text_tokens": [200, 100], "input_audio_tokens": [10000, 5000], "input_audio_duration_ms": [1200000, 600000], "output_text_tokens": [7200, 3600], "output_audio_tokens": [0, 0], "output_audio_duration_ms": [0, 0], "duration_ms": [0, 0] }, { "model": "tts-rt-v2", "days": ["2026-04-01", "2026-04-02"], "total_cost_usd": "0.1953750000", "total_input_cost_usd": "0.0180000000", "total_output_cost_usd": "0.1773750000", "total_duration_cost_usd": "0.0000000000", "cost_usd": ["0.1302500000", "0.0651250000"], "input_cost_usd": ["0.0120000000", "0.0060000000"], "output_cost_usd": ["0.1182500000", "0.0591250000"], "duration_cost_usd": ["0.0000000000", "0.0000000000"], "total_num_requests": 225, "total_input_text_tokens": 4500, "total_input_audio_tokens": 0, "total_input_audio_duration_ms": 0, "total_output_text_tokens": 0, "total_output_audio_tokens": 8250, "total_output_audio_duration_ms": 990000, "total_duration_ms": 0, "num_requests": [150, 75], "input_text_tokens": [3000, 1500], "input_audio_tokens": [0, 0], "input_audio_duration_ms": [0, 0], "output_text_tokens": [0, 0], "output_audio_tokens": [5500, 2750], "output_audio_duration_ms": [660000, 330000], "duration_ms": [0, 0] } ] } ``` See [`GET /v1/usage/summary`](/api-reference/other/get_usage_summary) for the full schema and field types. Only successfully started requests are counted. Requests rejected before they start - for example, due to an invalid API key, insufficient funds, or a validation error - are not included. # Community integrations URL: /integrations/community-integrations Discover useful Soniox integrations built by the community import Image from "next/image"; Web and mobile demo apps screenshot ## Overview This page lists code repositories, SDKs and other useful tools built by our user community. We appreciate the investment in time and willingness to give back to the community from these developers. Note that the projects linked here are community-built integrations and are not officially supported by Soniox. For any questions or support requests, please contact the respective developers or project maintainers. *** ## Agent Voice Response (AVR) [Agent Voice Response](https://github.com/agentvoiceresponse) is the ultimate conversational AI platform for Asterisk PBX systems. Experience ultra-low latency speech-to-speech, advanced Voice Activity Detection, and intelligent noise suppression. Choose between cloud and local AI providers based on your needs. Perfect for FreePBX, Asterisk-based contact centers, and enterprise telephony solutions. Integration wiki: [https://wiki.agentvoiceresponse.com/en/avr-soniox-speech-to-text](https://wiki.agentvoiceresponse.com/en/avr-soniox-speech-to-text). *** ## Go SDK (unofficial) Unofficial Go SDK for the Soniox Speech-to-Text real-time WebSocket API. Enable Soniox real-time speech-to-text transcription and translation in your Go applications. GitHub repository: [https://github.com/moxierobots/soniox-stt-go](https://github.com/moxierobots/soniox-stt-go). *** ## C++ library (unofficial) SonioxPP is an unofficial C++17 library for the Soniox API. It covers real-time WebSocket and async REST speech-to-text, plus REST and real-time WebSocket text-to-speech, with support for speaker diarization, language identification, translation, custom vocabulary, and temporary API keys. GitHub repository: [https://github.com/fatehmtd/SonioxPP](https://github.com/fatehmtd/SonioxPP). *** # Integrations URL: /integrations Explore Soniox Speech-to-Text and Text-to-Speech integrations for real-time, multilingual voice applications. Connect Soniox with LiveKit, Pipecat, LangChain, Twilio, Vercel AI SDK, and more. import { LinkCards } from "@/components/link-card"; Soniox Speech-to-Text integrates seamlessly with leading real-time communication platforms, AI frameworks, automation tools, and developer SDKs. These integrations make it easy to add high-accuracy, low-latency, multilingual speech recognition to live audio, voice agents, call centers, and AI-powered applications, without building everything from scratch. Whether you’re streaming audio in real time, orchestrating AI workflows, or deploying speech recognition at scale, Soniox integrations let you move faster while maintaining enterprise-grade accuracy and performance. Explore the available integrations to quickly connect Soniox Speech-to-Text to your existing stack and start transcribing speech in real time. See also [community integrations](/integrations/community-integrations). # n8n URL: /integrations/n8n How to use Soniox Speech-to-Text AI with n8n import Image from "next/image"; import VideoPlayer from "@/components/video-player"; import { Callout } from "@/components/callout";
Soniox x n8n
## Overview Soniox Speech-to-Text AI turns audio into highly accurate text. Paired with n8n, you can build powerful automation workflows that transcribe audio from any source. Use the Soniox node in your n8n workflows to: * Transcribe audio files uploaded to cloud storage * Process voice messages from messaging platforms * Build automated transcription pipelines at scale * Combine speech-to-text with other n8n integrations All with enterprise-grade accuracy. *** ## Getting started To use Soniox with n8n, you'll need: * An [n8n](https://n8n.io/) instance (self-hosted or cloud) * A [Soniox account](https://console.soniox.com/) with an API key *** ## Installation Soniox provides a first-party verified node in the n8n marketplace. Search for "Soniox" in the node panel to find it.
Searching for the Soniox verified node in the n8n UI
Alternatively, you can install via npm: ```bash npm install @soniox/n8n-nodes-soniox ``` ## Credentials The Soniox node requires an API key to authenticate. ### Get your API key 1. Sign in to the [Soniox Console](https://console.soniox.com/) 2. Navigate to **API Keys** 3. Create a new key or copy an existing one ### Add credentials in n8n 1. In n8n, go to **Credentials** > **Add Credential** 2. Search for **Soniox API** 3. Enter your API key 4. Click **Save** The credentials will be tested automatically. If successful, you're ready to use the Soniox node. *** ## Operations The Soniox node supports three operations: ### Create transcription Creates a new transcription job from an audio source. You can choose to wait for completion or receive results asynchronously via webhook. **Audio sources:** | Source | Description | | ----------- | ------------------------------------------------------------------------ | | Binary File | Upload audio from a previous node (e.g., HTTP Request, Read Binary File) | | Audio URL | Provide a publicly accessible URL to the audio file | | File ID | Use a file previously uploaded to Soniox | ### Get results Retrieves the status and transcript for an existing transcription job. Use this when processing transcriptions asynchronously. ### Delete Deletes a transcription and its associated file from Soniox. Simply provide the transcription ID — the node automatically fetches the file ID and deletes both resources. Use this to clean up after async workflows. You must ensure the file is deleted after the transcription is completed. Files and transcriptions stored via the Async API are automatically deleted after 30 days. Check [limits and quotas](https://soniox.com/docs/stt/async/limits-and-quotas) for more information. *** ## Basic usage ### Transcribe from URL The simplest way to transcribe audio is from a public URL: 1. Add the **Soniox** node to your workflow 2. Select **Create** operation 3. Set **Audio Source** to **Audio URL** 4. Enter the URL to your audio file 5. Execute the workflow The node will wait for the transcription to complete and return the full transcript. ### Transcribe from binary data To transcribe audio from another node (like HTTP Request or Read Binary File): 1. Connect the source node to the Soniox node 2. Select **Create** operation 3. Set **Audio Source** to **Binary File** 4. Set **Binary Property Name** to the property containing your audio (default: `data`) 5. Execute the workflow ### Polling settings When **Wait for Completion** is enabled, you can configure: | Setting | Default | Description | | ------------------- | ------- | -------------------------------------- | | Poll Interval (Sec) | 1 | How often to check for completion | | Max Wait (Sec) | 300 | Maximum time to wait before timing out | If the **Max Wait (Sec)** value is too large, the n8n cloud platform may timeout before the transcription completes. For long audio files, consider using [async processing with webhooks](#async-processing-with-webhooks) instead. *** ## Advanced usage ### Language hints The model automatically detects and transcribes any supported language. It also handles multilingual audio, even when multiple languages appear within the same conversation. If you know which languages are likely to be spoken, you can provide language hints to improve accuracy: 1. In the Soniox node, find **Language Hints** 2. Click **Add Language** 3. Enter the language code (e.g., `en`, `es`, `fr`) 4. Repeat for additional languages See [list of supported languages](/stt/concepts/supported-languages) for all available language codes. Learn more about [language hints](/stt/concepts/language-hints). ### Speaker diarization Enable **Enable Speaker Diarization** to identify and separate different speakers in the audio. The transcript will include speaker labels for each segment. ### Language identification Enable **Enable Language Identification** to include detected language information in the transcript output. ### Customization with context Provide context to help the model better understand domain-specific terminology, names, or phrases. **Simple text context:** 1. Set **Context Mode** to **Text** 2. Enter relevant terms or phrases in the **Context Text** field ``` Celebrex, Zyrtec, Xanax, Prilosec, Amoxicillin ``` **Structured JSON context:** For more control, use structured context: 1. Set **Context Mode** to **Structured JSON** 2. Enter a JSON object in the **Context JSON** field ```json { "general": [ {"key": "domain", "value": "Healthcare"} ], "text": "Medical consultation recording", "terms": ["Celebrex", "Zyrtec", "Xanax"], "translation_terms": [ {"source": "Dr. Smith", "target": "Dr. Smith"} ] } ``` Learn more about [customizing with context](/stt/concepts/context). ### Translation Soniox can translate the transcript to another language during transcription. **One-way translation:** Translate the transcript to a single target language: 1. Set **Translation Type** to **One Way** 2. Enter the **Target Language** code (e.g., `es` for Spanish) **Two-way translation:** For conversations between speakers of two languages, translate each speaker to the other's language: 1. Set **Translation Type** to **Two Way** 2. Enter **Language A** (e.g., `en`) 3. Enter **Language B** (e.g., `es`) *** ## Async processing with webhooks For long audio files or high-volume processing, you can use webhooks instead of waiting for completion: 1. Set **Wait for Completion** to **false** 2. Enter your **Webhook URL** 3. Optionally set **Webhook Auth Header Name** and **Webhook Auth Header Value** for authentication The node will immediately return the transcription ID. Soniox will send the results to your webhook when processing completes.
To fetch results later, use the **Get Results** operation with the transcription ID. *** ## Output options ### Output mode Choose what data to return when the transcription completes: | Mode | Description | | ------------- | -------------------------------------------------------------------------------------- | | Full Response | Returns the complete transcript with all metadata, timestamps, and speaker information | | Text Only | Returns only the transcribed text as a simple string | *** ## Cleanup and resource management When you upload binary files, Soniox stores them temporarily. To avoid accumulating unused files, use the cleanup features: ### Automatic cleanup (recommended) When **Wait for Completion** is enabled, the **Auto Delete** option is available (enabled by default). When enabled, the node automatically deletes: * The **transcription** — always deleted regardless of audio source * The **uploaded file** — only deleted when using Binary File as the audio source (since that's when a file is uploaded) This keeps your Soniox account clean without extra workflow steps. ### Manual cleanup For async workflows (when not waiting for completion), use the **Delete** operation to clean up after processing: 1. Add a new **Soniox** node after receiving the webhook callback 2. Select **Delete** operation 3. Enter the **Transcription ID** 4. Execute The Delete operation automatically fetches the transcription details to find the associated file ID, then deletes both the transcription and its file (if one exists). **Example async cleanup workflow:** 1. **Soniox** (Create) — Upload file, don't wait. 2. **Webhook** trigger — Receives completion callback from Soniox with `id` (transcription ID) 3. Process the transcript as needed 4. **Soniox** (Delete) — Clean up using the transcription ID from step 2 *** ## Resources * [Soniox API Reference](/api-reference) * [Supported Languages](/stt/concepts/supported-languages) * [n8n Documentation](https://docs.n8n.io/) * [GitHub Repository](https://github.com/soniox/n8n-nodes-soniox) # TanStack AI SDK URL: /integrations/tanstack-ai-sdk Soniox transcription adapter for the TanStack AI SDK. import Image from "next/image"; import { LinkCards } from "@/components/link-card";
Soniox x Tanstack AI SDK
## Overview [TanStack AI](https://tanstack.com/ai) is a TypeScript toolkit for building AI applications. It provides a unified API that abstracts away the differences between various AI providers, allowing developers to switch models with just a few lines of code. This package (`@soniox/tanstack-ai-adapter`) implements the SDK's transcription adapter, enabling you to use Soniox's Speech-to-Text models directly within the standard TanStack AI workflow. ## Installation ```bash npm install @soniox/tanstack-ai-adapter ``` ## Authentication Set `SONIOX_API_KEY` in your environment or pass `apiKey` when creating the adapter. Get your API key from the [Soniox Console](https://console.soniox.com). ## Example ```ts import { generateTranscription } from '@tanstack/ai'; import { sonioxTranscription } from '@soniox/tanstack-ai-adapter'; const result = await generateTranscription({ adapter: sonioxTranscription('stt-async-v4'), audio: new URL( 'https://soniox.com/media/examples/coffee_shop.mp3', ), modelOptions: { enableLanguageIdentification: true, enableSpeakerDiarization: true, }, }); console.log(result.text); console.log(result.segments); // Timestamped segments with speaker info ``` ## Adapter configuration Use `createSonioxTranscription` to customize the adapter instance: ```ts import { createSonioxTranscription } from '@soniox/tanstack-ai-adapter'; const adapter = createSonioxTranscription('stt-async-v4', process.env.SONIOX_API_KEY!, { baseUrl: 'https://api.soniox.com', pollingIntervalMs: 1000, timeout: 180000, }); ``` Options: * `apiKey`: override `SONIOX_API_KEY` (required when using `createSonioxTranscription`). * `baseUrl`: custom API base URL. See list of regional API endpoints [here](/data-residency#regional-endpoints). Default is `https://api.soniox.com`. * `headers`: additional request headers. * `timeout`: transcription timeout in milliseconds. Default is 180000ms (3 minutes). * `pollingIntervalMs`: transcription polling interval in milliseconds. Default is 1000ms. ## Transcription options Per-request options are passed via `modelOptions`: ```ts const result = await generateTranscription({ adapter: sonioxTranscription('stt-async-v4'), audio, modelOptions: { languageHints: ['en', 'es'], enableLanguageIdentification: true, enableSpeakerDiarization: true, context: { terms: ['Soniox', 'TanStack'], }, }, }); ``` Available options: * `languageHints` - Array of ISO language codes to bias recognition. If you pass the TanStack `language` option, this adapter will merge it into `languageHints` for convenience. * `languageHintsStrict` - When true, rely more heavily on language hints (note: not supported by all models) * `enableLanguageIdentification` - Automatically detect spoken language * `enableSpeakerDiarization` - Identify and separate different speakers * `context` - Additional context to improve accuracy * `clientReferenceId` - Optional client-defined reference ID * `webhookUrl` - Webhook URL for transcription completion notifications * `webhookAuthHeaderName` - Webhook authentication header name * `webhookAuthHeaderValue` - Webhook authentication header value * `translation` - Translation configuration For more information on the available options, see the [Speech-to-Text API reference](/api-reference/stt/transcriptions/create_transcription). ## Accessing raw tokens When using translation or working with multilingual audio, you may need access to raw tokens with per-token language information and translation status. The adapter attaches a non-standard `providerMetadata` field at runtime: ```ts const result = await generateTranscription({ adapter: sonioxTranscription('stt-async-v4'), audio, modelOptions: { translation: { type: 'one_way', targetLanguage: 'es' }, }, }); // Access raw Soniox tokens with full metadata const rawTokens = (result as any).providerMetadata?.soniox?.tokens; if (rawTokens) { rawTokens.forEach((token) => { // token.text - token text // token.start_ms - start time in milliseconds // token.end_ms - end time in milliseconds // token.language - detected language for this token // token.translation_status - translation status (if translation enabled) // token.speaker - speaker identifier // token.confidence - confidence score }); } ``` **Note:** When using translation, the API returns both transcription tokens (original) and translation tokens. The `segments` array always includes only transcription tokens. To access translation tokens, filter by `translation_status === 'translation'`. # Twilio URL: /integrations/twilio Stream Twilio call audio to Soniox Speech-to-Text API and get real-time transcriptions. import Image from "next/image"; import { LinkCards } from "@/components/link-card";
Soniox x Twilio
## Overview This guide demonstrates how to stream live Twilio call audio to the Soniox Speech-to-Text API and receive real-time transcription via WebSockets. If you want to see a complete example, check out this GitHub repository: ## Preparation ### Create a Twilio account To get started, you'll need a Twilio account. If you don't have one, you can [sign up](https://www.twilio.com/try-twilio). You will also need two phone numbers to test the integration: * One from a phone number you own, which needs to be verified by Twilio. * The other one is a Twilio-owned phone number that you can use for testing. ### Get your Soniox API key To use Soniox Speech-to-Text API in your application, you'll need to obtain an API key. You can get one by signing up at [Soniox Console](https://console.soniox.com/). ## Running the example ### Clone the repository Clone the repository and install the dependencies: ```bash git clone https://github.com/soniox/soniox-twilio-realtime-transcription.git cd soniox-twilio-realtime-transcription pip install -r requirements.txt ``` ### Configure server environment Copy the `.env.example` file to `.env` and update the values with your Twilio account credentials and Soniox API key: ```bash cp .env.example .env ``` ### Run the server and expose it to Twilio Run the server: ```bash python server.py ``` This will start the server and listen for incoming Twilio calls. You will specify where phone call recording is streamed later. To expose the server to Twilio, you can use [ngrok](https://ngrok.com/). ```bash ngrok http 5000 ``` Note the forwarding URL that ngrok provides. It should look like `https://.ngrok.io` or `https://.ngrok-free.app`. ### Run the client Edit `client.html` and set `WEBSOCKET_URL` to your ngrok URL with `/client` at the end, e.g. `wss://xxxxx.ngrok.io/client`. Open `client.html` in your browser to view live call transcriptions. ### Start a Twilio call You can configure Twilio calls with `TwiML Bin` files. More information about streaming can be found in the [Twilio documentation](https://www.twilio.com/docs/voice/twiml/stream). Here is an example `TwiML Bin` file calls you phone number and streams the audio to your websocket server: ```xml USER_PHONE_NUMBER "Hello, this is a test call. How are you?" "Thank you, bye!" ``` To start, we recommend using the provided `call_me.py` script to start a Twilio call. Simply set the following environment variables: * `TWILIO_ACCOUNT_SID` and `TWILIO_AUTH_TOKEN` (from Twilio) * `TWILIO_PHONE_NUMBER` (your Twilio number, rented on Twilio) * `WEBSOCKET_URL` with your ngrok URL with `/twilio` at the end, e.g. `wss://xxxxx.ngrok.io/twilio`. * `USER_PHONE_NUMBER` with your Twilio-verified phone number. ```bash python call_me.py ``` You should hear a voice message saying "Hello, this is a test call. How are you?" and then a message saying "Thank you, bye!". Simultaneously, you should see the transcription in the browser. # Vercel AI SDK URL: /integrations/vercel-ai-sdk Soniox transcription provider for the Vercel AI SDK. import Image from "next/image"; import { LinkCards } from "@/components/link-card";
Soniox x Vercel AI SDK
## Overview [Vercel AI SDK](https://sdk.vercel.ai/) is a TypeScript toolkit for building AI applications. It provides a unified API that abstracts away the differences between various AI providers, allowing developers to switch models with just a few lines of code. The [`@soniox/vercel-ai-sdk-provider`](https://www.npmjs.com/package/@soniox/vercel-ai-sdk-provider) package implements the SDK's Transcription Interface, enabling you to use Soniox's Speech-to-Text models directly within the standard Vercel AI workflow. Learn more about the Soniox provider in the [Vercel AI SDK Community Providers documentation](https://ai-sdk.dev/providers/community-providers/soniox). ## Installation ```bash npm install @soniox/vercel-ai-sdk-provider ``` ## Authentication Set `SONIOX_API_KEY` in your environment or pass `apiKey` when creating the provider. ## Example ```ts import { soniox } from '@soniox/vercel-ai-sdk-provider'; import { experimental_transcribe as transcribe } from 'ai'; const { text } = await transcribe({ model: soniox.transcription('stt-async-v4'), audio: new URL( 'https://soniox.com/media/examples/coffee_shop.mp3', ), }); ``` ## Provider options Use `createSoniox` to customize the provider instance: ```ts import { createSoniox } from '@soniox/vercel-ai-sdk-provider'; const soniox = createSoniox({ apiKey: process.env.SONIOX_API_KEY, apiBaseUrl: 'https://api.soniox.com', }); ``` Options: * `apiKey`: override `SONIOX_API_KEY`. * `apiBaseUrl`: custom API base URL. See list of regional API endpoints [here](/data-residency#regional-endpoints). * `headers`: additional request headers. * `fetch`: custom fetch implementation. * `pollingIntervalMs`: transcription polling interval in milliseconds. Default is 1000ms. ## Transcription options Per-request options are passed via `providerOptions`: ```ts const { text } = await transcribe({ model: soniox.transcription('stt-async-v4'), audio, providerOptions: { soniox: { languageHints: ['en', 'es'], enableLanguageIdentification: true, enableSpeakerDiarization: true, context: { terms: ["Soniox", "Vercel"] }, }, }, }); ``` Available options: * `languageHints` - Array of ISO language codes to bias recognition * `languageHintsStrict` - When true, rely more heavily on language hints (note: not supported by all models) * `enableLanguageIdentification` - Automatically detect spoken language * `enableSpeakerDiarization` - Identify and separate different speakers * `context` - Additional context to improve accuracy * `clientReferenceId` - Optional client-defined reference ID * `webhookUrl` - Webhook URL for transcription completion notifications * `webhookAuthHeaderName` - Webhook authentication header name * `webhookAuthHeaderValue` - Webhook authentication header value * `translation` - Translation configuration For more information on the available options, see the [Speech-to-Text API reference](/api-reference/stt/transcriptions/create_transcription). # Get started URL: /stt/get-started Learn how to use Soniox Speech-to-Text API. ## Learn how to use Soniox Speech-to-Text API in minutes Soniox Speech-to-Text is a **universal speech AI** that lets you transcribe speech in 60+ languages — from recorded files (async) or live audio streams (real-time). Languages can be freely mixed within the same conversation, and Soniox will handle them seamlessly with high accuracy and low latency. In just a few steps, you can run your first transcription. The examples also cover real-time and async transcription flows through the same simple API. ### Get API key Create a [Soniox account](https://console.soniox.com/signup) and log in to the [Console](https://console.soniox.com) to get your API key. API keys are created per project. In the Console, go to **My First Project** and click **API Keys** to generate one. Export it as an environment variable (replace with your key): ```sh title="Terminal" export SONIOX_API_KEY= ``` ### Get examples Clone the official examples repo: ```sh title="Terminal" git clone https://github.com/soniox/soniox_examples cd soniox_examples/speech_to_text ``` ### Run examples Choose your language and run the ready-to-use examples below. {/* TABLE START */} {/* NOTE: Width is set so that we have maximum of 2 lines in 'Example' column. */} {/* NOTE: Font size is set so the table doesn't look "too big". */}
|
Example
| What it does | Output | | ------------------------------------------ | --------------------------------------------------------- | ------------------------------- | | **Real-time
transcription** | Transcribes speech in any language in
real-time. | Transcript streamed to console. | | **Transcribe file from URL** | Transcribes an audio file directly from a public URL. | Transcript printed to console. | | **Transcribe local file** | Uploads and transcribes an audio file from your computer. | Transcript printed to console. |
{/* TABLE END */} {/* NOTE: Empty tag is needed so code block renders correctly */}
```sh title="Terminal" cd python_sdk # Set up environment python3 -m venv venv source venv/bin/activate pip install -r requirements.txt # Real-time examples python soniox_sdk_realtime.py --audio_path ../assets/coffee_shop.mp3 # Async examples python soniox_sdk_async.py --audio_url "https://soniox.com/media/examples/coffee_shop.mp3" python soniox_sdk_async.py --audio_path ../assets/coffee_shop.mp3 ``` {/* NOTE: Empty tag is needed so code block renders correctly */}
```sh title="Terminal" cd nodejs_sdk # Install dependencies npm install # Real-time examples node soniox_sdk_realtime.js --audio_path ../assets/coffee_shop.mp3 # Async examples node soniox_sdk_async.js --audio_url "https://soniox.com/media/examples/coffee_shop.mp3" node soniox_sdk_async.js --audio_path ../assets/coffee_shop.mp3 ``` {/* NOTE: Empty tag is needed so code block renders correctly */}
```sh title="Terminal" cd python # Set up environment python3 -m venv venv source venv/bin/activate pip install -r requirements.txt # Real-time examples python soniox_realtime.py --audio_path ../assets/coffee_shop.mp3 # Async examples python soniox_async.py --audio_url "https://soniox.com/media/examples/coffee_shop.mp3" python soniox_async.py --audio_path ../assets/coffee_shop.mp3 ``` {/* NOTE: Empty tag is needed so code block renders correctly */}
```sh title="Terminal" cd nodejs # Install dependencies npm install # Real-time examples node soniox_realtime.js --audio_path ../assets/coffee_shop.mp3 # Async examples node soniox_async.js --audio_url "https://soniox.com/media/examples/coffee_shop.mp3" node soniox_async.js --audio_path ../assets/coffee_shop.mp3 ``` ## Next steps * **Dive into the [Real-time API](/stt/rt/real-time-transcription)** → Run live transcription and endpoint detection. * **Explore the [Async API](/stt/async/async-transcription)** → Transcribe recorded files at scale and integrate with webhooks. # Models URL: /stt/models Learn about latest models, changelog, and deprecations. Soniox Speech-to-Text AI provides multiple models for real-time and asynchronous transcription and translation. This page lists the currently available models, their capabilities, and important updates. *** ## Current models {/*TABLE START */} {/* NOTE: Width is set so that we have maximum of 2 lines in 'Example' column. */} {/* NOTE: Font size is set so the table doesn't look "too big".*/}
|
Model
| {" "}
Type
| Status | | -------------------------------------------- | ----------------------------------------------- | ---------- | | **stt-rt-v5** | Real-time | **Active** | | **stt-async-v5** | Async | **Active** |
{/* TABLE END */} *** ## Aliases Aliases provide a stable reference so you don't need to change your code when newer versions are released. | Alias | Points to | Notes | | ---------------- | -------------- | ----- | | **stt-rt-v4** | `stt-rt-v5` | | | **stt-async-v4** | `stt-async-v5` | | *** ## Changelog ### June 16, 2026 **New models:** stt-rt-v5 **Replaces:** stt-rt-v4 #### Overview stt-rt-v5 is the new Soniox real-time speech-to-text model for live audio. It delivers higher accuracy, reinvented speaker separation, improved spoken language identification, higher-quality real-time translation, faster semantic endpointing, better context handling, and more reliable recognition and formatting of structured speech data such as numbers, dates, emails, names, addresses, and codes. #### Key improvements * Higher real-time transcription accuracy across 60+ languages * Better robustness on noisy audio, telephony, far-field microphones, accents, interruptions, overlapping speech, and mixed-language conversations * Reinvented speaker separation for identifying who said what in live conversations * Improved spoken language identification across multilingual and accented speech * Higher-quality real-time translation across 3,600+ language pairs * Faster and more reliable semantic endpointing for voice agents, dictation, command systems, and conversational apps * Better alphanumeric recognition and formatting for numbers, dates, times, emails, IDs, codes, names, and addresses * More robust context usage for names, domain terms, product names, translation preferences, and custom vocabulary #### API compatibility The stt-rt-v5 model is fully compatible with the existing stt-rt-v4 model and Soniox Real-Time API. To upgrade, simply replace the model name in your API request: `{ "model": "stt-rt-v5" }` #### Deprecation notice The stt-rt-v4 model will be removed on June 30, 2026. After June 30, 2026, requests using stt-rt-v4 will automatically route to stt-rt-v5 with no service interruption and no API changes required. ### June 11, 2026 **New models:** stt-async-v5 **Replaces:** stt-async-v4 #### Overview **stt-async-v5** is the new Soniox async speech-to-text model for processing recorded audio. It delivers higher accuracy, stronger speaker separation, improved language identification, better context handling, and more reliable formatting of structured speech data such as numbers, dates, emails, names, and codes. #### Key improvements * Higher transcription accuracy across 60+ languages * Better robustness on noisy audio, telephony, accents, and mixed-language speech * Completely reengineered speaker separation for identifying who said what * Improved spoken language identification across multilingual conversations * Better alphanumeric recognition and formatting for numbers, dates, times, emails, IDs, and codes * More robust context usage for names, domain terms, product names, and custom vocabulary #### API compatibility * The stt-async-v5 model is fully compatible with the existing stt-async-v4 model and Soniox API * To upgrade, simply replace the model name in your API request to `{ "model": "stt-async-v5" }` #### Deprecation notice * The stt-async-v4 model will be removed on June 30, 2026 * After June 30, 2026, requests using stt-async-v4 will automatically route to stt-async-v5 with no service interruption and no API changes required ### February 5, 2026 **New models:** stt-rt-v4 **Replaces:** stt-rt-v3 #### Overview **Soniox v4 Real-Time** is a next-generation real-time speech recognition model built for low-latency voice interactions. It delivers speaker-native accuracy across 60+ languages with improved latency, reliability, and conversational behavior. The model is production-ready and fully backward-compatible with v3 Real-Time. #### Key improvements * Higher accuracy across all supported languages * Better multilingual detection and mid-sentence language switching * Lower endpoint latency with faster final transcription * Improved semantic endpointing for more natural turn-taking * Lower manual finalization latency with faster final transcription * More stable, higher-quality transcription on long and multi-hour recordings * Stronger use of provided context for domain-specific accuracy * More fluent, accurate, and consistent translation across all supported languages * Added `max_endpoint_delay_ms` for controlling end-of-speech endpoint delay #### API compatibility * The stt-rt-v4 model is fully compatible with the existing stt-rt-v3 model and Soniox API * To upgrade, simply replace the model name in your API request: * `{ "model": "stt-rt-v4" }` for real-time #### Deprecation notice * The stt-rt-v3 model will be removed on February 28, 2026 * After February 28, 2026, requests will automatically route to stt-rt-v4 with no service interruption. No API changes required ### January 29, 2026 **New models:** stt-async-v4 **Replaces:** stt-async-v3 #### Overview **Soniox v4 Async** is the latest generation of Soniox’s asynchronous speech recognition and translation model. This release delivers a significant improvement in accuracy, robustness, and multilingual performance across more than 60 languages. v4 Async reaches human-parity transcription quality in real-world scenarios, while also introducing stronger long-form processing, improved speaker diarization, richer context handling, and higher-quality translation output. The model is designed for production-scale workloads and consistent, high-fidelity results across diverse acoustic environments and language mixes. #### Key improvements * Higher transcription accuracy across all languages, reaching speaker-native quality in many domains * More robust performance in noise, accents, overlapping speech, and poor audio * Better language identification and smoother mid-sentence language switching * Improved speaker separation and more consistent labeling in multi-speaker audio * Better normalization of dates, numbers, phone/email addresses, and other structured content * More stable, higher-quality transcription on long and multi-hour recordings * Stronger use of provided context for domain-specific accuracy * More fluent, accurate, and consistent translation across all supported languages #### API compatibility * The stt-async-v4 model is fully compatible with the existing stt-async-v3 model and Soniox API * To upgrade, simply replace the model name in your API request: * `{ "model": "stt-async-v4" }` for async #### Deprecation notice * The stt-async-v3 model will be removed on February 28, 2026 * After February 28, 2026, requests will automatically route to stt-async-v4 with no service interruption. No API changes required ### October 31, 2025 #### Model retirement and upgrade We have accelerated the retirement of older models following the overwhelmingly positive response to the new v3 models. The following models have been retired: * stt-async-preview-v1 * stt-rt-preview-v2 Both models have been **aliased to the new Soniox v3 models.** This means all existing requests using the old model names are now automatically served with v3, giving every user our most accurate, capable, and intelligent voice AI experience, without any code changes required. #### Context compatibility The context feature is now backward compatible with v3 models, ensuring smooth migration from older versions. However, we **strongly recommend updating to the new context** structure for best results and future flexibility. Learn more about [context](/stt/concepts/context). ### October 29, 2025 **Model update:** v3 enhancements **Applies to:** stt-rt-v3, stt-async-v3 #### New features * **Extended audio duration support:** both real-time (stt-rt-v3) and asynchronous (stt-async-v3) models now support **audio up to 5 hours** in a single request. #### Quality improvements * **Higher transcription accuracy** across challenging audio conditions and diverse languages. #### Notes * No API changes are required; existing integrations continue to work seamlessly. * For asynchronous processing, large files up to 5 hours can now be uploaded directly without chunking. * For real-time streaming, sessions up to 5 hours are supported under the same WebSocket connection. ### October 21, 2025 **New models:** stt-rt-v3, stt-async-v3 **Replaces:** stt-rt-preview-v2, stt-async-preview-v1 #### Overview The **v3 models** introduce major improvements across recognition, translation, and reasoning — making Soniox faster, more accurate, and more capable than ever before. These models power real-time and asynchronous speech processing in 60+ languages, with enhanced accuracy, robustness, and context understanding. #### Key improvements * Higher transcription accuracy across 60+ languages * Improved multilingual switching — seamless recognition when speakers change language mid-sentence * Significantly higher translation quality, especially for languages such as German and Korean * The async model now also supports translation * Support for new advanced structured context, enabling richer domain- and task-specific adaptation * Enhanced alphanumeric accuracy (addresses, IDs, codes, serials) * More accurate speaker diarization, even in overlapping speech * Extended maximum audio duration to 5 hours for both async and real-time models #### API compatibility * The v3 models are fully compatible with the existing Soniox API, if you are not using the context feature. * To upgrade, simply replace the model name in your API request: * `{ "model": "stt-rt-v3" }` for real-time * `{ "model": "stt-async-v3" }` for async * If you are using the context feature, update to the new structured [context](/stt/concepts/context) for improved accuracy. #### Deprecation notice The following preview models are **deprecated** and will be retired on **November 30, 2025:** * stt-async-preview-v1 * stt-rt-preview-v2 Please migrate to the v3 models before that date to ensure uninterrupted service. ### August 15, 2025 * Deprecated `stt-rt-preview-v1` ### August 5, 2025 * Released `stt-rt-preview-v2` * Higher transcription accuracy * Improved translation quality * Expanded to support all translation pairs * More reliable automatic language switching * **Replaces:** stt-rt-preview-v2, stt-async-preview-v1 # Get started URL: /translation/get-started Translate live speech across supported languages with low-latency streaming. import { LinkCards } from "@/components/link-card"; Soniox lets you translate live speech in real time across 60+ languages and 3,600+ language pairs. You can stream translated text as people speak, or combine Soniox STT and TTS to build full speech-to-speech translation. Speech translation is built into the Soniox real-time API, so you can start with transcription and enable translation with a simple configuration change. *** ## Two ways to build speech translation Translate live speech into written text.

Use the Soniox STT API to stream both the original transcript and translated text in real time. This is ideal for captions, subtitles, meeting translation, accessibility, agent assist, live dashboards, and multilingual transcription workflows. ), href: "/translation/stt-translation", showArrow: false, }, { title: "Speech-to-speech translation", description: ( <> Translate live speech into spoken output.

Combine Soniox STT and Soniox TTS to recognize speech, translate it in real time, and speak the result in the target language with low latency. This is ideal for live interpreters, bilingual voice agents, travel assistants, customer support, and real-time multilingual communication. ), href: "/translation/sts-translation", showArrow: false, }, ]} /> *** ## Two ways to translate live speech Translate speech from any supported language into one target language.

Use one-way translation when many speakers or many languages need to be understood in a single output language, such as live captions, lectures, broadcasts, meetings, events, and customer calls. ), href: "/translation/stt-translation#one-way-translation", showArrow: false, }, { title: "Two-way translation", description: ( <> Translate between two languages for live bilingual conversation.

Use two-way translation when both sides should speak naturally in their own language and understand each other in real time. ), href: "/translation/stt-translation#two-way-translation", showArrow: false, }, ]} /> *** ## Language coverage Speech translation uses the Soniox supported-language set and supports more than 3,600 language pairs. See [Supported languages](/translation/supported-languages) for the language list and coverage details. *** ## Run speech-to-text translation examples Use the official examples repo to try real-time one-way and two-way translation with sample audio. ### Get API key Create a [Soniox account](https://console.soniox.com/signup) and log in to the [Console](https://console.soniox.com) to get your API key. API keys are created per project. In the Console, go to **My First Project** and click **API Keys** to generate one. Export it as an environment variable: ```sh title="Terminal" export SONIOX_API_KEY= ``` ### Get examples Clone the official examples repo: ```sh title="Terminal" git clone https://github.com/soniox/soniox_examples cd soniox_examples/speech_to_text ``` ### Run speech-to-text translation examples Choose your language and run the translation examples below.
|
Example
| What it does | Output | | ------------------------------------------ | ------------------------------------------------------------------------------------------------ | ---------------------------------------------------------- | | **Real-time
one-way translation** | Transcribes speech in any detected language and translates it into Spanish in real time. | Transcript + Spanish translation streamed together. | | **Real-time
two-way translation** | Transcribes and translates English ↔ Spanish in real time. Spanish → English, English → Spanish. | Transcript + bidirectional translations streamed together. |
```sh title="Terminal" cd python_sdk # Set up environment python3 -m venv venv source venv/bin/activate pip install -r requirements.txt # One-way translation of a live audio stream python soniox_sdk_realtime.py --audio_path ../assets/coffee_shop.mp3 --translation one_way # Two-way translation of a live audio stream python soniox_sdk_realtime.py --audio_path ../assets/two_way_translation.mp3 --translation two_way ```
```sh title="Terminal" cd nodejs_sdk # Install dependencies npm install # One-way translation of a live audio stream node soniox_sdk_realtime.js --audio_path ../assets/coffee_shop.mp3 --translation one_way # Two-way translation of a live audio stream node soniox_sdk_realtime.js --audio_path ../assets/two_way_translation.mp3 --translation two_way ```
```sh title="Terminal" cd python # Set up environment python3 -m venv venv source venv/bin/activate pip install -r requirements.txt # One-way translation of a live audio stream python soniox_realtime.py --audio_path ../assets/coffee_shop.mp3 --translation one_way # Two-way translation of a live audio stream python soniox_realtime.py --audio_path ../assets/two_way_translation.mp3 --translation two_way ```
```sh title="Terminal" cd nodejs # Install dependencies npm install # One-way translation of a live audio stream node soniox_realtime.js --audio_path ../assets/coffee_shop.mp3 --translation one_way # Two-way translation of a live audio stream node soniox_realtime.js --audio_path ../assets/two_way_translation.mp3 --translation two_way ``` # Real-time speech-to-speech translation URL: /translation/sts-translation Build a real-time spoken translation pipeline by combining Soniox STT with translation and TTS. ## Overview **Soniox real-time speech-to-speech translation** takes spoken audio in one language and plays it back as spoken audio in another in real time. Build the pipeline by chaining two Soniox products: 1. **[Real-time speech-to-text translation](/translation/stt-translation/rt-translation):** Soniox recognizes speech and translates it in real time across supported languages. 2. **[Real-time text-to-speech](/tts/rt/real-time-generation):** Soniox speaks the translated text in the target language with a chosen [voice](/tts/concepts/voices), with streaming output. Both APIs are streaming and low-latency by design, so audio for the first translated words can play before the speaker finishes their sentence. Typical use cases: * **Live interpreters** for meetings, conversations, and business communication. * **Bilingual voice agents** for support, sales, scheduling, healthcare, and other multilingual workflows. * **Travel assistants** and customer support that translate calls while preserving names, numbers, and verification codes. * **Real-time multilingual communication:** anywhere two people who don't share a language need to speak naturally. Need translated **text** output instead? See [Speech-to-text translation](/translation/stt-translation). *** ## How it works The pipeline has two Soniox WebSocket connections and a small piece of application logic between them: 1. **STT + translation** receives audio streams transcription and translation tokens back. Translation tokens are sent to TTS. 2. **TTS** receives the translated text chunks and streams audio chunks back. Your app decodes and plays those chunks as they arrive. Because both APIs stream, your app can start sending translated text to TTS before the speaker has finished the full utterance. Check out the **[Soniox speech-to-speech translation demo](/demo-apps/soniox-speech-to-speech-translation)**, a FastAPI backend and vanilla JS frontend that wires the STT with translation to the TTS. *** ## Pipeline configuration You combine two configs, one for each API. Pick a translation mode for the STT side and a [voice](/tts/concepts/voices) and [audio output format](/tts/concepts/audio-formats) for the TTS side. Real-time STT with one-way translation into Spanish: ```json { "model": "stt-rt-v5", "audio_format": "auto", "enable_endpoint_detection": true, "max_endpoint_delay_ms": 500, "translation": { "type": "one_way", "target_language": "es" } } ``` For two-way conversations (e.g. English ⟷ Spanish), use `{"type": "two_way", "language_a": "en", "language_b": "es"}` so each speaker hears the other's language back. Real-time TTS opens one stream per utterance in the target language: ```json { "model": "tts-rt-v2", "voice": "Daniel", "audio_format": "pcm_s16le", "sample_rate": 24000 } ``` Voices are multilingual, so the same voice ID works across supported languages. *** ## Things to consider * **Latency:** total end-to-end latency is roughly STT translation latency + TTS time-to-first-audio. Keep TTS streams short (one utterance each) and start them eagerly. * **Utterance boundaries:** enable [endpoint detection](/stt/rt/endpoint-detection) on the STT side and use the final `` token to close the current TTS stream. * **Voice consistency:** Soniox voices work with [all 60+ supported languages](/tts/concepts/supported-languages), so you can keep the same voice across translation targets. * **Two-way mode:** for bilingual conversations, you can maintain separate TTS streams per direction and pick which to play based on the translation token's `language` field. # Supported languages URL: /translation/supported-languages Languages supported by Soniox speech translation and how to check coverage for your model. ## Overview Soniox speech translation uses the same supported-language set as Soniox Speech-to-Text and Text-to-Speech APIs. This language list is relevant to each translation workflow: * **Real-time speech-to-text translation:** live audio over WebSocket. * **Async speech-to-text translation:** recorded audio via REST. * **Real-time speech-to-speech translation:** STT + TTS pipeline. *** ## Translation pairs Translation supports one-way and two-way modes across 3600+ Soniox-supported language pairs. | Language | ISO Code | | ----------- | -------- | | Afrikaans | af | | Albanian | sq | | Arabic | ar | | Azerbaijani | az | | Basque | eu | | Belarusian | be | | Bengali | bn | | Bosnian | bs | | Bulgarian | bg | | Catalan | ca | | Chinese | zh | | Croatian | hr | | Czech | cs | | Danish | da | | Dutch | nl | | English | en | | Estonian | et | | Finnish | fi | | French | fr | | Galician | gl | | German | de | | Greek | el | | Gujarati | gu | | Hebrew | he | | Hindi | hi | | Hungarian | hu | | Indonesian | id | | Italian | it | | Japanese | ja | | Kannada | kn | | Kazakh | kk | | Korean | ko | | Latvian | lv | | Lithuanian | lt | | Macedonian | mk | | Malay | ms | | Malayalam | ml | | Marathi | mr | | Norwegian | no | | Persian | fa | | Polish | pl | | Portuguese | pt | | Punjabi | pa | | Romanian | ro | | Russian | ru | | Serbian | sr | | Slovak | sk | | Slovenian | sl | | Spanish | es | | Swahili | sw | | Swedish | sv | | Tagalog | tl | | Tamil | ta | | Telugu | te | | Thai | th | | Turkish | tr | | Ukrainian | uk | | Urdu | ur | | Vietnamese | vi | | Welsh | cy | # Get started URL: /tts/get-started Learn how to generate speech with the Soniox Text-to-Speech API. ## Learn how to use the Soniox Text-to-Speech API in minutes Soniox Text-to-Speech is built for the hardest parts of speech generation. It delivers native-speaker-quality speech in 60+ languages, with hallucination-free output and accurate pronunciation of alphanumerics such as phone numbers, email addresses, and IDs. Soniox TTS is optimized for ultra-low latency and can start generating speech from the first few words, before the full sentence is available. It is available through WebSocket streaming and request-response generation over REST. Use this guide to run your first Text-to-Speech request. ### Get API key Create a [Soniox account](https://console.soniox.com/signup) and log in to the [Console](https://console.soniox.com) to get your API key. API keys are created per project. In the Console, go to **My First Project** and click **API Keys** to generate one. Export it as an environment variable (replace with your key): ```sh title="Terminal" export SONIOX_API_KEY= ``` ### Get examples Clone the official examples repo: ```sh title="Terminal" git clone https://github.com/soniox/soniox_examples cd soniox_examples/text_to_speech ``` ### Run examples {/* NOTE: Empty tag is needed so code block renders correctly */}
```sh title="Terminal" cd python_sdk # Set up environment python3 -m venv venv source venv/bin/activate pip install -r requirements.txt # Real-time TTS example (WebSocket) python soniox_sdk_realtime.py --line "Hello from Soniox realtime Text-to-Speech." # REST TTS example python soniox_sdk_rest.py --text "Hello from Soniox REST Text-to-Speech." ``` {/* NOTE: Empty tag is needed so code block renders correctly */}
```sh title="Terminal" cd nodejs_sdk # Install dependencies npm install # Real-time TTS example (WebSocket) node soniox_sdk_realtime.js --line "Hello from Soniox realtime Text-to-Speech." # REST TTS example node soniox_sdk_rest.js --text "Hello from Soniox REST Text-to-Speech." ``` {/* NOTE: Empty tag is needed so code block renders correctly */}
```sh title="Terminal" cd python # Set up environment python3 -m venv venv source venv/bin/activate pip install -r requirements.txt # Real-time TTS example (WebSocket) python soniox_realtime.py --line "Hello from Soniox websocket Text-to-Speech." # REST TTS example python soniox_rest.py --text "Hello from Soniox REST Text-to-Speech." ``` {/* NOTE: Empty tag is needed so code block renders correctly */}
```sh title="Terminal" cd nodejs # Install dependencies npm install # Real-time TTS example (WebSocket) node soniox_realtime.js --line "Hello from Soniox websocket Text-to-Speech." # REST TTS example node soniox_rest.js --text "Hello from Soniox REST Text-to-Speech." ``` ## Next steps * **Dive into the [Real-time API](/tts/rt/real-time-generation)** → Stream audio as text arrives. Ideal for voice agents and LLM-driven applications. * **Explore the [REST API](/tts/rest-api/generate-speech)** → Generate full audio files in a single request. Ideal for server-side and batch generation. # Models URL: /tts/models Learn about latest Text-to-Speech models, changelog, and deprecations. Soniox Text-to-Speech is built for the hardest parts of speech generation. It delivers native-speaker-quality speech in 60+ languages, with hallucination-free output and accurate pronunciation of alphanumerics such as phone numbers, email addresses, and IDs. This page lists the currently available models, along with release notes and important updates. *** ## Current models {/*TABLE START */} {/* NOTE: Width is set so that we have maximum of 2 lines in 'Example' column. */} {/* NOTE: Font size is set so the table doesn't look "too big".*/}
|
Model
| {" "}
Type
| Status | | -------------------------------------------- | ----------------------------------------------- | ------------------------------------------------ | | **tts-rt-v2** | Real-time | **Active** | | **tts-rt-v1** | Real-time | **Deprecated**
Will be removed Aug 31, 2026 |
{/* TABLE END */} *** ## Aliases Aliases provide a stable reference so you don’t need to change your code when newer versions are released. | Alias | Points to | Notes | | --------------------- | ----------- | ----- | | **tts-rt-v1-preview** | `tts-rt-v1` | | *** ## Changelog ### Aug 11, 2026 #### Overview Soniox TTS v2 is now generally available. * `tts-rt-v2` is available to all API customers and deployed in all Soniox regions: US, EU, JP, and IN. #### Key improvements * Natural, expressive speech across more than 60 languages * Direct control over emotion, delivery, and vocal reactions through audio tags * High-fidelity voice cloning * Natural language switching within the same sentence * Even more precise pronunciation of names, terminology, numbers, codes, addresses, and identifiers * Reduced silence between sentences and punctuation for more responsive conversations #### API compatibility The `tts-rt-v2` model is fully compatible with the existing `tts-rt-v1` model and Soniox Text-to-Speech API. To upgrade, simply replace the model name in your API request: `{ "model": "tts-rt-v2" }` #### Deprecation notice The `tts-rt-v1` model will be removed on August 31, 2026. After August 31, 2026, requests using `tts-rt-v1` will automatically route to `tts-rt-v2` with no service interruption and no API changes required. ### April 29, 2026 #### Overview Soniox TTS is now generally available. * The preview model `tts-rt-v1-preview` is now available as the production model `tts-rt-v1`. * `tts-rt-v1` is available to all API customers and deployed in all Soniox regions: US, EU, and JP. * For backward compatibility, `tts-rt-v1-preview` now points to `tts-rt-v1` with no service interruption. We recommend updating your API requests to use `tts-rt-v1`. ### April 23, 2026 #### Overview `tts-rt-v1-preview` is the first Soniox Text-to-Speech model, released in preview to gather developer feedback and guide further improvements before general availability. #### Key capabilities * Native-speaker-quality speech in 60+ languages * Hallucination-free generation, with no invented words, dropped content, or unexpected substitutions * Accurate rendering of alphanumerics such as email addresses, phone numbers, street addresses, IDs, and codes * Streaming generation before the sentence ends for ultra-low-latency voice systems * Multiple voices that work across all supported languages * Configurable audio formats, sample rates, and bitrates * Support for both WebSocket and REST APIs # SDKs URL: /sdk Official Soniox SDKs for Speech-to-Text and Text-to-Speech across Python, Node.js, browser, React, and React Native. import { LinkCards } from "@/components/link-card"; Soniox SDKs give you fully typed access to our STT and TTS REST and real-time APIs across the languages and runtimes you already use. Pick the SDK that matches your stack to get started quickly with transcription, translation, and speech generation, without writing low-level WebSocket or REST plumbing. # React Native SDK URL: /sdk/react-native-SDK Build speech-to-text workflows in React Native with real-time API. Soniox [React SDK](/sdk/react-SDK) works with React Native and Expo out of the box, providing the same hooks for real-time speech-to-text. It lets you: * Capture audio from the device microphone with a single hook * Stream audio to Soniox in real time * Receive transcription and translation results as reactive state ## Quickstart ### Install Install via your preferred package manager: ```bash tab npm install @soniox/react @soniox/client ``` ```bash tab yarn add @soniox/react @soniox/client ``` ```bash tab pnpm add @soniox/react @soniox/client ``` ```bash tab bun add @soniox/react @soniox/client ``` ### Set up your temporary API key endpoint In client environments (browser, mobile app, React Native, etc.), you don't want to expose your API key. Create a temporary API key endpoint on your server and use it to issue short-lived keys for the client. Read more about using temporary API keys with the [React SDK](/sdk/react-SDK#set-up-your-temporary-api-key-endpoint). ### Create a custom audio source Wrap any RN audio streaming library (e.g. `@siteed/expo-audio-studio`) with the `AudioSource` interface to stream PCM audio chunks to Soniox ```ts import type { AudioSource, AudioSourceHandlers } from "@soniox/client"; class MyAudioSource implements AudioSource { private handlers: AudioSourceHandlers | null = null; async start(handlers: AudioSourceHandlers): Promise { this.handlers = handlers; // Start your audio capture here. // Call handlers.onData(chunk) with each audio chunk as an ArrayBuffer. // Call handlers.onError(error) if something goes wrong. // Call handlers.onMuted?.() / handlers.onUnmuted?.() when the mic is // muted or unmuted externally (e.g. OS-level, hardware switch). } stop(): void { // Stop audio capture and release resources. this.handlers = null; } } ``` ### Create your first real-time session The core hooks (e.g. [`useRecording`](/sdk/react-SDK/stt/realtime-transcription#userecording)) are platform-agnostic. To use them in React Native, provide a custom [`AudioSource`](/sdk/web-SDK/reference/types#audiosource) that streams PCM audio chunks ```tsx import { useRef } from "react"; import { SonioxProvider, useRecording } from "@soniox/react"; import { MyAudioSource } from "./MyAudioSource"; // Fetch a temporary API key from your server endpoint. async function fetchConfig() { const res = await fetch("/api/soniox-temporary-key", { method: "POST" }); const { api_key } = await res.json(); return { api_key }; } function App() { return ( // Wrap your app with a SonioxProvider. `permissions={null}` disables the // default browser permission resolver — not applicable on React Native. ); } function Transcription() { // Instantiate the audio source const sourceRef = useRef(null); if (sourceRef.current === null) { sourceRef.current = new MyAudioSource(); } // Create a recording session const { state, isActive, finalText, partialText, start, stop } = useRecording({ model: "stt-rt-v5", audio_format: "pcm_s16le", sample_rate: 16000, num_channels: 1, source: sourceRef.current, }); return ( {finalText} {partialText} {isActive ? ( ) : ( )}
); } ``` Learn more about [Real-time transcription](/sdk/react-SDK/stt/realtime-transcription) ### Generate your first speech ```tsx import { useRef } from "react"; import { useTts } from "@soniox/react"; async function fetchTtsConfig() { // Your server should issue a temporary key with usage_type: 'tts_rt' const res = await fetch("/tts-tmp-key"); const { api_key } = await res.json(); return { api_key }; } function SpeakButton() { const audioRef = useRef([]); const { speak, isSpeaking } = useTts({ config: fetchTtsConfig, voice: "Adrian", audio_format: "wav", onAudio: (chunk) => audioRef.current.push(chunk), onTerminated: () => { const blob = new Blob(audioRef.current, { type: "audio/wav" }); new Audio(URL.createObjectURL(blob)).play(); audioRef.current = []; }, }); return ( ); } ``` Learn more about [Real-time speech generation](/sdk/react-SDK/tts/realtime-speech-generation). ## Next steps * [Real-time transcription](/sdk/react-SDK/stt/realtime-transcription) * [Real-time speech generation](/sdk/react-SDK/tts/realtime-speech-generation) * [Full SDK reference](/sdk/react-SDK/reference) ## Package links * [GitHub repository](https://github.com/soniox/soniox-js) * [NPM package](https://www.npmjs.com/package/@soniox/react) # Web SDK URL: /sdk/web-SDK Build speech-to-text and text-to-speech workflows in browser with real-time APIs. import { LinkCards } from "@/components/link-card"; Soniox [Web SDK](https://www.npmjs.com/package/@soniox/client) is the official JavaScript/TypeScript SDK for using the Soniox [Real-time API](/api-reference/stt/websocket-api) and [Text-to-Speech API](/api-reference/tts/websocket-api) directly in the browser. It lets you: * Capture audio from the user's microphone * Stream audio to Soniox in real time * Receive transcription and translation results instantly * Generate speech from text over HTTP or WebSocket ## Quickstart ### Install Install via your preferred package manager: ```bash tab npm install @soniox/client ``` ```bash tab yarn add @soniox/client ``` ```bash tab pnpm add @soniox/client ``` ```bash tab bun add @soniox/client ``` ### Set up your temporary API key endpoint In client environment (browser, mobile app, React Native, etc.), you don't want to expose your API key to the client. For this reason, you can create a temporary API key endpoint on your server and use it to issue temporary API keys for the client. For example, you can use our [Node SDK](/sdk/node-SDK) to create a temporary API key endpoint. ```ts import express from 'express'; import { SonioxNodeClient } from '@soniox/node'; const app = express(); const client = new SonioxNodeClient(); // reads SONIOX_API_KEY from env // Create a temporary API key endpoint app.post('/tmp-key', async (_req, res) => { try { const { api_key, expires_at } = await client.auth.createTemporaryKey({ usage_type: 'transcribe_websocket', expires_in_seconds: 300, // 1..3600 }); res.json({ api_key, expires_at }); } catch (err) { res.status(500).json({ error: err instanceof Error ? err.message : 'Failed to create temporary key' }); } }); app.listen(3000, () => { console.log('Server listening on http://localhost:3000'); }); ``` Read more about our [Node SDK](/sdk/node-SDK) and [Temporary API keys](/guides/temporary-api-keys). ### Create your first real-time session ```ts import { SonioxClient } from "@soniox/client"; // Create a Soniox client const client = new SonioxClient({ // Pass a function that fetches a temporary API key (and optional region / URL overrides) // from your server for each new session. config: async () => { const res = await fetch("/tmp-key", { method: "POST" }); const { api_key } = await res.json(); return { api_key }; }, }); // Create a recording session const recording = client.realtime.record({ model: "stt-rt-v5" }); // Listen for transcription results recording.on("result", (result) => { const text = result.tokens.map((t) => t.text).join(""); if (text) console.log(text); }); // Listen for errors recording.on("error", (err) => console.error("Error:", err)); // Call this from your UI (e.g. a Stop button) to end gracefully and wait for final results. async function stopRecording() { await recording.stop(); } ``` Learn more about [Real-time transcription](/sdk/web-SDK/stt/realtime-transcription) ### Generate your first speech See [Real-time speech generation](/sdk/web-SDK/tts/realtime-speech-generation#set-up-your-temporary-api-key-endpoint) for an example server endpoint that issues a temporary key with `usage_type: 'tts_rt'`. ```ts import { SonioxClient } from "@soniox/client"; const client = new SonioxClient({ config: async () => { const res = await fetch("/tts-tmp-key"); const { api_key } = await res.json(); return { api_key }; // temporary key with usage_type: 'tts_rt' }, }); const stream = await client.realtime.tts({ voice: "Adrian", audio_format: "wav" }); stream.sendText("Hello from Soniox Web SDK text-to-speech.", { end: true }); const chunks: Uint8Array[] = []; for await (const chunk of stream) chunks.push(chunk); const blob = new Blob(chunks, { type: "audio/wav" }); await new Audio(URL.createObjectURL(blob)).play(); ``` Learn more about [Real-time speech generation](/sdk/web-SDK/tts/realtime-speech-generation) and [REST speech generation](/sdk/web-SDK/tts/rest-speech-generation). ## Next steps * [Real-time transcription](/sdk/web-SDK/stt/realtime-transcription) * [Real-time speech generation](/sdk/web-SDK/tts/realtime-speech-generation) * [REST speech generation](/sdk/web-SDK/tts/rest-speech-generation) * [Full SDK reference](/sdk/web-SDK/reference) ## Package links * [GitHub repository](https://github.com/soniox/soniox-js) * [NPM package](https://www.npmjs.com/package/@soniox/client) # Delete file URL: /api-reference/stt/files/delete_file Permanently deletes specified file. If a transcription that has not started processing yet still references the file, that transcription fails with `file_not_found`, so delete the file only after the transcription reaches `completed` or `error`. ## Delete file **Endpoint:** `DELETE /v1/files/{file_id}` Permanently deletes specified file. If a transcription that has not started processing yet still references the file, that transcription fails with `file_not_found`, so delete the file only after the transcription reaches `completed` or `error`. ### Parameters * `file_id` (path) (Required): ### Responses * **204**: File deleted. * **401**: Authentication error. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **404**: File not found. Error types: * `file_not_found`: No file with this ID exists in the project the API key is scoped to (it may have been deleted, the ID may be wrong, or it may belong to a different project). Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **429**: Rate / capacity limit exceeded. Error types: * `limit_exceeded`: The caller hit a per-minute request rate or other capacity limit. The `message` describes which limit was hit. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **500**: Internal server error. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` # Get file URL: /api-reference/stt/files/get_file Retrieve metadata for an uploaded file. ## Get file **Endpoint:** `GET /v1/files/{file_id}` Retrieve metadata for an uploaded file. ### Parameters * `file_id` (path) (Required): ### Responses * **200**: File metadata. Example (JSON): ```json { "client_reference_id": "some_internal_id", "created_at": "2024-11-26T00:00:00Z", "filename": "example.mp3", "id": "84c32fc6-4fb5-4e7a-b656-b5ec70493753", "size": 123456 } ``` Schema (YAML Structural Definition): ```yaml description: File metadata. properties: id: description: Unique identifier of the file. format: uuid type: string filename: description: Name of the file. type: string size: description: Size of the file in bytes. type: integer created_at: description: UTC timestamp indicating when the file was uploaded. format: date-time type: string client_reference_id: anyOf: - type: string - type: 'null' description: Tracking identifier string. required: - id - filename - size - created_at type: object ``` * **401**: Authentication error. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **404**: File not found. Error types: * `file_not_found`: No file with this ID exists in the project the API key is scoped to (it may have been deleted, the ID may be wrong, or it may belong to a different project). Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **429**: Rate / capacity limit exceeded. Error types: * `limit_exceeded`: The caller hit a per-minute request rate or other capacity limit. The `message` describes which limit was hit. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **500**: Internal server error. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` # Get files URL: /api-reference/stt/files/get_files Retrieves list of uploaded files. ## Get files **Endpoint:** `GET /v1/files` Retrieves list of uploaded files. ### Parameters * `limit` (query): Maximum number of files to return. * `cursor` (query): Pagination cursor for the next page of results. ### Responses * **200**: List of files. Example (JSON): ```json { "files": [ { "created_at": "2024-11-26T00:00:00Z", "filename": "example.mp3", "id": "84c32fc6-4fb5-4e7a-b656-b5ec70493753", "size": 123456 } ], "next_page_cursor": "cursor_or_null" } ``` Schema (YAML Structural Definition): ```yaml description: A list of files. properties: files: description: List of uploaded files. items: description: File metadata. example: client_reference_id: some_internal_id created_at: '2024-11-26T00:00:00Z' filename: example.mp3 id: 84c32fc6-4fb5-4e7a-b656-b5ec70493753 size: 123456 properties: id: description: Unique identifier of the file. format: uuid type: string filename: description: Name of the file. type: string size: description: Size of the file in bytes. type: integer created_at: description: UTC timestamp indicating when the file was uploaded. format: date-time type: string client_reference_id: anyOf: - type: string - type: 'null' description: Tracking identifier string. required: - id - filename - size - created_at type: object type: array next_page_cursor: anyOf: - type: string - type: 'null' description: >- A pagination token that references the next page of results. When more data is available, this field contains a value to pass in the cursor parameter of a subsequent request. When null, no additional results are available. required: - files type: object ``` * **400**: Invalid request. Error types: * `invalid_cursor`: The `cursor` parameter is invalid. Omit `cursor` to start pagination from the beginning. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **401**: Authentication error. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **429**: Rate / capacity limit exceeded. Error types: * `limit_exceeded`: The caller hit a per-minute request rate, total file count, or other capacity limit. The `message` describes which limit was hit. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **500**: Internal server error. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` # Get files count URL: /api-reference/stt/files/get_files_count Returns the total number of files, split by source. ## Get files count **Endpoint:** `GET /v1/files/count` Returns the total number of files, split by source. ### Responses * **200**: Total number of files, split by source. Schema (YAML Structural Definition): ```yaml properties: playground: description: Number of files uploaded via the Playground. type: integer public_api: description: Number of files uploaded via Public API. type: integer total: description: Total number of files across all sources. type: integer required: - total - public_api - playground type: object ``` * **401**: Authentication error. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **429**: Rate / capacity limit exceeded. Error types: * `limit_exceeded`: The caller hit a per-minute request rate or other capacity limit. The `message` describes which limit was hit. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **500**: Internal server error. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` # Upload file URL: /api-reference/stt/files/upload_file Uploads a new file. ## Upload file **Endpoint:** `POST /v1/files` Uploads a new file. ### Request Body Content-Type: `multipart/form-data` (Required) Schema (YAML Structural Definition): ```yaml type: object properties: client_reference_id: anyOf: - maxLength: 256 type: string - type: 'null' description: Optional tracking identifier string. Does not need to be unique. file: description: >- The file to upload. Original file name will be used unless a custom filename is provided. format: binary type: string required: - file ``` ### Responses * **201**: Uploaded file. Example (JSON): ```json { "client_reference_id": "some_internal_id", "created_at": "2024-11-26T00:00:00Z", "filename": "example.mp3", "id": "84c32fc6-4fb5-4e7a-b656-b5ec70493753", "size": 123456 } ``` Schema (YAML Structural Definition): ```yaml description: File metadata. properties: id: description: Unique identifier of the file. format: uuid type: string filename: description: Name of the file. type: string size: description: Size of the file in bytes. type: integer created_at: description: UTC timestamp indicating when the file was uploaded. format: date-time type: string client_reference_id: anyOf: - type: string - type: 'null' description: Tracking identifier string. required: - id - filename - size - created_at type: object ``` * **400**: Invalid request. Error types: * `invalid_request`: One or more parts of the multipart body are missing or invalid (missing `file`, filename or `client_reference_id` too long, malformed multipart body), or the uploaded file exceeds the per-file upload limit. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **401**: Authentication error. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **429**: Rate / capacity limit exceeded. Error types: * `limit_exceeded`: The caller hit a per-minute request rate, the total file count or total file size cap (organization or project), or this upload would exceed those caps. The `message` describes which limit was hit. Delete unused files via `DELETE /v1/files/{id}` or request a higher limit in the Soniox Console. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **500**: Internal server error. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` # Create transcription URL: /api-reference/stt/transcriptions/create_transcription Creates a new transcription. ## Create transcription **Endpoint:** `POST /v1/transcriptions` Creates a new transcription. ### Request Body Content-Type: `application/json` (Required) Schema (YAML Structural Definition): ```yaml properties: model: description: Speech-to-text model to use for the transcription. maxLength: 32 type: string audio_url: anyOf: - maxLength: 4096 pattern: ^https?://[^\s]+$ type: string - type: 'null' description: >- URL of the audio file to transcribe. Cannot be specified if `file_id` is specified. file_id: anyOf: - format: uuid type: string - type: 'null' description: >- ID of the uploaded file to transcribe. Cannot be specified if `audio_url` is specified. Keep the file until the transcription reaches `completed` or `error`; deleting it earlier fails the transcription with `file_not_found`. language_hints: anyOf: - items: maxLength: 10 type: string maxItems: 100 type: array - type: 'null' description: >- Expected languages in the audio. If not specified, languages are automatically detected. language_hints_strict: anyOf: - type: boolean - type: 'null' description: When `true`, the model will rely more on language hints. enable_speaker_diarization: anyOf: - type: boolean - type: 'null' description: >- When `true`, speakers are identified and separated in the transcription output. enable_language_identification: anyOf: - type: boolean - type: 'null' description: When `true`, language is detected for each part of the transcription. translation: anyOf: - properties: type: enum: - one_way - two_way type: string target_language: anyOf: - type: string - type: 'null' language_a: anyOf: - type: string - type: 'null' language_b: anyOf: - type: string - type: 'null' required: - type type: object - type: 'null' description: Translation configuration. context: anyOf: - properties: general: anyOf: - items: properties: key: description: Item key (e.g. "Domain"). type: string value: description: Item value (e.g. "medicine"). type: string required: - key - value type: object type: array - type: 'null' description: General context items. text: anyOf: - type: string - type: 'null' description: Text context. terms: anyOf: - items: type: string type: array - type: 'null' description: Terms that might occur in speech. translation_terms: anyOf: - items: properties: source: description: Source term. type: string target: description: Target term to translate to. type: string required: - source - target type: object type: array - type: 'null' description: >- Hints how to translate specific terms. Ignored if translation is not enabled. type: object - type: string - type: 'null' description: >- Additional context to improve transcription accuracy and formatting of specialized terms. webhook_url: anyOf: - maxLength: 256 pattern: ^https?://[^\s]+$ type: string - type: 'null' description: >- URL to receive webhook notifications when transcription is completed or fails. webhook_auth_header_name: anyOf: - maxLength: 256 type: string - type: 'null' description: Name of the authentication header sent with webhook notifications. webhook_auth_header_value: anyOf: - maxLength: 256 type: string - type: 'null' description: Authentication header value sent with webhook notifications. client_reference_id: anyOf: - maxLength: 256 type: string - type: 'null' description: Optional tracking identifier string. Does not need to be unique. required: - model type: object ``` ### Responses * **201**: Created transcription. Example (JSON): ```json { "audio_duration_ms": 0, "audio_url": "https://soniox.com/media/examples/coffee_shop.mp3", "client_reference_id": "some_internal_id", "created_at": "2024-11-26T00:00:00Z", "error_message": null, "error_type": null, "file_id": null, "filename": "coffee_shop.mp3", "id": "73d4357d-cad2-4338-a60d-ec6f2044f721", "language_hints": [ "en", "fr" ], "model": "stt-async-preview", "status": "queued", "webhook_auth_header_name": "Authorization", "webhook_auth_header_value": "******************", "webhook_status_code": null, "webhook_url": "https://example.com/webhook" } ``` Schema (YAML Structural Definition): ```yaml description: A transcription. properties: id: description: Unique identifier for the transcription request. format: uuid type: string status: description: Transcription status. enum: - queued - processing - completed - error type: string created_at: description: UTC timestamp indicating when the transcription was created. format: date-time type: string model: description: Speech-to-text model used for the transcription. type: string audio_url: anyOf: - type: string - type: 'null' description: URL of the file being transcribed. file_id: anyOf: - format: uuid type: string - type: 'null' description: ID of the file being transcribed. filename: description: Name of the file being transcribed. type: string language_hints: anyOf: - items: type: string type: array - type: 'null' description: >- Expected languages in the audio. If not specified, languages are automatically detected. enable_speaker_diarization: description: >- When `true`, speakers are identified and separated in the transcription output. type: boolean enable_language_identification: description: When `true`, language is detected for each part of the transcription. type: boolean audio_duration_ms: anyOf: - type: integer - type: 'null' description: >- Duration of the audio in milliseconds. Only available after processing begins. error_type: anyOf: - type: string - type: 'null' description: >- Error type if transcription failed. `null` for successful or in-progress transcriptions. error_message: anyOf: - type: string - type: 'null' description: >- Error message if transcription failed. `null` for successful or in-progress transcriptions. webhook_url: anyOf: - type: string - type: 'null' description: >- URL to receive webhook notifications when transcription is completed or fails. webhook_auth_header_name: anyOf: - type: string - type: 'null' description: Name of the authentication header sent with webhook notifications. webhook_auth_header_value: anyOf: - type: string - type: 'null' description: >- Authentication header value. Always returned masked as `******************`. webhook_status_code: anyOf: - type: integer - type: 'null' description: >- HTTP status code received from your server when webhook was delivered. `null` if not yet sent. client_reference_id: anyOf: - type: string - type: 'null' description: Tracking identifier string. required: - id - status - created_at - model - filename - enable_speaker_diarization - enable_language_identification type: object ``` * **400**: Invalid request. Error types: * `invalid_request`: One or more request body fields are missing or invalid (model, audio source, language hints, translation config, webhook config, etc.). Inspect `validation_errors`. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **401**: Authentication error. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **402**: Balance or budget exhausted. Error types: * `organization_balance_exhausted`: The organization's prepaid balance has dropped to zero. Top up at [https://console.soniox.com/org/billing/overview](https://console.soniox.com/org/billing/overview) or enable autopay. * `organization_monthly_budget_exhausted`: The organization has hit its configured monthly budget cap. Raise the cap at [https://console.soniox.com/org/limits](https://console.soniox.com/org/limits), or wait for the month to roll over. * `project_monthly_budget_exhausted`: The project has hit its configured monthly budget cap. Raise the cap at [https://console.soniox.com/org/projects/limits](https://console.soniox.com/org/projects/limits), or wait for the month to roll over. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **429**: Rate / capacity limit exceeded. Error types: * `limit_exceeded`: The caller hit a per-minute request rate, total transcription count, or pending transcription count limit (organization or project). The `message` describes which limit was hit. Delete completed transcriptions via `DELETE /v1/transcriptions/{id}` or request a higher limit in the Soniox Console. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **500**: Internal server error. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` # Delete transcription URL: /api-reference/stt/transcriptions/delete_transcription Permanently deletes a transcription. Files uploaded through the Files API are not deleted; use the delete file endpoint to remove them. Cannot delete transcriptions that are currently processing. ## Delete transcription **Endpoint:** `DELETE /v1/transcriptions/{transcription_id}` Permanently deletes a transcription. Files uploaded through the Files API are not deleted; use the delete file endpoint to remove them. Cannot delete transcriptions that are currently processing. ### Parameters * `transcription_id` (path) (Required): ### Responses * **204**: Transcription deleted. * **401**: Authentication error. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **404**: Transcription not found. Error types: * `transcription_not_found`: No transcription with this ID exists in the project the API key is scoped to (it may have been deleted, the ID may be wrong, or it may belong to a different project). Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **409**: Invalid transcription state. Error types: * `transcription_invalid_state`: The transcription cannot be deleted in its current state — it is still processing. Wait until `status` reaches `completed` or `error` and retry. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **429**: Rate / capacity limit exceeded. Error types: * `limit_exceeded`: The caller hit a per-minute request rate or other capacity limit. The `message` describes which limit was hit. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **500**: Internal server error. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` # Get transcription URL: /api-reference/stt/transcriptions/get_transcription Retrieves detailed information about a specific transcription. ## Get transcription **Endpoint:** `GET /v1/transcriptions/{transcription_id}` Retrieves detailed information about a specific transcription. ### Parameters * `transcription_id` (path) (Required): ### Responses * **200**: Transcription details. Example (JSON): ```json { "audio_duration_ms": 0, "audio_url": "https://soniox.com/media/examples/coffee_shop.mp3", "client_reference_id": "some_internal_id", "created_at": "2024-11-26T00:00:00Z", "error_message": null, "error_type": null, "file_id": null, "filename": "coffee_shop.mp3", "id": "73d4357d-cad2-4338-a60d-ec6f2044f721", "language_hints": [ "en", "fr" ], "model": "stt-async-preview", "status": "queued", "webhook_auth_header_name": "Authorization", "webhook_auth_header_value": "******************", "webhook_status_code": null, "webhook_url": "https://example.com/webhook" } ``` Schema (YAML Structural Definition): ```yaml description: A transcription. properties: id: description: Unique identifier for the transcription request. format: uuid type: string status: description: Transcription status. enum: - queued - processing - completed - error type: string created_at: description: UTC timestamp indicating when the transcription was created. format: date-time type: string model: description: Speech-to-text model used for the transcription. type: string audio_url: anyOf: - type: string - type: 'null' description: URL of the file being transcribed. file_id: anyOf: - format: uuid type: string - type: 'null' description: ID of the file being transcribed. filename: description: Name of the file being transcribed. type: string language_hints: anyOf: - items: type: string type: array - type: 'null' description: >- Expected languages in the audio. If not specified, languages are automatically detected. enable_speaker_diarization: description: >- When `true`, speakers are identified and separated in the transcription output. type: boolean enable_language_identification: description: When `true`, language is detected for each part of the transcription. type: boolean audio_duration_ms: anyOf: - type: integer - type: 'null' description: >- Duration of the audio in milliseconds. Only available after processing begins. error_type: anyOf: - type: string - type: 'null' description: >- Error type if transcription failed. `null` for successful or in-progress transcriptions. error_message: anyOf: - type: string - type: 'null' description: >- Error message if transcription failed. `null` for successful or in-progress transcriptions. webhook_url: anyOf: - type: string - type: 'null' description: >- URL to receive webhook notifications when transcription is completed or fails. webhook_auth_header_name: anyOf: - type: string - type: 'null' description: Name of the authentication header sent with webhook notifications. webhook_auth_header_value: anyOf: - type: string - type: 'null' description: >- Authentication header value. Always returned masked as `******************`. webhook_status_code: anyOf: - type: integer - type: 'null' description: >- HTTP status code received from your server when webhook was delivered. `null` if not yet sent. client_reference_id: anyOf: - type: string - type: 'null' description: Tracking identifier string. required: - id - status - created_at - model - filename - enable_speaker_diarization - enable_language_identification type: object ``` * **401**: Authentication error. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **404**: Transcription not found. Error types: * `transcription_not_found`: No transcription with this ID exists in the project the API key is scoped to (it may have been deleted, the ID may be wrong, or it may belong to a different project). Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **429**: Rate / capacity limit exceeded. Error types: * `limit_exceeded`: The caller hit a per-minute request rate or other capacity limit. The `message` describes which limit was hit. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **500**: Internal server error. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` # Get transcription transcript URL: /api-reference/stt/transcriptions/get_transcription_transcript Retrieves the full transcript text and detailed tokens for a completed transcription. Only available for successfully completed transcriptions. ## Get transcription transcript **Endpoint:** `GET /v1/transcriptions/{transcription_id}/transcript` Retrieves the full transcript text and detailed tokens for a completed transcription. Only available for successfully completed transcriptions. ### Parameters * `transcription_id` (path) (Required): ### Responses * **200**: Transcription transcript. Example (JSON): ```json { "id": "19b6d61d-02db-4c25-bc71-b4094dc310c8", "text": "Hello", "tokens": [ { "confidence": 0.95, "end_ms": 90, "start_ms": 10, "text": "Hel" }, { "confidence": 0.98, "end_ms": 160, "start_ms": 110, "text": "lo" } ] } ``` Schema (YAML Structural Definition): ```yaml description: The transcription text. properties: id: description: Unique identifier of the transcription this transcript belongs to. format: uuid type: string text: description: Complete transcribed text content. type: string tokens: description: List of detailed token information with timestamps and metadata. items: description: The transcript token. example: confidence: 0.95 end_ms: 90 start_ms: 10 text: Hel properties: text: description: Token text content. type: string start_ms: description: Start time of the token in milliseconds. type: integer end_ms: description: End time of the token in milliseconds. type: integer confidence: description: Confidence score of the token, between 0.0 and 1.0. type: number speaker: anyOf: - type: string - type: 'null' description: >- Speaker identifier. Only present when speaker diarization is enabled. language: anyOf: - type: string - type: 'null' description: >- Detected language code for this token. Only present when language identification is enabled. is_audio_event: anyOf: - type: boolean - type: 'null' description: >- Boolean indicating if this token represents an audio event. Only present when audio event detection is enabled. translation_status: anyOf: - type: string - type: 'null' description: >- Translation status ("none", "original" or "translation"). Only when if translation is enabled. required: - text - start_ms - end_ms - confidence type: object type: array required: - id - text - tokens type: object ``` * **401**: Authentication error. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **404**: Transcription not found. Error types: * `transcription_not_found`: No transcription with this ID exists in the project the API key is scoped to (it may have been deleted, the ID may be wrong, or it may belong to a different project). Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **409**: Invalid transcription state. The transcription exists but the transcript cannot be returned in its current state. Error types: * `transcription_invalid_state`. The `message` indicates which sub-case applies: * **Not completed yet** — transcription is still queued, downloading, or transcribing. Poll `GET /transcriptions/{id}` until `status` is `completed`, or configure a webhook on the transcription. * **Failed** — transcription ended in `failed` state. Inspect `error_type` / `error_message` on `GET /transcriptions/{id}` for the failure reason. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **429**: Rate / capacity limit exceeded. Error types: * `limit_exceeded`: The caller hit a per-minute request rate or other capacity limit. The `message` describes which limit was hit. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **500**: Internal server error. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` # Get transcriptions URL: /api-reference/stt/transcriptions/get_transcriptions Retrieves list of transcriptions. ## Get transcriptions **Endpoint:** `GET /v1/transcriptions` Retrieves list of transcriptions. ### Parameters * `limit` (query): Maximum number of transcriptions to return. * `cursor` (query): Pagination cursor for the next page of results. ### Responses * **200**: A list of transcriptions. Schema (YAML Structural Definition): ```yaml properties: transcriptions: description: List of transcriptions. items: description: A transcription. example: audio_duration_ms: 0 audio_url: https://soniox.com/media/examples/coffee_shop.mp3 client_reference_id: some_internal_id created_at: '2024-11-26T00:00:00Z' error_message: null error_type: null file_id: null filename: coffee_shop.mp3 id: 73d4357d-cad2-4338-a60d-ec6f2044f721 language_hints: - en - fr model: stt-async-preview status: queued webhook_auth_header_name: Authorization webhook_auth_header_value: '******************' webhook_status_code: null webhook_url: https://example.com/webhook properties: id: description: Unique identifier for the transcription request. format: uuid type: string status: description: Transcription status. enum: - queued - processing - completed - error type: string created_at: description: UTC timestamp indicating when the transcription was created. format: date-time type: string model: description: Speech-to-text model used for the transcription. type: string audio_url: anyOf: - type: string - type: 'null' description: URL of the file being transcribed. file_id: anyOf: - format: uuid type: string - type: 'null' description: ID of the file being transcribed. filename: description: Name of the file being transcribed. type: string language_hints: anyOf: - items: type: string type: array - type: 'null' description: >- Expected languages in the audio. If not specified, languages are automatically detected. enable_speaker_diarization: description: >- When `true`, speakers are identified and separated in the transcription output. type: boolean enable_language_identification: description: >- When `true`, language is detected for each part of the transcription. type: boolean audio_duration_ms: anyOf: - type: integer - type: 'null' description: >- Duration of the audio in milliseconds. Only available after processing begins. error_type: anyOf: - type: string - type: 'null' description: >- Error type if transcription failed. `null` for successful or in-progress transcriptions. error_message: anyOf: - type: string - type: 'null' description: >- Error message if transcription failed. `null` for successful or in-progress transcriptions. webhook_url: anyOf: - type: string - type: 'null' description: >- URL to receive webhook notifications when transcription is completed or fails. webhook_auth_header_name: anyOf: - type: string - type: 'null' description: Name of the authentication header sent with webhook notifications. webhook_auth_header_value: anyOf: - type: string - type: 'null' description: >- Authentication header value. Always returned masked as `******************`. webhook_status_code: anyOf: - type: integer - type: 'null' description: >- HTTP status code received from your server when webhook was delivered. `null` if not yet sent. client_reference_id: anyOf: - type: string - type: 'null' description: Tracking identifier string. required: - id - status - created_at - model - filename - enable_speaker_diarization - enable_language_identification type: object type: array next_page_cursor: anyOf: - type: string - type: 'null' description: >- A pagination token that references the next page of results. When more data is available, this field contains a value to pass in the cursor parameter of a subsequent request. When null, no additional results are available. required: - transcriptions type: object ``` * **400**: Invalid request. Error types: * `invalid_cursor`: The `cursor` parameter is invalid. Omit `cursor` to start pagination from the beginning. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **401**: Authentication error. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **429**: Rate / capacity limit exceeded. Error types: * `limit_exceeded`: The caller hit a per-minute request rate or other capacity limit. The `message` describes which limit was hit. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **500**: Internal server error. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` # Get transcriptions count URL: /api-reference/stt/transcriptions/get_transcriptions_count Returns the total number of transcriptions, split by request scope. ## Get transcriptions count **Endpoint:** `GET /v1/transcriptions/count` Returns the total number of transcriptions, split by request scope. ### Responses * **200**: Total number of transcriptions, split by request scope. Schema (YAML Structural Definition): ```yaml properties: playground: description: Number of transcriptions created via the Playground. type: integer public_api: description: Number of transcriptions created via Public API. type: integer total: description: Total number of transcriptions across all scopes. type: integer required: - total - public_api - playground type: object ``` * **401**: Authentication error. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **429**: Rate / capacity limit exceeded. Error types: * `limit_exceeded`: The caller hit a per-minute request rate or other capacity limit. The `message` describes which limit was hit. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **500**: Internal server error. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` # Create voice URL: /api-reference/tts/voices/create_voice Uploads a reference audio clip and creates a new voice. ## Create voice **Endpoint:** `POST /v1/voices` Uploads a reference audio clip and creates a new voice. ### Request Body Content-Type: `multipart/form-data` (Required) Schema (YAML Structural Definition): ```yaml type: object properties: name: description: A name for the voice, unique within your project. maxLength: 128 minLength: 1 type: string file: description: The reference audio clip for the voice. format: binary type: string required: - name - file ``` ### Responses * **201**: Created Schema (YAML Structural Definition): ```yaml properties: id: description: Unique identifier of the voice. format: uuid type: string name: description: Name of the voice. type: string filename: description: Original file name of the uploaded audio clip. type: string created_at: description: UTC timestamp indicating when the voice was created. format: date-time type: string models: description: >- Voice status for each available model. A model with status 'not_computed' is not prepared yet (e.g. it was released after the voice was created); call recompute to prepare the voice for it. items: properties: model: description: Name of the model. type: string status: description: Has to be 'ready' for the voice to be usable with this model. enum: - not_computed - processing - ready - failed type: string error_type: anyOf: - type: string - type: 'null' description: >- Machine-readable error category when status is 'failed'. Stable across releases — safe to use in control flow. `null` otherwise. error_message: anyOf: - type: string - type: 'null' description: >- Human-readable error message when status is 'failed' (e.g. the reference audio is too long). `null` otherwise. required: - model - status type: object type: array required: - id - name - filename - created_at - models type: object ``` * **400**: Bad Request Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **401**: Unauthorized Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **409**: Conflict Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **429**: Too Many Requests Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **500**: Internal Server Error Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` # Delete voice URL: /api-reference/tts/voices/delete_voice Permanently deletes the specified voice and its embeddings. ## Delete voice **Endpoint:** `DELETE /v1/voices/{voice_id}` Permanently deletes the specified voice and its embeddings. ### Parameters * `voice_id` (path) (Required): ### Responses * **204**: No Content * **401**: Unauthorized Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **404**: Not Found Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **429**: Too Many Requests Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **500**: Internal Server Error Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` # Get voice URL: /api-reference/tts/voices/get_voice Retrieve metadata for a voice. ## Get voice **Endpoint:** `GET /v1/voices/{voice_id}` Retrieve metadata for a voice. ### Parameters * `voice_id` (path) (Required): ### Responses * **200**: OK Schema (YAML Structural Definition): ```yaml properties: id: description: Unique identifier of the voice. format: uuid type: string name: description: Name of the voice. type: string filename: description: Original file name of the uploaded audio clip. type: string created_at: description: UTC timestamp indicating when the voice was created. format: date-time type: string models: description: >- Voice status for each available model. A model with status 'not_computed' is not prepared yet (e.g. it was released after the voice was created); call recompute to prepare the voice for it. items: properties: model: description: Name of the model. type: string status: description: Has to be 'ready' for the voice to be usable with this model. enum: - not_computed - processing - ready - failed type: string error_type: anyOf: - type: string - type: 'null' description: >- Machine-readable error category when status is 'failed'. Stable across releases — safe to use in control flow. `null` otherwise. error_message: anyOf: - type: string - type: 'null' description: >- Human-readable error message when status is 'failed' (e.g. the reference audio is too long). `null` otherwise. required: - model - status type: object type: array required: - id - name - filename - created_at - models type: object ``` * **401**: Unauthorized Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **404**: Not Found Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **429**: Too Many Requests Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **500**: Internal Server Error Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` # Get voices URL: /api-reference/tts/voices/get_voices Retrieves the list of voices in your project. ## Get voices **Endpoint:** `GET /v1/voices` Retrieves the list of voices in your project. ### Parameters * `limit` (query): Maximum number of voices to return. * `cursor` (query): Pagination cursor for the next page of results. ### Responses * **200**: OK Schema (YAML Structural Definition): ```yaml properties: voices: description: List of voices. items: properties: id: description: Unique identifier of the voice. format: uuid type: string name: description: Name of the voice. type: string filename: description: Original file name of the uploaded audio clip. type: string created_at: description: UTC timestamp indicating when the voice was created. format: date-time type: string models: description: >- Voice status for each available model. A model with status 'not_computed' is not prepared yet (e.g. it was released after the voice was created); call recompute to prepare the voice for it. items: properties: model: description: Name of the model. type: string status: description: Has to be 'ready' for the voice to be usable with this model. enum: - not_computed - processing - ready - failed type: string error_type: anyOf: - type: string - type: 'null' description: >- Machine-readable error category when status is 'failed'. Stable across releases — safe to use in control flow. `null` otherwise. error_message: anyOf: - type: string - type: 'null' description: >- Human-readable error message when status is 'failed' (e.g. the reference audio is too long). `null` otherwise. required: - model - status type: object type: array required: - id - name - filename - created_at - models type: object type: array next_page_cursor: anyOf: - type: string - type: 'null' description: >- A pagination token that references the next page of results. When more data is available, this field contains a value to pass in the cursor parameter of a subsequent request. When null, no additional results are available. required: - voices type: object ``` * **400**: Bad Request Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **401**: Unauthorized Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **429**: Too Many Requests Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **500**: Internal Server Error Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` # Get voices count URL: /api-reference/tts/voices/get_voices_count Returns the total number of voices in your project. ## Get voices count **Endpoint:** `GET /v1/voices/count` Returns the total number of voices in your project. ### Responses * **200**: OK Schema (YAML Structural Definition): ```yaml properties: total: description: Total number of voices in your project. type: integer required: - total type: object ``` * **401**: Unauthorized Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **429**: Too Many Requests Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **500**: Internal Server Error Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` # Recompute voice URL: /api-reference/tts/voices/recompute_voice Prepares the voice for use with available models it is not ready for yet. Use this after a new model is released to make an existing voice usable with it. Models the voice is already prepared for are left unchanged. ## Recompute voice **Endpoint:** `POST /v1/voices/{voice_id}/recompute` Prepares the voice for use with available models it is not ready for yet. Use this after a new model is released to make an existing voice usable with it. Models the voice is already prepared for are left unchanged. ### Request Body Content-Type: `application/json` (Required) Schema (YAML Structural Definition): ```yaml properties: model: anyOf: - type: string - type: 'null' description: >- The model to prepare this voice for. If omitted, the voice is prepared for every available model it is not ready for yet. type: object ``` ### Parameters * `voice_id` (path) (Required): ### Responses * **200**: OK Schema (YAML Structural Definition): ```yaml properties: id: description: Unique identifier of the voice. format: uuid type: string name: description: Name of the voice. type: string filename: description: Original file name of the uploaded audio clip. type: string created_at: description: UTC timestamp indicating when the voice was created. format: date-time type: string models: description: >- Voice status for each available model. A model with status 'not_computed' is not prepared yet (e.g. it was released after the voice was created); call recompute to prepare the voice for it. items: properties: model: description: Name of the model. type: string status: description: Has to be 'ready' for the voice to be usable with this model. enum: - not_computed - processing - ready - failed type: string error_type: anyOf: - type: string - type: 'null' description: >- Machine-readable error category when status is 'failed'. Stable across releases — safe to use in control flow. `null` otherwise. error_message: anyOf: - type: string - type: 'null' description: >- Human-readable error message when status is 'failed' (e.g. the reference audio is too long). `null` otherwise. required: - model - status type: object type: array required: - id - name - filename - created_at - models type: object ``` * **400**: Bad Request Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **401**: Unauthorized Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **404**: Not Found Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **429**: Too Many Requests Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **500**: Internal Server Error Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` # Create temporary API key URL: /api-reference/auth/create_temporary_api_key Creates a short-lived API key for specific temporary use cases. The key will automatically expire after the specified duration. Use `single_use` and `max_session_duration_seconds` to limit how the key can be used by a client. See the [Temporary API keys guide](https://soniox.com/docs/guides/temporary-api-keys) for details. ## Create temporary API key **Endpoint:** `POST /v1/auth/temporary-api-key` Creates a short-lived API key for specific temporary use cases. The key will automatically expire after the specified duration. Use `single_use` and `max_session_duration_seconds` to limit how the key can be used by a client. See the [Temporary API keys guide](https://soniox.com/docs/guides/temporary-api-keys) for details. ### Request Body Content-Type: `application/json` (Required) Example (JSON): ```json { "client_reference_id": "reference_id", "expires_in_seconds": 1800, "max_session_duration_seconds": 120, "single_use": true, "usage_type": "transcribe_websocket" } ``` Schema (YAML Structural Definition): ```yaml properties: usage_type: description: Intended usage of the temporary API key. enum: - transcribe_websocket - tts_rt type: string expires_in_seconds: description: Duration in seconds until the temporary API key expires. maximum: 3600 minimum: 1 type: integer client_reference_id: anyOf: - maxLength: 256 type: string - type: 'null' description: Optional tracking identifier string. Does not need to be unique. single_use: anyOf: - type: boolean - type: 'null' description: If true, the temporary API key can be used only once. max_session_duration_seconds: anyOf: - maximum: 18000 minimum: 1 type: integer - type: 'null' description: >- Maximum connection duration in seconds for WebSocket and TTS HTTP streaming endpoints. If exceeded, the connection will be dropped. If not set, no limit is applied. required: - usage_type - expires_in_seconds type: object ``` ### Responses * **201**: Created temporary API key. Example (JSON): ```json { "api_key": "snx_temp_AcsDGHvigal7tHRzzzqJI7EdJ5CFwk9C0PtXN_s_cUKJ.Oo1TyosFa7b3rgAcXA2bayqBFO7667gXROEu0mH0U4vgvlNzCqVGgTzitabbXlK7FKH-sSy0F1NKI1OOJzQaAw.YPj4oA", "expires_at": "2025-02-22T22:47:37.150Z" } ``` Schema (YAML Structural Definition): ```yaml properties: api_key: description: Created temporary API key. type: string expires_at: description: UTC timestamp indicating when generated temporary API key will expire. format: date-time type: string required: - api_key - expires_at type: object ``` * **400**: Invalid request. Error types: * `invalid_request`: One or more body fields are missing or invalid (`usage_type`, `expires_in_seconds` out of range, `client_reference_id` too long, `max_session_duration_seconds` out of range, etc.). Inspect `validation_errors`. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **401**: Authentication error. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **429**: Rate / capacity limit exceeded. Error types: * `limit_exceeded`: The caller hit a per-minute request rate or other capacity limit. The `message` describes which limit was hit. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` * **500**: Internal server error. Schema (YAML Structural Definition): ```yaml properties: status_code: description: HTTP status code of the response. type: integer error_type: description: > Machine-readable error category. Examples: `invalid_request`, `unauthenticated`, `limit_exceeded`, `model_not_available`, `internal_error`. type: string message: description: Human-readable error message. type: string validation_errors: description: >- List of per-field validation errors. Populated only when `error_type` is `invalid_request` and the failure came from request-body validation. items: properties: error_type: type: string location: type: string message: type: string required: - error_type - location - message type: object type: array request_id: description: >- Unique identifier for this request. Include it when contacting support at support@soniox.com so we can look up server-side logs. type: string more_info: anyOf: - type: string - type: 'null' description: > Optional URL with additional information about this error. Points to the Soniox documentation for errors a developer can resolve via code or configuration. required: - status_code - error_type - message - validation_errors - request_id type: object ``` # Connection keepalive URL: /tts/rt/connection-keepalive Learn how to keep a WebSocket connection alive during idle periods in Soniox Text-to-Speech. ## Overview In real-time Text-to-Speech, there may be periods when you are not sending any text — for example, while waiting on an upstream LLM, between turns in a conversational agent, or between streams on the same connection. If the WebSocket stays idle for too long, the server will close the connection as inactive. To prevent this, send a **keepalive control message**: ```json {"keep_alive": true} ``` This is a lightweight signal that tells the server the client is still present and the connection should stay open. **Keepalive only works after you start your first stream.** A freshly opened connection must authenticate by sending a [config message](/api-reference/tts/websocket-api#configuration) with a valid API key shortly after connecting (within about 10 seconds). Keepalive messages do not authenticate the connection, so a connection that only sends keepalives before its first stream will be closed. *** ## When to use Send a keepalive message whenever: * You are not sending text for an extended period. * You want to keep the WebSocket connection open between streams so you can start a new `stream_id` without reconnecting. This ensures that: * The connection stays open across idle gaps. * You avoid the latency cost of reopening a WebSocket for the next stream. You can also send them on a fixed interval even while sending text. *** ## Key points * The keepalive message does **not** trigger speech generation and has no server-side effect other than maintaining the connection. * **Send at least once every 20–30 seconds** when the connection is idle to prevent timeouts. * Keepalive applies to the whole WebSocket connection, not to a specific `stream_id`, one message keeps every stream on the connection alive. * Keepalive works only **after** you start your first stream. Before that, send your config message with a valid API key within about 10 seconds of connecting, or the connection is closed. *** ## Keepalive does not keep an unused connection open forever Keepalive prevents idle timeouts, but a connection that generates **no audio for more than 3 minutes** is closed even if keepalives are still flowing. Keepalives keep the socket from looking idle, they do not count as work. To keep a connection open across long gaps, generate audio at least once within every 3-minute window. If your application may go longer than that without speaking, let the connection close and open a new one when you have text to send. # Limits & quotas URL: /tts/rt/limits-and-quotas Learn about real-time Text-to-Speech API limits and quotas. ## WebSocket API limits Soniox applies default limits to real-time WebSocket connections and streams to ensure stability and fair use. Make sure your application respects these constraints and implements graceful recovery when a limit is reached. You can request higher limits (except for streams per connnection and stream duration) in the [Soniox Console](https://console.soniox.com/org/limits). | Limit | Value | Notes | | ---------------------- | ------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Requests per minute | **100** | Exceeding this may result in rate limiting. | | Concurrent streams | **3** | Maximum number of TTS streams running concurrently across all your WebSocket connections. | | Streams per connection | **5** | Maximum number of streams within a single WebSocket connection.
**This limit is fixed and cannot be increased.** | | Stream duration | **2 minutes** | Each real-time stream is capped at 2 minutes of generated audio. When the cap is reached, the stream ends with a [`max_audio_duration_reached`](/api-reference/errors#max-audio-duration-reached) error and the output is truncated. To continue beyond this, start a new stream.
**This limit is fixed and cannot be increased.** | Learn how to [track usage](/guides/usage-logs) and [monitor concurrency](/guides/concurrency-limits). # Real-time generation URL: /tts/rt/real-time-generation Learn how to generate speech in real time with the Soniox Text-to-Speech WebSocket API, including configuration, streaming, responses, and errors. ## Overview Soniox Text-to-Speech AI generates **native-speaker-quality speech** in over 60 languages with **ultra-low latency**, hallucination-free output, and accurate pronunciation of alphanumerics like phone numbers, email addresses, and IDs. This is ideal for use cases like **voice agents, conversational AI, live narration, and interactive assistants.** Real-time generation is provided through our [WebSocket API](/api-reference/tts/websocket-api), which streams audio back to you as text arrives. A single connection can host [multiple concurrent streams](/tts/rt/streams) for serving many users or voices at once. *** ## How real-time generation works Each stream is opened with a one-time config that picks the [voice](/tts/concepts/voices) and [audio output format](/tts/concepts/audio-formats). After that, Text-to-Speech is a **two-way stream:** you send text in small chunks as it becomes available, and Soniox streams audio back without waiting for the full input. * **Text in (chunks)** → You send text incrementally over the WebSocket. Each message carries a `text` fragment and a `text_end` flag marking the last chunk. * **Audio out (chunks)** → Soniox returns base64-encoded `audio` payloads as they are produced. You can decode and play or buffer each chunk on arrival. * **Interleaved** → Audio for the first words starts flowing back while you are still writing later chunks. You never have to wait for a full sentence before audio begins. ### Pairing with an LLM This is what makes real-time TTS the right fit for LLM-driven workloads. Pipe each LLM token (or token batch) straight into a Soniox text chunk. Because audio generation begins with the first words, the user hears speech *while the LLM is still generating the rest of the response*. No buffering the full reply before speaking, no awkward "thinking" pause between the model finishing and the voice starting. ### Example streaming timeline Here is what a single stream looks like on the wire. Outbound text chunks and inbound audio chunks are **interleaved**, they do not happen in two separate phases. **Client → Server** (first text chunk): ```json { "text": "Hello there,", "text_end": false, "stream_id": "stream-001" } ``` **Server → Client** (audio already streaming back for "Hello"): ```json { "audio": "", "audio_end": false, "stream_id": "stream-001" } ``` **Client → Server** (next text chunk sent while audio for earlier text keeps arriving): ```json { "text": " this is Soniox", "text_end": false, "stream_id": "stream-001" } ``` **Server → Client** (more audio chunks): ```json { "audio": "", "audio_end": false, "stream_id": "stream-001" } ``` **Client → Server** (final text chunk, `text_end: true`): ```json { "text": " speaking live.", "text_end": true, "stream_id": "stream-001" } ``` **Server → Client** (last audio chunk, `audio_end: true`): ```json { "audio": "", "audio_end": true, "stream_id": "stream-001" } ``` **Server → Client** (stream terminated): ```json { "terminated": true, "stream_id": "stream-001" } ``` **Bottom line:** You don't wait for a full sentence before you hear anything. Audio for the first words starts streaming back while the rest of the text is still being written. Treat the stream as complete only after [`terminated: true`](/tts/rt/termination). *** ## Connection keepalive During idle periods, for example, while waiting on an upstream LLM or between agent turns. Send keepalive messages to prevent the WebSocket connection from timing out. For details, see [Connection keepalive](/tts/rt/connection-keepalive). *** ## API reference For the full message schema, configuration parameters, cancellation, and error codes, see the [WebSocket API reference](/api-reference/tts/websocket-api). *** ## Code example **Prerequisite:** Complete the steps in [Get started](/tts/get-started). See on GitHub: [soniox\_sdk\_realtime.py](https://github.com/soniox/soniox_examples/blob/master/text_to_speech/python_sdk/soniox_sdk_realtime.py). ``` import argparse import os import threading import time from pathlib import Path from uuid import uuid4 from soniox import SonioxClient from soniox.errors import SonioxRealtimeError from soniox.types import RealtimeTTSConfig from soniox.utils import output_file_for_audio_format VALID_SAMPLE_RATES = [8000, 16000, 24000, 44100, 48000] VALID_BITRATES = [32000, 64000, 96000, 128000, 192000, 256000, 320000] VALID_AUDIO_FORMATS = [ "pcm_f32le", "pcm_s16le", "pcm_mulaw", "pcm_alaw", "wav", "aac", "mp3", "opus", "flac", ] DEFAULT_LINES = [ "Welcome to Soniox real-time Text-to-Speech. ", "As text is streamed in, audio streams back in parallel with high accuracy, ", "so your application can start playing speech ", "within milliseconds of the first word.", ] def get_config( model: str, language: str, voice: str, audio_format: str, sample_rate: int | None, bitrate: int | None, stream_id: str | None, ) -> RealtimeTTSConfig: config = RealtimeTTSConfig( # Stream id for this realtime TTS session. # If omitted, a random id is generated. stream_id=stream_id or f"tts-{uuid4()}", # # Select the model to use. # See: soniox.com/docs/tts/models model=model, # # Set the language of the input text. # See: soniox.com/docs/tts/concepts/supported-languages language=language, # # Select the voice to use. # See: soniox.com/docs/tts/concepts/voices voice=voice, # # Set output audio format and optional encoding parameters. # See: soniox.com/docs/api-reference/tts/websocket-api audio_format=audio_format, sample_rate=sample_rate, bitrate=bitrate, ) return config def run_session( client: SonioxClient, lines: list[str], model: str, language: str, voice: str, audio_format: str, sample_rate: int | None, bitrate: int | None, stream_id: str | None, output_path: str | None, ) -> None: # Build a realtime Text-to-Speech session configuration. config = get_config( model=model, language=language, voice=voice, audio_format=audio_format, sample_rate=sample_rate, bitrate=bitrate, stream_id=stream_id, ) sanitized_lines = [line.strip() for line in lines if line.strip()] if not sanitized_lines: raise ValueError("Text is empty after parsing.") destination = ( Path(output_path) if output_path else output_file_for_audio_format(audio_format, "tts_realtime") ) print("Connecting to Soniox...") audio_chunks: list[bytes] = [] try: with client.realtime.tts.connect(config=config) as session: print("Session started.") send_errors: list[Exception] = [] def send_worker() -> None: try: for line in sanitized_lines: session.send_text_chunk(line, text_end=False) time.sleep(0.1) session.finish() except Exception as exc: send_errors.append(exc) threading.Thread(target=send_worker, daemon=True).start() # Receive streamed audio chunks from the websocket. for audio_chunk in session.receive_audio_chunks(): audio_chunks.append(audio_chunk) if send_errors: raise RuntimeError(f"Failed to send realtime text: {send_errors[0]}") print("Session finished.") finally: audio = b"".join(audio_chunks) if audio: destination.write_bytes(audio) print(f"Wrote {len(audio)} bytes to {destination.resolve()}") else: print("No audio file was written.") def main() -> None: parser = argparse.ArgumentParser() parser.add_argument( "--line", action="append", default=None, help="Line to send to realtime TTS (repeat --line for multiple lines).", ) parser.add_argument("--model", default="tts-rt-v2") parser.add_argument("--language", default="en") parser.add_argument("--voice", default="Adrian") parser.add_argument("--audio_format", default="wav") parser.add_argument("--sample_rate", type=int) parser.add_argument("--bitrate", type=int) parser.add_argument("--stream_id", help="Optional stream id.") parser.add_argument( "--output_path", help="Optional output file path. If omitted, a timestamped path is generated.", ) args = parser.parse_args() if args.audio_format not in VALID_AUDIO_FORMATS: raise ValueError(f"audio_format must be one of {VALID_AUDIO_FORMATS}") if args.sample_rate is not None and args.sample_rate not in VALID_SAMPLE_RATES: raise ValueError(f"sample_rate must be None or one of {VALID_SAMPLE_RATES}") if args.bitrate is not None and args.bitrate not in VALID_BITRATES: raise ValueError(f"bitrate must be None or one of {VALID_BITRATES}") api_key = os.environ.get("SONIOX_API_KEY") if not api_key: raise RuntimeError( "Missing SONIOX_API_KEY.\n" "1. Get your API key at https://console.soniox.com\n" "2. Run: export SONIOX_API_KEY=" ) client = SonioxClient(api_key=api_key) try: run_session( client=client, lines=args.line or DEFAULT_LINES, model=args.model, language=args.language, voice=args.voice, audio_format=args.audio_format, sample_rate=args.sample_rate, bitrate=args.bitrate, stream_id=args.stream_id, output_path=args.output_path, ) except SonioxRealtimeError as exc: print("Soniox realtime error:", exc) finally: client.close() if __name__ == "__main__": main() ``` ```sh title="Terminal" # Generate speech with default settings (wav output) python soniox_sdk_realtime.py --line "Hello from Soniox realtime Text-to-Speech." # Generate raw PCM output python soniox_sdk_realtime.py --audio_format pcm_s16le --sample_rate 24000 --output_path tts-output.pcm ``` See on GitHub: [soniox\_sdk\_realtime.js](https://github.com/soniox/soniox_examples/blob/master/text_to_speech/nodejs_sdk/soniox_sdk_realtime.js). ``` import { RealtimeError, SonioxNodeClient } from "@soniox/node"; import fs from "fs"; import path from "path"; import { parseArgs } from "node:util"; import process from "process"; const VALID_SAMPLE_RATES = [8000, 16000, 24000, 44100, 48000]; const VALID_BITRATES = [32000, 64000, 96000, 128000, 192000, 256000, 320000]; const VALID_AUDIO_FORMATS = [ "pcm_f32le", "pcm_s16le", "pcm_mulaw", "pcm_alaw", "wav", "aac", "mp3", "opus", "flac", ]; const RAW_PCM_FORMATS = ["pcm_s16le", "pcm_f32le", "pcm_mulaw", "pcm_alaw"]; const DEFAULT_LINES = [ "Welcome to Soniox real-time Text-to-Speech. ", "As text is streamed in, audio streams back in parallel with high accuracy, ", "so your application can start playing speech ", "within milliseconds of the first word.", ]; // Initialize the client. // The API key is read from the SONIOX_API_KEY environment variable. const client = new SonioxNodeClient(); // Resolve a concrete output file path. // If the provided path has no extension, derive one from audio_format: // * pcm_s16le -> .wav (we wrap the bytes in a WAV container below) // * other pcm_* -> .pcm (raw, no container) // * anything else -> the format name (e.g. .flac, .mp3, .opus) function resolveOutputPath(outputPath, audioFormat) { if (outputPath && path.extname(outputPath)) { return outputPath; } const ext = audioFormat === "pcm_s16le" ? "wav" : RAW_PCM_FORMATS.includes(audioFormat) ? "pcm" : audioFormat; const base = outputPath || "tts_realtime"; return `${base}.${ext}`; } function pcmS16leToWav(pcm, { sampleRate, numChannels = 1 }) { const bitsPerSample = 16; const byteRate = sampleRate * numChannels * (bitsPerSample / 8); const blockAlign = numChannels * (bitsPerSample / 8); const dataSize = pcm.byteLength; const header = Buffer.alloc(44); header.write("RIFF", 0, "ascii"); header.writeUInt32LE(36 + dataSize, 4); header.write("WAVE", 8, "ascii"); header.write("fmt ", 12, "ascii"); header.writeUInt32LE(16, 16); header.writeUInt16LE(1, 20); header.writeUInt16LE(numChannels, 22); header.writeUInt32LE(sampleRate, 24); header.writeUInt32LE(byteRate, 28); header.writeUInt16LE(blockAlign, 32); header.writeUInt16LE(bitsPerSample, 34); header.write("data", 36, "ascii"); header.writeUInt32LE(dataSize, 40); return Buffer.concat([header, Buffer.from(pcm)]); } // Build a realtime TTS stream config. function getStreamConfig({ model, language, voice, audioFormat, sampleRate, bitrate, streamId, }) { const config = { // Client-defined stream id (auto-generated if omitted). ...(streamId && { stream_id: streamId }), // Select the model to use. // See: soniox.com/docs/tts/models model, // Set the language of the input text. // See: soniox.com/docs/tts/concepts/supported-languages language, // Select the voice to use. // See: soniox.com/docs/tts/concepts/voices voice, // Set output audio format and optional encoding parameters. // See: soniox.com/docs/api-reference/tts/websocket-api audio_format: audioFormat, }; if (sampleRate !== undefined) config.sample_rate = sampleRate; if (bitrate !== undefined) config.bitrate = bitrate; return config; } async function runSession({ lines, model, language, voice, audioFormat, sampleRate, bitrate, streamId, outputPath, }) { const sanitizedLines = lines .map((line) => line.trim()) .filter((line) => line.length > 0); if (sanitizedLines.length === 0) { throw new Error("Text is empty after parsing."); } const destination = resolveOutputPath(outputPath, audioFormat); const config = getStreamConfig({ model, language, voice, audioFormat, sampleRate, bitrate, streamId, }); console.log("Connecting to Soniox..."); const stream = await client.realtime.tts(config); console.log("Session started."); // Send text chunks in the background while receiving audio. let sendError = null; const sendPromise = (async () => { try { for (const line of sanitizedLines) { stream.sendText(line); // Sleep for 100 ms to simulate real-time streaming. await new Promise((res) => setTimeout(res, 100)); } stream.finish(); } catch (err) { sendError = err; } })(); // Collect streamed audio chunks. const audioChunks = []; try { for await (const chunk of stream) { audioChunks.push(chunk); } } finally { await sendPromise; stream.close(); } if (sendError) { throw new Error(`Failed to send realtime text: ${sendError.message}`); } console.log("Session finished."); const audio = Buffer.concat(audioChunks.map((c) => Buffer.from(c))); if (audio.length > 0) { // Wrap raw pcm_s16le in a WAV container so the .wav file plays everywhere. const bytes = audioFormat === "pcm_s16le" && path.extname(destination).toLowerCase() === ".wav" ? pcmS16leToWav(audio, { sampleRate }) : audio; fs.writeFileSync(destination, bytes); console.log(`Wrote ${bytes.length} bytes to ${path.resolve(destination)}`); } else { console.log("No audio file was written."); } } async function main() { const { values: argv } = parseArgs({ options: { line: { type: "string", multiple: true }, model: { type: "string", default: "tts-rt-v2" }, language: { type: "string", default: "en" }, voice: { type: "string", default: "Adrian" }, audio_format: { type: "string", default: "pcm_s16le" }, sample_rate: { type: "string" }, bitrate: { type: "string" }, stream_id: { type: "string" }, output_path: { type: "string" }, }, }); if (!VALID_AUDIO_FORMATS.includes(argv.audio_format)) { throw new Error( `audio_format must be one of ${VALID_AUDIO_FORMATS.join(", ")}`, ); } let sampleRate = argv.sample_rate !== undefined ? Number(argv.sample_rate) : undefined; if (sampleRate === undefined && RAW_PCM_FORMATS.includes(argv.audio_format)) { sampleRate = 24000; } if (sampleRate !== undefined && !VALID_SAMPLE_RATES.includes(sampleRate)) { throw new Error( `sample_rate must be one of ${VALID_SAMPLE_RATES.join(", ")}`, ); } const bitrate = argv.bitrate !== undefined ? Number(argv.bitrate) : undefined; if (bitrate !== undefined && !VALID_BITRATES.includes(bitrate)) { throw new Error(`bitrate must be one of ${VALID_BITRATES.join(", ")}`); } try { await runSession({ lines: argv.line && argv.line.length > 0 ? argv.line : DEFAULT_LINES, model: argv.model, language: argv.language, voice: argv.voice, audioFormat: argv.audio_format, sampleRate, bitrate, streamId: argv.stream_id, outputPath: argv.output_path, }); } catch (err) { if (err instanceof RealtimeError) { console.error("Soniox realtime error:", err.message); } else { throw err; } } } main().catch((err) => { console.error("Error:", err.message); process.exit(1); }); ``` ```sh title="Terminal" # Generate speech with default settings (wav output) node soniox_sdk_realtime.js --line "Hello from Soniox realtime Text-to-Speech." # Generate raw PCM output node soniox_sdk_realtime.js --audio_format pcm_s16le --sample_rate 24000 --output_path tts-output.pcm ``` {/* NOTE: Empty tag is needed so code block renders correctly */}
See on GitHub: [soniox\_realtime.py](https://github.com/soniox/soniox_examples/blob/master/text_to_speech/python/soniox_realtime.py). ``` import argparse import base64 import json import os import threading import time from typing import Any from websockets import ConnectionClosedOK from websockets.sync.client import connect SONIOX_TTS_WEBSOCKET_URL = "wss://tts-rt.soniox.com/tts-websocket" MODEL = "tts-rt-v2" VALID_SAMPLE_RATES = [8000, 16000, 24000, 44100, 48000] VALID_BITRATES = [32000, 64000, 96000, 128000, 192000, 256000, 320000] VALID_AUDIO_FORMATS = [ "pcm_f32le", "pcm_s16le", "pcm_mulaw", "pcm_alaw", "wav", "aac", "mp3", "opus", "flac", ] DEFAULT_LINES = [ "Welcome to Soniox real-time Text-to-Speech. ", "As text is streamed in, audio streams back in parallel with high accuracy, ", "so your application can start playing speech ", "within milliseconds of the first word.", ] def get_output_path(*, output_path: str, audio_format: str) -> str: """ Generates the resulting output path for the given audio format. """ if "." in os.path.basename(output_path): return output_path ext = "pcm" if audio_format in ("pcm_s16le", "pcm_s16be") else audio_format return f"{output_path}.{ext}" # Get Soniox TTS config. def get_config( api_key: str, stream_id: str, language: str, voice: str, audio_format: str, sample_rate: int | None, bitrate: int | None, ) -> dict: config: dict[str, Any] = { # Get your API key at console.soniox.com, then run: export SONIOX_API_KEY= "api_key": api_key, # # Client-defined stream id to identify this realtime request. "stream_id": stream_id, # # Select the model to use. # See: soniox.com/docs/tts/models "model": MODEL, # # Set the language of the input text. # See: soniox.com/docs/tts/concepts/supported-languages "language": language, # # Select the voice to use. # See: soniox.com/docs/tts/concepts/voices "voice": voice, # # Audio format. # See: soniox.com/docs/tts/concepts/audio-formats "audio_format": audio_format, } if sample_rate is not None: config["sample_rate"] = sample_rate if bitrate is not None: config["bitrate"] = bitrate return config def get_text_request(text: str, stream_id: str, text_end: bool) -> dict: return { "text": text, "text_end": text_end, "stream_id": stream_id, } # Stream text lines to the websocket. def stream_text(lines: list[str], stream_id: str, ws) -> None: for line in lines: clean_line = line.strip() if not clean_line: continue ws.send(json.dumps(get_text_request(clean_line, stream_id, text_end=False))) # Sleep for 100 ms to simulate real-time streaming. time.sleep(0.1) # Send text_end=true after the last chunk. ws.send(json.dumps(get_text_request("", stream_id, text_end=True))) def send_requests( ws, api_key: str, lines: list[str], language: str, voice: str, audio_format: str, sample_rate: int | None, bitrate: int | None, stream_id: str, ) -> None: config = get_config( api_key=api_key, stream_id=stream_id, language=language, voice=voice, audio_format=audio_format, sample_rate=sample_rate, bitrate=bitrate, ) ws.send(json.dumps(config)) stream_text(lines, stream_id, ws) def run_session( api_key: str, lines: list[str], language: str, voice: str, audio_format: str, sample_rate: int | None, bitrate: int | None, stream_id: str, output_path: str, ) -> None: print("Connecting to Soniox...") with connect(SONIOX_TTS_WEBSOCKET_URL) as ws: send_errors: list[Exception] = [] def send_worker() -> None: try: send_requests( ws, api_key, lines, language, voice, audio_format, sample_rate, bitrate, stream_id, ) except Exception as exc: send_errors.append(exc) # Send config and text in the background while receiving responses. threading.Thread( target=send_worker, daemon=True, ).start() print("Session started.") audio_chunks: list[bytes] = [] try: while True: if send_errors: raise RuntimeError(f"Failed to send realtime requests: {send_errors[0]}") message = ws.recv() res = json.loads(message) # Error from server. if res.get("error_code") is not None: print(f"Error: {res['error_code']} - {res['error_message']}") break # Collect audio bytes from base64-encoded chunks. audio_b64 = res.get("audio") if audio_b64: audio_chunks.append(base64.b64decode(audio_b64)) # Session finished. if res.get("terminated"): break except ConnectionClosedOK: # Normal, server closed after finished. pass except KeyboardInterrupt: print("\nInterrupted by user.") except Exception as e: print(f"Error: {e}") finally: audio_data = b"".join(audio_chunks) if audio_data: destination = get_output_path(output_path=output_path, audio_format=audio_format) with open(destination, "wb") as fh: fh.write(audio_data) print(f"Wrote {len(audio_data)} bytes to {destination}") else: print("No audio file was written.") def main(): parser = argparse.ArgumentParser() parser.add_argument( "--line", action="append", default=None, help="Line to send to realtime TTS (repeat --line for multiple lines).", ) parser.add_argument("--language", default="en") parser.add_argument("--voice", default="Adrian") parser.add_argument("--audio_format", default="wav") parser.add_argument("--stream_id", default="stream-1") parser.add_argument("--output_path", default="tts-ws") parser.add_argument("--sample_rate", type=int) parser.add_argument("--bitrate", type=int) args = parser.parse_args() if args.audio_format not in VALID_AUDIO_FORMATS: raise ValueError(f"audio_format must be one of {VALID_AUDIO_FORMATS}") if args.sample_rate is not None and args.sample_rate not in VALID_SAMPLE_RATES: raise ValueError(f"sample_rate must be None or one of {VALID_SAMPLE_RATES}") if args.bitrate is not None and args.bitrate not in VALID_BITRATES: raise ValueError(f"bitrate must be None or one of {VALID_BITRATES}") api_key = os.environ.get("SONIOX_API_KEY") if not api_key: raise RuntimeError( "Missing SONIOX_API_KEY.\n" "1. Get your API key at https://console.soniox.com\n" "2. Run: export SONIOX_API_KEY=" ) run_session( api_key=api_key, lines=args.line or DEFAULT_LINES, language=args.language, voice=args.voice, audio_format=args.audio_format, sample_rate=args.sample_rate, bitrate=args.bitrate, stream_id=args.stream_id, output_path=args.output_path, ) if __name__ == "__main__": main() ``` ```sh title="Terminal" # Generate speech with default settings (wav output) python soniox_realtime.py --line "Hello from Soniox websocket Text-to-Speech." # Generate raw PCM output python soniox_realtime.py --audio_format pcm_s16le --sample_rate 24000 --output_path tts-output ``` See on GitHub: [soniox\_realtime.js](https://github.com/soniox/soniox_examples/blob/master/text_to_speech/nodejs/soniox_realtime.js). ``` import fs from "fs"; import path from "path"; import WebSocket from "ws"; import { parseArgs } from "node:util"; import process from "process"; const SONIOX_TTS_WEBSOCKET_URL = "wss://tts-rt.soniox.com/tts-websocket"; const MODEL = "tts-rt-v2"; const VALID_SAMPLE_RATES = [8000, 16000, 24000, 44100, 48000]; const VALID_BITRATES = [32000, 64000, 96000, 128000, 192000, 256000, 320000]; const VALID_AUDIO_FORMATS = [ "pcm_f32le", "pcm_s16le", "pcm_mulaw", "pcm_alaw", "wav", "aac", "mp3", "opus", "flac", ]; const RAW_PCM_FORMATS = ["pcm_s16le", "pcm_f32le", "pcm_mulaw", "pcm_alaw"]; const DEFAULT_LINES = [ "Welcome to Soniox real-time Text-to-Speech. ", "As text is streamed in, audio streams back in parallel with high accuracy, ", "so your application can start playing speech ", "within milliseconds of the first word.", ]; // Resolve a concrete output file path. // If the provided path has no extension, derive one from audio_format: // * pcm_s16le -> .wav (we wrap the bytes in a WAV container below) // * other pcm_* -> .pcm (raw, no container) // * anything else -> the format name (e.g. .flac, .mp3, .opus) function resolveOutputPath(outputPath, audioFormat) { if (outputPath && path.extname(outputPath)) { return outputPath; } const ext = audioFormat === "pcm_s16le" ? "wav" : RAW_PCM_FORMATS.includes(audioFormat) ? "pcm" : audioFormat; const base = outputPath || "tts-ws"; return `${base}.${ext}`; } function pcmS16leToWav(pcm, { sampleRate, numChannels = 1 }) { const bitsPerSample = 16; const byteRate = sampleRate * numChannels * (bitsPerSample / 8); const blockAlign = numChannels * (bitsPerSample / 8); const dataSize = pcm.byteLength; const header = Buffer.alloc(44); header.write("RIFF", 0, "ascii"); header.writeUInt32LE(36 + dataSize, 4); header.write("WAVE", 8, "ascii"); header.write("fmt ", 12, "ascii"); header.writeUInt32LE(16, 16); header.writeUInt16LE(1, 20); header.writeUInt16LE(numChannels, 22); header.writeUInt32LE(sampleRate, 24); header.writeUInt32LE(byteRate, 28); header.writeUInt16LE(blockAlign, 32); header.writeUInt16LE(bitsPerSample, 34); header.write("data", 36, "ascii"); header.writeUInt32LE(dataSize, 40); return Buffer.concat([header, Buffer.from(pcm)]); } // Get Soniox TTS config. function getConfig({ apiKey, streamId, language, voice, audioFormat, sampleRate, bitrate, }) { const config = { // Get your API key at console.soniox.com, then run: export SONIOX_API_KEY= api_key: apiKey, // Client-defined stream id to identify this realtime request. stream_id: streamId, // Select the model to use. // See: soniox.com/docs/tts/models model: MODEL, // Set the language of the input text. // See: soniox.com/docs/tts/concepts/supported-languages language, // Select the voice to use. // See: soniox.com/docs/tts/concepts/voices voice, // Audio format. // See: soniox.com/docs/tts/concepts/audio-formats audio_format: audioFormat, }; if (sampleRate !== undefined) config.sample_rate = sampleRate; if (bitrate !== undefined) config.bitrate = bitrate; return config; } function getTextRequest(text, streamId, textEnd) { return { text, text_end: textEnd, stream_id: streamId, }; } // Stream text lines to the websocket. async function streamText(lines, streamId, ws) { for (const line of lines) { const cleanLine = line.trim(); if (!cleanLine) continue; ws.send(JSON.stringify(getTextRequest(cleanLine, streamId, false))); // Sleep for 100 ms to simulate real-time streaming. await new Promise((res) => setTimeout(res, 100)); } // Send text_end=true after the last chunk. ws.send(JSON.stringify(getTextRequest("", streamId, true))); } function runSession({ apiKey, lines, language, voice, audioFormat, sampleRate, bitrate, streamId, outputPath, }) { return new Promise((resolve, reject) => { console.log("Connecting to Soniox..."); const ws = new WebSocket(SONIOX_TTS_WEBSOCKET_URL); const audioChunks = []; const finalize = (err) => { const destination = resolveOutputPath(outputPath, audioFormat); if (audioChunks.length > 0) { const audio = Buffer.concat(audioChunks); // Wrap raw pcm_s16le in a WAV container so the .wav file plays everywhere. const bytes = audioFormat === "pcm_s16le" && path.extname(destination).toLowerCase() === ".wav" ? pcmS16leToWav(audio, { sampleRate }) : audio; fs.writeFileSync(destination, bytes); console.log(`Wrote ${bytes.length} bytes to ${path.resolve(destination)}`); } else { console.log("No audio file was written."); } if (err) reject(err); else resolve(); }; ws.on("open", () => { const config = getConfig({ apiKey, streamId, language, voice, audioFormat, sampleRate, bitrate, }); // Send first request with config. ws.send(JSON.stringify(config)); // Start streaming text in the background. streamText(lines, streamId, ws).catch((err) => { console.error("Text stream error:", err); }); console.log("Session started."); }); ws.on("message", (msg) => { let res; try { res = JSON.parse(msg.toString()); } catch { return; } // Error from server. // See: https://soniox.com/docs/api-reference/tts/websocket-api#error-response if (res.error_code) { console.error(`Error: ${res.error_code} - ${res.error_message}`); ws.close(); return; } // Collect audio bytes from base64-encoded chunks. if (res.audio) { audioChunks.push(Buffer.from(res.audio, "base64")); } // Session finished. if (res.terminated) { console.log("Session finished."); ws.close(); } }); ws.on("close", () => { finalize(null); }); ws.on("error", (err) => { console.error("WebSocket error:", err.message); finalize(err); }); }); } async function main() { const { values: argv } = parseArgs({ options: { line: { type: "string", multiple: true }, language: { type: "string", default: "en" }, voice: { type: "string", default: "Adrian" }, audio_format: { type: "string", default: "pcm_s16le" }, stream_id: { type: "string", default: "stream-1" }, output_path: { type: "string", default: "tts-ws" }, sample_rate: { type: "string" }, bitrate: { type: "string" }, }, }); if (!VALID_AUDIO_FORMATS.includes(argv.audio_format)) { throw new Error( `audio_format must be one of ${VALID_AUDIO_FORMATS.join(", ")}`, ); } let sampleRate = argv.sample_rate !== undefined ? Number(argv.sample_rate) : undefined; if (sampleRate === undefined && RAW_PCM_FORMATS.includes(argv.audio_format)) { sampleRate = 24000; } if (sampleRate !== undefined && !VALID_SAMPLE_RATES.includes(sampleRate)) { throw new Error( `sample_rate must be one of ${VALID_SAMPLE_RATES.join(", ")}`, ); } const bitrate = argv.bitrate !== undefined ? Number(argv.bitrate) : undefined; if (bitrate !== undefined && !VALID_BITRATES.includes(bitrate)) { throw new Error(`bitrate must be one of ${VALID_BITRATES.join(", ")}`); } const apiKey = process.env.SONIOX_API_KEY; if (!apiKey) { throw new Error( "Missing SONIOX_API_KEY.\n" + "1. Get your API key at https://console.soniox.com\n" + "2. Run: export SONIOX_API_KEY=", ); } await runSession({ apiKey, lines: argv.line && argv.line.length > 0 ? argv.line : DEFAULT_LINES, language: argv.language, voice: argv.voice, audioFormat: argv.audio_format, sampleRate, bitrate, streamId: argv.stream_id, outputPath: argv.output_path, }); } main().catch((err) => { console.error("Error:", err.message); process.exit(1); }); ``` ```sh title="Terminal" # Generate speech with default settings (wav output) node soniox_realtime.js --line "Hello from Soniox websocket Text-to-Speech." # Generate raw PCM output node soniox_realtime.js --audio_format pcm_s16le --sample_rate 24000 --output_path tts-output ``` # Streams URL: /tts/rt/streams Run multiple concurrent Text-to-Speech streams over a single WebSocket connection using stream_id. ## Overview A single WebSocket connection can host up to **5 concurrent streams**, each an independent Text-to-Speech generation identified by a unique `stream_id`. This lets you serve multiple users, voices, or languages from one connection without reopening a socket per request. Multiple streams are useful when you need to: * Generate speech for multiple users or conversations over one connection. * Produce audio in different voices or languages simultaneously. * Reduce connection overhead by reusing a single WebSocket instead of opening one per request. Each stream has its own lifecycle: it is started with a config message, receives its own text chunks, and produces its own audio output. Streams are fully isolated — an error in one stream does not affect the others. *** ## Core rules * Start each stream with a config message that includes a unique `stream_id`. See the [WebSocket API reference](/api-reference/tts/websocket-api#configuration) for the full payload. * Use the same `stream_id` for all text chunks belonging to that stream. * You can run up to 5 active streams per connection. * A stream error does not terminate other streams on the same connection. * Keep the WebSocket connection alive with [keepalive](/tts/rt/connection-keepalive) messages during idle periods. * A stream can be [canceled](/tts/rt/termination#client-initiated-cancellation) at any time. *** ## Starting multiple streams Send a config message for each stream before sending text for that stream. Stream configs are independent — each one picks its own `voice`, `language`, and audio output settings. ### Example: two streams on one connection Stream A config (English, voice `Adrian`): ```json { "api_key": "", "model": "tts-rt-v2", "language": "en", "voice": "Adrian", "audio_format": "pcm_s16le", "sample_rate": 24000, "stream_id": "stream-A" } ``` Stream B config (Italian, voice `Adrian`) — same connection, different `stream_id`, any other fields may differ too: ```json { "api_key": "", "model": "tts-rt-v2", "language": "it", "voice": "Adrian", "audio_format": "pcm_s16le", "sample_rate": 24000, "stream_id": "stream-B" } ``` **Important:** Always send a stream's config message before sending any text for that stream. After configuring, send text chunks with the matching `stream_id`. You can interleave chunks across streams freely — only the `stream_id` field determines routing: ```json {"text": "This audio belongs to stream A.", "text_end": false, "stream_id": "stream-A"} ``` ```json {"text": "Questo audio appartiene allo stream B.", "text_end": false, "stream_id": "stream-B"} ``` *** ## Route messages by stream ID The server sends audio for all active streams over the same connection. Use the `stream_id` field to route each incoming message to the correct output: 1. Inspect `stream_id` on every incoming message. 2. Route `audio` chunks to the matching output buffer or player. 3. Track termination state separately for each stream. 4. Mark a stream as finished only after receiving `terminated: true` for that `stream_id`. Audio message: ```json { "audio": "", "audio_end": false, "stream_id": "stream-A" } ``` Termination message: ```json { "terminated": true, "stream_id": "stream-A" } ``` *** ## Error isolation * **Stream-level error** — The server returns an error message for that `stream_id`, followed by a `terminated: true` event. Other streams on the connection continue normally. * **Validation or malformed-message error** — An incoming message was malformed or invalid (bad JSON, missing fields, unknown `stream_id`, duplicate start). The server returns an error response but does **not** send a `terminated` event, because no stream was actually running. The connection stays open for other valid streams and messages. * **Connection-level failure** — If the WebSocket itself closes (network failure, server error), all active streams on that connection end immediately. For the full list of error codes and messages, see the [WebSocket API reference](/api-reference/tts/websocket-api#full-list-of-possible-error-codes-and-messages). *** ## Code example **Prerequisite:** Complete the steps in [Get started](/tts/get-started). See on GitHub: [soniox\_sdk\_realtime.py](https://github.com/soniox/soniox_examples/blob/master/text_to_speech/python_sdk/soniox_sdk_realtime.py). ``` import argparse import os import threading import time from pathlib import Path from uuid import uuid4 from soniox import SonioxClient from soniox.errors import SonioxRealtimeError from soniox.types import RealtimeTTSConfig from soniox.utils import output_file_for_audio_format VALID_SAMPLE_RATES = [8000, 16000, 24000, 44100, 48000] VALID_BITRATES = [32000, 64000, 96000, 128000, 192000, 256000, 320000] VALID_AUDIO_FORMATS = [ "pcm_f32le", "pcm_s16le", "pcm_mulaw", "pcm_alaw", "wav", "aac", "mp3", "opus", "flac", ] DEFAULT_LINES = [ "Welcome to Soniox real-time Text-to-Speech. ", "As text is streamed in, audio streams back in parallel with high accuracy, ", "so your application can start playing speech ", "within milliseconds of the first word.", ] def get_config( model: str, language: str, voice: str, audio_format: str, sample_rate: int | None, bitrate: int | None, stream_id: str | None, ) -> RealtimeTTSConfig: config = RealtimeTTSConfig( # Stream id for this realtime TTS session. # If omitted, a random id is generated. stream_id=stream_id or f"tts-{uuid4()}", # # Select the model to use. # See: soniox.com/docs/tts/models model=model, # # Set the language of the input text. # See: soniox.com/docs/tts/concepts/supported-languages language=language, # # Select the voice to use. # See: soniox.com/docs/tts/concepts/voices voice=voice, # # Set output audio format and optional encoding parameters. # See: soniox.com/docs/api-reference/tts/websocket-api audio_format=audio_format, sample_rate=sample_rate, bitrate=bitrate, ) return config def run_session( client: SonioxClient, lines: list[str], model: str, language: str, voice: str, audio_format: str, sample_rate: int | None, bitrate: int | None, stream_id: str | None, output_path: str | None, ) -> None: # Build a realtime Text-to-Speech session configuration. config = get_config( model=model, language=language, voice=voice, audio_format=audio_format, sample_rate=sample_rate, bitrate=bitrate, stream_id=stream_id, ) sanitized_lines = [line.strip() for line in lines if line.strip()] if not sanitized_lines: raise ValueError("Text is empty after parsing.") destination = ( Path(output_path) if output_path else output_file_for_audio_format(audio_format, "tts_realtime") ) print("Connecting to Soniox...") audio_chunks: list[bytes] = [] try: with client.realtime.tts.connect(config=config) as session: print("Session started.") send_errors: list[Exception] = [] def send_worker() -> None: try: for line in sanitized_lines: session.send_text_chunk(line, text_end=False) time.sleep(0.1) session.finish() except Exception as exc: send_errors.append(exc) threading.Thread(target=send_worker, daemon=True).start() # Receive streamed audio chunks from the websocket. for audio_chunk in session.receive_audio_chunks(): audio_chunks.append(audio_chunk) if send_errors: raise RuntimeError(f"Failed to send realtime text: {send_errors[0]}") print("Session finished.") finally: audio = b"".join(audio_chunks) if audio: destination.write_bytes(audio) print(f"Wrote {len(audio)} bytes to {destination.resolve()}") else: print("No audio file was written.") def main() -> None: parser = argparse.ArgumentParser() parser.add_argument( "--line", action="append", default=None, help="Line to send to realtime TTS (repeat --line for multiple lines).", ) parser.add_argument("--model", default="tts-rt-v2") parser.add_argument("--language", default="en") parser.add_argument("--voice", default="Adrian") parser.add_argument("--audio_format", default="wav") parser.add_argument("--sample_rate", type=int) parser.add_argument("--bitrate", type=int) parser.add_argument("--stream_id", help="Optional stream id.") parser.add_argument( "--output_path", help="Optional output file path. If omitted, a timestamped path is generated.", ) args = parser.parse_args() if args.audio_format not in VALID_AUDIO_FORMATS: raise ValueError(f"audio_format must be one of {VALID_AUDIO_FORMATS}") if args.sample_rate is not None and args.sample_rate not in VALID_SAMPLE_RATES: raise ValueError(f"sample_rate must be None or one of {VALID_SAMPLE_RATES}") if args.bitrate is not None and args.bitrate not in VALID_BITRATES: raise ValueError(f"bitrate must be None or one of {VALID_BITRATES}") api_key = os.environ.get("SONIOX_API_KEY") if not api_key: raise RuntimeError( "Missing SONIOX_API_KEY.\n" "1. Get your API key at https://console.soniox.com\n" "2. Run: export SONIOX_API_KEY=" ) client = SonioxClient(api_key=api_key) try: run_session( client=client, lines=args.line or DEFAULT_LINES, model=args.model, language=args.language, voice=args.voice, audio_format=args.audio_format, sample_rate=args.sample_rate, bitrate=args.bitrate, stream_id=args.stream_id, output_path=args.output_path, ) except SonioxRealtimeError as exc: print("Soniox realtime error:", exc) finally: client.close() if __name__ == "__main__": main() ``` ```sh title="Terminal" # Generate speech with default settings (wav output) python soniox_sdk_realtime.py --line "Hello from Soniox realtime Text-to-Speech." # Generate raw PCM output python soniox_sdk_realtime.py --audio_format pcm_s16le --sample_rate 24000 --output_path tts-output.pcm ``` See on GitHub: [soniox\_sdk\_realtime.js](https://github.com/soniox/soniox_examples/blob/master/text_to_speech/nodejs_sdk/soniox_sdk_realtime.js). ``` import { RealtimeError, SonioxNodeClient } from "@soniox/node"; import fs from "fs"; import path from "path"; import { parseArgs } from "node:util"; import process from "process"; const VALID_SAMPLE_RATES = [8000, 16000, 24000, 44100, 48000]; const VALID_BITRATES = [32000, 64000, 96000, 128000, 192000, 256000, 320000]; const VALID_AUDIO_FORMATS = [ "pcm_f32le", "pcm_s16le", "pcm_mulaw", "pcm_alaw", "wav", "aac", "mp3", "opus", "flac", ]; const RAW_PCM_FORMATS = ["pcm_s16le", "pcm_f32le", "pcm_mulaw", "pcm_alaw"]; const DEFAULT_LINES = [ "Welcome to Soniox real-time Text-to-Speech. ", "As text is streamed in, audio streams back in parallel with high accuracy, ", "so your application can start playing speech ", "within milliseconds of the first word.", ]; // Initialize the client. // The API key is read from the SONIOX_API_KEY environment variable. const client = new SonioxNodeClient(); // Resolve a concrete output file path. // If the provided path has no extension, derive one from audio_format: // * pcm_s16le -> .wav (we wrap the bytes in a WAV container below) // * other pcm_* -> .pcm (raw, no container) // * anything else -> the format name (e.g. .flac, .mp3, .opus) function resolveOutputPath(outputPath, audioFormat) { if (outputPath && path.extname(outputPath)) { return outputPath; } const ext = audioFormat === "pcm_s16le" ? "wav" : RAW_PCM_FORMATS.includes(audioFormat) ? "pcm" : audioFormat; const base = outputPath || "tts_realtime"; return `${base}.${ext}`; } function pcmS16leToWav(pcm, { sampleRate, numChannels = 1 }) { const bitsPerSample = 16; const byteRate = sampleRate * numChannels * (bitsPerSample / 8); const blockAlign = numChannels * (bitsPerSample / 8); const dataSize = pcm.byteLength; const header = Buffer.alloc(44); header.write("RIFF", 0, "ascii"); header.writeUInt32LE(36 + dataSize, 4); header.write("WAVE", 8, "ascii"); header.write("fmt ", 12, "ascii"); header.writeUInt32LE(16, 16); header.writeUInt16LE(1, 20); header.writeUInt16LE(numChannels, 22); header.writeUInt32LE(sampleRate, 24); header.writeUInt32LE(byteRate, 28); header.writeUInt16LE(blockAlign, 32); header.writeUInt16LE(bitsPerSample, 34); header.write("data", 36, "ascii"); header.writeUInt32LE(dataSize, 40); return Buffer.concat([header, Buffer.from(pcm)]); } // Build a realtime TTS stream config. function getStreamConfig({ model, language, voice, audioFormat, sampleRate, bitrate, streamId, }) { const config = { // Client-defined stream id (auto-generated if omitted). ...(streamId && { stream_id: streamId }), // Select the model to use. // See: soniox.com/docs/tts/models model, // Set the language of the input text. // See: soniox.com/docs/tts/concepts/supported-languages language, // Select the voice to use. // See: soniox.com/docs/tts/concepts/voices voice, // Set output audio format and optional encoding parameters. // See: soniox.com/docs/api-reference/tts/websocket-api audio_format: audioFormat, }; if (sampleRate !== undefined) config.sample_rate = sampleRate; if (bitrate !== undefined) config.bitrate = bitrate; return config; } async function runSession({ lines, model, language, voice, audioFormat, sampleRate, bitrate, streamId, outputPath, }) { const sanitizedLines = lines .map((line) => line.trim()) .filter((line) => line.length > 0); if (sanitizedLines.length === 0) { throw new Error("Text is empty after parsing."); } const destination = resolveOutputPath(outputPath, audioFormat); const config = getStreamConfig({ model, language, voice, audioFormat, sampleRate, bitrate, streamId, }); console.log("Connecting to Soniox..."); const stream = await client.realtime.tts(config); console.log("Session started."); // Send text chunks in the background while receiving audio. let sendError = null; const sendPromise = (async () => { try { for (const line of sanitizedLines) { stream.sendText(line); // Sleep for 100 ms to simulate real-time streaming. await new Promise((res) => setTimeout(res, 100)); } stream.finish(); } catch (err) { sendError = err; } })(); // Collect streamed audio chunks. const audioChunks = []; try { for await (const chunk of stream) { audioChunks.push(chunk); } } finally { await sendPromise; stream.close(); } if (sendError) { throw new Error(`Failed to send realtime text: ${sendError.message}`); } console.log("Session finished."); const audio = Buffer.concat(audioChunks.map((c) => Buffer.from(c))); if (audio.length > 0) { // Wrap raw pcm_s16le in a WAV container so the .wav file plays everywhere. const bytes = audioFormat === "pcm_s16le" && path.extname(destination).toLowerCase() === ".wav" ? pcmS16leToWav(audio, { sampleRate }) : audio; fs.writeFileSync(destination, bytes); console.log(`Wrote ${bytes.length} bytes to ${path.resolve(destination)}`); } else { console.log("No audio file was written."); } } async function main() { const { values: argv } = parseArgs({ options: { line: { type: "string", multiple: true }, model: { type: "string", default: "tts-rt-v2" }, language: { type: "string", default: "en" }, voice: { type: "string", default: "Adrian" }, audio_format: { type: "string", default: "pcm_s16le" }, sample_rate: { type: "string" }, bitrate: { type: "string" }, stream_id: { type: "string" }, output_path: { type: "string" }, }, }); if (!VALID_AUDIO_FORMATS.includes(argv.audio_format)) { throw new Error( `audio_format must be one of ${VALID_AUDIO_FORMATS.join(", ")}`, ); } let sampleRate = argv.sample_rate !== undefined ? Number(argv.sample_rate) : undefined; if (sampleRate === undefined && RAW_PCM_FORMATS.includes(argv.audio_format)) { sampleRate = 24000; } if (sampleRate !== undefined && !VALID_SAMPLE_RATES.includes(sampleRate)) { throw new Error( `sample_rate must be one of ${VALID_SAMPLE_RATES.join(", ")}`, ); } const bitrate = argv.bitrate !== undefined ? Number(argv.bitrate) : undefined; if (bitrate !== undefined && !VALID_BITRATES.includes(bitrate)) { throw new Error(`bitrate must be one of ${VALID_BITRATES.join(", ")}`); } try { await runSession({ lines: argv.line && argv.line.length > 0 ? argv.line : DEFAULT_LINES, model: argv.model, language: argv.language, voice: argv.voice, audioFormat: argv.audio_format, sampleRate, bitrate, streamId: argv.stream_id, outputPath: argv.output_path, }); } catch (err) { if (err instanceof RealtimeError) { console.error("Soniox realtime error:", err.message); } else { throw err; } } } main().catch((err) => { console.error("Error:", err.message); process.exit(1); }); ``` ```sh title="Terminal" # Generate speech with default settings (wav output) node soniox_sdk_realtime.js --line "Hello from Soniox realtime Text-to-Speech." # Generate raw PCM output node soniox_sdk_realtime.js --audio_format pcm_s16le --sample_rate 24000 --output_path tts-output.pcm ``` {/* NOTE: Empty tag is needed so code block renders correctly */}
See on GitHub: [soniox\_realtime.py](https://github.com/soniox/soniox_examples/blob/master/text_to_speech/python/soniox_realtime.py). ``` import argparse import base64 import json import os import threading import time from typing import Any from websockets import ConnectionClosedOK from websockets.sync.client import connect SONIOX_TTS_WEBSOCKET_URL = "wss://tts-rt.soniox.com/tts-websocket" MODEL = "tts-rt-v2" VALID_SAMPLE_RATES = [8000, 16000, 24000, 44100, 48000] VALID_BITRATES = [32000, 64000, 96000, 128000, 192000, 256000, 320000] VALID_AUDIO_FORMATS = [ "pcm_f32le", "pcm_s16le", "pcm_mulaw", "pcm_alaw", "wav", "aac", "mp3", "opus", "flac", ] DEFAULT_LINES = [ "Welcome to Soniox real-time Text-to-Speech. ", "As text is streamed in, audio streams back in parallel with high accuracy, ", "so your application can start playing speech ", "within milliseconds of the first word.", ] def get_output_path(*, output_path: str, audio_format: str) -> str: """ Generates the resulting output path for the given audio format. """ if "." in os.path.basename(output_path): return output_path ext = "pcm" if audio_format in ("pcm_s16le", "pcm_s16be") else audio_format return f"{output_path}.{ext}" # Get Soniox TTS config. def get_config( api_key: str, stream_id: str, language: str, voice: str, audio_format: str, sample_rate: int | None, bitrate: int | None, ) -> dict: config: dict[str, Any] = { # Get your API key at console.soniox.com, then run: export SONIOX_API_KEY= "api_key": api_key, # # Client-defined stream id to identify this realtime request. "stream_id": stream_id, # # Select the model to use. # See: soniox.com/docs/tts/models "model": MODEL, # # Set the language of the input text. # See: soniox.com/docs/tts/concepts/supported-languages "language": language, # # Select the voice to use. # See: soniox.com/docs/tts/concepts/voices "voice": voice, # # Audio format. # See: soniox.com/docs/tts/concepts/audio-formats "audio_format": audio_format, } if sample_rate is not None: config["sample_rate"] = sample_rate if bitrate is not None: config["bitrate"] = bitrate return config def get_text_request(text: str, stream_id: str, text_end: bool) -> dict: return { "text": text, "text_end": text_end, "stream_id": stream_id, } # Stream text lines to the websocket. def stream_text(lines: list[str], stream_id: str, ws) -> None: for line in lines: clean_line = line.strip() if not clean_line: continue ws.send(json.dumps(get_text_request(clean_line, stream_id, text_end=False))) # Sleep for 100 ms to simulate real-time streaming. time.sleep(0.1) # Send text_end=true after the last chunk. ws.send(json.dumps(get_text_request("", stream_id, text_end=True))) def send_requests( ws, api_key: str, lines: list[str], language: str, voice: str, audio_format: str, sample_rate: int | None, bitrate: int | None, stream_id: str, ) -> None: config = get_config( api_key=api_key, stream_id=stream_id, language=language, voice=voice, audio_format=audio_format, sample_rate=sample_rate, bitrate=bitrate, ) ws.send(json.dumps(config)) stream_text(lines, stream_id, ws) def run_session( api_key: str, lines: list[str], language: str, voice: str, audio_format: str, sample_rate: int | None, bitrate: int | None, stream_id: str, output_path: str, ) -> None: print("Connecting to Soniox...") with connect(SONIOX_TTS_WEBSOCKET_URL) as ws: send_errors: list[Exception] = [] def send_worker() -> None: try: send_requests( ws, api_key, lines, language, voice, audio_format, sample_rate, bitrate, stream_id, ) except Exception as exc: send_errors.append(exc) # Send config and text in the background while receiving responses. threading.Thread( target=send_worker, daemon=True, ).start() print("Session started.") audio_chunks: list[bytes] = [] try: while True: if send_errors: raise RuntimeError(f"Failed to send realtime requests: {send_errors[0]}") message = ws.recv() res = json.loads(message) # Error from server. if res.get("error_code") is not None: print(f"Error: {res['error_code']} - {res['error_message']}") break # Collect audio bytes from base64-encoded chunks. audio_b64 = res.get("audio") if audio_b64: audio_chunks.append(base64.b64decode(audio_b64)) # Session finished. if res.get("terminated"): break except ConnectionClosedOK: # Normal, server closed after finished. pass except KeyboardInterrupt: print("\nInterrupted by user.") except Exception as e: print(f"Error: {e}") finally: audio_data = b"".join(audio_chunks) if audio_data: destination = get_output_path(output_path=output_path, audio_format=audio_format) with open(destination, "wb") as fh: fh.write(audio_data) print(f"Wrote {len(audio_data)} bytes to {destination}") else: print("No audio file was written.") def main(): parser = argparse.ArgumentParser() parser.add_argument( "--line", action="append", default=None, help="Line to send to realtime TTS (repeat --line for multiple lines).", ) parser.add_argument("--language", default="en") parser.add_argument("--voice", default="Adrian") parser.add_argument("--audio_format", default="wav") parser.add_argument("--stream_id", default="stream-1") parser.add_argument("--output_path", default="tts-ws") parser.add_argument("--sample_rate", type=int) parser.add_argument("--bitrate", type=int) args = parser.parse_args() if args.audio_format not in VALID_AUDIO_FORMATS: raise ValueError(f"audio_format must be one of {VALID_AUDIO_FORMATS}") if args.sample_rate is not None and args.sample_rate not in VALID_SAMPLE_RATES: raise ValueError(f"sample_rate must be None or one of {VALID_SAMPLE_RATES}") if args.bitrate is not None and args.bitrate not in VALID_BITRATES: raise ValueError(f"bitrate must be None or one of {VALID_BITRATES}") api_key = os.environ.get("SONIOX_API_KEY") if not api_key: raise RuntimeError( "Missing SONIOX_API_KEY.\n" "1. Get your API key at https://console.soniox.com\n" "2. Run: export SONIOX_API_KEY=" ) run_session( api_key=api_key, lines=args.line or DEFAULT_LINES, language=args.language, voice=args.voice, audio_format=args.audio_format, sample_rate=args.sample_rate, bitrate=args.bitrate, stream_id=args.stream_id, output_path=args.output_path, ) if __name__ == "__main__": main() ``` ```sh title="Terminal" # Generate speech with default settings (wav output) python soniox_realtime.py --line "Hello from Soniox websocket Text-to-Speech." # Generate raw PCM output python soniox_realtime.py --audio_format pcm_s16le --sample_rate 24000 --output_path tts-output ``` See on GitHub: [soniox\_realtime.js](https://github.com/soniox/soniox_examples/blob/master/text_to_speech/nodejs/soniox_realtime.js). ``` import fs from "fs"; import path from "path"; import WebSocket from "ws"; import { parseArgs } from "node:util"; import process from "process"; const SONIOX_TTS_WEBSOCKET_URL = "wss://tts-rt.soniox.com/tts-websocket"; const MODEL = "tts-rt-v2"; const VALID_SAMPLE_RATES = [8000, 16000, 24000, 44100, 48000]; const VALID_BITRATES = [32000, 64000, 96000, 128000, 192000, 256000, 320000]; const VALID_AUDIO_FORMATS = [ "pcm_f32le", "pcm_s16le", "pcm_mulaw", "pcm_alaw", "wav", "aac", "mp3", "opus", "flac", ]; const RAW_PCM_FORMATS = ["pcm_s16le", "pcm_f32le", "pcm_mulaw", "pcm_alaw"]; const DEFAULT_LINES = [ "Welcome to Soniox real-time Text-to-Speech. ", "As text is streamed in, audio streams back in parallel with high accuracy, ", "so your application can start playing speech ", "within milliseconds of the first word.", ]; // Resolve a concrete output file path. // If the provided path has no extension, derive one from audio_format: // * pcm_s16le -> .wav (we wrap the bytes in a WAV container below) // * other pcm_* -> .pcm (raw, no container) // * anything else -> the format name (e.g. .flac, .mp3, .opus) function resolveOutputPath(outputPath, audioFormat) { if (outputPath && path.extname(outputPath)) { return outputPath; } const ext = audioFormat === "pcm_s16le" ? "wav" : RAW_PCM_FORMATS.includes(audioFormat) ? "pcm" : audioFormat; const base = outputPath || "tts-ws"; return `${base}.${ext}`; } function pcmS16leToWav(pcm, { sampleRate, numChannels = 1 }) { const bitsPerSample = 16; const byteRate = sampleRate * numChannels * (bitsPerSample / 8); const blockAlign = numChannels * (bitsPerSample / 8); const dataSize = pcm.byteLength; const header = Buffer.alloc(44); header.write("RIFF", 0, "ascii"); header.writeUInt32LE(36 + dataSize, 4); header.write("WAVE", 8, "ascii"); header.write("fmt ", 12, "ascii"); header.writeUInt32LE(16, 16); header.writeUInt16LE(1, 20); header.writeUInt16LE(numChannels, 22); header.writeUInt32LE(sampleRate, 24); header.writeUInt32LE(byteRate, 28); header.writeUInt16LE(blockAlign, 32); header.writeUInt16LE(bitsPerSample, 34); header.write("data", 36, "ascii"); header.writeUInt32LE(dataSize, 40); return Buffer.concat([header, Buffer.from(pcm)]); } // Get Soniox TTS config. function getConfig({ apiKey, streamId, language, voice, audioFormat, sampleRate, bitrate, }) { const config = { // Get your API key at console.soniox.com, then run: export SONIOX_API_KEY= api_key: apiKey, // Client-defined stream id to identify this realtime request. stream_id: streamId, // Select the model to use. // See: soniox.com/docs/tts/models model: MODEL, // Set the language of the input text. // See: soniox.com/docs/tts/concepts/supported-languages language, // Select the voice to use. // See: soniox.com/docs/tts/concepts/voices voice, // Audio format. // See: soniox.com/docs/tts/concepts/audio-formats audio_format: audioFormat, }; if (sampleRate !== undefined) config.sample_rate = sampleRate; if (bitrate !== undefined) config.bitrate = bitrate; return config; } function getTextRequest(text, streamId, textEnd) { return { text, text_end: textEnd, stream_id: streamId, }; } // Stream text lines to the websocket. async function streamText(lines, streamId, ws) { for (const line of lines) { const cleanLine = line.trim(); if (!cleanLine) continue; ws.send(JSON.stringify(getTextRequest(cleanLine, streamId, false))); // Sleep for 100 ms to simulate real-time streaming. await new Promise((res) => setTimeout(res, 100)); } // Send text_end=true after the last chunk. ws.send(JSON.stringify(getTextRequest("", streamId, true))); } function runSession({ apiKey, lines, language, voice, audioFormat, sampleRate, bitrate, streamId, outputPath, }) { return new Promise((resolve, reject) => { console.log("Connecting to Soniox..."); const ws = new WebSocket(SONIOX_TTS_WEBSOCKET_URL); const audioChunks = []; const finalize = (err) => { const destination = resolveOutputPath(outputPath, audioFormat); if (audioChunks.length > 0) { const audio = Buffer.concat(audioChunks); // Wrap raw pcm_s16le in a WAV container so the .wav file plays everywhere. const bytes = audioFormat === "pcm_s16le" && path.extname(destination).toLowerCase() === ".wav" ? pcmS16leToWav(audio, { sampleRate }) : audio; fs.writeFileSync(destination, bytes); console.log(`Wrote ${bytes.length} bytes to ${path.resolve(destination)}`); } else { console.log("No audio file was written."); } if (err) reject(err); else resolve(); }; ws.on("open", () => { const config = getConfig({ apiKey, streamId, language, voice, audioFormat, sampleRate, bitrate, }); // Send first request with config. ws.send(JSON.stringify(config)); // Start streaming text in the background. streamText(lines, streamId, ws).catch((err) => { console.error("Text stream error:", err); }); console.log("Session started."); }); ws.on("message", (msg) => { let res; try { res = JSON.parse(msg.toString()); } catch { return; } // Error from server. // See: https://soniox.com/docs/api-reference/tts/websocket-api#error-response if (res.error_code) { console.error(`Error: ${res.error_code} - ${res.error_message}`); ws.close(); return; } // Collect audio bytes from base64-encoded chunks. if (res.audio) { audioChunks.push(Buffer.from(res.audio, "base64")); } // Session finished. if (res.terminated) { console.log("Session finished."); ws.close(); } }); ws.on("close", () => { finalize(null); }); ws.on("error", (err) => { console.error("WebSocket error:", err.message); finalize(err); }); }); } async function main() { const { values: argv } = parseArgs({ options: { line: { type: "string", multiple: true }, language: { type: "string", default: "en" }, voice: { type: "string", default: "Adrian" }, audio_format: { type: "string", default: "pcm_s16le" }, stream_id: { type: "string", default: "stream-1" }, output_path: { type: "string", default: "tts-ws" }, sample_rate: { type: "string" }, bitrate: { type: "string" }, }, }); if (!VALID_AUDIO_FORMATS.includes(argv.audio_format)) { throw new Error( `audio_format must be one of ${VALID_AUDIO_FORMATS.join(", ")}`, ); } let sampleRate = argv.sample_rate !== undefined ? Number(argv.sample_rate) : undefined; if (sampleRate === undefined && RAW_PCM_FORMATS.includes(argv.audio_format)) { sampleRate = 24000; } if (sampleRate !== undefined && !VALID_SAMPLE_RATES.includes(sampleRate)) { throw new Error( `sample_rate must be one of ${VALID_SAMPLE_RATES.join(", ")}`, ); } const bitrate = argv.bitrate !== undefined ? Number(argv.bitrate) : undefined; if (bitrate !== undefined && !VALID_BITRATES.includes(bitrate)) { throw new Error(`bitrate must be one of ${VALID_BITRATES.join(", ")}`); } const apiKey = process.env.SONIOX_API_KEY; if (!apiKey) { throw new Error( "Missing SONIOX_API_KEY.\n" + "1. Get your API key at https://console.soniox.com\n" + "2. Run: export SONIOX_API_KEY=", ); } await runSession({ apiKey, lines: argv.line && argv.line.length > 0 ? argv.line : DEFAULT_LINES, language: argv.language, voice: argv.voice, audioFormat: argv.audio_format, sampleRate, bitrate, streamId: argv.stream_id, outputPath: argv.output_path, }); } main().catch((err) => { console.error("Error:", err.message); process.exit(1); }); ``` ```sh title="Terminal" # Generate speech with default settings (wav output) node soniox_realtime.js --line "Hello from Soniox websocket Text-to-Speech." # Generate raw PCM output node soniox_realtime.js --audio_format pcm_s16le --sample_rate 24000 --output_path tts-output ``` # Stream termination URL: /tts/rt/termination Learn how to terminate a stream and handle finalization and errors in the Soniox Text-to-Speech WebSocket API. ## Overview A real-time stream ends with a well-defined **three-step handshake:** the client signals the end of input with `text_end: true`, the server emits a final audio chunk with `audio_end: true`, and then the server confirms cleanup with `terminated: true`. Only after `terminated: true` is it safe to release stream state or reuse the `stream_id`. Stream termination does not close the WebSocket connection. Other streams on the same connection continue normally, and you can start new streams on the same socket. *** ## Normal termination sequence 1. **Client** sends final text chunk with `text_end: true`: ```json { "text": "Final chunk of text.", "text_end": true, "stream_id": "stream-001" } ``` 2. **Server** sends the last audio chunk with `audio_end: true`: ```json { "audio": "", "audio_end": true, "stream_id": "stream-001" } ``` 3. **Server** sends the final stream event: ```json { "terminated": true, "stream_id": "stream-001" } ``` * `audio_end: true` indicates the audio payload sequence is complete, no more `audio` chunks will arrive for this stream. * `terminated: true` indicates stream lifecycle cleanup is complete on the server, it is now safe to release stream state and reuse the `stream_id`. *** ## Client-initiated cancellation If you need to stop a stream early, for example, user interruption in a voice agent or a timeout on the calling side — send a cancel message. The server finalizes the stream without sending further audio chunks. Cancel request: ```json { "stream_id": "stream-001", "cancel": true } ``` Finalization response: ```json { "terminated": true, "stream_id": "stream-001" } ``` *** ## Error termination If a stream fails, the server responds with an error message followed by `terminated: true` for that `stream_id`: ```json { "stream_id": "stream-001", "error_code": 400, "error_type": "invalid_request", "error_message": "Missing model", "more_info": "https://soniox.com/docs/api-reference/errors#invalid-request", "request_id": "3d37a3bd-5078-47ee-a369-b204e3bbedda" } ``` ```json { "terminated": true, "stream_id": "stream-001" } ``` Branch your client code on `error_type` (stable) rather than `error_message` (human-readable, may change). Log the `request_id` so it can be passed to [support@soniox.com](mailto:support@soniox.com) when needed. One error follows a slightly different sequence: when generated audio reaches the [2-minute cap per stream](/tts/rt/limits-and-quotas), the final audio chunk is delivered without `audio_end: true`, followed by a [`max_audio_duration_reached`](/api-reference/errors#max-audio-duration-reached) error and `terminated: true`. Generate the remaining text on a new stream. The WebSocket connection stays open unless there is a connection-level failure, so other streams continue. For the full catalog of `error_type` values and recovery steps, see the [Errors reference](/api-reference/errors). For the per-endpoint message list, see the [WebSocket API reference](/api-reference/tts/websocket-api#full-list-of-possible-error-codes-and-messages). *** # Timestamps URL: /tts/rt/timestamps Learn how to receive character-level audio timestamps from the Soniox Text-to-Speech WebSocket API for subtitle highlighting and agent-interruption flows. ## Overview The Text-to-Speech WebSocket API can return **character-level timestamps** alongside the generated audio. For each character of the spoken text, you receive the start and end time (in seconds) of the audio that pronounces it. This makes it possible to line up the audio you play with the exact text that produced it, in real time, as the stream arrives. Timestamps are available on the [WebSocket API](/api-reference/tts/websocket-api) only. The [REST endpoint](/tts/rest-api/generate-speech) streams raw audio bytes with no JSON envelope, so it has nowhere to carry alignment data. *** ## Why use timestamps * **Subtitle and caption highlighting.** Drive karaoke-style highlighting that follows the voice character by character, or word by word. Because timestamps stream back with the audio, you can highlight live instead of waiting for the full clip. * **Agent-interruption flows.** When a user interrupts a voice agent mid-sentence, the timestamps tell you how far into the text the audio actually reached. You can compute the exact interruption point and feed the spoken-so-far text back to the LLM, so the agent knows what the user did and did not hear. *** ## Quick start Set `return_timestamps: true` in the configuration message when you start a stream: ```json { "api_key": "", "model": "tts-rt-v2", "language": "en", "voice": "Adrian", "audio_format": "pcm_s16le", "sample_rate": 24000, "stream_id": "stream-001", "return_timestamps": true } ``` Audio responses then include a `timestamps` object: ```json { "stream_id": "stream-001", "audio": "", "timestamps": { "characters": ["H", "e", "l", "l", "o"], "character_start_times_seconds": [0.0, 0.1, 0.2, 0.3, 0.4], "character_end_times_seconds": [0.1, 0.2, 0.3, 0.4, 0.5] } } ``` *** ## Request `return_timestamps` is an optional field on the configuration (start) message. It defaults to `false` and has no effect on `text`, `keep_alive`, or `cancel` messages. Request character-level timestamps in the responses. Defaults to `false`. *** ## Response When timestamps are active, a `timestamps` object is attached to audio response frames. It contains three equal-length, parallel arrays: One entry per character (Unicode codepoint) of the spoken text. Start time of each character, in seconds. End time of each character, in seconds. ### How timestamps are delivered * **Streaming and chunked.** Timestamps arrive incrementally, interleaved with audio. Each `timestamps` object covers only the characters in that frame. Concatenating the `characters` arrays across all frames, in order, reconstructs the full spoken text. * **Timestamps always ride with audio.** A `timestamps` object is always attached to an `audio` frame, never sent on its own. The service may, however, emit an `audio` frame that carries no `timestamps` (audio-only), so clients must handle audio frames both with and without alignment data. * **Monotonic timing.** Times are non-decreasing across the whole stream, including across chunk boundaries, and never exceed the duration of audio generated so far. * **Omitted when empty.** The `timestamps` object is left out of frames that carry no alignment data, such as the terminal `audio_end` and `terminated` frames. ### Preprocessed text Timestamps map to the **preprocessed** text, not the raw input you sent. Before generation, the model normalizes the input: for example, normalizing whitespace and removing characters it cannot pronounce, such as emojis. The `characters` you receive reflect that normalized form, so they may differ from your original string. There is currently no mapping back to the original input text, only the preprocessed alignment is returned. For agent-interruption use cases this is usually sufficient: feed the normalized text, up to the interruption point, back to the LLM. ## End-to-end example **Client → Server** (start with `return_timestamps: true`): ```json { "api_key": "", "model": "tts-rt-v2", "language": "en", "voice": "Adrian", "audio_format": "mp3", "bitrate": 128000, "stream_id": "stream-001", "return_timestamps": true } ``` **Client → Server** (text, then end of text): ```json { "text": "Hi", "text_end": true, "stream_id": "stream-001" } ``` **Server → Client** (audio with timestamps): ```json { "stream_id": "stream-001", "audio": "", "timestamps": { "characters": ["H", "i"], "character_start_times_seconds": [0.0, 0.1], "character_end_times_seconds": [0.1, 0.25] } } ``` **Server → Client** (last audio chunk, carrying the trailing audio with no new characters, so no `timestamps`): ```json { "stream_id": "stream-001", "audio": "", "audio_end": true } ``` **Server → Client** (stream terminated): ```json { "stream_id": "stream-001", "terminated": true } ``` *** ## REST API The [REST endpoint](/tts/rest-api/generate-speech) returns raw audio bytes as the HTTP response body (for example, `audio/mpeg`), with no JSON wrapper. There is nowhere to place alignment data without corrupting the audio stream, so **timestamps are not available over REST**, and `return_timestamps` is ignored there. Use the [WebSocket API](/api-reference/tts/websocket-api) if you need timestamps. *** ## API reference For the full message schema, configuration parameters, and error codes, see the [WebSocket API reference](/api-reference/tts/websocket-api). # Classes URL: /sdk/node-SDK/reference/classes Soniox Node SDK — Class Reference ## SonioxNodeClient Soniox Node Client ### Example ```typescript import { SonioxNodeClient } from '@soniox/node'; // Default (US) region const client = new SonioxNodeClient({ api_key: 'your-api-key' }); // EU region const client = new SonioxNodeClient({ api_key: 'your-api-key', region: 'eu' }); // REST TTS const audio = await client.tts.generate({ text: 'Hello', voice: 'Adrian', language: 'en', }); // WebSocket TTS const stream = await client.realtime.tts({ model: 'tts-rt-v2', voice: 'Adrian', language: 'en', audio_format: 'wav', }); ``` ### Constructor ```ts new SonioxNodeClient(options): SonioxNodeClient; ``` **Parameters** | Parameter | Type | | --------- | ---------------------------------------------------------- | | `options` | [`SonioxNodeClientOptions`](types#sonioxnodeclientoptions) | **Returns** `SonioxNodeClient` ### Properties | Property | Type | | ------------------- | ------------------------------------------------------------------ | | `auth` | [`SonioxAuthAPI`](classes#sonioxauthapi) | | `concurrencyLimits` | [`SonioxConcurrencyLimitsAPI`](classes#sonioxconcurrencylimitsapi) | | `files` | [`SonioxFilesAPI`](classes#sonioxfilesapi) | | `models` | [`SonioxModelsAPI`](classes#sonioxmodelsapi) | | `realtime` | [`SonioxRealtimeApi`](classes#sonioxrealtimeapi) | | `stt` | [`SonioxSttApi`](classes#sonioxsttapi) | | `tts` | [`SonioxTtsApi`](classes#sonioxttsapi) | | `usageLogs` | [`SonioxUsageLogsAPI`](classes#sonioxusagelogsapi) | | `webhooks` | [`SonioxWebhooksAPI`](classes#sonioxwebhooksapi) | *** ## SonioxFilesAPI ### count() ```ts count(options): Promise; ``` Returns the total number of files, split by source. **Parameters** | Parameter | Type | Description | | ----------------- | ------------------------------ | --------------------------------- | | `options` | \{ `signal?`: `AbortSignal`; } | Optional cancellation parameters. | | `options.signal?` | `AbortSignal` | - | **Returns** `Promise`\<[`FilesCountResponse`](types#filescountresponse)> Total file counts for Playground, Public API, and all sources. **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors. **Example** ```typescript const counts = await client.files.count(); console.log(counts.total); ``` *** ### delete() ```ts delete(file, signal?): Promise; ``` Permanently deletes a file. This operation is idempotent - succeeds even if the file doesn't exist. **Parameters** | Parameter | Type | Description | | --------- | ---------------------------------------- | --------------------------------------------- | | `file` | [`FileIdentifier`](types#fileidentifier) | The UUID of the file or a SonioxFile instance | | `signal?` | `AbortSignal` | Optional AbortSignal for cancellation | **Returns** `Promise`\<`void`> **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors (except 404) **Example** ```typescript // Delete by ID await client.files.delete('550e8400-e29b-41d4-a716-446655440000'); // Or delete a file instance const file = await client.files.get('550e8400-e29b-41d4-a716-446655440000'); if (file) { await client.files.delete(file); } // Or just use the instance method await file.delete(); ``` *** ### delete\_all() ```ts delete_all(options): Promise; ``` Permanently deletes all uploaded files. Iterates through all pages of files and deletes each one. **Parameters** | Parameter | Type | Description | | --------- | ------------------------------------------------------ | -------------------------------------- | | `options` | [`DeleteAllFilesOptions`](types#deleteallfilesoptions) | Optional signal and progress callback. | **Returns** `Promise`\<`void`> The number of files deleted. **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors. **Throws** `Error` If the operation is aborted via signal. **Example** ```typescript // Delete all files await client.files.delete_all(); console.log(`Deleted all files.`); // With cancellation const controller = new AbortController(); await client.files.delete_all({ signal: controller.signal }); ``` *** ### get() ```ts get(file, signal?): Promise; ``` Retrieve metadata for an uploaded file. **Parameters** | Parameter | Type | Description | | --------- | ---------------------------------------- | --------------------------------------------- | | `file` | [`FileIdentifier`](types#fileidentifier) | The UUID of the file or a SonioxFile instance | | `signal?` | `AbortSignal` | Optional AbortSignal for cancellation | **Returns** `Promise`\<[`SonioxFile`](classes#sonioxfile) | `null`> The file instance, or null if not found **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors (except 404) **Example** ```typescript const file = await client.files.get('550e8400-e29b-41d4-a716-446655440000'); if (file) { console.log(file.filename, file.size); } ``` *** ### list() ```ts list(options): Promise; ``` Retrieves list of uploaded files The returned result is async iterable - use `for await...of` **Parameters** | Parameter | Type | Description | | --------- | -------------------------------------------- | ----------------------------------------------- | | `options` | [`ListFilesOptions`](types#listfilesoptions) | Optional pagination and cancellation parameters | **Returns** `Promise`\<[`FileListResult`](classes#filelistresult)> FileListResult **Throws** [SonioxHttpError](classes#sonioxhttperror) **Example** ```typescript const result = await client.files.list(); // Automatic paging - iterates through ALL files across all pages for await (const file of result) { console.log(file.filename, file.size); } // Or access just the first page for (const file of result.files) { console.log(file.filename); } // Check if there are more pages if (result.isPaged()) { console.log('More pages available'); } // Manual paging using cursor const page1 = await client.files.list({ limit: 10 }); if (page1.next_page_cursor) { const page2 = await client.files.list({ cursor: page1.next_page_cursor }); } // With cancellation const controller = new AbortController(); const result = await client.files.list({ signal: controller.signal }); ``` *** ### upload() ```ts upload(file, options): Promise; ``` Uploads a file to Soniox for transcription **Parameters** | Parameter | Type | Description | | --------- | ---------------------------------------------- | ------------------------------------------- | | `file` | [`UploadFileInput`](types#uploadfileinput) | Buffer, Uint8Array, Blob, or ReadableStream | | `options` | [`UploadFileOptions`](types#uploadfileoptions) | Upload options | **Returns** `Promise`\<[`SonioxFile`](classes#sonioxfile)> The uploaded file metadata **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors **Throws** `Error` On validation errors (file too large, invalid input) **Examples** ```typescript import * as fs from 'node:fs'; const buffer = await fs.promises.readFile('/path/to/audio.mp3'); const file = await client.files.upload(buffer, { filename: 'audio.mp3' }); ``` ```typescript const file = await client.files.upload(Bun.file('/path/to/audio.mp3')); ``` ```typescript const file = await client.files.upload(buffer, { filename: 'audio.mp3', client_reference_id: 'order-12345', }); ``` ```typescript const controller = new AbortController(); setTimeout(() => controller.abort(), 30000); const file = await client.files.upload(buffer, { filename: 'audio.mp3', signal: controller.signal, }); ``` *** ## SonioxSttApi ### count() ```ts count(signal?): Promise; ``` Returns the total number of transcriptions, split by request scope. **Parameters** | Parameter | Type | Description | | --------- | ------------- | ---------------------------------------------- | | `signal?` | `AbortSignal` | Optional AbortSignal for request cancellation. | **Returns** `Promise`\<[`TranscriptionsCountResponse`](types#transcriptionscountresponse)> Total transcription counts for Playground, Public API, and all scopes. **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors. **Example** ```typescript const counts = await client.stt.count(); console.log(counts.total); ``` *** ### create() ```ts create(options, signal?): Promise; ``` Creates a new transcription from audio\_url or file\_id **Parameters** | Parameter | Type | Description | | --------- | ---------------------------------------------------------------- | ------------------------------------------------------- | | `options` | [`CreateTranscriptionOptions`](types#createtranscriptionoptions) | Transcription options including model and audio source. | | `signal?` | `AbortSignal` | - | **Returns** `Promise`\<[`SonioxTranscription`](classes#sonioxtranscription)> The created transcription. **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors. **Example** ```typescript // Transcribe from URL const transcription = await client.stt.create({ model: 'stt-async-v5', audio_url: 'https://soniox.com/media/examples/coffee_shop.mp3', }); // Transcribe from uploaded file const file = await client.files.upload(buffer); const transcription = await client.stt.create({ model: 'stt-async-v5', file_id: file.id, }); // With speaker diarization const transcription = await client.stt.create({ model: 'stt-async-v5', audio_url: 'https://soniox.com/media/examples/coffee_shop.mp3', enable_speaker_diarization: true, }); ``` *** ### delete() ```ts delete(id, signal?): Promise; ``` Permanently deletes a transcription. This operation is idempotent - succeeds even if the transcription doesn't exist. **Parameters** | Parameter | Type | Description | | --------- | ---------------------------------------------------------- | --------------------------------------------------------------- | | `id` | [`TranscriptionIdentifier`](types#transcriptionidentifier) | The UUID of the transcription or a SonioxTranscription instance | | `signal?` | `AbortSignal` | - | **Returns** `Promise`\<`void`> **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors (except 404) **Example** ```typescript // Delete by ID await client.stt.delete('550e8400-e29b-41d4-a716-446655440000'); // Or delete a transcription instance const transcription = await client.stt.get('550e8400-e29b-41d4-a716-446655440000'); if (transcription) { await client.stt.delete(transcription); } ``` *** ### delete\_all() ```ts delete_all(options): Promise; ``` Permanently deletes all transcriptions. Iterates through all pages of transcriptions and deletes each one. **Parameters** | Parameter | Type | Description | | --------- | ------------------------------------------------------------------------ | ---------------- | | `options` | [`DeleteAllTranscriptionsOptions`](types#deletealltranscriptionsoptions) | Optional signal. | **Returns** `Promise`\<`void`> **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors. **Throws** `Error` If the operation is aborted via signal. **Example** ```typescript // Delete all transcriptions await client.stt.delete_all(); console.log(`Deleted all transcriptions.`); // With cancellation const controller = new AbortController(); await client.stt.delete_all({ signal: controller.signal }); ``` *** ### destroy() ```ts destroy(id): Promise; ``` Permanently deletes a transcription and its associated file (if any). This operation is idempotent - succeeds even if resources don't exist. **Parameters** | Parameter | Type | Description | | --------- | ---------------------------------------------------------- | --------------------------------------------------------------- | | `id` | [`TranscriptionIdentifier`](types#transcriptionidentifier) | The UUID of the transcription or a SonioxTranscription instance | **Returns** `Promise`\<`void`> **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors (except 404) **Example** ```typescript // Clean up both transcription and uploaded file const transcription = await client.stt.transcribe({ model: 'stt-async-v5', file: buffer, wait: true, }); // ... use transcription ... await client.stt.destroy(transcription); // Deletes both // Or by ID await client.stt.destroy('550e8400-e29b-41d4-a716-446655440000'); ``` *** ### destroy\_all() ```ts destroy_all(options): Promise; ``` Permanently deletes all transcriptions and their associated files. Iterates through all pages of transcriptions and calls [destroy](#destroy) on each one, removing both the transcription and its uploaded file. **Parameters** | Parameter | Type | Description | | --------- | ------------------------------------------------------------------------ | -------------------------------------- | | `options` | [`DeleteAllTranscriptionsOptions`](types#deletealltranscriptionsoptions) | Optional signal and progress callback. | **Returns** `Promise`\<`void`> The number of transcriptions destroyed. **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors. **Throws** `Error` If the operation is aborted via signal. **Example** ```typescript // Destroy all transcriptions and their files await client.stt.destroy_all(); console.log(`Destroyed all transcriptions and their files.`); // With cancellation const controller = new AbortController(); await client.stt.destroy_all({ signal: controller.signal }); ``` *** ### get() ```ts get(id, signal?): Promise; ``` Retrieves a transcription by ID **Parameters** | Parameter | Type | Description | | --------- | ---------------------------------------------------------- | ---------------------------------------------------------------- | | `id` | [`TranscriptionIdentifier`](types#transcriptionidentifier) | The UUID of the transcription or a SonioxTranscription instance. | | `signal?` | `AbortSignal` | - | **Returns** `Promise`\<[`SonioxTranscription`](classes#sonioxtranscription) | `null`> The transcription, or null if not found. **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors (except 404). **Example** ```typescript const transcription = await client.stt.get('550e8400-e29b-41d4-a716-446655440000'); if (transcription) { console.log(transcription.status, transcription.model); } ``` *** ### getTranscript() ```ts getTranscript(id, signal?): Promise; ``` Retrieves the full transcript text and tokens for a completed transcription. Only available for successfully completed transcriptions. **Parameters** | Parameter | Type | Description | | --------- | ---------------------------------------------------------- | --------------------------------------------------------------- | | `id` | [`TranscriptionIdentifier`](types#transcriptionidentifier) | The UUID of the transcription or a SonioxTranscription instance | | `signal?` | `AbortSignal` | - | **Returns** `Promise`\<[`SonioxTranscript`](classes#sonioxtranscript) | `null`> The transcript with text and detailed tokens, or null if not found **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors (except 404) **Example** ```typescript const transcript = await client.stt.getTranscript('550e8400-e29b-41d4-a716-446655440000'); if (transcript) { console.log(transcript.text); for (const token of transcript.tokens) { console.log(token.text, token.start_ms, token.end_ms, token.confidence); } } ``` *** ### list() ```ts list(options, signal?): Promise; ``` Retrieves list of transcriptions The returned result is async iterable - use `for await...of` to iterate through all pages **Parameters** | Parameter | Type | Description | | --------- | -------------------------------------------------------------- | ------------------------------------------ | | `options` | [`ListTranscriptionsOptions`](types#listtranscriptionsoptions) | Optional pagination and filter parameters. | | `signal?` | `AbortSignal` | - | **Returns** `Promise`\<[`TranscriptionListResult`](classes#transcriptionlistresult)> TranscriptionListResult with async iteration support. **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors. **Example** ```typescript const result = await client.stt.list(); // Automatic paging - iterates through ALL transcriptions across all pages for await (const transcription of result) { console.log(transcription.id, transcription.status); } // Or access just the first page for (const transcription of result.transcriptions) { console.log(transcription.id); } // Check if there are more pages if (result.isPaged()) { console.log('More pages available'); } ``` *** ### transcribe() ```ts transcribe(options): Promise; ``` Unified transcribe method - supports direct file upload When `file` is provided, uploads it first then creates a transcription When `wait: true`, waits for completion before returning When `cleanup` is specified (requires `wait: true`), cleans up resources after completion or on error/timeout **Parameters** | Parameter | Type | Description | | --------- | ---------------------------------------------- | -------------------------------------------------------------------- | | `options` | [`TranscribeOptions`](types#transcribeoptions) | Transcribe options including model, audio source, and wait settings. | **Returns** `Promise`\<[`SonioxTranscription`](classes#sonioxtranscription)> The transcription (completed if wait=true, otherwise in queued/processing state). **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors. **Throws** `Error` On validation errors or wait timeout. **Example** ```typescript // Transcribe from URL and wait for completion const result = await client.stt.transcribe({ model: 'stt-async-v5', audio_url: 'https://soniox.com/media/examples/coffee_shop.mp3', wait: true, }); // Upload file and transcribe in one call const result = await client.stt.transcribe({ model: 'stt-async-v5', file: buffer, // or Blob, ReadableStream filename: 'meeting.mp3', enable_speaker_diarization: true, wait: true, }); // With wait progress callback const result = await client.stt.transcribe({ model: 'stt-async-v5', file: buffer, wait: true, wait_options: { interval_ms: 2000, on_status_change: (status) => console.log(`Status: ${status}`), }, }); // Auto-cleanup uploaded file after transcription const result = await client.stt.transcribe({ model: 'stt-async-v5', file: buffer, wait: true, cleanup: ['file'], // Deletes uploaded file, keeps transcription record }); // Auto-cleanup everything after transcription const result = await client.stt.transcribe({ model: 'stt-async-v5', file: buffer, wait: true, cleanup: ['file', 'transcription'], // Deletes both file and transcription record }); ``` *** ### transcribeFromFile() ```ts transcribeFromFile(file, options): Promise; ``` Wrapper to transcribe from raw file data. **Parameters** | Parameter | Type | Description | | --------- | -------------------------------------------------------------- | ------------------------------------------- | | `file` | [`UploadFileInput`](types#uploadfileinput) | Buffer, Uint8Array, Blob, or ReadableStream | | `options` | [`TranscribeFromFileOptions`](types#transcribefromfileoptions) | Transcription options (excluding file) | **Returns** `Promise`\<[`SonioxTranscription`](classes#sonioxtranscription)> The transcription (completed if wait=true, otherwise in queued/processing state). *** ### transcribeFromFileId() ```ts transcribeFromFileId(file_id, options): Promise; ``` Wrapper to transcribe from an uploaded file ID. **Parameters** | Parameter | Type | Description | | --------- | ------------------------------------------------------------------ | ------------------------------------------ | | `file_id` | `string` | ID of a previously uploaded file | | `options` | [`TranscribeFromFileIdOptions`](types#transcribefromfileidoptions) | Transcription options (excluding file\_id) | **Returns** `Promise`\<[`SonioxTranscription`](classes#sonioxtranscription)> The transcription (completed if wait=true, otherwise in queued/processing state). *** ### transcribeFromUrl() ```ts transcribeFromUrl(audio_url, options): Promise; ``` Wrapper to transcribe from a URL. **Parameters** | Parameter | Type | Description | | ----------- | ------------------------------------------------------------ | -------------------------------------------- | | `audio_url` | `string` | Publicly accessible audio URL | | `options` | [`TranscribeFromUrlOptions`](types#transcribefromurloptions) | Transcription options (excluding audio\_url) | **Returns** `Promise`\<[`SonioxTranscription`](classes#sonioxtranscription)> The transcription (completed if wait=true, otherwise in queued/processing state). *** ### translate() ```ts translate(options): Promise; ``` Starts a translation job. Sugar around [SonioxSttApi.transcribe](#transcribe): configures translation from the supplied [TranslateMode](types#translatemode) (`{ to }`, `{ to, from }`, or `{ between }`) and returns a [SonioxTranslationJob](classes#sonioxtranslationjob). With `wait: true`, waits for completion before returning, mirroring [SonioxSttApi.transcribe](#transcribe). Method-owned fields (cannot be set by the caller): * `translation` — derived from `to` / `between`. * `language_hints` + `language_hints_strict: true` — derived from `from` / `between`. * `enable_language_identification: true` — always. The reshape step relies on every token carrying a `language` field to group originals with their translations. * `fetch_transcript` — derived from `fetch_translation` when `wait=true`. **Parameters** | Parameter | Type | Description | | --------- | -------------------------------------------- | --------------------------------------------- | | `options` | [`TranslateOptions`](types#translateoptions) | Mode, audio source, and pass-through options. | **Returns** `Promise`\<[`SonioxTranslationJob`](classes#sonioxtranslationjob)> The translation job. **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors. **Throws** `Error` On validation errors or wait timeout. **Example** ```typescript // One-way: detect any source language, translate to Spanish. const job = await client.stt.translate({ audio_url: 'https://soniox.com/media/examples/coffee_shop.mp3', to: 'es', }); const completed = await job.wait(); const result = await completed.getTranslation(); console.log(result?.translation_text); // One-way with explicit source language. const explicitJob = await client.stt.translate({ file: buffer, filename: 'meeting.mp3', from: 'en', to: 'es', wait: true, }); console.log(explicitJob.translation?.translation_text); // Two-way: bidirectional translation between two languages. const twoWayJob = await client.stt.translate({ audio_url: 'https://soniox.com/media/examples/coffee_shop.mp3', between: ['en', 'es'], enable_speaker_diarization: true, wait: true, }); for (const seg of twoWayJob.translation?.segments ?? []) { console.log(`[${seg.from}] ${seg.original_text}`); if (seg.translation_text) { console.log(` → [${seg.to}] ${seg.translation_text}`); } } ``` *** ### wait() ```ts wait(id, options?): Promise; ``` Waits for a transcription to complete **Parameters** | Parameter | Type | Description | | ---------- | ---------------------------------------------------------- | ---------------------------------------------------------------- | | `id` | [`TranscriptionIdentifier`](types#transcriptionidentifier) | The UUID of the transcription or a SonioxTranscription instance. | | `options?` | [`WaitOptions`](types#waitoptions) | Wait options including polling interval, timeout, and callbacks. | **Returns** `Promise`\<[`SonioxTranscription`](classes#sonioxtranscription)> The completed or errored transcription. **Throws** `Error` If the wait times out or is aborted. **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors. **Example** ```typescript const completed = await client.stt.wait('550e8400-e29b-41d4-a716-446655440000'); // With progress callback const completed = await client.stt.wait('id', { interval_ms: 2000, on_status_change: (status) => console.log(`Status: ${status}`), }); ``` *** ## SonioxModelsAPI ### list() ```ts list(signal?): Promise; ``` List of available models and their attributes. **Parameters** | Parameter | Type | Description | | --------- | ------------- | ------------------------------------- | | `signal?` | `AbortSignal` | Optional AbortSignal for cancellation | **Returns** `Promise`\<[`SonioxModel`](types#sonioxmodel)\[]> List of available models and their attributes. **See** [https://soniox.com/docs/stt/api-reference/models/get\_models](https://soniox.com/docs/stt/api-reference/models/get_models) *** ## SonioxWebhooksAPI Webhook utilities API accessible via client.webhooks Provides methods for handling incoming Soniox webhook requests. When used via the client, results include lazy fetch helpers for transcripts. ### getAuthFromEnv() ```ts getAuthFromEnv(): WebhookAuthConfig | undefined; ``` Get webhook authentication configuration from environment variables. Reads `SONIOX_API_WEBHOOK_HEADER` and `SONIOX_API_WEBHOOK_SECRET` environment variables. Returns undefined if either variable is not set (both are required for authentication). **Returns** [`WebhookAuthConfig`](types#webhookauthconfig) | `undefined` *** ### handle() ```ts handle(options): WebhookHandlerResultWithFetch; ``` Framework-agnostic webhook handler **Parameters** | Parameter | Type | | --------- | ---------------------------------------------------- | | `options` | [`HandleWebhookOptions`](types#handlewebhookoptions) | **Returns** [`WebhookHandlerResultWithFetch`](types#webhookhandlerresultwithfetch) *** ### handleExpress() ```ts handleExpress(req, auth?): WebhookHandlerResultWithFetch; ``` Handle a webhook from an Express-like request **Parameters** | Parameter | Type | | --------- | ------------------------------------------------ | | `req` | [`ExpressLikeRequest`](types#expresslikerequest) | | `auth?` | [`WebhookAuthConfig`](types#webhookauthconfig) | **Returns** [`WebhookHandlerResultWithFetch`](types#webhookhandlerresultwithfetch) **Example** ```typescript app.post('/webhook', async (req, res) => { const result = soniox.webhooks.handleExpress(req); if (result.ok && result.event.status === 'completed') { const transcript = await result.fetchTranscript(); console.log(transcript?.text); } res.status(result.status).json({ received: true }); }); ``` *** ### handleFastify() ```ts handleFastify(req, auth?): WebhookHandlerResultWithFetch; ``` Handle a webhook from a Fastify request **Parameters** | Parameter | Type | | --------- | ------------------------------------------------ | | `req` | [`FastifyLikeRequest`](types#fastifylikerequest) | | `auth?` | [`WebhookAuthConfig`](types#webhookauthconfig) | **Returns** [`WebhookHandlerResultWithFetch`](types#webhookhandlerresultwithfetch) *** ### handleHono() ```ts handleHono(c, auth?): Promise; ``` Handle a webhook from a Hono context **Parameters** | Parameter | Type | | --------- | ---------------------------------------------- | | `c` | [`HonoLikeContext`](types#honolikecontext) | | `auth?` | [`WebhookAuthConfig`](types#webhookauthconfig) | **Returns** `Promise`\<[`WebhookHandlerResultWithFetch`](types#webhookhandlerresultwithfetch)> *** ### handleNestJS() ```ts handleNestJS(req, auth?): WebhookHandlerResultWithFetch; ``` Handle a webhook from a NestJS request **Parameters** | Parameter | Type | | --------- | ---------------------------------------------- | | `req` | [`NestJSLikeRequest`](types#nestjslikerequest) | | `auth?` | [`WebhookAuthConfig`](types#webhookauthconfig) | **Returns** [`WebhookHandlerResultWithFetch`](types#webhookhandlerresultwithfetch) *** ### handleRequest() ```ts handleRequest(request, auth?): Promise; ``` Handle a webhook from a Fetch API Request **Parameters** | Parameter | Type | | --------- | ---------------------------------------------- | | `request` | `Request` | | `auth?` | [`WebhookAuthConfig`](types#webhookauthconfig) | **Returns** `Promise`\<[`WebhookHandlerResultWithFetch`](types#webhookhandlerresultwithfetch)> *** ### isEvent() ```ts isEvent(payload): payload is WebhookEvent; ``` Type guard to check if a value is a valid WebhookEvent **Parameters** | Parameter | Type | | --------- | --------- | | `payload` | `unknown` | **Returns** `payload is WebhookEvent` *** ### parseEvent() ```ts parseEvent(payload): WebhookEvent; ``` Parse and validate a webhook event payload **Parameters** | Parameter | Type | | --------- | --------- | | `payload` | `unknown` | **Returns** [`WebhookEvent`](types#webhookevent) *** ### verifyAuth() ```ts verifyAuth(headers, auth): boolean; ``` Verify webhook authentication header **Parameters** | Parameter | Type | | --------- | ---------------------------------------------- | | `headers` | [`WebhookHeaders`](types#webhookheaders) | | `auth` | [`WebhookAuthConfig`](types#webhookauthconfig) | **Returns** `boolean` *** ## SonioxAuthAPI ### createTemporaryKey() ```ts createTemporaryKey(request, signal?): Promise; ``` Creates a temporary API key for client-side use. **Parameters** | Parameter | Type | Description | | --------- | -------------------------------------------------------- | ---------------------------------------- | | `request` | [`TemporaryApiKeyRequest`](types#temporaryapikeyrequest) | Request parameters for the temporary key | | `signal?` | `AbortSignal` | Optional AbortSignal for cancellation | **Returns** `Promise`\<[`TemporaryApiKeyResponse`](types#temporaryapikeyresponse)> The temporary API key response **Example** ```typescript const sttKey = await client.auth.createTemporaryKey({ usage_type: 'transcribe_websocket', expires_in_seconds: 300, }); const ttsKey = await client.auth.createTemporaryKey({ usage_type: 'tts_rt', expires_in_seconds: 300, single_use: true, max_session_duration_seconds: 600, }); ``` *** ## SonioxConcurrencyLimitsAPI ### get() ```ts get(signal?): Promise; ``` Retrieves current concurrency counts and configured limits. Values are region-scoped according to the client's configured REST API endpoint. **Parameters** | Parameter | Type | Description | | --------- | ------------- | -------------------------------------- | | `signal?` | `AbortSignal` | Optional AbortSignal for cancellation. | **Returns** `Promise`\<[`ConcurrencyLimitsResponse`](types#concurrencylimitsresponse)> Current counts and configured limits for project and organization scopes. **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors. **Example** ```typescript const limits = await client.concurrencyLimits.get(); console.log(limits.project.current.transcribe_concurrent); console.log(limits.project.limits.transcribe_concurrent); ``` *** ### getHistory() ```ts getHistory(options): Promise; ``` Retrieves historical concurrent stream aggregates for the project. Returns every aggregation period in the requested window (no gaps). Periods with no recorded activity have every numeric field set to `0`. **Parameters** | Parameter | Type | Description | | --------- | -------------------------------------------------------------------------------- | --------------------------------------------------------------------------------- | | `options` | [`GetConcurrentStreamsHistoryOptions`](types#getconcurrentstreamshistoryoptions) | Required time window, aggregation period, stream kind, and optional cancellation. | **Returns** `Promise`\<[`ConcurrentStreamsHistoryResponse`](types#concurrentstreamshistoryresponse)> Concurrent streams history for the requested kind. **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors. **Example** ```typescript const history = await client.concurrencyLimits.getHistory({ start_time: '2026-04-28T09:00:00Z', end_time: '2026-04-28T10:00:00Z', period_sec: 60, kind: 'stt', }); for (const entry of history.entries) { console.log(entry.period_start, entry.sample_max); } ``` *** ## SonioxUsageLogsAPI ### getSummary() ```ts getSummary(options): Promise; ``` Retrieves an aggregated usage summary for the project over a time window. Returns totals across all models plus one entry per model that recorded usage. Per-day series are aligned to each entry's `days` array. **Parameters** | Parameter | Type | Description | | --------- | -------------------------------------------------------- | ------------------------------------------------ | | `options` | [`GetUsageSummaryOptions`](types#getusagesummaryoptions) | Required time window plus optional cancellation. | **Returns** `Promise`\<[`UsageSummaryResponse`](types#usagesummaryresponse)> Aggregated usage summary. **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors. **Example** ```typescript const summary = await client.usageLogs.getSummary({ start_time: '2026-04-01T00:00:00Z', end_time: '2026-04-03T00:00:00Z', }); console.log(summary.total.total_cost_usd); for (const model of summary.models) { console.log(model.model, model.total_num_requests); } ``` *** ### list() ```ts list(options): Promise; ``` Retrieves per-request usage log entries for the project. The returned result is async iterable. Use `for await...of` to iterate through all pages. **Parameters** | Parameter | Type | Description | | --------- | ---------------------------------------------------- | ------------------------------------------------------------------------- | | `options` | [`ListUsageLogsOptions`](types#listusagelogsoptions) | Required time window plus optional pagination, sorting, and cancellation. | **Returns** `Promise`\<[`UsageLogListResult`](classes#usageloglistresult)> UsageLogListResult with async iteration support. **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors. **Example** ```typescript const result = await client.usageLogs.list({ start_time: '2026-04-28T09:00:00Z', end_time: '2026-04-29T09:00:00Z', sort: 'end_time_desc', }); for await (const log of result) { console.log(log.model, log.cost_usd); } ``` *** ## SonioxRealtimeApi Real-time API factory for creating STT sessions and TTS connections. ### Examples ```typescript const session = client.realtime.stt({ model: 'stt-rt-v5' }); await session.connect(); ``` ```typescript const stream = await client.realtime.tts({ model: 'tts-rt-v2', voice: 'Adrian', language: 'en', audio_format: 'wav', }); stream.sendText("Hello"); stream.finish(); for await (const chunk of stream) { ... } ``` ```typescript const conn = await client.realtime.tts.multiStream(); const stream = await conn.stream({ model: 'tts-rt-v2', voice: 'Adrian', language: 'en', audio_format: 'wav', }); ``` ### stt() ```ts stt(config, options?): RealtimeSttSession; ``` Create a new Speech-to-Text session. `config` is shallow-merged on top of `stt_defaults` from the client options; caller-provided fields override the defaults. **Parameters** | Parameter | Type | | ---------- | ---------------------------------------------- | | `config` | [`SttSessionConfig`](types#sttsessionconfig) | | `options?` | [`SttSessionOptions`](types#sttsessionoptions) | **Returns** [`RealtimeSttSession`](classes#realtimesttsession) ### Properties | Property | Type | | -------- | ------------ | | `tts` | `TtsFactory` | *** ## SonioxFile Uploaded file ### Constructor ```ts new SonioxFile(data, _http): SonioxFile; ``` **Parameters** | Parameter | Type | | --------- | ---------------------------------------- | | `data` | [`SonioxFileData`](types#sonioxfiledata) | | `_http` | [`HttpClient`](types#httpclient) | **Returns** `SonioxFile` ### delete() ```ts delete(signal?): Promise; ``` Permanently deletes this file. This operation is idempotent - succeeds even if the file doesn't exist. **Parameters** | Parameter | Type | Description | | --------- | ------------- | ------------------------------------- | | `signal?` | `AbortSignal` | Optional AbortSignal for cancellation | **Returns** `Promise`\<`void`> **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors (except 404) **Example** ```typescript const file = await client.files.get('550e8400-e29b-41d4-a716-446655440000'); if (file) { await file.delete(); } ``` *** ### toJSON() ```ts toJSON(): SonioxFileData; ``` Returns the raw data for this file. **Returns** [`SonioxFileData`](types#sonioxfiledata) ### Properties | Property | Type | | --------------------- | ----------------------- | | `client_reference_id` | `string` \| `undefined` | | `created_at` | `string` | | `filename` | `string` | | `id` | `string` | | `size` | `number` | *** ## SonioxTranscription A Transcription instance ### Extended by * [`SonioxTranslationJob`](classes#sonioxtranslationjob) ### Constructor ```ts new SonioxTranscription( data, _http, transcript?): SonioxTranscription; ``` **Parameters** | Parameter | Type | | ------------- | ---------------------------------------------------------- | | `data` | [`SonioxTranscriptionData`](types#sonioxtranscriptiondata) | | `_http` | [`HttpClient`](types#httpclient) | | `transcript?` | [`SonioxTranscript`](classes#sonioxtranscript) \| `null` | **Returns** `SonioxTranscription` ### delete() ```ts delete(): Promise; ``` Permanently deletes this transcription. This operation is idempotent - succeeds even if the transcription doesn't exist. **Returns** `Promise`\<`void`> **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors (except 404) **Example** ```typescript const transcription = await client.stt.get('550e8400-e29b-41d4-a716-446655440000'); await transcription.delete(); ``` *** ### destroy() ```ts destroy(): Promise; ``` Permanently deletes this transcription and its associated file (if any). This operation is idempotent - succeeds even if resources don't exist. **Returns** `Promise`\<`void`> **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors (except 404) **Example** ```typescript // Clean up both transcription and uploaded file const transcription = await client.stt.transcribe({ model: 'stt-async-v5', file: buffer, wait: true, }); // ... use transcription ... await transcription.destroy(); // Deletes both transcription and file ``` *** ### getTranscript() ```ts getTranscript(options?): Promise; ``` Retrieves the full transcript text and tokens for this transcription. Only available for successfully completed transcriptions. Returns cached transcript if available (when using `transcribe()` with `wait: true`). Use `force: true` to bypass the cache and fetch fresh data from the API. **Parameters** | Parameter | Type | Description | | ----------------- | --------------------------------------------------- | -------------------------------------------------------- | | `options?` | \{ `force?`: `boolean`; `signal?`: `AbortSignal`; } | Optional settings | | `options.force?` | `boolean` | If true, bypasses cached transcript and fetches from API | | `options.signal?` | `AbortSignal` | Optional AbortSignal for request cancellation | **Returns** `Promise`\<[`SonioxTranscript`](classes#sonioxtranscript) | `null`> The transcript with text and detailed tokens, or null if not found. **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors (except 404). **Example** ```typescript const transcription = await client.stt.get('550e8400-e29b-41d4-a716-446655440000'); if (transcription) { const transcript = await transcription.getTranscript(); if (transcript) { console.log(transcript.text); } } // Force re-fetch from API const freshTranscript = await transcription.getTranscript({ force: true }); ``` *** ### refresh() ```ts refresh(signal?): Promise; ``` Re-fetches this transcription to get the latest status. **Parameters** | Parameter | Type | Description | | --------- | ------------- | ---------------------------------------------- | | `signal?` | `AbortSignal` | Optional AbortSignal for request cancellation. | **Returns** `Promise`\<`SonioxTranscription`> A new SonioxTranscription instance with updated data. **Throws** [SonioxHttpError](classes#sonioxhttperror) **Example** ```typescript let transcription = await client.stt.get('550e8400-e29b-41d4-a716-446655440000'); transcription = await transcription.refresh(); console.log(transcription.status); ``` *** ### toJSON() ```ts toJSON(): SonioxTranscriptionData; ``` Returns the raw data for this transcription. **Returns** [`SonioxTranscriptionData`](types#sonioxtranscriptiondata) *** ### wait() ```ts wait(options): Promise; ``` Waits for the transcription to complete or fail. Polls the API at the specified interval until the status is 'completed' or 'error'. **Parameters** | Parameter | Type | Description | | --------- | ---------------------------------- | ---------------------------------------------------------------- | | `options` | [`WaitOptions`](types#waitoptions) | Wait options including polling interval, timeout, and callbacks. | **Returns** `Promise`\<`SonioxTranscription`> The completed or errored transcription. **Throws** `Error` If the wait times out or is aborted. **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors. **Example** ```typescript const transcription = await client.stt.create({ model: 'stt-async-v5', audio_url: 'https://soniox.com/media/examples/coffee_shop.mp3', }); // Simple wait const completed = await transcription.wait(); // Wait with progress callback const completed = await transcription.wait({ interval_ms: 2000, on_status_change: (status) => console.log(`Status: ${status}`), }); ``` ### Properties | Property | Type | Description | | -------------------------------- | -------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `audio_duration_ms` | `number` \| `null` \| `undefined` | Duration of the audio in milliseconds. Only available after processing begins. | | `audio_url` | `string` \| `null` \| `undefined` | URL of the audio file being transcribed. | | `client_reference_id` | `string` \| `null` \| `undefined` | Optional tracking identifier. | | `context` | \| [`TranscriptionContext`](types#transcriptioncontext) \| `null` \| `undefined` | Additional context provided for the transcription. | | `created_at` | `string` | UTC timestamp when the transcription was created. | | `enable_language_identification` | `boolean` | When true, language is detected for each part of the transcription. | | `enable_speaker_diarization` | `boolean` | When true, speakers are identified and separated in the transcription output. | | `error_message` | `string` \| `null` \| `undefined` | Error message if transcription failed. | | `error_type` | `string` \| `null` \| `undefined` | Error type if transcription failed. | | `file_id` | `string` \| `null` \| `undefined` | ID of the uploaded file being transcribed. | | `filename` | `string` | Name of the file being transcribed. | | `id` | `string` | Unique identifier of the transcription. | | `language_hints` | `string`\[] \| `undefined` | Expected languages in the audio. | | `model` | `string` | Speech-to-text model used. | | `status` | [`TranscriptionStatus`](types#transcriptionstatus) | Current status of the transcription. | | `transcript` | [`SonioxTranscript`](classes#sonioxtranscript) \| `null` \| `undefined` | Pre-fetched transcript. Only available when using `transcribe()` with `wait: true`, `fetch_transcript !== false`, and the transcription completed successfully | | `webhook_auth_header_name` | `string` \| `null` \| `undefined` | Name of the authentication header sent with webhook notifications. | | `webhook_auth_header_value` | `string` \| `null` \| `undefined` | Authentication header value (masked). | | `webhook_status_code` | `number` \| `null` \| `undefined` | HTTP status code received when webhook was delivered. | | `webhook_url` | `string` \| `null` \| `undefined` | URL to receive webhook notifications. | *** ## SonioxTranscript A Transcript result containing the transcribed text and tokens. ### Constructor ```ts new SonioxTranscript(data): SonioxTranscript; ``` **Parameters** | Parameter | Type | | --------- | ------------------------------------------------ | | `data` | [`TranscriptResponse`](types#transcriptresponse) | **Returns** `SonioxTranscript` ### segments() ```ts segments(options?): TranscriptSegment[]; ``` Groups tokens into segments based on specified grouping keys. A new segment starts when any of the `group_by` fields changes. **Parameters** | Parameter | Type | Description | | ---------- | ------------------------------------------------------------ | -------------------- | | `options?` | [`SegmentTranscriptOptions`](types#segmenttranscriptoptions) | Segmentation options | **Returns** [`TranscriptSegment`](types#transcriptsegment)\[] Array of segments with combined text and timing **Example** ```typescript const transcript = await transcription.getTranscript(); // Group by both speaker and language (default) const segments = transcript.segments(); // Group by speaker only const bySpeaker = transcript.segments({ group_by: ['speaker'] }); for (const s of segments) { console.log(`[Speaker ${s.speaker}] ${s.text}`); } ``` ### Properties | Property | Type | Description | | -------- | --------------------------------------------- | ------------------------------------------------------------------ | | `id` | `string` | Unique identifier of the transcription this transcript belongs to. | | `text` | `string` | Complete transcribed text content. | | `tokens` | [`TranscriptToken`](types#transcripttoken)\[] | List of detailed token information with timestamps and metadata. | *** ## FileListResult Result set for file listing ### Constructor ```ts new FileListResult( initialResponse, _http, _limit, _signal): FileListResult; ``` **Parameters** | Parameter | Type | Default value | | ----------------- | ------------------------------------------------------------------------------------------ | ------------- | | `initialResponse` | [`ListFilesResponse`](types#listfilesresponset)\<[`SonioxFileData`](types#sonioxfiledata)> | `undefined` | | `_http` | [`HttpClient`](types#httpclient) | `undefined` | | `_limit` | `number` \| `undefined` | `undefined` | | `_signal` | `AbortSignal` \| `undefined` | `undefined` | **Returns** `FileListResult` ### \[asyncIterator]\() ```ts asyncIterator: AsyncIterator; ``` Async iterator that automatically fetches all pages Use with `for await...of` to iterate through all files **Returns** `AsyncIterator`\<[`SonioxFile`](classes#sonioxfile)> *** ### isPaged() ```ts isPaged(): boolean; ``` Returns true if there are more pages of results beyond the first page **Returns** `boolean` *** ### toJSON() ```ts toJSON(): ListFilesResponse; ``` Returns the raw data for this list result. Also used by JSON.stringify() to prevent serialization of internal HTTP client. **Returns** [`ListFilesResponse`](types#listfilesresponset)\<[`SonioxFileData`](types#sonioxfiledata)> ### Properties | Property | Type | Description | | ------------------ | ------------------------------------- | ---------------------------------------------------------- | | `files` | [`SonioxFile`](classes#sonioxfile)\[] | Files from the first page of results | | `next_page_cursor` | `string` \| `null` | Pagination cursor for the next page. Null if no more pages | *** ## TranscriptionListResult Result set for transcription listing. ### Constructor ```ts new TranscriptionListResult( initialResponse, _http, _options, _signal?): TranscriptionListResult; ``` **Parameters** | Parameter | Type | | ----------------- | ------------------------------------------------------------------------------------------------------------------------------ | | `initialResponse` | [`ListTranscriptionsResponse`](types#listtranscriptionsresponset)\<[`SonioxTranscriptionData`](types#sonioxtranscriptiondata)> | | `_http` | [`HttpClient`](types#httpclient) | | `_options` | [`ListTranscriptionsOptions`](types#listtranscriptionsoptions) | | `_signal?` | `AbortSignal` | **Returns** `TranscriptionListResult` ### \[asyncIterator]\() ```ts asyncIterator: AsyncIterator; ``` Async iterator that automatically fetches all pages. Use with `for await...of` to iterate through all transcriptions. **Returns** `AsyncIterator`\<[`SonioxTranscription`](classes#sonioxtranscription)> *** ### isPaged() ```ts isPaged(): boolean; ``` Returns true if there are more pages of results beyond the first page. **Returns** `boolean` *** ### toJSON() ```ts toJSON(): ListTranscriptionsResponse; ``` Returns the raw data for this list result **Returns** [`ListTranscriptionsResponse`](types#listtranscriptionsresponset)\<[`SonioxTranscriptionData`](types#sonioxtranscriptiondata)> ### Properties | Property | Type | Description | | ------------------ | ------------------------------------------------------- | ----------------------------------------------------------- | | `next_page_cursor` | `string` \| `null` | Pagination cursor for the next page. Null if no more pages. | | `transcriptions` | [`SonioxTranscription`](classes#sonioxtranscription)\[] | Transcriptions from the first page of results. | *** ## UsageLogListResult Result set for usage log listing. ### Constructor ```ts new UsageLogListResult( initialResponse, _http, _options): UsageLogListResult; ``` **Parameters** | Parameter | Type | | ----------------- | ------------------------------------------------------ | | `initialResponse` | [`ListUsageLogsResponse`](types#listusagelogsresponse) | | `_http` | [`HttpClient`](types#httpclient) | | `_options` | [`ListUsageLogsOptions`](types#listusagelogsoptions) | **Returns** `UsageLogListResult` ### \[asyncIterator]\() ```ts asyncIterator: AsyncIterator; ``` Async iterator that automatically fetches all pages. Use with `for await...of` to iterate through all usage logs. **Returns** `AsyncIterator`\<[`SonioxUsageLog`](types#sonioxusagelog)> *** ### isPaged() ```ts isPaged(): boolean; ``` Returns true if there are more pages of results beyond the first page. **Returns** `boolean` *** ### toJSON() ```ts toJSON(): ListUsageLogsResponse; ``` Returns the raw data for this list result. **Returns** [`ListUsageLogsResponse`](types#listusagelogsresponse) ### Properties | Property | Type | Description | | ------------------ | ------------------------------------------- | ----------------------------------------------------------- | | `next_page_cursor` | `string` \| `null` | Pagination cursor for the next page. Null if no more pages. | | `usage_logs` | [`SonioxUsageLog`](types#sonioxusagelog)\[] | Usage logs from the first page of results. | *** ## RealtimeSttSession Real-time Speech-to-Text session Provides WebSocket-based streaming transcription with support for: * Event-based and async iterator consumption * Pause/resume with automatic keepalive while paused * AbortSignal cancellation ### Example ```typescript const session = new RealtimeSttSession(apiKey, wsUrl, { model: 'stt-rt-v5' }); session.on('result', (result) => { console.log(result.tokens.map(t => t.text).join('')); }); await session.connect(); session.sendAudio(audioChunk); await session.finish(); ``` ### paused ```ts get paused(): boolean; ``` Whether the session is currently paused. **Returns** `boolean` *** ### state ```ts get state(): SttSessionState; ``` Current session state. **Returns** [`SttSessionState`](types#sttsessionstate) ### Constructor ```ts new RealtimeSttSession( apiKey, wsBaseUrl, config, options?): RealtimeSttSession; ``` **Parameters** | Parameter | Type | | ----------- | ---------------------------------------------- | | `apiKey` | `string` | | `wsBaseUrl` | `string` | | `config` | [`SttSessionConfig`](types#sttsessionconfig) | | `options?` | [`SttSessionOptions`](types#sttsessionoptions) | **Returns** `RealtimeSttSession` ### \[asyncIterator]\() ```ts asyncIterator: AsyncIterator; ``` Async iterator for consuming events. The returned iterator's `return()` resets the internal iterator-attach flag and drops any buffered events, so consumers that exit `for await` early (via `break` etc.) stop accruing memory while the session keeps running. **Returns** `AsyncIterator`\<[`RealtimeEvent`](types#realtimeevent)> *** ### close() ```ts close(): void; ``` Close (cancel) the session immediately without waiting **Returns** `void` *** ### connect() ```ts connect(): Promise; ``` Connect to the Soniox WebSocket API. **Returns** `Promise`\<`void`> **Throws** [AbortError](classes#aborterror) If aborted **Throws** [ConnectionError](classes#connectionerror) If connection fails **Throws** [StateError](classes#stateerror) If already connected *** ### finalize() ```ts finalize(options?): void; ``` Requests the server to finalize current transcription **Parameters** | Parameter | Type | | ------------------------------ | -------------------------------------- | | `options?` | \{ `trailing_silence_ms?`: `number`; } | | `options.trailing_silence_ms?` | `number` | **Returns** `void` *** ### finish() ```ts finish(): Promise; ``` Gracefully finish the session **Returns** `Promise`\<`void`> *** ### keepAlive() ```ts keepAlive(): void; ``` Send a keepalive message **Returns** `void` *** ### off() ```ts off(event, handler): this; ``` Remove an event handler **Type Parameters** | Type Parameter | | ---------------------------------------------------------------- | | `E` *extends* keyof [`SttSessionEvents`](types#sttsessionevents) | **Parameters** | Parameter | Type | | --------- | -------------------------------------------------- | | `event` | `E` | | `handler` | [`SttSessionEvents`](types#sttsessionevents)\[`E`] | **Returns** `this` *** ### on() ```ts on(event, handler): this; ``` Register an event handler **Type Parameters** | Type Parameter | | ---------------------------------------------------------------- | | `E` *extends* keyof [`SttSessionEvents`](types#sttsessionevents) | **Parameters** | Parameter | Type | | --------- | -------------------------------------------------- | | `event` | `E` | | `handler` | [`SttSessionEvents`](types#sttsessionevents)\[`E`] | **Returns** `this` *** ### once() ```ts once(event, handler): this; ``` Register a one-time event handler **Type Parameters** | Type Parameter | | ---------------------------------------------------------------- | | `E` *extends* keyof [`SttSessionEvents`](types#sttsessionevents) | **Parameters** | Parameter | Type | | --------- | -------------------------------------------------- | | `event` | `E` | | `handler` | [`SttSessionEvents`](types#sttsessionevents)\[`E`] | **Returns** `this` *** ### pause() ```ts pause(): void; ``` Pause audio transmission and starts automatic keepalive messages **Returns** `void` *** ### resume() ```ts resume(): void; ``` Resume audio transmission **Returns** `void` *** ### sendAudio() ```ts sendAudio(data): void; ``` Send audio data to the server **Parameters** | Parameter | Type | Description | | --------- | ----------- | --------------------------------------- | | `data` | `AudioData` | Audio data as Uint8Array or ArrayBuffer | **Returns** `void` **Throws** [AbortError](classes#aborterror) If aborted **Throws** [StateError](classes#stateerror) If not connected *** ### sendStream() ```ts sendStream(stream, options?): Promise; ``` Stream audio data from an async iterable source. **Parameters** | Parameter | Type | Description | | ---------- | ---------------------------------------------- | ---------------------------------------- | | `stream` | `AsyncIterable`\<`AudioData`> | Async iterable yielding audio chunks | | `options?` | [`SendStreamOptions`](types#sendstreamoptions) | Optional pacing and auto-finish settings | **Returns** `Promise`\<`void`> **Throws** [AbortError](classes#aborterror) If aborted during streaming **Throws** [StateError](classes#stateerror) If not connected *** ## RealtimeSegmentBuffer Rolling buffer for turning real-time results into stable segments. ### size ```ts get size(): number; ``` Number of tokens currently buffered. **Returns** `number` ### Constructor ```ts new RealtimeSegmentBuffer(options?): RealtimeSegmentBuffer; ``` **Parameters** | Parameter | Type | | ---------- | -------------------------------------------------------------------- | | `options?` | [`RealtimeSegmentBufferOptions`](types#realtimesegmentbufferoptions) | **Returns** `RealtimeSegmentBuffer` ### add() ```ts add(result): RealtimeSegment[]; ``` Add a real-time result and return stable segments. **Parameters** | Parameter | Type | | --------- | ---------------------------------------- | | `result` | [`RealtimeResult`](types#realtimeresult) | **Returns** [`RealtimeSegment`](types#realtimesegment)\[] *** ### flushAll() ```ts flushAll(): RealtimeSegment[]; ``` Flush all buffered tokens into segments and clear the buffer. Includes tokens that are not yet stable by final\_audio\_proc\_ms. **Returns** [`RealtimeSegment`](types#realtimesegment)\[] *** ### reset() ```ts reset(): void; ``` Clear all buffered tokens. **Returns** `void` *** ## RealtimeUtteranceBuffer Collects real-time results into utterances for endpoint-driven workflows. ### Constructor ```ts new RealtimeUtteranceBuffer(options?): RealtimeUtteranceBuffer; ``` **Parameters** | Parameter | Type | | ---------- | ------------------------------------------------------------------------ | | `options?` | [`RealtimeUtteranceBufferOptions`](types#realtimeutterancebufferoptions) | **Returns** `RealtimeUtteranceBuffer` ### addResult() ```ts addResult(result): RealtimeSegment[]; ``` Add a real-time result and collect stable segments. **Parameters** | Parameter | Type | | --------- | ---------------------------------------- | | `result` | [`RealtimeResult`](types#realtimeresult) | **Returns** [`RealtimeSegment`](types#realtimesegment)\[] *** ### markEndpoint() ```ts markEndpoint(): RealtimeUtterance | undefined; ``` Mark an endpoint and flush the current utterance. **Returns** [`RealtimeUtterance`](types#realtimeutterance) | `undefined` *** ### reset() ```ts reset(): void; ``` Clear buffered segments and tokens. **Returns** `void` *** ## SonioxError ### Extends * `Error` ### Extended by * [`SonioxHttpError`](classes#sonioxhttperror) * [`RealtimeError`](classes#realtimeerror) ### toJSON() ```ts toJSON(): Record; ``` Converts to a plain object for logging/serialization **Returns** `Record`\<`string`, `unknown`> *** ### toString() ```ts toString(): string; ``` Creates a human-readable string representation **Returns** `string` ### Properties | Property | Type | Description | | ------------ | ---------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- | | `cause` | `unknown` | The underlying error that caused this error, if any. | | `code` | \| `string` & \{ } \| [`SonioxErrorCode`](types#sonioxerrorcode) | Error code describing the type of error. Typed as `string` at the base level to allow subclasses (e.g. HTTP errors) to use their own error code unions. | | `statusCode` | `number` \| `undefined` | HTTP status code when applicable (e.g., 401 for auth errors, 500 for server errors). | *** ## SonioxHttpError HTTP error class for all HTTP-related failures (REST API). Thrown when HTTP requests fail due to network issues, timeouts, server errors, or response parsing failures. ### Extends * [`SonioxError`](classes#sonioxerror) ### toJSON() ```ts toJSON(): Record; ``` Converts to a plain object for logging/serialization **Returns** `Record`\<`string`, `unknown`> **Overrides** [`SonioxError`](classes#sonioxerror).[`toJSON`](classes#sonioxerror-tojson) *** ### toString() ```ts toString(): string; ``` Creates a human-readable string representation **Returns** `string` **Overrides** [`SonioxError`](classes#sonioxerror).[`toString`](classes#sonioxerror-tostring) ### Properties | Property | Type | Description | | ------------ | -------------------------------------------- | ------------------------------------------------------------------------------------ | | `bodyText` | `string` \| `undefined` | Response body text, capped at 4KB (only for http\_error/parse\_error) | | `cause` | `unknown` | The underlying error that caused this error, if any. | | `code` | [`HttpErrorCode`](types#httperrorcode) | Categorized HTTP error code | | `headers` | `Record`\<`string`, `string`> \| `undefined` | Response headers (only for http\_error) | | `method` | [`HttpMethod`](types#httpmethod) | HTTP method | | `statusCode` | `number` \| `undefined` | HTTP status code when applicable (e.g., 401 for auth errors, 500 for server errors). | | `url` | `string` | Request URL | *** ## RealtimeError Base error class for all real-time (WebSocket) SDK errors ### Extends * [`SonioxError`](classes#sonioxerror) ### Extended by * [`AuthError`](classes#autherror) * [`BadRequestError`](classes#badrequesterror) * [`QuotaError`](classes#quotaerror) * [`ConnectionError`](classes#connectionerror) * [`NetworkError`](classes#networkerror) * [`AbortError`](classes#aborterror) * [`StateError`](classes#stateerror) ### toJSON() ```ts toJSON(): Record; ``` Converts to a plain object for logging/serialization **Returns** `Record`\<`string`, `unknown`> **Overrides** [`SonioxError`](classes#sonioxerror).[`toJSON`](classes#sonioxerror-tojson) *** ### toString() ```ts toString(): string; ``` Creates a human-readable string representation **Returns** `string` **Overrides** [`SonioxError`](classes#sonioxerror).[`toString`](classes#sonioxerror-tostring) ### Properties | Property | Type | Description | | ------------ | ---------------------------------------------- | -------------------------------------------------------------------------------------------------- | | `cause` | `unknown` | The underlying error that caused this error, if any. | | `code` | [`RealtimeErrorCode`](types#realtimeerrorcode) | Real-time error code | | `raw` | `unknown` | Original response payload for debugging. Contains the raw WebSocket message that caused the error. | | `statusCode` | `number` \| `undefined` | HTTP status code when applicable (e.g., 401 for auth errors, 500 for server errors). | *** ## AuthError Authentication error (401). Thrown when the API key is invalid or expired. ### Extends * [`RealtimeError`](classes#realtimeerror) ### toJSON() ```ts toJSON(): Record; ``` Converts to a plain object for logging/serialization **Returns** `Record`\<`string`, `unknown`> **Inherited from** [`RealtimeError`](classes#realtimeerror).[`toJSON`](classes#realtimeerror-tojson) *** ### toString() ```ts toString(): string; ``` Creates a human-readable string representation **Returns** `string` **Inherited from** [`RealtimeError`](classes#realtimeerror).[`toString`](classes#realtimeerror-tostring) ### Properties | Property | Type | Description | | ------------ | ---------------------------------------------- | -------------------------------------------------------------------------------------------------- | | `cause` | `unknown` | The underlying error that caused this error, if any. | | `code` | [`RealtimeErrorCode`](types#realtimeerrorcode) | Real-time error code | | `raw` | `unknown` | Original response payload for debugging. Contains the raw WebSocket message that caused the error. | | `statusCode` | `number` \| `undefined` | HTTP status code when applicable (e.g., 401 for auth errors, 500 for server errors). | *** ## BadRequestError Bad request error (400). Thrown for invalid configuration or parameters. ### Extends * [`RealtimeError`](classes#realtimeerror) ### toJSON() ```ts toJSON(): Record; ``` Converts to a plain object for logging/serialization **Returns** `Record`\<`string`, `unknown`> **Inherited from** [`RealtimeError`](classes#realtimeerror).[`toJSON`](classes#realtimeerror-tojson) *** ### toString() ```ts toString(): string; ``` Creates a human-readable string representation **Returns** `string` **Inherited from** [`RealtimeError`](classes#realtimeerror).[`toString`](classes#realtimeerror-tostring) ### Properties | Property | Type | Description | | ------------ | ---------------------------------------------- | -------------------------------------------------------------------------------------------------- | | `cause` | `unknown` | The underlying error that caused this error, if any. | | `code` | [`RealtimeErrorCode`](types#realtimeerrorcode) | Real-time error code | | `raw` | `unknown` | Original response payload for debugging. Contains the raw WebSocket message that caused the error. | | `statusCode` | `number` \| `undefined` | HTTP status code when applicable (e.g., 401 for auth errors, 500 for server errors). | *** ## QuotaError Quota error (402, 429). Thrown when rate limits are exceeded or quota is exhausted. ### Extends * [`RealtimeError`](classes#realtimeerror) ### toJSON() ```ts toJSON(): Record; ``` Converts to a plain object for logging/serialization **Returns** `Record`\<`string`, `unknown`> **Inherited from** [`RealtimeError`](classes#realtimeerror).[`toJSON`](classes#realtimeerror-tojson) *** ### toString() ```ts toString(): string; ``` Creates a human-readable string representation **Returns** `string` **Inherited from** [`RealtimeError`](classes#realtimeerror).[`toString`](classes#realtimeerror-tostring) ### Properties | Property | Type | Description | | ------------ | ---------------------------------------------- | -------------------------------------------------------------------------------------------------- | | `cause` | `unknown` | The underlying error that caused this error, if any. | | `code` | [`RealtimeErrorCode`](types#realtimeerrorcode) | Real-time error code | | `raw` | `unknown` | Original response payload for debugging. Contains the raw WebSocket message that caused the error. | | `statusCode` | `number` \| `undefined` | HTTP status code when applicable (e.g., 401 for auth errors, 500 for server errors). | *** ## ConnectionError Connection error. Thrown for WebSocket connection failures and transport errors. ### Extends * [`RealtimeError`](classes#realtimeerror) ### toJSON() ```ts toJSON(): Record; ``` Converts to a plain object for logging/serialization **Returns** `Record`\<`string`, `unknown`> **Inherited from** [`RealtimeError`](classes#realtimeerror).[`toJSON`](classes#realtimeerror-tojson) *** ### toString() ```ts toString(): string; ``` Creates a human-readable string representation **Returns** `string` **Inherited from** [`RealtimeError`](classes#realtimeerror).[`toString`](classes#realtimeerror-tostring) ### Properties | Property | Type | Description | | ------------ | ---------------------------------------------- | -------------------------------------------------------------------------------------------------- | | `cause` | `unknown` | The underlying error that caused this error, if any. | | `code` | [`RealtimeErrorCode`](types#realtimeerrorcode) | Real-time error code | | `raw` | `unknown` | Original response payload for debugging. Contains the raw WebSocket message that caused the error. | | `statusCode` | `number` \| `undefined` | HTTP status code when applicable (e.g., 401 for auth errors, 500 for server errors). | *** ## NetworkError Network error. Thrown for server-side network issues (408, 500, 503). ### Extends * [`RealtimeError`](classes#realtimeerror) ### toJSON() ```ts toJSON(): Record; ``` Converts to a plain object for logging/serialization **Returns** `Record`\<`string`, `unknown`> **Inherited from** [`RealtimeError`](classes#realtimeerror).[`toJSON`](classes#realtimeerror-tojson) *** ### toString() ```ts toString(): string; ``` Creates a human-readable string representation **Returns** `string` **Inherited from** [`RealtimeError`](classes#realtimeerror).[`toString`](classes#realtimeerror-tostring) ### Properties | Property | Type | Description | | ------------ | ---------------------------------------------- | -------------------------------------------------------------------------------------------------- | | `cause` | `unknown` | The underlying error that caused this error, if any. | | `code` | [`RealtimeErrorCode`](types#realtimeerrorcode) | Real-time error code | | `raw` | `unknown` | Original response payload for debugging. Contains the raw WebSocket message that caused the error. | | `statusCode` | `number` \| `undefined` | HTTP status code when applicable (e.g., 401 for auth errors, 500 for server errors). | *** ## AbortError Abort error. Thrown when an operation is cancelled via AbortSignal. ### Extends * [`RealtimeError`](classes#realtimeerror) ### toJSON() ```ts toJSON(): Record; ``` Converts to a plain object for logging/serialization **Returns** `Record`\<`string`, `unknown`> **Inherited from** [`RealtimeError`](classes#realtimeerror).[`toJSON`](classes#realtimeerror-tojson) *** ### toString() ```ts toString(): string; ``` Creates a human-readable string representation **Returns** `string` **Inherited from** [`RealtimeError`](classes#realtimeerror).[`toString`](classes#realtimeerror-tostring) ### Properties | Property | Type | Description | | ------------ | ---------------------------------------------- | -------------------------------------------------------------------------------------------------- | | `cause` | `unknown` | The underlying error that caused this error, if any. | | `code` | [`RealtimeErrorCode`](types#realtimeerrorcode) | Real-time error code | | `raw` | `unknown` | Original response payload for debugging. Contains the raw WebSocket message that caused the error. | | `statusCode` | `number` \| `undefined` | HTTP status code when applicable (e.g., 401 for auth errors, 500 for server errors). | *** ## StateError State error. Thrown when an operation is attempted in an invalid state. ### Extends * [`RealtimeError`](classes#realtimeerror) ### toJSON() ```ts toJSON(): Record; ``` Converts to a plain object for logging/serialization **Returns** `Record`\<`string`, `unknown`> **Inherited from** [`RealtimeError`](classes#realtimeerror).[`toJSON`](classes#realtimeerror-tojson) *** ### toString() ```ts toString(): string; ``` Creates a human-readable string representation **Returns** `string` **Inherited from** [`RealtimeError`](classes#realtimeerror).[`toString`](classes#realtimeerror-tostring) ### Properties | Property | Type | Description | | ------------ | ---------------------------------------------- | -------------------------------------------------------------------------------------------------- | | `cause` | `unknown` | The underlying error that caused this error, if any. | | `code` | [`RealtimeErrorCode`](types#realtimeerrorcode) | Real-time error code | | `raw` | `unknown` | Original response payload for debugging. Contains the raw WebSocket message that caused the error. | | `statusCode` | `number` \| `undefined` | HTTP status code when applicable (e.g., 401 for auth errors, 500 for server errors). | *** ## RealtimeTtsConnection WebSocket connection for real-time Text-to-Speech. Supports up to 5 concurrent streams multiplexed by `stream_id`. The connection automatically sends keepalive messages while open. ### Example ```typescript const conn = new RealtimeTtsConnection(apiKey, wsUrl, ttsDefaults); await conn.connect(); const s1 = conn.stream({ model, voice, language, audio_format }); s1.sendText("Hello"); s1.finish(); for await (const chunk of s1) { ... } conn.close(); ``` ### Extends * `TypedEmitter`\<[`TtsConnectionEvents`](types#ttsconnectionevents)> ### isConnected ```ts get isConnected(): boolean; ``` Whether the WebSocket is connected. **Returns** `boolean` ### Constructor ```ts new RealtimeTtsConnection( apiKey, wsUrl, ttsDefaults?, options?): RealtimeTtsConnection; ``` **Parameters** | Parameter | Type | | -------------- | ------------------------------------------------------ | | `apiKey` | `string` | | `wsUrl` | `string` | | `ttsDefaults?` | `Partial`\<[`TtsStreamConfig`](types#ttsstreamconfig)> | | `options?` | [`TtsConnectionOptions`](types#ttsconnectionoptions) | **Returns** `RealtimeTtsConnection` **Overrides** ```ts TypedEmitter.constructor ``` ### close() ```ts close(): void; ``` Close the WebSocket connection and terminate all active streams. **Returns** `void` *** ### connect() ```ts connect(): Promise; ``` Open the WebSocket connection and start keepalive. Called automatically by [stream](#stream) if not yet connected. **Returns** `Promise`\<`void`> *** ### emit() ```ts emit(event, ...args): void; ``` Emit an event to all registered handlers. Handler errors do not prevent other handlers from running. Errors are reported to an `error` event if present, otherwise rethrown async. **Type Parameters** | Type Parameter | | ---------------------------------------------------------------------- | | `E` *extends* keyof [`TtsConnectionEvents`](types#ttsconnectionevents) | **Parameters** | Parameter | Type | | --------- | ----------------------------------------------------------------------- | | `event` | `E` | | ...`args` | `Parameters`\<[`TtsConnectionEvents`](types#ttsconnectionevents)\[`E`]> | **Returns** `void` **Inherited from** ```ts TypedEmitter.emit ``` *** ### off() ```ts off(event, handler): this; ``` Remove an event handler. **Type Parameters** | Type Parameter | | ---------------------------------------------------------------------- | | `E` *extends* keyof [`TtsConnectionEvents`](types#ttsconnectionevents) | **Parameters** | Parameter | Type | | --------- | -------------------------------------------------------- | | `event` | `E` | | `handler` | [`TtsConnectionEvents`](types#ttsconnectionevents)\[`E`] | **Returns** `this` **Inherited from** ```ts TypedEmitter.off ``` *** ### on() ```ts on(event, handler): this; ``` Register an event handler. **Type Parameters** | Type Parameter | | ---------------------------------------------------------------------- | | `E` *extends* keyof [`TtsConnectionEvents`](types#ttsconnectionevents) | **Parameters** | Parameter | Type | | --------- | -------------------------------------------------------- | | `event` | `E` | | `handler` | [`TtsConnectionEvents`](types#ttsconnectionevents)\[`E`] | **Returns** `this` **Inherited from** ```ts TypedEmitter.on ``` *** ### once() ```ts once(event, handler): this; ``` Register a one-time event handler. **Type Parameters** | Type Parameter | | ---------------------------------------------------------------------- | | `E` *extends* keyof [`TtsConnectionEvents`](types#ttsconnectionevents) | **Parameters** | Parameter | Type | | --------- | -------------------------------------------------------- | | `event` | `E` | | `handler` | [`TtsConnectionEvents`](types#ttsconnectionevents)\[`E`] | **Returns** `this` **Inherited from** ```ts TypedEmitter.once ``` *** ### removeAllListeners() ```ts removeAllListeners(event?): void; ``` Remove all event handlers. **Parameters** | Parameter | Type | | --------- | ------------------------- | | `event?` | keyof TtsConnectionEvents | **Returns** `void` **Inherited from** ```ts TypedEmitter.removeAllListeners ``` *** ### stream() ```ts stream(input?): Promise; ``` Open a new TTS stream on this connection. Auto-connects if the WebSocket is not yet open. **Parameters** | Parameter | Type | Description | | --------- | ---------------------------------------- | ------------------------------------------------ | | `input?` | [`TtsStreamInput`](types#ttsstreaminput) | Stream configuration (merged with tts\_defaults) | **Returns** `Promise`\<[`RealtimeTtsStream`](classes#realtimettsstream)> A ready-to-use stream handle *** ## RealtimeTtsStream Handle for one TTS stream on a WebSocket connection. Emits typed events and supports async iteration over decoded audio chunks. ### Examples ```typescript stream.on('audio', (chunk) => process(chunk)); stream.on('terminated', () => console.log('done')); stream.sendText("Hello world"); stream.finish(); ``` ```typescript stream.sendText("Hello world"); stream.finish(); for await (const chunk of stream) { process(chunk); } ``` ### Extends * `TypedEmitter`\<[`TtsStreamEvents`](types#ttsstreamevents)> ### state ```ts get state(): TtsStreamState; ``` Current stream lifecycle state. **Returns** [`TtsStreamState`](types#ttsstreamstate) ### \[asyncIterator]\() ```ts asyncIterator: AsyncIterator>; ``` Async iterator that yields decoded audio chunks. The returned iterator's `return()` resets the internal iterator-attach flag and drops any buffered audio, so consumers that exit `for await` early (via `break` etc.) stop accruing memory while the stream keeps receiving server audio. **Returns** `AsyncIterator`\<`Uint8Array`\<`ArrayBufferLike`>> *** ### cancel() ```ts cancel(): void; ``` Cancel this stream. The server will stop generating and send `terminated`. **Returns** `void` *** ### close() ```ts close(): void; ``` Close this stream. For single-stream usage (created via `tts(input)`), also closes the underlying WebSocket connection. **Returns** `void` *** ### emit() ```ts emit(event, ...args): void; ``` Emit an event to all registered handlers. Handler errors do not prevent other handlers from running. Errors are reported to an `error` event if present, otherwise rethrown async. **Type Parameters** | Type Parameter | | -------------------------------------------------------------- | | `E` *extends* keyof [`TtsStreamEvents`](types#ttsstreamevents) | **Parameters** | Parameter | Type | | --------- | --------------------------------------------------------------- | | `event` | `E` | | ...`args` | `Parameters`\<[`TtsStreamEvents`](types#ttsstreamevents)\[`E`]> | **Returns** `void` **Inherited from** ```ts TypedEmitter.emit ``` *** ### finish() ```ts finish(): void; ``` Signal that no more text will be sent for this stream. The server will finish generating audio and send `terminated`. **Returns** `void` *** ### off() ```ts off(event, handler): this; ``` Remove an event handler. **Type Parameters** | Type Parameter | | -------------------------------------------------------------- | | `E` *extends* keyof [`TtsStreamEvents`](types#ttsstreamevents) | **Parameters** | Parameter | Type | | --------- | ------------------------------------------------ | | `event` | `E` | | `handler` | [`TtsStreamEvents`](types#ttsstreamevents)\[`E`] | **Returns** `this` **Inherited from** ```ts TypedEmitter.off ``` *** ### on() ```ts on(event, handler): this; ``` Register an event handler. **Type Parameters** | Type Parameter | | -------------------------------------------------------------- | | `E` *extends* keyof [`TtsStreamEvents`](types#ttsstreamevents) | **Parameters** | Parameter | Type | | --------- | ------------------------------------------------ | | `event` | `E` | | `handler` | [`TtsStreamEvents`](types#ttsstreamevents)\[`E`] | **Returns** `this` **Inherited from** ```ts TypedEmitter.on ``` *** ### once() ```ts once(event, handler): this; ``` Register a one-time event handler. **Type Parameters** | Type Parameter | | -------------------------------------------------------------- | | `E` *extends* keyof [`TtsStreamEvents`](types#ttsstreamevents) | **Parameters** | Parameter | Type | | --------- | ------------------------------------------------ | | `event` | `E` | | `handler` | [`TtsStreamEvents`](types#ttsstreamevents)\[`E`] | **Returns** `this` **Inherited from** ```ts TypedEmitter.once ``` *** ### removeAllListeners() ```ts removeAllListeners(event?): void; ``` Remove all event handlers. **Parameters** | Parameter | Type | | --------- | --------------------- | | `event?` | keyof TtsStreamEvents | **Returns** `void` **Inherited from** ```ts TypedEmitter.removeAllListeners ``` *** ### sendStream() ```ts sendStream(source): Promise; ``` Pipe an async iterable of text chunks into the stream. Automatically calls [finish](#finish) when the iterable completes. Designed for concurrent use: call `sendStream()` and consume audio via `for await` or events simultaneously. **Parameters** | Parameter | Type | | --------- | -------------------------- | | `source` | `AsyncIterable`\<`string`> | **Returns** `Promise`\<`void`> **Example** ```typescript stream.sendStream(llmTokenStream); for await (const audio of stream) { forward(audio); } ``` *** ### sendText() ```ts sendText(text, options?): void; ``` Send one text chunk to the TTS stream. **Parameters** | Parameter | Type | Description | | -------------- | ----------------------- | --------------------------------------------- | | `text` | `string` | Text to synthesize | | `options?` | \{ `end?`: `boolean`; } | - | | `options.end?` | `boolean` | If true, signals this is the final text chunk | **Returns** `void` ### Properties | Property | Type | | ---------- | -------- | | `streamId` | `string` | *** ## SonioxTranslationJob A translation job. Mirrors [SonioxTranscription](classes#sonioxtranscription), with an additional translation mode and helpers that reshape the completed transcript into a [SonioxTranslation](types#sonioxtranslation). ### Extends * [`SonioxTranscription`](classes#sonioxtranscription) ### Constructor ```ts new SonioxTranslationJob( data, http, _mode, transcript?, translation?): SonioxTranslationJob; ``` **Parameters** | Parameter | Type | | -------------- | ------------------------------------------------------------------ | | `data` | [`SonioxTranscriptionData`](types#sonioxtranscriptiondata) | | `http` | [`HttpClient`](types#httpclient) | | `_mode` | [`TranslateFromTranscriptMode`](types#translatefromtranscriptmode) | | `transcript?` | [`SonioxTranscript`](classes#sonioxtranscript) \| `null` | | `translation?` | [`SonioxTranslation`](types#sonioxtranslation) \| `null` | **Returns** `SonioxTranslationJob` **Overrides** [`SonioxTranscription`](classes#sonioxtranscription).[`constructor`](classes#sonioxtranscription-constructor) ### delete() ```ts delete(): Promise; ``` Permanently deletes this transcription. This operation is idempotent - succeeds even if the transcription doesn't exist. **Returns** `Promise`\<`void`> **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors (except 404) **Example** ```typescript const transcription = await client.stt.get('550e8400-e29b-41d4-a716-446655440000'); await transcription.delete(); ``` **Inherited from** [`SonioxTranscription`](classes#sonioxtranscription).[`delete`](classes#sonioxtranscription-delete) *** ### destroy() ```ts destroy(): Promise; ``` Permanently deletes this transcription and its associated file (if any). This operation is idempotent - succeeds even if resources don't exist. **Returns** `Promise`\<`void`> **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors (except 404) **Example** ```typescript // Clean up both transcription and uploaded file const transcription = await client.stt.transcribe({ model: 'stt-async-v5', file: buffer, wait: true, }); // ... use transcription ... await transcription.destroy(); // Deletes both transcription and file ``` **Inherited from** [`SonioxTranscription`](classes#sonioxtranscription).[`destroy`](classes#sonioxtranscription-destroy) *** ### fetchTranslation() ```ts fetchTranslation(options?): Promise; ``` Alias for [SonioxTranslationJob.getTranslation](#gettranslation). **Parameters** | Parameter | Type | | ----------------- | --------------------------------------------------- | | `options?` | \{ `force?`: `boolean`; `signal?`: `AbortSignal`; } | | `options.force?` | `boolean` | | `options.signal?` | `AbortSignal` | **Returns** `Promise`\<[`SonioxTranslation`](types#sonioxtranslation) | `null`> *** ### getTranscript() ```ts getTranscript(options?): Promise; ``` Retrieves the full transcript text and tokens for this transcription. Only available for successfully completed transcriptions. Returns cached transcript if available (when using `transcribe()` with `wait: true`). Use `force: true` to bypass the cache and fetch fresh data from the API. **Parameters** | Parameter | Type | Description | | ----------------- | --------------------------------------------------- | -------------------------------------------------------- | | `options?` | \{ `force?`: `boolean`; `signal?`: `AbortSignal`; } | Optional settings | | `options.force?` | `boolean` | If true, bypasses cached transcript and fetches from API | | `options.signal?` | `AbortSignal` | Optional AbortSignal for request cancellation | **Returns** `Promise`\<[`SonioxTranscript`](classes#sonioxtranscript) | `null`> The transcript with text and detailed tokens, or null if not found. **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors (except 404). **Example** ```typescript const transcription = await client.stt.get('550e8400-e29b-41d4-a716-446655440000'); if (transcription) { const transcript = await transcription.getTranscript(); if (transcript) { console.log(transcript.text); } } // Force re-fetch from API const freshTranscript = await transcription.getTranscript({ force: true }); ``` **Inherited from** [`SonioxTranscription`](classes#sonioxtranscription).[`getTranscript`](classes#sonioxtranscription-gettranscript) *** ### getTranslation() ```ts getTranslation(options?): Promise; ``` Retrieves and reshapes the completed transcript into a structured translation result. Returns cached translation if available. Use `force: true` to bypass the cached translation and fetch a fresh transcript from the API. **Parameters** | Parameter | Type | | ----------------- | --------------------------------------------------- | | `options?` | \{ `force?`: `boolean`; `signal?`: `AbortSignal`; } | | `options.force?` | `boolean` | | `options.signal?` | `AbortSignal` | **Returns** `Promise`\<[`SonioxTranslation`](types#sonioxtranslation) | `null`> *** ### refresh() ```ts refresh(signal?): Promise; ``` Re-fetches this translation job to get the latest status. **Parameters** | Parameter | Type | | --------- | ------------- | | `signal?` | `AbortSignal` | **Returns** `Promise`\<`SonioxTranslationJob`> **Overrides** [`SonioxTranscription`](classes#sonioxtranscription).[`refresh`](classes#sonioxtranscription-refresh) *** ### toJSON() ```ts toJSON(): SonioxTranscriptionData; ``` Returns the raw data for this transcription. **Returns** [`SonioxTranscriptionData`](types#sonioxtranscriptiondata) **Inherited from** [`SonioxTranscription`](classes#sonioxtranscription).[`toJSON`](classes#sonioxtranscription-tojson) *** ### wait() ```ts wait(options): Promise; ``` Waits for the translation job to complete or fail. **Parameters** | Parameter | Type | | --------- | ---------------------------------- | | `options` | [`WaitOptions`](types#waitoptions) | **Returns** `Promise`\<`SonioxTranslationJob`> **Overrides** [`SonioxTranscription`](classes#sonioxtranscription).[`wait`](classes#sonioxtranscription-wait) ### Properties | Property | Type | Description | | -------------------------------- | -------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `audio_duration_ms` | `number` \| `null` \| `undefined` | Duration of the audio in milliseconds. Only available after processing begins. | | `audio_url` | `string` \| `null` \| `undefined` | URL of the audio file being transcribed. | | `client_reference_id` | `string` \| `null` \| `undefined` | Optional tracking identifier. | | `context` | \| [`TranscriptionContext`](types#transcriptioncontext) \| `null` \| `undefined` | Additional context provided for the transcription. | | `created_at` | `string` | UTC timestamp when the transcription was created. | | `enable_language_identification` | `boolean` | When true, language is detected for each part of the transcription. | | `enable_speaker_diarization` | `boolean` | When true, speakers are identified and separated in the transcription output. | | `error_message` | `string` \| `null` \| `undefined` | Error message if transcription failed. | | `error_type` | `string` \| `null` \| `undefined` | Error type if transcription failed. | | `file_id` | `string` \| `null` \| `undefined` | ID of the uploaded file being transcribed. | | `filename` | `string` | Name of the file being transcribed. | | `id` | `string` | Unique identifier of the transcription. | | `language_hints` | `string`\[] \| `undefined` | Expected languages in the audio. | | `model` | `string` | Speech-to-text model used. | | `status` | [`TranscriptionStatus`](types#transcriptionstatus) | Current status of the transcription. | | `transcript` | [`SonioxTranscript`](classes#sonioxtranscript) \| `null` \| `undefined` | Pre-fetched transcript. Only available when using `transcribe()` with `wait: true`, `fetch_transcript !== false`, and the transcription completed successfully | | `translation` | [`SonioxTranslation`](types#sonioxtranslation) \| `null` \| `undefined` | Pre-fetched translation result. Only available when using `translate()` with `wait: true`, `fetch_translation !== false`, and the job completed successfully. | | `webhook_auth_header_name` | `string` \| `null` \| `undefined` | Name of the authentication header sent with webhook notifications. | | `webhook_auth_header_value` | `string` \| `null` \| `undefined` | Authentication header value (masked). | | `webhook_status_code` | `number` \| `null` \| `undefined` | HTTP status code received when webhook was delivered. | | `webhook_url` | `string` \| `null` \| `undefined` | URL to receive webhook notifications. | *** ## SonioxTtsApi REST API for Text-to-Speech generation and TTS model listing. Accessed via `client.tts` on [SonioxNodeClient](classes#sonioxnodeclient). Inherits browser-safe `generate()` and `generateStream()` from `TtsRestClient` in `@soniox/core`, and adds Node-specific methods `generateToFile()` and `listModels()`. ### Extends * `TtsRestClient` ### Constructor ```ts new SonioxTtsApi( apiKey, ttsApiUrl, http): SonioxTtsApi; ``` **Parameters** | Parameter | Type | | ----------- | -------------------------------- | | `apiKey` | `string` | | `ttsApiUrl` | `string` | | `http` | [`HttpClient`](types#httpclient) | **Returns** `SonioxTtsApi` **Overrides** ```ts TtsRestClient.constructor ``` ### generate() ```ts generate(options): Promise>; ``` Generate speech audio from text. Returns the full audio as a `Uint8Array`. **Parameters** | Parameter | Type | | --------- | ------------------------------------------------------ | | `options` | [`GenerateSpeechOptions`](types#generatespeechoptions) | **Returns** `Promise`\<`Uint8Array`\<`ArrayBufferLike`>> **Throws** [SonioxHttpError](classes#sonioxhttperror) on non-2xx responses, network failures, or aborted requests. **Inherited from** ```ts TtsRestClient.generate ``` *** ### generateStream() ```ts generateStream(options): AsyncIterable>; ``` Generate speech audio from text as a streaming async iterable. Yields `Uint8Array` chunks as they arrive from the server response body. Lower time-to-first-audio than [generate](#generate). **Known limitation:** Mid-stream server errors (reported via HTTP trailers) cannot be detected through the `fetch` API. The iterator may end early without an explicit error. Use WebSocket TTS for reliable error detection. **Parameters** | Parameter | Type | | --------- | ------------------------------------------------------ | | `options` | [`GenerateSpeechOptions`](types#generatespeechoptions) | **Returns** `AsyncIterable`\<`Uint8Array`\<`ArrayBufferLike`>> **Throws** [SonioxHttpError](classes#sonioxhttperror) on non-2xx responses, network failures, or aborted requests (before the stream starts). **Inherited from** ```ts TtsRestClient.generateStream ``` *** ### generateToFile() ```ts generateToFile(output, options): Promise; ``` Generate speech audio and write to a file or writable stream. **Parameters** | Parameter | Type | Description | | --------- | --------------------------------------------------------------- | ---------------------------------------------------- | | `output` | `string` \| `WritableStream`\<`Uint8Array`\<`ArrayBufferLike`>> | File path (string) or a `WritableStream` | | `options` | [`GenerateSpeechOptions`](types#generatespeechoptions) | Generation options | **Returns** `Promise`\<`number`> Number of bytes written **Examples** ```typescript const bytes = await client.tts.generateToFile('output.wav', { text: 'Hello world', voice: 'Adrian', language: 'en', }); ``` ```typescript const bytes = await client.tts.generateToFile(writableStream, { text: 'Hello world', voice: 'Adrian', language: 'en', }); ``` *** ### listModels() ```ts listModels(signal?): Promise; ``` List available TTS models and their voices. **Parameters** | Parameter | Type | | --------- | ------------- | | `signal?` | `AbortSignal` | **Returns** `Promise`\<[`TtsModel`](types#ttsmodel)\[]> **Example** ```typescript const models = await client.tts.listModels(); for (const model of models) { console.log(model.id, model.voices.map(v => v.id)); } ``` ### Properties | Property | Type | Description | | -------- | -------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `voices` | [`SonioxVoicesAPI`](classes#sonioxvoicesapi) | Manage custom (cloned) voices: create, list, get, recompute, and delete. **Example** `const voice = await client.tts.voices.create({ name: 'My voice', file: buffer });` | *** ## SonioxVoice A custom (cloned) Text-to-Speech voice. Use a `ready` voice by passing its [id](#id) as the `voice` value in any TTS request (REST or realtime). ### Constructor ```ts new SonioxVoice(data, _http): SonioxVoice; ``` **Parameters** | Parameter | Type | | --------- | ------------------------------------------ | | `data` | [`SonioxVoiceData`](types#sonioxvoicedata) | | `_http` | [`HttpClient`](types#httpclient) | **Returns** `SonioxVoice` ### delete() ```ts delete(signal?): Promise; ``` Permanently deletes this voice and its embeddings. This operation is idempotent - succeeds even if the voice doesn't exist. **Parameters** | Parameter | Type | Description | | --------- | ------------- | -------------------------------------- | | `signal?` | `AbortSignal` | Optional AbortSignal for cancellation. | **Returns** `Promise`\<`void`> **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors (except 404). **Example** ```typescript const voice = await client.tts.voices.get(voiceId); if (voice) { await voice.delete(); } ``` *** ### isReady() ```ts isReady(model): boolean; ``` Returns true if the voice is `ready` for the given model. **Parameters** | Parameter | Type | Description | | --------- | -------- | --------------------------- | | `model` | `string` | Name of the model to check. | **Returns** `boolean` **Example** ```typescript const voice = await client.tts.voices.get(voiceId); if (voice?.isReady('tts-rt-v2')) { const audio = await client.tts.generate({ text: 'Hi', voice: voice.id, language: 'en', model: 'tts-rt-v2' }); } ``` *** ### recompute() ```ts recompute(options): Promise; ``` Prepares this voice for use with available models it is not ready for yet. Models the voice is already prepared for are left unchanged. **Parameters** | Parameter | Type | Description | | --------- | ------------------------------------------------------ | ------------------------------------------------- | | `options` | [`RecomputeVoiceOptions`](types#recomputevoiceoptions) | Optional model to target and cancellation signal. | **Returns** `Promise`\<`SonioxVoice`> The updated voice. **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors. **Example** ```typescript const voice = await client.tts.voices.get(voiceId); const updated = await voice?.recompute(); ``` *** ### toJSON() ```ts toJSON(): SonioxVoiceData; ``` Returns the raw data for this voice. **Returns** [`SonioxVoiceData`](types#sonioxvoicedata) ### Properties | Property | Type | | ------------ | --------------------------------------------------------- | | `created_at` | `string` | | `filename` | `string` | | `id` | `string` | | `models` | [`VoiceModelStatusEntry`](types#voicemodelstatusentry)\[] | | `name` | `string` | *** ## SonioxVoicesAPI REST API for managing custom (cloned) Text-to-Speech voices. Accessed via `client.tts.voices` on `SonioxNodeClient`. Create a voice from a short reference clip, wait until it is `ready` for the model you intend to use, then pass its `id` as the `voice` value in any TTS request. ### Constructor ```ts new SonioxVoicesAPI(http): SonioxVoicesAPI; ``` **Parameters** | Parameter | Type | | --------- | -------------------------------- | | `http` | [`HttpClient`](types#httpclient) | **Returns** `SonioxVoicesAPI` ### count() ```ts count(options): Promise; ``` Returns the total number of voices in your project. **Parameters** | Parameter | Type | Description | | ----------------- | ------------------------------ | --------------------------------- | | `options` | \{ `signal?`: `AbortSignal`; } | Optional cancellation parameters. | | `options.signal?` | `AbortSignal` | - | **Returns** `Promise`\<[`VoicesCountResponse`](types#voicescountresponse)> The total voice count. **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors. **Example** ```typescript const { total } = await client.tts.voices.count(); ``` *** ### create() ```ts create(options): Promise; ``` Creates a new voice by uploading a reference audio clip. The voice begins processing in the background per model. Poll [SonioxVoicesAPI.get](#get) (or check [SonioxVoice.isReady](classes#sonioxvoice-isready)) until the model you need reports `ready`. **Parameters** | Parameter | Type | Description | | --------- | ------------------------------------------------ | ---------------------------------- | | `options` | [`CreateVoiceOptions`](types#createvoiceoptions) | The voice name and reference clip. | **Returns** `Promise`\<[`SonioxVoice`](classes#sonioxvoice)> The created voice. **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors. **Throws** `Error` On validation errors (file too large, invalid input). **Example** ```typescript import * as fs from 'node:fs'; const buffer = await fs.promises.readFile('/path/to/reference.wav'); const voice = await client.tts.voices.create({ name: 'My narrator', file: buffer, filename: 'reference.wav', }); ``` *** ### delete() ```ts delete(voice, signal?): Promise; ``` Permanently deletes a voice and its embeddings. This operation is idempotent - succeeds even if the voice doesn't exist. **Parameters** | Parameter | Type | Description | | --------- | ------------------------------------------ | ------------------------------------------------ | | `voice` | [`VoiceIdentifier`](types#voiceidentifier) | The UUID of the voice or a SonioxVoice instance. | | `signal?` | `AbortSignal` | Optional AbortSignal for cancellation. | **Returns** `Promise`\<`void`> **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors (except 404). **Example** ```typescript await client.tts.voices.delete('550e8400-e29b-41d4-a716-446655440000'); ``` *** ### delete\_all() ```ts delete_all(options): Promise; ``` Permanently deletes all voices in your project. Iterates through all pages of voices and deletes each one. **Parameters** | Parameter | Type | Description | | --------- | -------------------------------------------------------- | ----------------------------- | | `options` | [`DeleteAllVoicesOptions`](types#deleteallvoicesoptions) | Optional cancellation signal. | **Returns** `Promise`\<`void`> **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors. **Throws** `Error` If the operation is aborted via signal. **Example** ```typescript await client.tts.voices.delete_all(); ``` *** ### get() ```ts get(voice, signal?): Promise; ``` Retrieve metadata for a voice. **Parameters** | Parameter | Type | Description | | --------- | ------------------------------------------ | ------------------------------------------------ | | `voice` | [`VoiceIdentifier`](types#voiceidentifier) | The UUID of the voice or a SonioxVoice instance. | | `signal?` | `AbortSignal` | Optional AbortSignal for cancellation. | **Returns** `Promise`\<[`SonioxVoice`](classes#sonioxvoice) | `null`> The voice instance, or null if not found. **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors (except 404). **Example** ```typescript const voice = await client.tts.voices.get('550e8400-e29b-41d4-a716-446655440000'); if (voice) { console.log(voice.name, voice.models); } ``` *** ### list() ```ts list(options): Promise; ``` Retrieves the list of voices in your project. The returned result is async iterable - use `for await...of`. **Parameters** | Parameter | Type | Description | | --------- | ---------------------------------------------- | ------------------------------------------------ | | `options` | [`ListVoicesOptions`](types#listvoicesoptions) | Optional pagination and cancellation parameters. | **Returns** `Promise`\<[`VoiceListResult`](classes#voicelistresult)> VoiceListResult **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors. **Example** ```typescript const result = await client.tts.voices.list(); // Automatic paging - iterates through ALL voices across all pages for await (const voice of result) { console.log(voice.id, voice.name); } ``` *** ### recompute() ```ts recompute(voice, options): Promise; ``` Prepares a voice for use with available models it is not ready for yet. Use this after a new model is released to make an existing voice usable with it. Models the voice is already prepared for are left unchanged. **Parameters** | Parameter | Type | Description | | --------- | ------------------------------------------------------ | ------------------------------------------------- | | `voice` | [`VoiceIdentifier`](types#voiceidentifier) | The UUID of the voice or a SonioxVoice instance. | | `options` | [`RecomputeVoiceOptions`](types#recomputevoiceoptions) | Optional model to target and cancellation signal. | **Returns** `Promise`\<[`SonioxVoice`](classes#sonioxvoice)> The updated voice. **Throws** [SonioxHttpError](classes#sonioxhttperror) On API errors. **Example** ```typescript // Prepare for every model the voice is not ready for yet await client.tts.voices.recompute(voiceId); // Or target a specific model await client.tts.voices.recompute(voiceId, { model: 'tts-rt-v2' }); ``` *** ## VoiceListResult Result set for voice listing. ### Constructor ```ts new VoiceListResult( initialResponse, _http, _limit, _signal): VoiceListResult; ``` **Parameters** | Parameter | Type | Default value | | ----------------- | ---------------------------------------------------------------------------------------------- | ------------- | | `initialResponse` | [`ListVoicesResponse`](types#listvoicesresponset)\<[`SonioxVoiceData`](types#sonioxvoicedata)> | `undefined` | | `_http` | [`HttpClient`](types#httpclient) | `undefined` | | `_limit` | `number` \| `undefined` | `undefined` | | `_signal` | `AbortSignal` \| `undefined` | `undefined` | **Returns** `VoiceListResult` ### \[asyncIterator]\() ```ts asyncIterator: AsyncIterator; ``` Async iterator that automatically fetches all pages. Use with `for await...of` to iterate through all voices. **Returns** `AsyncIterator`\<[`SonioxVoice`](classes#sonioxvoice)> *** ### isPaged() ```ts isPaged(): boolean; ``` Returns true if there are more pages of results beyond the first page. **Returns** `boolean` *** ### toJSON() ```ts toJSON(): ListVoicesResponse; ``` Returns the raw data for this list result. Also used by JSON.stringify() to prevent serialization of internal HTTP client. **Returns** [`ListVoicesResponse`](types#listvoicesresponset)\<[`SonioxVoiceData`](types#sonioxvoicedata)> ### Properties | Property | Type | Description | | ------------------ | --------------------------------------- | ----------------------------------------------------------- | | `next_page_cursor` | `string` \| `null` | Pagination cursor for the next page. Null if no more pages. | | `voices` | [`SonioxVoice`](classes#sonioxvoice)\[] | Voices from the first page of results. | # Full Node SDK reference URL: /sdk/node-SDK/reference Full SDK reference for the Node SDK ## Environment variables Environment variables configure the client when no explicit option is passed. Every variable has a matching option on `new SonioxNodeClient({ ... })` — see [`SonioxNodeClientOptions`](/sdk/node-SDK/reference/types#sonioxnodeclientoptions). | Variable | Maps to option | Default | | --------------------------- | --------------------------------------------------------------------- | ---------------------------------------------- | | `SONIOX_API_KEY` | `api_key` | - | | `SONIOX_REGION` | `region` | - (US root) | | `SONIOX_BASE_DOMAIN` | `base_domain` — derives all `{api,stt-rt,tts-rt}.{base_domain}` hosts | - | | `SONIOX_API_BASE_URL` | `base_url` (REST API) | `https://api.soniox.com` | | `SONIOX_WS_URL` | `realtime.ws_base_url` (STT WebSocket) | `wss://stt-rt.soniox.com/transcribe-websocket` | | `SONIOX_TTS_API_URL` | `tts_api_url` (TTS REST) | `https://tts-rt.soniox.com` | | `SONIOX_TTS_WS_URL` | `realtime.tts_ws_url` (TTS WebSocket) | `wss://tts-rt.soniox.com/tts-websocket` | | `SONIOX_API_WEBHOOK_HEADER` | Webhook auth header name | - | | `SONIOX_API_WEBHOOK_SECRET` | Webhook auth header value | - | ### Resolution precedence For every configurable host, the SDK resolves in this order: 1. Explicit option on `new SonioxNodeClient({ ... })` (e.g. `base_url`, `tts_api_url`, `realtime.ws_base_url`, `realtime.tts_ws_url`). 2. Matching environment variable (`SONIOX_API_BASE_URL`, `SONIOX_TTS_API_URL`, `SONIOX_WS_URL`, `SONIOX_TTS_WS_URL`). 3. `base_domain` / `SONIOX_BASE_DOMAIN` — derives all service endpoints from `{service}.{base_domain}`. 4. `region` / `SONIOX_REGION` — shorthand for `base_domain: '{region}.soniox.com'`. 5. Root US default (no region prefix). Setting `region: 'eu'`, `region: 'jp'`, or `region: 'in'` resolves every host family — `api.*`, `stt-rt.*`, and `tts-rt.*` — to the matching regional variant. Use `SONIOX_TTS_API_URL` / `SONIOX_TTS_WS_URL` (or the corresponding explicit options) only when you need to override just the TTS hosts while leaving the rest of the client at its default region. ## Client ### Available client methods | Method | Description | | ----------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------- | | [`client.files.upload(file, options)`](/sdk/node-SDK/reference/classes#sonioxfilesapi-upload) | Upload file | | [`client.files.list(options)`](/sdk/node-SDK/reference/classes#sonioxfilesapi-list) | List files | | [`client.files.count(options)`](/sdk/node-SDK/reference/classes#sonioxfilesapi-count) | Count uploaded files by source | | [`client.files.get(id)`](/sdk/node-SDK/reference/classes#sonioxfilesapi-get) | Get file | | [`client.files.delete(id)`](/sdk/node-SDK/reference/classes#sonioxfilesapi-delete) | Delete file | | [`client.files.delete_all()`](/sdk/node-SDK/reference/classes#sonioxfilesapi-delete_all) | Delete all files | | --- | --- | | [`client.stt.create(options)`](/sdk/node-SDK/reference/classes#sonioxsttapi-create) | Create transcription | | [`client.stt.list(options)`](/sdk/node-SDK/reference/classes#sonioxsttapi-list) | List transcriptions | | [`client.stt.count(signal?)`](/sdk/node-SDK/reference/classes#sonioxsttapi-count) | Count transcriptions by request scope | | [`client.stt.get(id)`](/sdk/node-SDK/reference/classes#sonioxsttapi-get) | Get transcription | | [`client.stt.getTranscript(id)`](/sdk/node-SDK/reference/classes#sonioxsttapi-gettranscript) | Get transcription transcript | | [`client.stt.delete(id)`](/sdk/node-SDK/reference/classes#sonioxsttapi-delete) | Delete transcription | | [`client.stt.destroy(id)`](/sdk/node-SDK/reference/classes#sonioxsttapi-destroy) | Delete transcription and its file | | [`client.stt.wait(id)`](/sdk/node-SDK/reference/classes#sonioxsttapi-wait) | Wait for transcription to complete | | [`client.stt.transcribe(options)`](/sdk/node-SDK/reference/classes#sonioxsttapi-transcribe) | Transcribe audio | | [`client.stt.translate(options)`](/sdk/node-SDK/reference/classes#sonioxsttapi-translate) | Translate audio asynchronously | | [`client.stt.transcribeFromUrl(url, options)`](/sdk/node-SDK/reference/classes#sonioxsttapi-transcribefromurl) | Transcribe audio from URL | | [`client.stt.transcribeFromFile(file, options)`](/sdk/node-SDK/reference/classes#sonioxsttapi-transcribefromfile) | Transcribe audio from file | | [`client.stt.transcribeFromFileId(id, options)`](/sdk/node-SDK/reference/classes#sonioxsttapi-transcribefromfileid) | Transcribe audio from file ID | | [`client.stt.delete_all()`](/sdk/node-SDK/reference/classes#sonioxsttapi-delete_all) | Delete all transcriptions | | [`client.stt.destroy_all()`](/sdk/node-SDK/reference/classes#sonioxsttapi-destroy_all) | Delete all transcriptions and their files | | --- | --- | | [`client.tts.generate(options)`](/sdk/node-SDK/reference/classes#sonioxttsapi-generate) | Generate speech audio from text | | [`client.tts.generateStream(options)`](/sdk/node-SDK/reference/classes#sonioxttsapi-generatestream) | Generate speech audio as a streaming async iterable | | [`client.tts.generateToFile(output, options)`](/sdk/node-SDK/reference/classes#sonioxttsapi-generatetofile) | Generate speech audio and write it to a file or writable stream | | [`client.tts.listModels()`](/sdk/node-SDK/reference/classes#sonioxttsapi-listmodels) | List available TTS models and voices | | [`client.tts.voices.create(options)`](/sdk/node-SDK/reference/classes#sonioxvoicesapi-create) | Create a custom (cloned) voice from a reference audio clip | | [`client.tts.voices.list(options)`](/sdk/node-SDK/reference/classes#sonioxvoicesapi-list) | List voices in your project | | [`client.tts.voices.count(options)`](/sdk/node-SDK/reference/classes#sonioxvoicesapi-count) | Count voices in your project | | [`client.tts.voices.get(voice, signal?)`](/sdk/node-SDK/reference/classes#sonioxvoicesapi-get) | Get voice metadata | | [`client.tts.voices.delete(voice, signal?)`](/sdk/node-SDK/reference/classes#sonioxvoicesapi-delete) | Delete a voice | | [`client.tts.voices.delete_all(options)`](/sdk/node-SDK/reference/classes#sonioxvoicesapi-delete_all) | Delete all voices | | [`client.tts.voices.recompute(voice, options)`](/sdk/node-SDK/reference/classes#sonioxvoicesapi-recompute) | Prepare a voice for additional TTS models | | --- | --- | | [`client.models.list()`](/sdk/node-SDK/reference/classes#sonioxmodelsapi-list) | List available models | | --- | --- | | [`client.usageLogs.list(options)`](/sdk/node-SDK/reference/classes#sonioxusagelogsapi-list) | List per-request usage and cost logs | | [`client.usageLogs.getSummary(options)`](/sdk/node-SDK/reference/classes#sonioxusagelogsapi-getsummary) | Get daily usage and cost aggregated per model | | --- | --- | | [`client.concurrencyLimits.get(signal?)`](/sdk/node-SDK/reference/classes#sonioxconcurrencylimitsapi-get) | Get region-scoped concurrency counts and limits | | [`client.concurrencyLimits.getHistory(options)`](/sdk/node-SDK/reference/classes#sonioxconcurrencylimitsapi-gethistory) | Get historical concurrent stream aggregates per period | | --- | --- | | [`client.webhooks.handle(options)`](/sdk/node-SDK/reference/classes#sonioxwebhooksapi-handle) | Handle webhook | | [`client.webhooks.handleRequest(request, options)`](/sdk/node-SDK/reference/classes#sonioxwebhooksapi-handlerequest) | Handle webhook with Fetch API | | [`client.webhooks.handleExpress(req, options)`](/sdk/node-SDK/reference/classes#sonioxwebhooksapi-handleexpress) | Handle webhook with Express | | [`client.webhooks.handleFastify(req, options)`](/sdk/node-SDK/reference/classes#sonioxwebhooksapi-handlefastify) | Handle webhook with Fastify | | [`client.webhooks.handleNestJS(req, options)`](/sdk/node-SDK/reference/classes#sonioxwebhooksapi-handlenestjs) | Handle webhook with NestJS | | [`client.webhooks.handleHono(c, options)`](/sdk/node-SDK/reference/classes#sonioxwebhooksapi-handlehono) | Handle webhook with Hono | | [`client.webhooks.getAuthFromEnv()`](/sdk/node-SDK/reference/classes#sonioxwebhooksapi-getauthfromenv) | Get webhook auth from environment variables | | [`client.webhooks.verifyAuth(headers, auth)`](/sdk/node-SDK/reference/classes#sonioxwebhooksapi-verifyauth) | Verify webhook auth | | [`client.webhooks.parseEvent(body)`](/sdk/node-SDK/reference/classes#sonioxwebhooksapi-parseevent) | Parse webhook event | | [`client.webhooks.isEvent(body)`](/sdk/node-SDK/reference/classes#sonioxwebhooksapi-isevent) | Check if body is a webhook event | | --- | --- | | [`client.auth.createTemporaryKey(options)`](/sdk/node-SDK/reference/classes#sonioxauthapi-createtemporarykey) | Create temporary API key | | --- | --- | | [`client.realtime.stt(options)`](/sdk/node-SDK/reference/classes#sonioxrealtimeapi-stt) | Create real-time STT session | | [`client.realtime.tts(input?)`](/sdk/node-SDK/reference/classes#realtimettsstream) | Create a single real-time TTS stream (opens its own WebSocket) | | [`client.realtime.tts.multiStream()`](/sdk/node-SDK/reference/classes#realtimettsconnection) | Open a TTS WebSocket connection that can host multiple streams | ## File ### Available file instance methods | Method | Description | | -------------------------------------------------------------------- | ----------- | | [`file.delete()`](/sdk/node-SDK/reference/classes#sonioxfile-delete) | Delete file | ## Transcription ### Available transcription instance methods | Method | Description | | ---------------------------------------------------------------------------------------------------- | ---------------------------------- | | [`transcription.getTranscript()`](/sdk/node-SDK/reference/classes#sonioxtranscription-gettranscript) | Get transcription transcript | | [`transcription.delete()`](/sdk/node-SDK/reference/classes#sonioxtranscription-delete) | Delete transcription | | [`transcription.destroy()`](/sdk/node-SDK/reference/classes#sonioxtranscription-destroy) | Delete transcription and its file | | [`transcription.wait()`](/sdk/node-SDK/reference/classes#sonioxtranscription-wait) | Wait for transcription to complete | | [`transcription.refresh()`](/sdk/node-SDK/reference/classes#sonioxtranscription-refresh) | Refresh transcription | ## Translation job ### Available translation job instance methods | Method | Description | | ------------------------------------------------------------------------------------------------------------ | -------------------------------------------------------------------- | | [`translationJob.getTranslation()`](/sdk/node-SDK/reference/classes#sonioxtranslationjob-gettranslation) | Fetch and reshape the completed transcript into a translation result | | [`translationJob.fetchTranslation()`](/sdk/node-SDK/reference/classes#sonioxtranslationjob-fetchtranslation) | Alias for `getTranslation()` | | [`translationJob.getTranscript()`](/sdk/node-SDK/reference/classes#sonioxtranslationjob-gettranscript) | Get the underlying transcription transcript | | [`translationJob.wait()`](/sdk/node-SDK/reference/classes#sonioxtranslationjob-wait) | Wait for the translation job to complete | | [`translationJob.refresh()`](/sdk/node-SDK/reference/classes#sonioxtranslationjob-refresh) | Refresh translation job status | ## Voice ### Available voice instance methods | Method | Description | | ----------------------------------------------------------------------------------- | ------------------------------------------------- | | [`voice.delete()`](/sdk/node-SDK/reference/classes#sonioxvoice-delete) | Permanently delete this voice | | [`voice.isReady(model)`](/sdk/node-SDK/reference/classes#sonioxvoice-isready) | Check if the voice is ready for a given TTS model | | [`voice.recompute(options)`](/sdk/node-SDK/reference/classes#sonioxvoice-recompute) | Prepare this voice for additional TTS models | | [`voice.toJSON()`](/sdk/node-SDK/reference/classes#sonioxvoice-tojson) | Get raw voice metadata | ## Transcript | Method | Description | | ------------------------------------------------------------------------------------ | ----------------------- | | [`transcript.segments()`](/sdk/node-SDK/reference/classes#sonioxtranscript-segments) | Get transcript segments | | [`transcript.text`](/sdk/node-SDK/reference/classes#sonioxtranscript) | Transcript text | | [`transcript.tokens`](/sdk/node-SDK/reference/classes#sonioxtranscript) | Transcript tokens | ## Real-time STT Session | Method | Description | | ----------------------------------------------------------------------------------------------------------- | ------------------------------------------------ | | [`realtime.stt.connect()`](/sdk/node-SDK/reference/classes#realtimesttsession-connect) | Establish websocket connection | | [`realtime.stt.close()`](/sdk/node-SDK/reference/classes#realtimesttsession-close) | Close websocket connection | | [`realtime.stt.finalize()`](/sdk/node-SDK/reference/classes#realtimesttsession-finalize) | Request server to finalize current transcription | | [`realtime.stt.finish()`](/sdk/node-SDK/reference/classes#realtimesttsession-finish) | Gracefully finish the session | | [`realtime.stt.keepAlive()`](/sdk/node-SDK/reference/classes#realtimesttsession-keepalive) | Send keepalive message | | [`realtime.stt.off()`](/sdk/node-SDK/reference/classes#realtimesttsession-off) | Remove event handler | | [`realtime.stt.on()`](/sdk/node-SDK/reference/classes#realtimesttsession-on) | Register event handler | | [`realtime.stt.once()`](/sdk/node-SDK/reference/classes#realtimesttsession-once) | Register one-time event handler | | [`realtime.stt.sendAudio(audio)`](/sdk/node-SDK/reference/classes#realtimesttsession-sendaudio) | Send audio chunk | | [`realtime.stt.sendStream(stream, options)`](/sdk/node-SDK/reference/classes#realtimesttsession-sendstream) | Send audio stream | | [`realtime.stt.pause()`](/sdk/node-SDK/reference/classes#realtimesttsession-pause) | Pause audio transmission | | [`realtime.stt.resume()`](/sdk/node-SDK/reference/classes#realtimesttsession-resume) | Resume audio transmission | | [`realtime.stt.paused`](/sdk/node-SDK/reference/classes#realtimesttsession-paused) | Whether the session is currently paused | | [`realtime.stt.state`](/sdk/node-SDK/reference/classes#realtimesttsession-state) | Current session state | ## Real-time TTS Stream | Method | Description | | ----------------------------------------------------------------------------------------------- | ------------------------------------------------------------ | | [`stream.sendText(text, options?)`](/sdk/node-SDK/reference/classes#realtimettsstream-sendtext) | Send one text chunk to the TTS stream | | [`stream.sendStream(source)`](/sdk/node-SDK/reference/classes#realtimettsstream-sendstream) | Pipe an async iterable of text chunks into the stream | | [`stream.finish()`](/sdk/node-SDK/reference/classes#realtimettsstream-finish) | Signal that no more text will be sent | | [`stream.cancel()`](/sdk/node-SDK/reference/classes#realtimettsstream-cancel) | Cancel generation and terminate the stream | | [`stream.close()`](/sdk/node-SDK/reference/classes#realtimettsstream-close) | Close the stream (and the WebSocket for single-stream usage) | | [`stream.on()`](/sdk/node-SDK/reference/classes#realtimettsstream-on) | Register event handler | | [`stream.once()`](/sdk/node-SDK/reference/classes#realtimettsstream-once) | Register one-time event handler | | [`stream.off()`](/sdk/node-SDK/reference/classes#realtimettsstream-off) | Remove event handler | | [`stream.state`](/sdk/node-SDK/reference/classes#realtimettsstream-state) | Current stream lifecycle state | | [`stream.streamId`](/sdk/node-SDK/reference/classes#realtimettsstream-properties) | Identifier of this TTS stream | ## Real-time TTS Connection | Method | Description | | --------------------------------------------------------------------------------------------- | --------------------------------------------------------------- | | [`connection.connect()`](/sdk/node-SDK/reference/classes#realtimettsconnection-connect) | Open the WebSocket connection | | [`connection.stream(input?)`](/sdk/node-SDK/reference/classes#realtimettsconnection-stream) | Open a new TTS stream on this connection | | [`connection.close()`](/sdk/node-SDK/reference/classes#realtimettsconnection-close) | Close the WebSocket connection and terminate all active streams | | [`connection.on()`](/sdk/node-SDK/reference/classes#realtimettsconnection-on) | Register event handler | | [`connection.once()`](/sdk/node-SDK/reference/classes#realtimettsconnection-once) | Register one-time event handler | | [`connection.off()`](/sdk/node-SDK/reference/classes#realtimettsconnection-off) | Remove event handler | | [`connection.isConnected`](/sdk/node-SDK/reference/classes#realtimettsconnection-isconnected) | Whether the WebSocket is currently connected | # Types URL: /sdk/node-SDK/reference/types Soniox Node SDK — Types Reference ## AudioData ```ts type AudioData = Buffer | Uint8Array | ArrayBuffer; ``` Audio data types accepted by sendAudio. In Node.js, Buffer is also accepted since Buffer extends Uint8Array. *** ## AudioFormat ```ts type AudioFormat = | "pcm_s8" | "pcm_s8le" | "pcm_s8be" | "pcm_s16le" | "pcm_s16be" | "pcm_s24le" | "pcm_s24be" | "pcm_s32le" | "pcm_s32be" | "pcm_u8" | "pcm_u8le" | "pcm_u8be" | "pcm_u16le" | "pcm_u16be" | "pcm_u24le" | "pcm_u24be" | "pcm_u32le" | "pcm_u32be" | "pcm_f32le" | "pcm_f32be" | "pcm_f64le" | "pcm_f64be" | "mulaw" | "alaw" | "aac" | "aiff" | "amr" | "asf" | "wav" | "mp3" | "flac" | "ogg" | "webm"; ``` Supported audio formats for real-time transcription. *** ## CleanupTarget ```ts type CleanupTarget = "file" | "transcription"; ``` Resource types that can be cleaned up after transcription completes. * `'file'` - The uploaded file * `'transcription'` - The transcription record *** ## ConcurrencyCurrentValues ```ts type ConcurrencyCurrentValues = { transcribe_concurrent: number; tts_concurrent: number; }; ``` Live concurrency counts. **Properties** | Property | Type | Description | | --------------------------------------------------------------------------------- | -------- | ---------------------------------------------------- | | `transcribe_concurrent` | `number` | Current number of concurrent transcription sessions. | | `tts_concurrent` | `number` | Current number of concurrent TTS sessions. | *** ## ConcurrencyLimitValues ```ts type ConcurrencyLimitValues = { transcribe_concurrent: number | null; tts_concurrent: number | null; }; ``` Configured concurrency limits. **Properties** | Property | Type | Description | | ------------------------------------------------------------------------------- | ------------------ | --------------------------------------------------------------------------- | | `transcribe_concurrent` | `number` \| `null` | Configured transcription concurrency limit. Null means no configured limit. | | `tts_concurrent` | `number` \| `null` | Configured TTS concurrency limit. Null means no configured limit. | *** ## ConcurrencyLimitsResponse ```ts type ConcurrencyLimitsResponse = { organization: ConcurrencyScopeValues; project: ConcurrencyScopeValues; }; ``` Current concurrent counts plus configured concurrency limits for the project and its organization. Values are region-scoped. **Properties** | Property | Type | Description | | ---------------------------------------------------------------- | -------------------------------------------------------- | ------------------------------------------------- | | `organization` | [`ConcurrencyScopeValues`](types#concurrencyscopevalues) | Organization-level concurrency counts and limits. | | `project` | [`ConcurrencyScopeValues`](types#concurrencyscopevalues) | Project-level concurrency counts and limits. | *** ## ConcurrencyScopeValues ```ts type ConcurrencyScopeValues = { current: ConcurrencyCurrentValues; limits: ConcurrencyLimitValues; }; ``` Current counts and configured limits for a concurrency scope. **Properties** | Property | Type | Description | | --------------------------------------------------- | ------------------------------------------------------------ | -------------------------------- | | `current` | [`ConcurrencyCurrentValues`](types#concurrencycurrentvalues) | Current live concurrency counts. | | `limits` | [`ConcurrencyLimitValues`](types#concurrencylimitvalues) | Configured concurrency limits. | *** ## ConcurrentStreamKind ```ts type ConcurrentStreamKind = "stt" | "tts"; ``` Stream kind for concurrent streams history. * `stt`: Speech-to-Text WebSocket sessions * `tts`: Text-to-Speech WebSocket streams and REST requests *** ## ConcurrentStreamsHistoryEntry ```ts type ConcurrentStreamsHistoryEntry = { period_sec: number; period_start: string; sample_count: number; sample_max: number; sample_min: number; sample_sum: number; total_count: number; }; ``` Per-period concurrent stream aggregate for the authenticated project. **Properties** | Property | Type | Description | | -------------------------------------------------------------------- | -------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `period_sec` | `number` | Aggregation period in seconds. | | `period_start` | `string` | Start of the aggregation period, UTC. Aligned to a multiple of `period_sec`. **Format** date-time | | `sample_count` | `number` | Number of values actually recorded in the period. For `period_sec=60` this is how many samples were taken during that minute, so it is usually larger than `total_count`. For hourly and daily periods it is the number of source periods that had data, at most `total_count`. `0` when the period had no activity. | | `sample_max` | `number` | Peak concurrent stream count in the period. Stays exact when periods are rolled up into hours and days. `0` when the period had no activity. | | `sample_min` | `number` | Lowest recorded concurrent stream count in the period. Always `0`, at every tier, because that is what the per-minute tier records. Use `sample_max` for the peak. | | `sample_sum` | `number` | Sum of the recorded concurrency values in the period. Divide by `sample_count` for the average while streams were active, or by `total_count` for the average across the whole period (idle slots as zero). | | `total_count` | `number` | Number of slots the period covers. `1` for `period_sec=60`, `60` for `3600` (minutes per hour), `24` for `86400` (hours per day). `0` when the period had no activity. | *** ## ConcurrentStreamsHistoryPeriodSec ```ts type ConcurrentStreamsHistoryPeriodSec = 60 | 3600 | 86400; ``` Aggregation period for concurrent streams history, in seconds. * `60`: per-minute * `3600`: hourly * `86400`: daily *** ## ConcurrentStreamsHistoryResponse ```ts type ConcurrentStreamsHistoryResponse = { entries: ConcurrentStreamsHistoryEntry[]; kind: ConcurrentStreamKind; }; ``` Concurrent streams history for the authenticated project. **Properties** | Property | Type | Description | | ------------------------------------------------------------- | ------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `entries` | [`ConcurrentStreamsHistoryEntry`](types#concurrentstreamshistoryentry)\[] | Per-period aggregates ordered by `period_start` ascending. Every aggregation period in the requested window is returned, with no gaps. Periods with no recorded activity have every numeric field set to `0`. | | `kind` | [`ConcurrentStreamKind`](types#concurrentstreamkind) | Stream kind these entries describe (`stt` or `tts`). | *** ## ContextGeneralEntry ```ts type ContextGeneralEntry = { key: string; value: string; }; ``` Key-value pair for general context information. **Properties** | Property | Type | Description | | -------------------------------------------- | -------- | ------------------------------------------------------------------------ | | `key` | `string` | The key describing the context type (e.g., "domain", "topic", "doctor"). | | `value` | `string` | The value for the context key. | *** ## ContextTranslationTerm ```ts type ContextTranslationTerm = { source: string; target: string; }; ``` Custom translation term mapping. **Properties** | Property | Type | Description | | ------------------------------------------------- | -------- | ------------------------------------ | | `source` | `string` | The source term to translate. | | `target` | `string` | The target translation for the term. | *** ## CreateTranscriptionOptions ```ts type CreateTranscriptionOptions = { audio_url?: string; client_reference_id?: string; context?: TranscriptionContext; enable_language_identification?: boolean; enable_speaker_diarization?: boolean; file_id?: string; language_hints?: string[]; language_hints_strict?: boolean; model: string; translation?: TranslationConfig; webhook_auth_header_name?: string; webhook_auth_header_value?: string; webhook_url?: string; }; ``` Options for creating a transcription. **Properties** | Property | Type | Description | | ------------------------------------------------------------------------------------------------------ | ---------------------------------------------------- | ------------------------------------------------------------------------------------------------- | | `audio_url?` | `string` | URL of a publicly accessible audio file. **Max Length** 4096 | | `client_reference_id?` | `string` | Optional tracking identifier. **Max Length** 256 | | `context?` | [`TranscriptionContext`](types#transcriptioncontext) | Additional context to improve transcription accuracy and formatting of specialized terms. | | `enable_language_identification?` | `boolean` | Enable automatic language identification. | | `enable_speaker_diarization?` | `boolean` | Enable speaker diarization to identify different speakers. | | `file_id?` | `string` | ID of a previously uploaded file. **Format** uuid | | `language_hints?` | `string`\[] | Array of expected ISO language codes to bias recognition. | | `language_hints_strict?` | `boolean` | When true, model relies more heavily on language hints. | | `model` | `string` | Speech-to-text model to use. **Max Length** 32 | | `translation?` | [`TranslationConfig`](types#translationconfig) | Translation configuration. | | `webhook_auth_header_name?` | `string` | Name of the authentication header sent with webhook notifications. **Max Length** 256 | | `webhook_auth_header_value?` | `string` | Authentication header value sent with webhook notifications. **Max Length** 256 | | `webhook_url?` | `string` | URL to receive webhook notifications when transcription is completed or fails. **Max Length** 256 | *** ## CreateVoiceInput ```ts type CreateVoiceInput = UploadFileInput; ``` Supported input types for the reference audio clip. *** ## CreateVoiceOptions ```ts type CreateVoiceOptions = { file: CreateVoiceInput; filename?: string; name: string; signal?: AbortSignal; timeout_ms?: number; }; ``` Options for creating a voice. **Properties** | Property | Type | Description | | ------------------------------------------------------ | -------------------------------------------- | ------------------------------------------------------------------------------------- | | `file` | [`CreateVoiceInput`](types#createvoiceinput) | The reference audio clip for the voice. Keep it up to 2 minutes and within 35 MB. | | `filename?` | `string` | Custom filename for the uploaded reference clip. | | `name` | `string` | A name for the voice, unique within your project. **Min Length** 1 **Max Length** 128 | | `signal?` | `AbortSignal` | AbortSignal for cancelling the request. | | `timeout_ms?` | `number` | Request timeout in milliseconds. | *** ## DeleteAllFilesOptions ```ts type DeleteAllFilesOptions = { signal?: AbortSignal; }; ``` Options for purging all files. **Properties** | Property | Type | Description | | ------------------------------------------------- | ------------- | ----------------------------------------------------- | | `signal?` | `AbortSignal` | AbortSignal for cancelling the delete\_all operation. | *** ## DeleteAllTranscriptionsOptions ```ts type DeleteAllTranscriptionsOptions = { on_progress?: (transcription, index) => void; signal?: AbortSignal; }; ``` Options for deleting all transcriptions. **Properties** | Property | Type | Description | | -------------------------------------------------------------------- | ------------------------------------ | ------------------------------------------------------------------------------------------------------------- | | `on_progress?` | (`transcription`, `index`) => `void` | Callback invoked before each transcription is deleted. Receives the transcription data and its 0-based index. | | `signal?` | `AbortSignal` | AbortSignal for cancelling the delete\_all operation. | *** ## DeleteAllVoicesOptions ```ts type DeleteAllVoicesOptions = { signal?: AbortSignal; }; ``` Options for deleting all voices. **Properties** | Property | Type | Description | | -------------------------------------------------- | ------------- | ----------------------------------------------------- | | `signal?` | `AbortSignal` | AbortSignal for cancelling the delete\_all operation. | *** ## ExpressLikeRequest ```ts type ExpressLikeRequest = { body?: unknown; headers: Record; method: string; }; ``` Express/Connect-style request object **Properties** | Property | Type | | ----------------------------------------------- | ----------------------------------------------------------- | | `body?` | `unknown` | | `headers` | `Record`\<`string`, `string` \| `string`\[] \| `undefined`> | | `method` | `string` | *** ## FastifyLikeRequest ```ts type FastifyLikeRequest = { body?: unknown; headers: Record; method: string; }; ``` Fastify-style request object **Properties** | Property | Type | | ----------------------------------------------- | ----------------------------------------------------------- | | `body?` | `unknown` | | `headers` | `Record`\<`string`, `string` \| `string`\[] \| `undefined`> | | `method` | `string` | *** ## FileIdentifier ```ts type FileIdentifier = | string | { id: string; }; ``` File identifier - either a string ID or an object with an id property. *** ## FilesCountResponse ```ts type FilesCountResponse = { playground: number; public_api: number; total: number; }; ``` Total number of files, split by source. **Properties** | Property | Type | Description | | ----------------------------------------------------- | -------- | -------------------------------------------- | | `playground` | `number` | Number of files uploaded via the Playground. | | `public_api` | `number` | Number of files uploaded via Public API. | | `total` | `number` | Total number of files across all sources. | *** ## GenerateSpeechOptions ```ts type GenerateSpeechOptions = { audio_format?: string; bitrate?: number; client_reference_id?: string; language?: string; model?: string; reduce_silence?: boolean; sample_rate?: number; signal?: AbortSignal; speed?: number; text: string; voice: string; }; ``` Options for REST TTS generation (`generate` / `generateStream`). **Properties** | Property | Type | Description | | --------------------------------------------------------------------------- | ------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `audio_format?` | `string` | Output audio format **Default** `'wav'` | | `bitrate?` | `number` | Codec bitrate in bps (for compressed formats). | | `client_reference_id?` | `string` | Optional tracking identifier. Does not need to be unique. Ignored if the request authenticates with a temporary API key. **Max Length** 256 | | `language?` | `string` | Language code. **Default** `'en'` | | `model?` | `string` | Text-to-Speech model to use. **Default** `'tts-rt-v2'` | | `reduce_silence?` | `boolean` | Shorten pauses between words in the generated speech. `false` (default) keeps the model's natural pacing; `true` tightens delivery by reducing silence between words. | | `sample_rate?` | `number` | Output sample rate in Hz. Required for raw PCM formats. | | `signal?` | `AbortSignal` | Optional AbortSignal for cancellation. | | `speed?` | `number` | Speaking rate. `1.0` is the normal rate; values below `1.0` slow speech down and values above `1.0` speed it up. Supported range is `0.7`-`1.3`. Defaults to `1.0` when omitted. | | `text` | `string` | Input text to generate as speech. | | `voice` | `string` | Voice identifier. | *** ## GetConcurrentStreamsHistoryOptions ```ts type GetConcurrentStreamsHistoryOptions = { end_time: string; kind: ConcurrentStreamKind; period_sec: ConcurrentStreamsHistoryPeriodSec; signal?: AbortSignal; start_time: string; }; ``` Options for retrieving concurrent streams history. **Properties** | Property | Type | Description | | --------------------------------------------------------------------- | ------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `end_time` | `string` | End of the time window (exclusive). Must be an ISO 8601 timestamp in UTC and strictly after `start_time`. Filters by `period_start`. **Example** `'2026-04-28T10:00:00Z'` | | `kind` | [`ConcurrentStreamKind`](types#concurrentstreamkind) | Stream kind to return. | | `period_sec` | [`ConcurrentStreamsHistoryPeriodSec`](types#concurrentstreamshistoryperiodsec) | Aggregation period in seconds. Also caps how long the requested window may be: for `period_sec=60` the window must not exceed 7 days. A window longer than the cap for its period, or one that would return more than 20000 entries, is rejected with a `400 invalid_request` error. | | `signal?` | `AbortSignal` | AbortSignal for cancelling the request. | | `start_time` | `string` | Start of the time window (inclusive). Must be an ISO 8601 timestamp in UTC. Filters by `period_start`. **Example** `'2026-04-28T09:00:00Z'` | *** ## GetUsageSummaryOptions ```ts type GetUsageSummaryOptions = { end_time: string; signal?: AbortSignal; start_time: string; }; ``` Options for retrieving an aggregated usage summary. Usage is aggregated by whole UTC day over the half-open window `[start_time, end_time)`: a day is included when the window covers any part of it. The window must not cover more than 366 UTC days. **Properties** | Property | Type | Description | | --------------------------------------------------------- | ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `end_time` | `string` | End of the window (exclusive). Must be an ISO 8601 timestamp in UTC and strictly after `start_time`. Its UTC day is included unless it falls exactly on midnight. **Example** `'2026-04-03T00:00:00Z'` | | `signal?` | `AbortSignal` | AbortSignal for cancelling the request. | | `start_time` | `string` | Start of the window (inclusive). Must be an ISO 8601 timestamp in UTC. Its UTC day is included. **Example** `'2026-04-01T00:00:00Z'` | *** ## HandleWebhookOptions ```ts type HandleWebhookOptions = { auth?: WebhookAuthConfig; body: unknown; headers: WebhookHeaders; method: string; }; ``` Options for the handleWebhook function **Properties** | Property | Type | Description | | ------------------------------------------------- | ---------------------------------------------- | ---------------------------------------- | | `auth?` | [`WebhookAuthConfig`](types#webhookauthconfig) | Optional authentication configuration | | `body` | `unknown` | Request body (parsed JSON or raw string) | | `headers` | [`WebhookHeaders`](types#webhookheaders) | Request headers | | `method` | `string` | HTTP method of the request | *** ## HonoLikeContext ```ts type HonoLikeContext = { req: { method: string; header: string | undefined; json: Promise; }; }; ``` Hono context object **Properties** | Property | Type | | ------------------------------------------------- | ------------------------------------------------------------------------------------------ | | `req` | \{ `method`: `string`; `header`: `string` \| `undefined`; `json`: `Promise`\<`unknown`>; } | | `req.method` | `string` | | `req.header` | `string` \| `undefined` | | `req.json` | `Promise`\<`unknown`> | *** ## HttpErrorCode ```ts type HttpErrorCode = "network_error" | "timeout" | "aborted" | "http_error" | "parse_error"; ``` Error codes for HTTP client errors *** ## HttpMethod ```ts type HttpMethod = "GET" | "POST" | "PUT" | "PATCH" | "DELETE" | "HEAD"; ``` HTTP methods supported by the client *** ## HttpRequestBody ```ts type HttpRequestBody = | string | Record | ArrayBuffer | Uint8Array | FormData | null; ``` Request body types *** ## HttpResponseType ```ts type HttpResponseType = "json" | "text" | "arrayBuffer"; ``` Response types *** ## ListFilesOptions ```ts type ListFilesOptions = { cursor?: string; limit?: number; signal?: AbortSignal; }; ``` Options for listing files. **Properties** | Property | Type | Description | | -------------------------------------------- | ------------- | ------------------------------------------------------------------------------------ | | `cursor?` | `string` | Pagination cursor for the next page of results. | | `limit?` | `number` | Maximum number of files to return. **Default** `1000` **Minimum** 1 **Maximum** 1000 | | `signal?` | `AbortSignal` | AbortSignal for cancelling the request | *** ## ListFilesResponse\ ```ts type ListFilesResponse = { files: T[]; next_page_cursor: string | null; }; ``` Response from listing files. **Type Parameters** | Type Parameter | | -------------- | | `T` | **Properties** | Property | Type | Description | | ----------------------------------------------------------------- | ------------------ | ------------------------------------------------------------------------------------------------------------ | | `files` | `T`\[] | List of uploaded files. | | `next_page_cursor` | `string` \| `null` | A pagination token that references the next page of results. When null, no additional results are available. | *** ## ListTranscriptionsOptions ```ts type ListTranscriptionsOptions = { cursor?: string; limit?: number; }; ``` Options for listing transcriptions **Properties** | Property | Type | Description | | ----------------------------------------------------- | -------- | --------------------------------------------------------------------------------------------- | | `cursor?` | `string` | Pagination cursor for the next page of results | | `limit?` | `number` | Maximum number of transcriptions to return. **Default** `1000` **Minimum** 1 **Maximum** 1000 | *** ## ListTranscriptionsResponse\ ```ts type ListTranscriptionsResponse = { next_page_cursor: string | null; transcriptions: T[]; }; ``` Response from listing transcriptions. **Type Parameters** | Type Parameter | | -------------- | | `T` | **Properties** | Property | Type | Description | | -------------------------------------------------------------------------- | ------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------ | | `next_page_cursor` | `string` \| `null` | A pagination token that references the next page of results. When null, no additional results are available. TODO: potentially can be undefined? | | `transcriptions` | `T`\[] | List of transcriptions. | *** ## ListUsageLogsOptions ```ts type ListUsageLogsOptions = { cursor?: string; end_time: string; limit?: number; signal?: AbortSignal; sort?: UsageLogsSort; start_time: string; }; ``` Options for listing usage logs. **Properties** | Property | Type | Description | | ------------------------------------------------------- | -------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- | | `cursor?` | `string` | Pagination cursor for the next page of results. | | `end_time` | `string` | End of the time window (exclusive), filtering by request end time. Must be an ISO 8601 timestamp in UTC. **Example** `'2026-04-29T09:00:00Z'` | | `limit?` | `number` | Maximum number of usage log entries to return. **Default** `1000` **Minimum** 1 **Maximum** 1000 | | `signal?` | `AbortSignal` | AbortSignal for cancelling the request. | | `sort?` | [`UsageLogsSort`](types#usagelogssort) | Sort order by end\_time. **Default** `'end_time_asc'` | | `start_time` | `string` | Start of the time window (inclusive), filtering by request end time. Must be an ISO 8601 timestamp in UTC. **Example** `'2026-04-28T09:00:00Z'` | *** ## ListUsageLogsResponse ```ts type ListUsageLogsResponse = { next_page_cursor: string | null; usage_logs: SonioxUsageLog[]; }; ``` Response from listing usage logs. **Properties** | Property | Type | Description | | -------------------------------------------------------------------- | ------------------------------------------- | ---------------------------------------------------------------------- | | `next_page_cursor` | `string` \| `null` | Pagination cursor for the next page of results. Null if no more pages. | | `usage_logs` | [`SonioxUsageLog`](types#sonioxusagelog)\[] | Per-request usage log entries ordered by end\_time and UUID. | *** ## ListVoicesOptions ```ts type ListVoicesOptions = { cursor?: string; limit?: number; signal?: AbortSignal; }; ``` Options for listing voices. **Properties** | Property | Type | Description | | --------------------------------------------- | ------------- | ----------------------------------------------- | | `cursor?` | `string` | Pagination cursor for the next page of results. | | `limit?` | `number` | Maximum number of voices to return. | | `signal?` | `AbortSignal` | AbortSignal for cancelling the request. | *** ## ListVoicesResponse\ ```ts type ListVoicesResponse = { next_page_cursor: string | null; voices: T[]; }; ``` Response from listing voices. **Type Parameters** | Type Parameter | | -------------- | | `T` | **Properties** | Property | Type | Description | | ------------------------------------------------------------------ | ------------------ | ------------------------------------------------------------------------------------------------------------ | | `next_page_cursor` | `string` \| `null` | A pagination token that references the next page of results. When null, no additional results are available. | | `voices` | `T`\[] | List of voices. | *** ## NestJSLikeRequest ```ts type NestJSLikeRequest = { body?: unknown; headers: Record; method: string; }; ``` NestJS-style request object (uses Express under the hood by default) **Properties** | Property | Type | | ---------------------------------------------- | ----------------------------------------------------------- | | `body?` | `unknown` | | `headers` | `Record`\<`string`, `string` \| `string`\[] \| `undefined`> | | `method` | `string` | *** ## OneWayTranslation ```ts type OneWayTranslation = { duration_ms: number; from?: string; mode: "one_way"; original_text: string; segments: TranslationSegment[]; to: string; translation_text: string; }; ``` Result of a one-way translation (`{ to }` or `{ to, from }` mode). `original_text` and `translation_text` flatten the per-segment content across the whole audio, which is useful when the caller just wants two parallel strings. **Properties** | Property | Type | Description | | ---------------------------------------------------------------- | --------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- | | `duration_ms` | `number` | Total audio duration in milliseconds. Equals the largest `end_ms` across all original tokens, or `0` when there are no original tokens. | | `from?` | `string` | Source language hint that was supplied via `from`. Undefined when only `to` was provided and the source language was auto-detected. | | `mode` | `"one_way"` | - | | `original_text` | `string` | Concatenated text of every original token across all segments. | | `segments` | [`TranslationSegment`](types#translationsegment)\[] | Per-utterance segments in audio order. | | `to` | `string` | Target language code (the `to` value passed in). | | `translation_text` | `string` | Concatenated text of every translation token across all segments. | *** ## OneWayTranslationConfig ```ts type OneWayTranslationConfig = { target_language: string; type: "one_way"; }; ``` One-way translation configuration. Translates all spoken languages into a single target language. **Properties** | Property | Type | Description | | -------------------------------------------------------------------- | ----------- | -------------------------------------------------------------- | | `target_language` | `string` | Target language code for translation (e.g., "fr", "es", "de"). | | `type` | `"one_way"` | Translation type. | *** ## QueryParams ```ts type QueryParams = Record; ``` Query parameters *** ## RealtimeClientOptions ```ts type RealtimeClientOptions = { api_key: string; default_session_options?: SttSessionOptions; stt_defaults?: Partial; tts_connection_options?: TtsConnectionOptions; tts_defaults?: Partial; tts_ws_url: string; ws_base_url: string; }; ``` Real-time API configuration options for the client. **Properties** | Property | Type | Description | | ----------------------------------------------------------------------------------- | -------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------- | | `api_key` | `string` | API key for real-time sessions. | | `default_session_options?` | [`SttSessionOptions`](types#sttsessionoptions) | Default session options applied to all real-time STT sessions. Can be overridden per-session. | | `stt_defaults?` | `Partial`\<[`SttSessionConfig`](types#sttsessionconfig)> | STT session config defaults. Merged as the base layer when opening STT sessions via `realtime.stt(config)`; caller fields override. | | `tts_connection_options?` | [`TtsConnectionOptions`](types#ttsconnectionoptions) | Default TTS connection options. | | `tts_defaults?` | `Partial`\<[`TtsStreamConfig`](types#ttsstreamconfig)> | TTS stream config defaults. Merged as the base layer when opening TTS streams via `realtime.tts(...)`; caller fields override. | | `tts_ws_url` | `string` | TTS WebSocket URL for real-time connections. **Default** `'wss://tts-rt.soniox.com/tts-websocket'` | | `ws_base_url` | `string` | STT WebSocket base URL for real-time connections. **Default** `'wss://stt-rt.soniox.com/transcribe-websocket'` | *** ## RealtimeErrorCode ```ts type RealtimeErrorCode = | "auth_error" | "bad_request" | "quota_exceeded" | "connection_error" | "network_error" | "aborted" | "state_error" | "realtime_error"; ``` Error codes for Real-time (WebSocket) API errors *** ## RealtimeEvent ```ts type RealtimeEvent = | { data: RealtimeResult; kind: "result"; } | { kind: "endpoint"; } | { kind: "finalized"; } | { kind: "finished"; }; ``` Typed event for async iterator consumption. *** ## RealtimeOptions ```ts type RealtimeOptions = { default_session_options?: SttSessionOptions; stt_defaults?: Partial; tts_connection_options?: TtsConnectionOptions; tts_defaults?: Partial; tts_ws_url?: string; ws_base_url?: string; }; ``` Real-time configuration options for the main client. **Properties** | Property | Type | Description | | ----------------------------------------------------------------------------- | -------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `default_session_options?` | [`SttSessionOptions`](types#sttsessionoptions) | Default session options applied to all real-time STT sessions. Can be overridden per-session. | | `stt_defaults?` | `Partial`\<[`SttSessionConfig`](types#sttsessionconfig)> | Default STT session config fields (model, language hints, context, etc.). Merged as the base layer when opening STT sessions via `client.realtime.stt(config)`. Fields on the caller-provided `config` override these defaults. Equivalent to SonioxConnectionConfig.stt\_defaults on the web/react clients. | | `tts_connection_options?` | [`TtsConnectionOptions`](types#ttsconnectionoptions) | Default TTS connection options (keepalive interval, connect timeout). | | `tts_defaults?` | `Partial`\<[`TtsStreamConfig`](types#ttsstreamconfig)> | Default TTS stream config fields (model, voice, language, audio\_format, etc.). Merged as the base layer when opening TTS streams via `client.realtime.tts(...)`. Fields on the caller-provided [TtsStreamInput](types#ttsstreaminput) override these defaults. Equivalent to SonioxConnectionConfig.tts\_defaults on the web/react clients. | | `tts_ws_url?` | `string` | TTS WebSocket URL for real-time connections. Falls back to SONIOX\_TTS\_WS\_URL environment variable, then to 'wss\://tts-rt.soniox.com/tts-websocket'. | | `ws_base_url?` | `string` | STT WebSocket base URL for real-time connections. Falls back to SONIOX\_WS\_URL environment variable, then to 'wss\://stt-rt.soniox.com/transcribe-websocket'. | *** ## RealtimeResult ```ts type RealtimeResult = { final_audio_proc_ms: number; finished?: boolean; tokens: RealtimeToken[]; total_audio_proc_ms: number; }; ``` A result message from the real-time WebSocket. **Properties** | Property | Type | Description | | ------------------------------------------------------------------- | ----------------------------------------- | -------------------------------------------------- | | `final_audio_proc_ms` | `number` | Milliseconds of audio that have been finalized. | | `finished?` | `boolean` | Whether this is the final result (session ending). | | `tokens` | [`RealtimeToken`](types#realtimetoken)\[] | Tokens in this result. | | `total_audio_proc_ms` | `number` | Total milliseconds of audio processed. | *** ## RealtimeSegment ```ts type RealtimeSegment = { end_ms?: number; language?: string; speaker?: string; start_ms?: number; text: string; tokens: RealtimeToken[]; }; ``` A segment of contiguous real-time tokens grouped by speaker/language. **Properties** | Property | Type | Description | | ----------------------------------------------- | ----------------------------------------- | ------------------------------------------------------------- | | `end_ms?` | `number` | End time of the segment in milliseconds (from last token). | | `language?` | `string` | Detected language code (if language identification enabled). | | `speaker?` | `string` | Speaker identifier (if diarization enabled). | | `start_ms?` | `number` | Start time of the segment in milliseconds (from first token). | | `text` | `string` | Concatenated text of all tokens in this segment. | | `tokens` | [`RealtimeToken`](types#realtimetoken)\[] | Original tokens in this segment. | *** ## RealtimeSegmentBufferOptions ```ts type RealtimeSegmentBufferOptions = { final_only?: boolean; group_by?: SegmentGroupKey[]; max_ms?: number; max_tokens?: number; }; ``` Options for rolling real-time segmentation buffers. **Properties** | Property | Type | Description | | ---------------------------------------------------------------- | --------------------------------------------- | --------------------------------------------------------------------------------------------------------------- | | `final_only?` | `boolean` | When true, only tokens marked as final are buffered. **Default** `true` | | `group_by?` | [`SegmentGroupKey`](types#segmentgroupkey)\[] | Fields to group by. A new segment starts when any of these fields changes **Default** `['speaker', 'language']` | | `max_ms?` | `number` | Maximum time window to keep in milliseconds (requires token timings). | | `max_tokens?` | `number` | Maximum number of tokens to keep in the buffer. **Default** `2000` | *** ## RealtimeSegmentOptions ```ts type RealtimeSegmentOptions = { final_only?: boolean; group_by?: SegmentGroupKey[]; }; ``` Options for segmenting real-time tokens. **Properties** | Property | Type | Description | | ---------------------------------------------------------- | --------------------------------------------- | --------------------------------------------------------------------------------------------------------------- | | `final_only?` | `boolean` | When true, only tokens marked as final are included. **Default** `false` | | `group_by?` | [`SegmentGroupKey`](types#segmentgroupkey)\[] | Fields to group by. A new segment starts when any of these fields changes **Default** `['speaker', 'language']` | *** ## RealtimeToken ```ts type RealtimeToken = { confidence: number; end_ms?: number; is_final: boolean; language?: string; source_language?: string; speaker?: string; start_ms?: number; text: string; translation_status?: "none" | "original" | "translation"; }; ``` A single token from the real-time transcription. **Properties** | Property | Type | Description | | ----------------------------------------------------------------- | ------------------------------------------- | ------------------------------------------------------------ | | `confidence` | `number` | Confidence score (0.0 to 1.0). | | `end_ms?` | `number` | End time in milliseconds relative to audio start. | | `is_final` | `boolean` | Whether this is a finalized token. | | `language?` | `string` | Detected language code (if language identification enabled). | | `source_language?` | `string` | Source language for translated tokens. | | `speaker?` | `string` | Speaker identifier (if diarization enabled). | | `start_ms?` | `number` | Start time in milliseconds relative to audio start. | | `text` | `string` | The transcribed text. | | `translation_status?` | `"none"` \| `"original"` \| `"translation"` | Translation status of this token. | *** ## RealtimeUtterance ```ts type RealtimeUtterance = { end_ms?: number; final_audio_proc_ms?: number; language?: string; segments: RealtimeSegment[]; speaker?: string; start_ms?: number; text: string; tokens: RealtimeToken[]; total_audio_proc_ms?: number; }; ``` A single utterance built from real-time segments. **Properties** | Property | Type | Description | | ----------------------------------------------------------------------- | --------------------------------------------- | ----------------------------------------------------------------- | | `end_ms?` | `number` | End time of the utterance in milliseconds (from last segment). | | `final_audio_proc_ms?` | `number` | Milliseconds of audio that have been finalized at flush time. | | `language?` | `string` | Detected language code when consistent across segments. | | `segments` | [`RealtimeSegment`](types#realtimesegment)\[] | Segments included in this utterance. | | `speaker?` | `string` | Speaker identifier when consistent across segments. | | `start_ms?` | `number` | Start time of the utterance in milliseconds (from first segment). | | `text` | `string` | Concatenated text of all segments in this utterance. | | `tokens` | [`RealtimeToken`](types#realtimetoken)\[] | Tokens included in this utterance. | | `total_audio_proc_ms?` | `number` | Total milliseconds of audio processed at flush time. | *** ## RealtimeUtteranceBufferOptions ```ts type RealtimeUtteranceBufferOptions = { final_only?: boolean; group_by?: SegmentGroupKey[]; max_ms?: number; max_tokens?: number; }; ``` Options for buffering real-time utterances. **Properties** | Property | Type | Description | | ------------------------------------------------------------------ | --------------------------------------------- | --------------------------------------------------------------------------------------------------------------- | | `final_only?` | `boolean` | When true, only tokens marked as final are buffered. **Default** `true` | | `group_by?` | [`SegmentGroupKey`](types#segmentgroupkey)\[] | Fields to group by. A new segment starts when any of these fields changes **Default** `['speaker', 'language']` | | `max_ms?` | `number` | Maximum time window to keep in milliseconds (requires token timings). | | `max_tokens?` | `number` | Maximum number of tokens to keep in the buffer. **Default** `2000` | *** ## RecomputeVoiceOptions ```ts type RecomputeVoiceOptions = { model?: string | null; signal?: AbortSignal; }; ``` Options for recomputing a voice. **Properties** | Property | Type | Description | | ------------------------------------------------- | ------------------ | ------------------------------------------------------------------------------------------------------------------------- | | `model?` | `string` \| `null` | The model to prepare this voice for. If omitted, the voice is prepared for every available model it is not ready for yet. | | `signal?` | `AbortSignal` | AbortSignal for cancelling the request. | *** ## SegmentGroupKey ```ts type SegmentGroupKey = "speaker" | "language"; ``` Fields that can be used to group tokens into segments *** ## SegmentTranscriptOptions ```ts type SegmentTranscriptOptions = { group_by?: SegmentGroupKey[]; }; ``` Options for segmenting a transcript **Properties** | Property | Type | Description | | -------------------------------------------------------- | --------------------------------------------- | --------------------------------------------------------------------------------------------------------------- | | `group_by?` | [`SegmentGroupKey`](types#segmentgroupkey)\[] | Fields to group by. A new segment starts when any of these fields changes **Default** `['speaker', 'language']` | *** ## SendStreamOptions ```ts type SendStreamOptions = { finish?: boolean; pace_ms?: number; }; ``` Options for streaming audio from an async iterable source. **Properties** | Property | Type | Description | | ----------------------------------------------- | --------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- | | `finish?` | `boolean` | When true, calls finish() automatically after the stream ends. **Default** `false` | | `pace_ms?` | `number` | Delay in milliseconds between sending chunks. Useful for simulating real-time pace when streaming pre-recorded files. Not needed for live audio sources. | *** ## SonioxErrorCode ```ts type SonioxErrorCode = | RealtimeErrorCode | "soniox_error" | HttpErrorCode; ``` All possible SDK error codes (real-time + HTTP-specific codes) *** ## SonioxFileData ```ts type SonioxFileData = { client_reference_id?: string | null; created_at: string; filename: string; id: string; size: number; }; ``` Raw file metadata from the API. **Properties** | Property | Type | Description | | -------------------------------------------------------------------- | ------------------ | ------------------------------------------------------------------------- | | `client_reference_id?` | `string` \| `null` | Optional tracking identifier string. | | `created_at` | `string` | UTC timestamp indicating when the file was uploaded. **Format** date-time | | `filename` | `string` | Name of the file. | | `id` | `string` | Unique identifier of the file. **Format** uuid | | `size` | `number` | Size of the file in bytes. | *** ## SonioxLanguage ```ts type SonioxLanguage = { code: string; name: string; }; ``` **Properties** | Property | Type | Description | | ------------------------------------- | -------- | ----------------------- | | `code` | `string` | 2-letter language code. | | `name` | `string` | Language name. | *** ## SonioxModel ```ts type SonioxModel = { aliased_model_id: string | null; context_version: number | null; endpoint_latency_adjustment_max_level: number; id: string; languages: SonioxLanguage[]; name: string; one_way_translation: string | null; supports_endpoint_latency_adjustment: boolean; supports_endpoint_sensitivity: boolean; supports_language_hints_strict: boolean; supports_max_endpoint_delay: boolean; transcription_mode: SonioxTranscriptionMode; translation_targets: SonioxTranslationTarget[]; two_way_translation: string | null; two_way_translation_pairs: string[]; }; ``` **Properties** | Property | Type | Description | | ---------------------------------------------------------------------------------------------------- | ------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `aliased_model_id` | `string` \| `null` | If this is an alias, the id of the aliased model. Null for non-alias models. | | `context_version` | `number` \| `null` | Version of context supported. | | `endpoint_latency_adjustment_max_level` | `number` | Maximum `endpoint_latency_adjustment_level` the model accepts. Valid levels are `0` (no adjustment) through this value; `0` means the feature is unsupported. | | `id` | `string` | Unique identifier of the model. | | `languages` | [`SonioxLanguage`](types#sonioxlanguage)\[] | List of languages supported by the model. | | `name` | `string` | Name of the model. | | `one_way_translation` | `string` \| `null` | When contains string 'all\_languages', any laguage from languages can be used | | `supports_endpoint_latency_adjustment` | `boolean` | Whether the model supports `endpoint_latency_adjustment_level` on real-time sessions. | | `supports_endpoint_sensitivity` | `boolean` | Whether the model supports endpoint sensitivity configuration. | | `supports_language_hints_strict` | `boolean` | Whether the model supports `language_hints_strict`. | | `supports_max_endpoint_delay` | `boolean` | Whether the model supports `max_endpoint_delay_ms` on real-time sessions. | | `transcription_mode` | [`SonioxTranscriptionMode`](types#sonioxtranscriptionmode) | Transcription mode of the model. | | `translation_targets` | [`SonioxTranslationTarget`](types#sonioxtranslationtarget)\[] | List of supported one-way translation targets. If list is empty, check for one\_way\_translation field | | `two_way_translation` | `string` \| `null` | When contains string 'all\_languages',' any laguage pair from languages can be used | | `two_way_translation_pairs` | `string`\[] | List of supported two-way translation pairs. If list is empty, check for two\_way\_translation field | *** ## SonioxNodeClientOptions ```ts type SonioxNodeClientOptions = { api_key?: string; base_domain?: string; base_url?: string; http_client?: HttpClient; realtime?: RealtimeOptions; region?: SonioxRegion; stt_defaults?: Partial; tts_api_url?: string; tts_defaults?: Partial; }; ``` **Properties** | Property | Type | Description | | --------------------------------------------------------------- | -------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `api_key?` | `string` | API key for authentication. Falls back to SONIOX\_API\_KEY environment variable if not provided. | | `base_domain?` | `string` | Base domain for all Soniox service URLs. A single override that derives all service endpoints from the pattern `{service}.{base_domain}`. Takes precedence over `region`. Falls back to SONIOX\_BASE\_DOMAIN environment variable. Individual URL fields (`base_url`, `tts_api_url`, `realtime.ws_base_url`, `realtime.tts_ws_url`) still take final precedence. **Example** `'eu.soniox.com'` | | `base_url?` | `string` | Base URL for the REST API. Falls back to SONIOX\_API\_BASE\_URL environment variable, then to the region-derived URL, then to '[https://api.soniox.com](https://api.soniox.com)'. | | `http_client?` | [`HttpClient`](types#httpclient) | Custom HTTP client implementation. | | `realtime?` | [`RealtimeOptions`](types#realtimeoptions) | Real-time API configuration options. | | `region?` | `SonioxRegion` | Deployment region. Determines which regional endpoints are used for both the REST API and real-time WebSocket connections. Leave `undefined` for the default (US) region. Shorthand for `base_domain: '{region}.soniox.com'`. `base_domain` takes precedence when both are provided. **See** [https://soniox.com/docs/stt/data-residency](https://soniox.com/docs/stt/data-residency) | | `stt_defaults?` | `Partial`\<[`SttSessionConfig`](types#sttsessionconfig)> | Default STT session config fields applied to every real-time STT session opened via `client.realtime.stt(config)`. Caller-provided fields override. Equivalent to SonioxConnectionConfig.stt\_defaults on the web/react clients. Prefer this when you want the same defaults across your whole Node process. | | `tts_api_url?` | `string` | TTS REST API URL. Falls back to SONIOX\_TTS\_API\_URL environment variable, then to the region-derived URL, then to '[https://tts-rt.soniox.com](https://tts-rt.soniox.com)'. | | `tts_defaults?` | `Partial`\<[`TtsStreamConfig`](types#ttsstreamconfig)> | Default TTS stream config fields applied to every real-time TTS stream opened via `client.realtime.tts(...)`. Caller-provided fields override. Equivalent to SonioxConnectionConfig.tts\_defaults on the web/react clients. | *** ## SonioxTranscriptionData ```ts type SonioxTranscriptionData = { audio_duration_ms?: number | null; audio_url?: string | null; client_reference_id?: string | null; context?: TranscriptionContext | null; created_at: string; enable_language_identification: boolean; enable_speaker_diarization: boolean; error_message?: string | null; error_type?: string | null; file_id?: string | null; filename: string; id: string; language_hints?: string[] | null; model: string; status: TranscriptionStatus; webhook_auth_header_name?: string | null; webhook_auth_header_value?: string | null; webhook_status_code?: number | null; webhook_url?: string | null; }; ``` Raw transcription metadata from the API. **Properties** | Property | Type | Description | | -------------------------------------------------------------------------------------------------- | -------------------------------------------------------------- | -------------------------------------------------------------------------------------------- | | `audio_duration_ms?` | `number` \| `null` | Duration of the audio in milliseconds. Only available after processing begins. | | `audio_url?` | `string` \| `null` | URL of the audio file being transcribed. | | `client_reference_id?` | `string` \| `null` | Optional tracking identifier. **Max Length** 256 | | `context?` | [`TranscriptionContext`](types#transcriptioncontext) \| `null` | Additional context provided for the transcription. | | `created_at` | `string` | UTC timestamp when the transcription was created. **Format** date-time | | `enable_language_identification` | `boolean` | When true, language is detected for each part of the transcription. | | `enable_speaker_diarization` | `boolean` | When true, speakers are identified and separated in the transcription output. | | `error_message?` | `string` \| `null` | Error message if transcription failed. Null for successful or in-progress transcriptions. | | `error_type?` | `string` \| `null` | Error type if transcription failed. Null for successful or in-progress transcriptions. | | `file_id?` | `string` \| `null` | ID of the uploaded file being transcribed. **Format** uuid | | `filename` | `string` | Name of the file being transcribed. | | `id` | `string` | Unique identifier of the transcription. **Format** uuid | | `language_hints?` | `string`\[] \| `null` | Expected languages in the audio. If not specified, languages are automatically detected. | | `model` | `string` | Speech-to-text model used. | | `status` | [`TranscriptionStatus`](types#transcriptionstatus) | Current status of the transcription. | | `webhook_auth_header_name?` | `string` \| `null` | Name of the authentication header sent with webhook notifications. | | `webhook_auth_header_value?` | `string` \| `null` | Authentication header value. Always returned masked. | | `webhook_status_code?` | `number` \| `null` | HTTP status code received from your server when webhook was delivered. Null if not yet sent. | | `webhook_url?` | `string` \| `null` | URL to receive webhook notifications when transcription is completed or fails. | *** ## SonioxTranscriptionMode ```ts type SonioxTranscriptionMode = "real_time" | "async"; ``` Transcription mode of the model. *** ## SonioxTranslation ```ts type SonioxTranslation = | OneWayTranslation | TwoWayTranslation; ``` Discriminated translation result returned by [SonioxTranslationJob.getTranslation](classes#sonioxtranslationjob-gettranslation), [SonioxTranslationJob.fetchTranslation](classes#sonioxtranslationjob-fetchtranslation), and [translateFromTranscript](types#translatefromtranscript). *** ## SonioxTranslationTarget ```ts type SonioxTranslationTarget = { exclude_source_languages: string[]; source_languages: string[]; target_language: string; }; ``` **Properties** | Property | Type | | -------------------------------------------------------------------------------------- | ----------- | | `exclude_source_languages` | `string`\[] | | `source_languages` | `string`\[] | | `target_language` | `string` | *** ## SonioxUsageLog ```ts type SonioxUsageLog = { client_reference_id?: string | null; cost_usd: string; end_time: string; input_audio_cost_usd: string; input_audio_duration_ms: number; input_audio_tokens: number; input_cost_usd: string; input_text_cost_usd: string; input_text_tokens: number; model: string; output_audio_cost_usd: string; output_audio_duration_ms: number; output_audio_tokens: number; output_cost_usd: string; output_text_cost_usd: string; output_text_tokens: number; request_scope: string; start_time: string; uuid: string; }; ``` Per-request usage log entry. **Properties** | Property | Type | Description | | ----------------------------------------------------------------------------- | ------------------ | ----------------------------------------------------------------------- | | `client_reference_id?` | `string` \| `null` | Optional tracking identifier provided by the caller. | | `cost_usd` | `string` | Total request cost in USD, represented as a decimal string. | | `end_time` | `string` | UTC timestamp indicating when the request ended. **Format** date-time | | `input_audio_cost_usd` | `string` | Input audio cost in USD, represented as a decimal string. | | `input_audio_duration_ms` | `number` | Input audio duration in milliseconds. | | `input_audio_tokens` | `number` | Number of input audio tokens. | | `input_cost_usd` | `string` | Input cost in USD, represented as a decimal string. | | `input_text_cost_usd` | `string` | Input text cost in USD, represented as a decimal string. | | `input_text_tokens` | `number` | Number of input text tokens. | | `model` | `string` | Model used for the request. | | `output_audio_cost_usd` | `string` | Output audio cost in USD, represented as a decimal string. | | `output_audio_duration_ms` | `number` | Output audio duration in milliseconds. | | `output_audio_tokens` | `number` | Number of output audio tokens. | | `output_cost_usd` | `string` | Output cost in USD, represented as a decimal string. | | `output_text_cost_usd` | `string` | Output text cost in USD, represented as a decimal string. | | `output_text_tokens` | `number` | Number of output text tokens. | | `request_scope` | `string` | Request scope. | | `start_time` | `string` | UTC timestamp indicating when the request started. **Format** date-time | | `uuid` | `string` | Unique identifier of the request. **Format** uuid | *** ## SonioxVoiceData ```ts type SonioxVoiceData = { created_at: string; filename: string; id: string; models: VoiceModelStatusEntry[]; name: string; }; ``` Raw voice metadata from the API. **Properties** | Property | Type | Description | | -------------------------------------------------- | --------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `created_at` | `string` | UTC timestamp indicating when the voice was created. **Format** date-time | | `filename` | `string` | Original file name of the uploaded audio clip. | | `id` | `string` | Unique identifier of the voice. **Format** uuid | | `models` | [`VoiceModelStatusEntry`](types#voicemodelstatusentry)\[] | Voice status for each available model. A model with status `not_computed` is not prepared yet (e.g. it was released after the voice was created); call recompute to prepare the voice for it. | | `name` | `string` | Name of the voice. | *** ## SttSessionConfig ```ts type SttSessionConfig = { audio_format?: "auto" | AudioFormat; client_reference_id?: string; context?: TranscriptionContext; enable_endpoint_detection?: boolean; enable_language_identification?: boolean; enable_speaker_diarization?: boolean; endpoint_latency_adjustment_level?: number; endpoint_sensitivity?: number; language_hints?: string[]; language_hints_strict?: boolean; max_endpoint_delay_ms?: number; model: string; num_channels?: number; sample_rate?: number; translation?: TranslationConfig; }; ``` Configuration sent to the Soniox WebSocket API when starting a session. **Properties** | Property | Type | Description | | -------------------------------------------------------------------------------------------------- | ---------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `audio_format?` | `"auto"` \| [`AudioFormat`](types#audioformat) | Audio format. Use 'auto' for automatic detection of container formats. For raw PCM formats, also set sample\_rate and num\_channels. **Default** `'auto'` | | `client_reference_id?` | `string` | Optional tracking identifier (max 256 chars). | | `context?` | [`TranscriptionContext`](types#transcriptioncontext) | Additional context to improve transcription accuracy. | | `enable_endpoint_detection?` | `boolean` | Enable endpoint detection for utterance boundaries. Useful for voice AI agents. | | `enable_language_identification?` | `boolean` | Enable automatic language detection. | | `enable_speaker_diarization?` | `boolean` | Enable speaker identification. | | `endpoint_latency_adjustment_level?` | `number` | Reduces endpoint latency compared to the default endpointing behavior. Higher values reduce endpoint latency more aggressively, which means endpoints are returned sooner and more endpoints may be emitted. This can split long speech into more segments and may slightly reduce word recognition accuracy because speech is finalized earlier. Allowed values are 0, 1, 2, and 3. The default value is 0 (default semantic endpointing behavior). | | `endpoint_sensitivity?` | `number` | Controls how aggressively endpoints are detected. Adjusts how likely the model is to emit an endpoint. Higher values make endpoints more likely, which can finalize segments sooner. Lower values make endpoints less likely, which can help the system wait longer before finalizing. Allowed values are between -1.0 and 1.0. The default value is 0.0. | | `language_hints?` | `string`\[] | Expected languages in the audio (ISO language codes). | | `language_hints_strict?` | `boolean` | When true, recognition is strongly biased toward language hints. Best-effort only, not a hard guarantee. | | `max_endpoint_delay_ms?` | `number` | Maximum delay between the end of speech and returned endpoint. Allowed values for maximum delay are between 500ms and 3000ms. The default value is 2000ms | | `model` | `string` | Speech-to-text model to use. | | `num_channels?` | `number` | Number of audio channels (required for raw audio formats). | | `sample_rate?` | `number` | Sample rate in Hz (required for PCM formats). | | `translation?` | [`TranslationConfig`](types#translationconfig) | Translation configuration. | *** ## SttSessionEvents ```ts type SttSessionEvents = { connected: () => void; disconnected: (reason?) => void; endpoint: () => void; error: (error) => void; finalized: () => void; finished: () => void; result: (result) => void; state_change: (update) => void; token: (token) => void; }; ``` Event handlers for the STT session. **Properties** | Property | Type | Description | | ------------------------------------------------------- | --------------------- | ------------------------------------------------- | | `connected` | () => `void` | Session connected and ready. | | `disconnected` | (`reason?`) => `void` | Session disconnected. | | `endpoint` | () => `void` | Endpoint detected (\ token). | | `error` | (`error`) => `void` | Error occurred. | | `finalized` | () => `void` | Finalization complete (\ token). | | `finished` | () => `void` | Session finished (server signaled end of stream). | | `result` | (`result`) => `void` | Parsed result received. | | `state_change` | (`update`) => `void` | Session state transition. | | `token` | (`token`) => `void` | Individual token received. | *** ## SttSessionOptions ```ts type SttSessionOptions = { connect_timeout_ms?: number; keepalive_interval_ms?: number; signal?: AbortSignal; }; ``` SDK-level session options (not sent to the server). **Properties** | Property | Type | Description | | --------------------------------------------------------------------------- | ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `connect_timeout_ms?` | `number` | Maximum time to wait for the WebSocket connection to open (milliseconds). If the connection is not established within this time, a [ConnectionError](classes#connectionerror) with message "Connection timed out" is thrown. **Default** `20000` | | `keepalive_interval_ms?` | `number` | Interval for sending keepalive messages while paused (milliseconds). **Default** `5000` | | `signal?` | `AbortSignal` | AbortSignal for cancellation. | *** ## SttSessionState ```ts type SttSessionState = | "idle" | "connecting" | "connected" | "finishing" | "finished" | "canceled" | "closed" | "error"; ``` Session lifecycle states. *** ## TemporaryApiKeyRequest ```ts type TemporaryApiKeyRequest = { client_reference_id?: string; expires_in_seconds: number; max_session_duration_seconds?: number; single_use?: boolean; usage_type: TemporaryApiKeyUsageType; }; ``` **Properties** | Property | Type | Description | | ---------------------------------------------------------------------------------------------- | ------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------- | | `client_reference_id?` | `string` | Optional tracking identifier string. Does not need to be unique **Max Length** 256 | | `expires_in_seconds` | `number` | Duration in seconds until the temporary API key expires **Minimum** 1 **Maximum** 3600 | | `max_session_duration_seconds?` | `number` | Maximum connection duration in seconds for WebSocket and TTS HTTP streaming endpoints. **Minimum** 1 **Maximum** 18000 | | `single_use?` | `boolean` | When true, restricts the temporary API key to a single use. | | `usage_type` | [`TemporaryApiKeyUsageType`](types#temporaryapikeyusagetype) | Intended usage of the temporary API key. | *** ## TemporaryApiKeyResponse ```ts type TemporaryApiKeyResponse = { api_key: string; expires_at: string; }; ``` **Properties** | Property | Type | Description | | ---------------------------------------------------------- | -------- | ------------------------------------------------------------------------------------------ | | `api_key` | `string` | Created temporary API key. | | `expires_at` | `string` | UTC timestamp indicating when generated temporary API key will expire **Format** date-time | *** ## TemporaryApiKeyUsageType ```ts type TemporaryApiKeyUsageType = "transcribe_websocket" | "tts_rt"; ``` *** ## TranscribeBaseOptions ```ts type TranscribeBaseOptions = { cleanup?: CleanupTarget[]; client_reference_id?: string; context?: TranscriptionContext; enable_language_identification?: boolean; enable_speaker_diarization?: boolean; fetch_transcript?: boolean; language_hints?: string[]; language_hints_strict?: boolean; model: string; signal?: AbortSignal; timeout_ms?: number; translation?: TranslationConfig; wait?: boolean; wait_options?: WaitOptions; webhook_auth_header_name?: string; webhook_auth_header_value?: string; webhook_query?: string | URLSearchParams | Record; webhook_url?: string; }; ``` Base options shared by all audio source variants. **Properties** | Property | Type | Description | | ------------------------------------------------------------------------------------------------- | -------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `cleanup?` | [`CleanupTarget`](types#cleanuptarget)\[] | Resources to clean up after transcription completes or on error/timeout. Only applies when `wait: true`. Cleanup runs in all cases when `wait: true`: - After successful completion - After transcription errors (status: 'error') - On timeout or abort This ensures no orphaned resources are left behind. **Example** `// Delete only the uploaded file cleanup: ['file'] // Delete only the transcription record cleanup: ['transcription'] // Delete both file and transcription cleanup: ['file', 'transcription']` | | `client_reference_id?` | `string` | Optional tracking identifier. **Max Length** 256 | | `context?` | [`TranscriptionContext`](types#transcriptioncontext) | Additional context to improve transcription accuracy and formatting of specialized terms. | | `enable_language_identification?` | `boolean` | Enable automatic language identification. | | `enable_speaker_diarization?` | `boolean` | Enable speaker diarization to identify different speakers. | | `fetch_transcript?` | `boolean` | When true (default), fetches the transcript and attaches it to the result when wait=true and the transcription completes successfully. Set to false to skip fetching the full transcript payload. **Default** `true` | | `language_hints?` | `string`\[] | Array of expected ISO language codes to bias recognition. | | `language_hints_strict?` | `boolean` | When true, model relies more heavily on language hints. | | `model` | `string` | Speech-to-text model to use. **Max Length** 32 | | `signal?` | `AbortSignal` | AbortSignal to cancel the operation | | `timeout_ms?` | `number` | Timeout in milliseconds | | `translation?` | [`TranslationConfig`](types#translationconfig) | Translation configuration. | | `wait?` | `boolean` | When true, waits for transcription to complete before returning. **Default** `false` | | `wait_options?` | [`WaitOptions`](types#waitoptions) | Options for waiting (only used when wait=true). | | `webhook_auth_header_name?` | `string` | Name of the authentication header sent with webhook notifications. **Max Length** 256 | | `webhook_auth_header_value?` | `string` | Authentication header value sent with webhook notifications. **Max Length** 256 | | `webhook_query?` | `string` \| `URLSearchParams` \| `Record`\<`string`, `string`> | Query parameters to append to the webhook URL. Useful for encoding metadata like transcription ID in the webhook callback. Can be a string, URLSearchParams, or Record\. | | `webhook_url?` | `string` | URL to receive webhook notifications when transcription is completed or fails. **Max Length** 256 | *** ## TranscribeFromFile ```ts type TranscribeFromFile = TranscribeBaseOptions & { audio_url?: never; file: UploadFileInput; file_id?: never; filename?: string; }; ``` Transcribe from a direct file upload (Buffer, Uint8Array, Blob, or ReadableStream) **Type Declaration** | Name | Type | Description | | ------------ | ------------------------------------------ | ----------------------------------- | | `audio_url?` | `never` | - | | `file` | [`UploadFileInput`](types#uploadfileinput) | File data to upload and transcribe. | | `file_id?` | `never` | - | | `filename?` | `string` | - | *** ## TranscribeFromFileId ```ts type TranscribeFromFileId = TranscribeBaseOptions & { audio_url?: never; file?: never; file_id: string; filename?: never; }; ``` Transcribe from a previously uploaded file **Type Declaration** | Name | Type | Description | | ------------ | -------- | ------------------------------------------------- | | `audio_url?` | `never` | - | | `file?` | `never` | - | | `file_id` | `string` | ID of a previously uploaded file. **Format** uuid | | `filename?` | `never` | - | *** ## TranscribeFromFileIdOptions ```ts type TranscribeFromFileIdOptions = Omit; ``` Options for transcribing from an uploaded file ID via `transcribeFromFileId`. *** ## TranscribeFromFileOptions ```ts type TranscribeFromFileOptions = Omit; ``` Options for transcribing from a file via `transcribeFromFile`. *** ## TranscribeFromUrl ```ts type TranscribeFromUrl = TranscribeBaseOptions & { audio_url: string; file?: never; file_id?: never; filename?: never; }; ``` Transcribe from a publicly accessible audio URL **Type Declaration** | Name | Type | Description | | ----------- | -------- | ------------------------------------------------------------ | | `audio_url` | `string` | URL of a publicly accessible audio file. **Max Length** 4096 | | `file?` | `never` | - | | `file_id?` | `never` | - | | `filename?` | `never` | - | *** ## TranscribeFromUrlOptions ```ts type TranscribeFromUrlOptions = Omit; ``` Options for transcribing from a URL via `transcribeFromUrl`. *** ## TranscribeOptions ```ts type TranscribeOptions = | TranscribeFromFile | TranscribeFromFileId | TranscribeFromUrl; ``` Options for the unified transcribe method Exactly one audio source must be provided: `file`, `file_id`, or `audio_url` *** ## TranscriptResponse ```ts type TranscriptResponse = { id: string; text: string; tokens: TranscriptToken[]; }; ``` Response from getting a transcription transcript. **Properties** | Property | Type | Description | | --------------------------------------------- | --------------------------------------------- | ---------------------------------------------------------------------------------- | | `id` | `string` | Unique identifier of the transcription this transcript belongs to. **Format** uuid | | `text` | `string` | Complete transcribed text content. | | `tokens` | [`TranscriptToken`](types#transcripttoken)\[] | List of detailed token information with timestamps and metadata. | *** ## TranscriptSegment ```ts type TranscriptSegment = { end_ms?: number; language?: string; speaker?: string; start_ms?: number; text: string; tokens: TranscriptToken[]; }; ``` A segment of contiguous tokens grouped by speaker and language **Properties** | Property | Type | Description | | ------------------------------------------------- | --------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------- | | `end_ms?` | `number` | End time of the segment in milliseconds (from last token). Absent for translation-only segments where the underlying tokens carry no timestamps. | | `language?` | `string` | Detected language code (if language identification was enabled). | | `speaker?` | `string` | Speaker identifier (if speaker diarization was enabled). | | `start_ms?` | `number` | Start time of the segment in milliseconds (from first token). Absent for translation-only segments where the underlying tokens carry no timestamps. | | `text` | `string` | Concatenated text of all tokens in this segment. | | `tokens` | [`TranscriptToken`](types#transcripttoken)\[] | Original tokens in this segment. | *** ## TranscriptToken ```ts type TranscriptToken = { confidence: number; end_ms?: number; is_audio_event?: boolean | null; language?: string | null; source_language?: string | null; speaker?: string | null; start_ms?: number; text: string; translation_status?: "none" | "original" | "translation" | null; }; ``` A single token from the transcript with timing and confidence information. **Properties** | Property | Type | Description | | ------------------------------------------------------------------- | ----------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `confidence` | `number` | Confidence score for this token (0.0 to 1.0). | | `end_ms?` | `number` | End time of the token in milliseconds. Present on original tokens (`translation_status` of `'original'` or `'none'`) and absent on translation tokens (`translation_status: 'translation'`), which do not carry timing. | | `is_audio_event?` | `boolean` \| `null` | Whether this token represents an audio event. | | `language?` | `string` \| `null` | Language code for this token. For original tokens (`translation_status` of `'original'` or `'none'`) this is the spoken language. For translation tokens (`translation_status: 'translation'`) this is the target language. Present on every token whenever language identification or translation is configured. | | `source_language?` | `string` \| `null` | Source language for translation tokens (`translation_status: 'translation'`). Identifies the language being translated from. Not set on original or `'none'` tokens; their language is in [TranscriptToken.language](#language). | | `speaker?` | `string` \| `null` | Speaker identifier (if speaker diarization was enabled). | | `start_ms?` | `number` | Start time of the token in milliseconds. Present on original tokens (`translation_status` of `'original'` or `'none'`) and absent on translation tokens (`translation_status: 'translation'`), which do not carry timing. | | `text` | `string` | The text content of this token. | | `translation_status?` | `"none"` \| `"original"` \| `"translation"` \| `null` | Translation status for this token. | *** ## TranscriptionContext ```ts type TranscriptionContext = { general?: ContextGeneralEntry[]; terms?: string[]; text?: string; translation_terms?: ContextTranslationTerm[]; }; ``` Additional context to improve transcription and translation accuracy. All sections are optional - include only what's relevant for your use case. **Properties** | Property | Type | Description | | ---------------------------------------------------------------------- | ----------------------------------------------------------- | --------------------------------------------------------------------------------------------------- | | `general?` | [`ContextGeneralEntry`](types#contextgeneralentry)\[] | Structured key-value pairs describing domain, topic, intent, participant names, etc. | | `terms?` | `string`\[] | Domain-specific or uncommon words to recognize. | | `text?` | `string` | Longer free-form background text, prior interaction history, reference documents, or meeting notes. | | `translation_terms?` | [`ContextTranslationTerm`](types#contexttranslationterm)\[] | Custom translations for ambiguous terms. | *** ## TranscriptionIdentifier ```ts type TranscriptionIdentifier = | string | { id: string; }; ``` Transcription identifier - either a string ID or an object with an id property. *** ## TranscriptionStatus ```ts type TranscriptionStatus = "queued" | "processing" | "completed" | "error"; ``` Status of a transcription request. *** ## TranscriptionsCountResponse ```ts type TranscriptionsCountResponse = { playground: number; public_api: number; total: number; }; ``` Total number of transcriptions, split by request scope. **Properties** | Property | Type | Description | | -------------------------------------------------------------- | -------- | ---------------------------------------------------- | | `playground` | `number` | Number of transcriptions created via the Playground. | | `public_api` | `number` | Number of transcriptions created via Public API. | | `total` | `number` | Total number of transcriptions across all scopes. | *** ## TranslateAudioSource ```ts type TranslateAudioSource = | { audio_url?: never; file: UploadFileInput; file_id?: never; filename?: string; } | { audio_url?: never; file?: never; file_id: string; filename?: never; } | { audio_url: string; file?: never; file_id?: never; filename?: never; }; ``` Audio source for [SonioxSttApi.translate](classes#sonioxsttapi-translate). Exactly one of `file`, `file_id`, or `audio_url` must be provided. **Type Declaration** ```ts { audio_url?: never; file: UploadFileInput; file_id?: never; filename?: string; } ``` | Name | Type | Description | | ------------ | ------------------------------------------ | ---------------------------------- | | `audio_url?` | `never` | - | | `file` | [`UploadFileInput`](types#uploadfileinput) | File data to upload and translate. | | `file_id?` | `never` | - | | `filename?` | `string` | - | ```ts { audio_url?: never; file?: never; file_id: string; filename?: never; } ``` | Name | Type | Description | | ------------ | -------- | ------------------------------------------------- | | `audio_url?` | `never` | - | | `file?` | `never` | - | | `file_id` | `string` | ID of a previously uploaded file. **Format** uuid | | `filename?` | `never` | - | ```ts { audio_url: string; file?: never; file_id?: never; filename?: never; } ``` | Name | Type | Description | | ----------- | -------- | ------------------------------------------------------------ | | `audio_url` | `string` | URL of a publicly accessible audio file. **Max Length** 4096 | | `file?` | `never` | - | | `file_id?` | `never` | - | | `filename?` | `never` | - | *** ## TranslateBaseOptions ```ts type TranslateBaseOptions = { cleanup?: CleanupTarget[]; client_reference_id?: string; context?: TranscriptionContext; enable_speaker_diarization?: boolean; fetch_translation?: boolean; model?: string; signal?: AbortSignal; timeout_ms?: number; wait?: boolean; wait_options?: WaitOptions; webhook_auth_header_name?: string; webhook_auth_header_value?: string; webhook_query?: string | URLSearchParams | Record; webhook_url?: string; }; ``` Common (non-mode, non-source) options shared by every translate call. **Properties** | Property | Type | Description | | ---------------------------------------------------------------------------------------- | -------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- | | `cleanup?` | [`CleanupTarget`](types#cleanuptarget)\[] | Resources to clean up after translation completes or on error/timeout. | | `client_reference_id?` | `string` | Optional tracking identifier. **Max Length** 256 | | `context?` | [`TranscriptionContext`](types#transcriptioncontext) | Additional context to improve transcription and translation accuracy. | | `enable_speaker_diarization?` | `boolean` | Enable speaker diarization to identify different speakers. | | `fetch_translation?` | `boolean` | When true (default), fetches and reshapes the translation result when `wait=true` and the job completes successfully. **Default** `true` | | `model?` | `string` | Speech-to-text model to use. **Default** `'stt-async-v5'` **Max Length** 32 | | `signal?` | `AbortSignal` | AbortSignal to cancel the operation. | | `timeout_ms?` | `number` | Timeout in milliseconds. | | `wait?` | `boolean` | When true, waits for translation to complete before returning. **Default** `false` | | `wait_options?` | [`WaitOptions`](types#waitoptions) | Options for waiting on completion. | | `webhook_auth_header_name?` | `string` | Name of the authentication header sent with webhook notifications. **Max Length** 256 | | `webhook_auth_header_value?` | `string` | Authentication header value sent with webhook notifications. **Max Length** 256 | | `webhook_query?` | `string` \| `URLSearchParams` \| `Record`\<`string`, `string`> | Query parameters to append to the webhook URL. | | `webhook_url?` | `string` | URL to receive webhook notifications when translation is completed or fails. **Max Length** 256 | *** ## TranslateFromTranscriptMode ```ts type TranslateFromTranscriptMode = | { from?: string; to: string; type: "one_way"; } | { language_a: string; language_b: string; type: "two_way"; }; ``` Mode parameter accepted by [translateFromTranscript](types#translatefromtranscript). The async `translate()` method stores this internally on the returned job; webhook handlers (and other callers that already have a transcript in hand) supply it directly. *** ## TranslateMode ```ts type TranslateMode = | { between?: never; from?: never; to: string; } | { between?: never; from: string; to: string; } | { between: [string, string]; from?: never; to?: never; }; ``` Shorthand specification of the translation direction(s) for [SonioxSttApi.translate](classes#sonioxsttapi-translate). Three mutually exclusive shapes: * `{ to }` — one-way translation into `to`. Source language(s) are detected automatically. * `{ to, from }` — one-way translation from `from` to `to`. The source language is hinted to the model. * `{ between: [a, b] }` — two-way translation between `a` and `b`. Each side is translated into the other; speech in any third language is passed through as-is. *** ## TranslateOptions ```ts type TranslateOptions = TranslateMode & TranslateAudioSource & TranslateBaseOptions; ``` Options for [SonioxSttApi.translate](classes#sonioxsttapi-translate). Combines a [TranslateMode](types#translatemode) (the translation direction shorthand), a [TranslateAudioSource](types#translateaudiosource) (file, file\_id, or audio\_url), and [TranslateBaseOptions](types#translatebaseoptions). *** ## TranslationConfig ```ts type TranslationConfig = | OneWayTranslationConfig | TwoWayTranslationConfig; ``` Translation configuration. *** ## TranslationSegment ```ts type TranslationSegment = { end_ms?: number; from: string; original_text: string; original_tokens: TranscriptToken[]; speaker?: string; start_ms?: number; to?: string; translation_text?: string; translation_tokens?: TranscriptToken[]; }; ``` A grouped pair of original speech and (optionally) its translation, derived from the underlying transcript tokens. In one-way mode every segment that originated from speech in the source language carries both `original_*` and `translation_*` fields. In two-way mode the same is true for the two configured languages; speech in a third language flows through with `translation_status: 'none'` and the translation fields are omitted. **Properties** | Property | Type | Description | | ---------------------------------------------------------------------- | --------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | | `end_ms?` | `number` | End time of the segment in milliseconds, taken from the last original token. Absent when the segment has no original tokens. | | `from` | `string` | Source language code. Derived from `original_tokens[0].language` when originals are present, otherwise from `translation_tokens[0].source_language`. | | `original_text` | `string` | Concatenated text of `original_tokens`. | | `original_tokens` | [`TranscriptToken`](types#transcripttoken)\[] | Original tokens (`translation_status` of `'original'` or `'none'`) for this segment, in order. | | `speaker?` | `string` | Speaker identifier (when speaker diarization is enabled). | | `start_ms?` | `number` | Start time of the segment in milliseconds, taken from the first original token. Absent when the segment has no original tokens. | | `to?` | `string` | Target language code. Omitted when there are no translation tokens (e.g. third-language pass-through under `between`). | | `translation_text?` | `string` | Concatenated text of `translation_tokens`. Omitted when there are no translation tokens. | | `translation_tokens?` | [`TranscriptToken`](types#transcripttoken)\[] | Translation tokens (`translation_status: 'translation'`) for this segment, in order. Omitted when there are no translation tokens. | *** ## TtsAudioFormat ```ts type TtsAudioFormat = | "pcm_f32le" | "pcm_s16le" | "pcm_s16be" | "pcm_mulaw" | "pcm_alaw" | "wav" | "aac" | "mp3" | "opus" | "flac" | string & { }; ``` Supported audio formats for Text-to-Speech output. *** ## TtsConnectionEvents ```ts type TtsConnectionEvents = { close: () => void; error: (error) => void; }; ``` Events emitted by a TTS WebSocket connection. **Properties** | Property | Type | Description | | -------------------------------------------- | ------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `close` | () => `void` | The WebSocket connection was closed. | | `error` | (`error`) => `void` | A connection-level error occurred. Always a [RealtimeError](classes#realtimeerror) subclass (e.g. [ConnectionError](classes#connectionerror), [NetworkError](classes#networkerror), [AuthError](classes#autherror)). | *** ## TtsConnectionOptions ```ts type TtsConnectionOptions = { connect_timeout_ms?: number; keepalive_interval_ms?: number; }; ``` Options for creating a TTS connection. **Properties** | Property | Type | Description | | ------------------------------------------------------------------------------ | -------- | --------------------------------------------------------------------------------------------- | | `connect_timeout_ms?` | `number` | Maximum time to wait for the WebSocket connection to open (milliseconds). **Default** `20000` | | `keepalive_interval_ms?` | `number` | Interval for sending keepalive messages (milliseconds). **Default** `5000` **Minimum** 1000 | *** ## TtsEvent ```ts type TtsEvent = { audio?: string; audio_end?: boolean; error_code?: number; error_message?: string; stream_id?: string; terminated?: boolean; timestamps?: TtsTimestamps; }; ``` Raw JSON event received from the TTS WebSocket server. **Properties** | Property | Type | | -------------------------------------------------- | --------------- | | `audio?` | `string` | | `audio_end?` | `boolean` | | `error_code?` | `number` | | `error_message?` | `string` | | `stream_id?` | `string` | | `terminated?` | `boolean` | | `timestamps?` | `TtsTimestamps` | *** ## TtsLanguage ```ts type TtsLanguage = { code: string; name: string; }; ``` A language supported by a Text-to-Speech model. **Properties** | Property | Type | Description | | ---------------------------------- | -------- | ----------------------------- | | `code` | `string` | ISO language code. | | `name` | `string` | Human-readable language name. | *** ## TtsModel ```ts type TtsModel = { aliased_model_id: string | null; id: string; languages: TtsLanguage[]; name: string; speed_max: number; speed_min: number; supports_silence_reduction: boolean; supports_speed_adjustment: boolean; supports_timestamps?: boolean; voices: TtsVoice[]; }; ``` A Text-to-Speech model. **Properties** | Property | Type | Description | | --------------------------------------------------------------------------- | ------------------------------------- | ---------------------------------------------------------------------------------------------- | | `aliased_model_id` | `string` \| `null` | If this is an alias, the id of the aliased model. Null for non-alias models. | | `id` | `string` | Unique identifier of the model. | | `languages` | [`TtsLanguage`](types#ttslanguage)\[] | Languages supported by this model. | | `name` | `string` | Name of the model. | | `speed_max` | `number` | Maximum supported speaking rate. | | `speed_min` | `number` | Minimum supported speaking rate. | | `supports_silence_reduction` | `boolean` | Whether the model supports shortening pauses between words via the `reduce_silence` parameter. | | `supports_speed_adjustment` | `boolean` | Whether the model supports adjusting the speaking rate via the `speed` parameter. | | `supports_timestamps?` | `boolean` | Whether the model can return character-level audio timestamps via `return_timestamps`. | | `voices` | [`TtsVoice`](types#ttsvoice)\[] | Voices supported by this model. | *** ## TtsStreamConfig ```ts type TtsStreamConfig = { audio_format: string; bitrate?: number; language: string; model: string; reduce_silence?: boolean; return_timestamps?: boolean; sample_rate?: number; speed?: number; stream_id: string; voice: string; }; ``` Fully resolved TTS stream config sent over the WebSocket. All required fields are present after merging input with defaults. **Properties** | Property | Type | | ----------------------------------------------------------------- | --------- | | `audio_format` | `string` | | `bitrate?` | `number` | | `language` | `string` | | `model` | `string` | | `reduce_silence?` | `boolean` | | `return_timestamps?` | `boolean` | | `sample_rate?` | `number` | | `speed?` | `number` | | `stream_id` | `string` | | `voice` | `string` | *** ## TtsStreamEvents ```ts type TtsStreamEvents = { audio: (chunk, timestamps?) => void; audioEnd: () => void; error: (error) => void; terminated: () => void; }; ``` Events emitted by a TTS stream. **Properties** | Property | Type | Description | | -------------------------------------------------- | ---------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `audio` | (`chunk`, `timestamps?`) => `void` | Decoded audio chunk received. When `return_timestamps` is enabled, the second argument carries the character-level alignment for this frame (it is `undefined` for audio-only frames). | | `audioEnd` | () => `void` | Server marked the final audio payload for this stream. | | `error` | (`error`) => `void` | A stream-level error occurred. Always a [RealtimeError](classes#realtimeerror) subclass mapped from the server `error_code` / `error_message`. | | `terminated` | () => `void` | Stream has been fully terminated by the server. | *** ## TtsStreamInput ```ts type TtsStreamInput = { audio_format?: TtsAudioFormat; bitrate?: number; language?: string; model?: string; reduce_silence?: boolean; return_timestamps?: boolean; sample_rate?: number; speed?: number; stream_id?: string; voice?: string; }; ``` Input for creating a TTS stream. All fields are optional and are merged with `tts_defaults` from the resolved connection config. After merging, `model`, `language`, `voice`, and `audio_format` must be present. **Properties** | Property | Type | Description | | ---------------------------------------------------------------- | ---------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `audio_format?` | [`TtsAudioFormat`](types#ttsaudioformat) | Output audio format **Example** `'wav'` | | `bitrate?` | `number` | Codec bitrate in bps (for compressed formats). | | `language?` | `string` | Language code for speech generation. **Example** `'en'` | | `model?` | `string` | Text-to-Speech model to use. **Example** `'tts-rt-v2'` | | `reduce_silence?` | `boolean` | Shorten pauses between words in the generated speech. `false` (default) keeps the model's natural pacing; `true` tightens delivery by reducing silence between words. | | `return_timestamps?` | `boolean` | Request character-level audio timestamps in the responses. When enabled, audio frames may carry a TtsTimestamps payload aligning each character of the spoken text to its start/end time in the audio. WebSocket (realtime) only — the REST endpoint streams raw audio bytes and ignores this flag. Timestamps map to the model's preprocessed text, not the raw input. Defaults to `false` when omitted. | | `sample_rate?` | `number` | Output sample rate in Hz. Required for raw PCM formats. | | `speed?` | `number` | Speaking rate. `1.0` is the normal rate; values below `1.0` slow speech down and values above `1.0` speed it up. Supported range is `0.7`-`1.3`. Defaults to `1.0` when omitted. | | `stream_id?` | `string` | Client-generated stream identifier. Must be unique among active streams on the same connection. Auto-generated if omitted. | | `voice?` | `string` | Voice identifier. **Example** `'Adrian'` | *** ## TtsStreamState ```ts type TtsStreamState = "active" | "finishing" | "ended" | "error"; ``` Lifecycle states for a TTS stream. *** ## TtsVoice ```ts type TtsVoice = { description: string; gender: TtsVoiceGender; id: string; }; ``` A Text-to-Speech voice. **Properties** | Property | Type | Description | | --------------------------------------------- | ---------------------------------------- | --------------------------------- | | `description` | `string` | Human-readable voice description. | | `gender` | [`TtsVoiceGender`](types#ttsvoicegender) | Voice gender metadata. | | `id` | `string` | Unique identifier of the voice. | *** ## TtsVoiceGender ```ts type TtsVoiceGender = "male" | "female" | "neutral"; ``` Voice gender metadata returned by the TTS models API. *** ## TwoWayTranslation ```ts type TwoWayTranslation = { duration_ms: number; language_a: string; language_b: string; mode: "two_way"; segments: TranslationSegment[]; }; ``` Result of a two-way translation (`{ between }` mode). No flat `original_text` / `translation_text` strings are exposed because which side is "original" depends on the segment. Read `segments` and filter / format per `from` / `to` as needed. **Properties** | Property | Type | Description | | ------------------------------------------------------ | --------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- | | `duration_ms` | `number` | Total audio duration in milliseconds. Equals the largest `end_ms` across all original tokens, or `0` when there are no original tokens. | | `language_a` | `string` | First configured language (the `between[0]` value). | | `language_b` | `string` | Second configured language (the `between[1]` value). | | `mode` | `"two_way"` | - | | `segments` | [`TranslationSegment`](types#translationsegment)\[] | Per-utterance segments in audio order. | *** ## TwoWayTranslationConfig ```ts type TwoWayTranslationConfig = { language_a: string; language_b: string; type: "two_way"; }; ``` Two-way translation configuration. Translates between two specified languages. **Properties** | Property | Type | Description | | ---------------------------------------------------------- | ----------- | --------------------- | | `language_a` | `string` | First language code. | | `language_b` | `string` | Second language code. | | `type` | `"two_way"` | Translation type. | *** ## UploadFileInput ```ts type UploadFileInput = | Buffer | Uint8Array | Blob | ReadableStream | NodeJS.ReadableStream; ``` Supported input types for file upload *** ## UploadFileOptions ```ts type UploadFileOptions = { client_reference_id?: string; filename?: string; signal?: AbortSignal; timeout_ms?: number; }; ``` Options for uploading a file **Properties** | Property | Type | Description | | ----------------------------------------------------------------------- | ------------- | ---------------------------------------------------------------------------------- | | `client_reference_id?` | `string` | Optional tracking identifier string. Does not need to be unique **Max Length** 256 | | `filename?` | `string` | Custom filename for the uploaded file | | `signal?` | `AbortSignal` | AbortSignal for cancelling the upload | | `timeout_ms?` | `number` | Request timeout in milliseconds | *** ## UsageLogsSort ```ts type UsageLogsSort = "end_time_asc" | "end_time_desc"; ``` Sort order for usage logs. *** ## UsageSummaryEntry ```ts type UsageSummaryEntry = { cost_usd: string[]; days: string[]; duration_cost_usd: string[]; duration_ms: number[]; input_audio_duration_ms: number[]; input_audio_tokens: number[]; input_cost_usd: string[]; input_text_tokens: number[]; model: string | null; num_requests: number[]; output_audio_duration_ms: number[]; output_audio_tokens: number[]; output_cost_usd: string[]; output_text_tokens: number[]; total_cost_usd: string; total_duration_cost_usd: string; total_duration_ms: number; total_input_audio_duration_ms: number; total_input_audio_tokens: number; total_input_cost_usd: string; total_input_text_tokens: number; total_num_requests: number; total_output_audio_duration_ms: number; total_output_audio_tokens: number; total_output_cost_usd: string; total_output_text_tokens: number; }; ``` Aggregated usage for a model (or the project total) over a time window. Per-day arrays are aligned to [UsageSummaryEntry.days](#days). **Properties** | Property | Type | Description | | -------------------------------------------------------------------------------------------- | ------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------- | | `cost_usd` | `string`\[] | Cost per day, in USD, aligned to `days`. | | `days` | `string`\[] | One UTC day (`YYYY-MM-DD`) per element, in ascending order. Every day in the requested window is present, including days with no usage. **Format** date | | `duration_cost_usd` | `string`\[] | Duration-billed cost per day, in USD, aligned to `days`. | | `duration_ms` | `number`\[] | Billed session duration per day, in milliseconds, aligned to `days`. | | `input_audio_duration_ms` | `number`\[] | - | | `input_audio_tokens` | `number`\[] | - | | `input_cost_usd` | `string`\[] | Cost of input tokens per day, in USD, aligned to `days`. | | `input_text_tokens` | `number`\[] | - | | `model` | `string` \| `null` | Model identifier. `null` on the [UsageSummaryResponse.total](types#usagesummaryresponse-total) entry. | | `num_requests` | `number`\[] | Number of requests per day, aligned to `days`. | | `output_audio_duration_ms` | `number`\[] | - | | `output_audio_tokens` | `number`\[] | - | | `output_cost_usd` | `string`\[] | Cost of output tokens per day, in USD, aligned to `days`. | | `output_text_tokens` | `number`\[] | - | | `total_cost_usd` | `string` | Total cost over the window, in USD. Equals `total_input_cost_usd` + `total_output_cost_usd` + `total_duration_cost_usd`. | | `total_duration_cost_usd` | `string` | Total cost over the window for models billed by session duration rather than by tokens, in USD. `0` for Speech-to-Text and Text-to-Speech models. | | `total_duration_ms` | `number` | Billed session duration over the window, in milliseconds, for models billed by duration. `0` for Speech-to-Text and Text-to-Speech models. | | `total_input_audio_duration_ms` | `number` | - | | `total_input_audio_tokens` | `number` | - | | `total_input_cost_usd` | `string` | Total cost of input tokens over the window, in USD. | | `total_input_text_tokens` | `number` | - | | `total_num_requests` | `number` | Number of requests over the window. | | `total_output_audio_duration_ms` | `number` | - | | `total_output_audio_tokens` | `number` | - | | `total_output_cost_usd` | `string` | Total cost of output tokens over the window, in USD. | | `total_output_text_tokens` | `number` | - | *** ## UsageSummaryResponse ```ts type UsageSummaryResponse = { models: UsageSummaryEntry[]; total: UsageSummaryEntry; }; ``` Aggregated usage summary for the authenticated project. **Properties** | Property | Type | Description | | ----------------------------------------------- | ------------------------------------------------- | ------------------------------------------------------------------------------------------- | | `models` | [`UsageSummaryEntry`](types#usagesummaryentry)\[] | One entry per model that recorded usage in the window. Empty when the project had no usage. | | `total` | [`UsageSummaryEntry`](types#usagesummaryentry) | Cost and activity across all models. Its `model` is `null`. | *** ## VoiceIdentifier ```ts type VoiceIdentifier = | string | { id: string; }; ``` Voice identifier - either a string ID or an object with an id property. *** ## VoiceModelStatus ```ts type VoiceModelStatus = "not_computed" | "processing" | "ready" | "failed"; ``` Processing status of a voice for a specific model. * `not_computed`: Not prepared for this model yet (e.g. the model was released after the voice was created). Call recompute to prepare it. * `processing`: Still being processed for this model. Wait and check again. * `ready`: Usable with this model. * `failed`: Processing failed permanently for this model. Fix the reference clip and create a new voice. *** ## VoiceModelStatusEntry ```ts type VoiceModelStatusEntry = { error_message?: string | null; error_type?: string | null; model: string; status: VoiceModelStatus; }; ``` Voice status for a single model. **Properties** | Property | Type | Description | | --------------------------------------------------------------- | -------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------- | | `error_message?` | `string` \| `null` | Human-readable error message when status is `failed` (e.g. the reference audio is too long). `null` otherwise. | | `error_type?` | `string` \| `null` | Machine-readable error category when status is `failed`. Stable across releases — safe to use in control flow. `null` otherwise. | | `model` | `string` | Name of the model. | | `status` | [`VoiceModelStatus`](types#voicemodelstatus) | Has to be `ready` for the voice to be usable with this model. | *** ## VoicesCountResponse ```ts type VoicesCountResponse = { total: number; }; ``` Total number of voices in your project. **Properties** | Property | Type | Description | | -------------------------------------------- | -------- | --------------------------------------- | | `total` | `number` | Total number of voices in your project. | *** ## WaitOptions ```ts type WaitOptions = { interval_ms?: number; on_status_change?: (status, transcription) => void; signal?: AbortSignal; timeout_ms?: number; }; ``` Options for polling/waiting for transcription completion. **Properties** | Property | Type | Description | | ----------------------------------------------------------- | ------------------------------------- | ---------------------------------------------------------------------- | | `interval_ms?` | `number` | Polling interval in milliseconds. **Default** `1000` **Minimum** 1000 | | `on_status_change?` | (`status`, `transcription`) => `void` | Callback invoked when status changes. | | `signal?` | `AbortSignal` | AbortSignal to cancel waiting. | | `timeout_ms?` | `number` | Maximum time to wait in milliseconds. **Default** `300000 (5 minutes)` | *** ## WebhookAuthConfig ```ts type WebhookAuthConfig = { name: string; value: string; }; ``` Authentication configuration for webhook verification **Properties** | Property | Type | Description | | ------------------------------------------ | -------- | -------------------------------------------------- | | `name` | `string` | Expected header name (case-insensitive comparison) | | `value` | `string` | Expected header value (exact match) | *** ## WebhookEvent ```ts type WebhookEvent = { id: string; status: WebhookEventStatus; }; ``` Webhook event payload sent by Soniox when a transcription completes or fails. **Properties** | Property | Type | Description | | --------------------------------------- | ------------------------------------------------ | -------------------------------- | | `id` | `string` | Transcription ID **Format** uuid | | `status` | [`WebhookEventStatus`](types#webhookeventstatus) | Transcription result status | *** ## WebhookEventStatus ```ts type WebhookEventStatus = "completed" | "error"; ``` Webhook event status values *** ## WebhookHandlerResult ```ts type WebhookHandlerResult = { error?: string; event?: WebhookEvent; ok: boolean; status: number; }; ``` Result of webhook handling **Properties** | Property | Type | Description | | ----------------------------------------------- | ------------------------------------ | ------------------------------------------------ | | `error?` | `string` | Error message (only present when ok=false) | | `event?` | [`WebhookEvent`](types#webhookevent) | Parsed webhook event (only present when ok=true) | | `ok` | `boolean` | Whether the webhook was handled successfully | | `status` | `number` | HTTP status code to return | *** ## WebhookHandlerResultWithFetch ```ts type WebhookHandlerResultWithFetch = WebhookHandlerResult & { fetchTranscript: | () => Promise | undefined; fetchTranscription: | () => Promise | undefined; }; ``` Result of webhook handling with lazy fetch capabilities. When using `client.webhooks.handleExpress()` (or other framework handlers), the result includes helper methods to fetch the transcript or transcription. **Type Declaration** | Name | Type | Description | | -------------------- | -------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `fetchTranscript` | \| () => `Promise`\<[`ISonioxTranscript`](types#isonioxtranscript) \| `null`> \| `undefined` | Fetch the transcript for a completed transcription. Only available when `ok=true` and `event.status='completed'`. **Example** `const result = soniox.webhooks.handleExpress(req); if (result.ok && result.event.status === 'completed') { const transcript = await result.fetchTranscript(); console.log(transcript?.text); }` | | `fetchTranscription` | \| () => `Promise`\<[`ISonioxTranscription`](types#isonioxtranscription) \| `null`> \| `undefined` | Fetch the full transcription object. Useful for both completed (metadata) and error (error details) statuses. **Example** `const result = soniox.webhooks.handleExpress(req); if (result.ok && result.event.status === 'error') { const transcription = await result.fetchTranscription(); console.log(transcription?.error_message); }` | *** ## WebhookHeaders ```ts type WebhookHeaders = | Headers | Record | { get: string | null; }; ``` Headers object type - supports both standard headers and record types *** ## HttpClient Pluggable HTTP client interface **Methods** **request()** ```ts request(request): Promise>; ``` Perform an HTTP request **Type Parameters** | Type Parameter | | -------------- | | `T` | **Parameters** | Parameter | Type | Description | | --------- | ---------------------------------- | --------------------- | | `request` | [`HttpRequest`](types#httprequest) | Request configuration | **Returns** `Promise`\<[`HttpResponse`](types#httpresponset)\<`T`>> Promise resolving to the response **Throws** [SonioxHttpError](classes#sonioxhttperror) On network errors, timeouts, HTTP errors, or parse errors *** ## HttpErrorDetails Error details for SonioxHttpError **Properties** | Property | Type | Description | | ---------------------------------------------------- | -------------------------------------- | ---------------------------------- | | `bodyText?` | `string` | Response body text (capped at 4KB) | | `cause?` | `unknown` | - | | `code` | [`HttpErrorCode`](types#httperrorcode) | - | | `headers?` | `Record`\<`string`, `string`> | - | | `message` | `string` | - | | `method` | [`HttpMethod`](types#httpmethod) | - | | `statusCode?` | `number` | - | | `url` | `string` | - | *** ## HttpRequest HTTP request configuration **Properties** | Property | Type | Description | | --------------------------------------------------- | -------------------------------------------- | ------------------------------------------------------------------------------------------------------ | | `body?` | [`HttpRequestBody`](types#httprequestbody) | Request body | | `headers?` | `Record`\<`string`, `string`> | Request headers | | `method` | [`HttpMethod`](types#httpmethod) | HTTP method | | `path` | `string` | URL path (relative to baseUrl) or absolute URL | | `query?` | [`QueryParams`](types#queryparams) | Query parameters (will be URL-encoded) | | `responseType?` | [`HttpResponseType`](types#httpresponsetype) | Expected response type **Default** `'json'` | | `signal?` | `AbortSignal` | Optional AbortSignal for request cancellation If provided along with timeoutMs, both will be respected | | `timeoutMs?` | `number` | Request timeout in milliseconds If not specified, uses the client's default timeout | *** ## HttpResponse\ HTTP response from the client **Type Parameters** | Type Parameter | | -------------- | | `T` | **Properties** | Property | Type | Description | | ------------------------------------------ | ----------------------------- | ----------------------------------------------- | | `data` | `T` | Parsed response data | | `headers` | `Record`\<`string`, `string`> | Response headers (normalized to lowercase keys) | | `status` | `number` | HTTP status code | *** ## ISonioxTranscript Type contract for SonioxTranscript class. **See** SonioxTranscript for full documentation. **Methods** **segments()** ```ts segments(options?): TranscriptSegment[]; ``` **Parameters** | Parameter | Type | | ---------- | ------------------------------------------------------------ | | `options?` | [`SegmentTranscriptOptions`](types#segmenttranscriptoptions) | **Returns** [`TranscriptSegment`](types#transcriptsegment)\[] **Properties** | Property | Type | | -------------------------------------------- | --------------------------------------------- | | `id` | `string` | | `text` | `string` | | `tokens` | [`TranscriptToken`](types#transcripttoken)\[] | *** ## ISonioxTranscription Type contract for SonioxTranscription class. **See** SonioxTranscription for full documentation. **Extended by** * [`ISonioxTranslationJob`](types#isonioxtranslationjob) **Methods** **delete()** ```ts delete(): Promise; ``` **Returns** `Promise`\<`void`> *** **destroy()** ```ts destroy(): Promise; ``` **Returns** `Promise`\<`void`> *** **getTranscript()** ```ts getTranscript(options?): Promise; ``` **Parameters** | Parameter | Type | | ----------------- | --------------------------------------------------- | | `options?` | \{ `force?`: `boolean`; `signal?`: `AbortSignal`; } | | `options.force?` | `boolean` | | `options.signal?` | `AbortSignal` | **Returns** `Promise`\<[`ISonioxTranscript`](types#isonioxtranscript) | `null`> *** **refresh()** ```ts refresh(signal?): Promise; ``` **Parameters** | Parameter | Type | | --------- | ------------- | | `signal?` | `AbortSignal` | **Returns** `Promise`\<`ISonioxTranscription`> *** **toJSON()** ```ts toJSON(): SonioxTranscriptionData; ``` **Returns** [`SonioxTranscriptionData`](types#sonioxtranscriptiondata) *** **wait()** ```ts wait(options?): Promise; ``` **Parameters** | Parameter | Type | | ---------- | ---------------------------------- | | `options?` | [`WaitOptions`](types#waitoptions) | **Returns** `Promise`\<`ISonioxTranscription`> **Properties** | Property | Type | | ----------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------- | | `audio_duration_ms` | `number` \| `null` \| `undefined` | | `audio_url` | `string` \| `null` \| `undefined` | | `client_reference_id` | `string` \| `null` \| `undefined` | | `context` | \| [`TranscriptionContext`](types#transcriptioncontext) \| `null` \| `undefined` | | `created_at` | `string` | | `enable_language_identification` | `boolean` | | `enable_speaker_diarization` | `boolean` | | `error_message` | `string` \| `null` \| `undefined` | | `error_type` | `string` \| `null` \| `undefined` | | `file_id` | `string` \| `null` \| `undefined` | | `filename` | `string` | | `id` | `string` | | `language_hints` | `string`\[] \| `undefined` | | `model` | `string` | | `status` | [`TranscriptionStatus`](types#transcriptionstatus) | | `transcript` | [`ISonioxTranscript`](types#isonioxtranscript) \| `null` \| `undefined` | | `webhook_auth_header_name` | `string` \| `null` \| `undefined` | | `webhook_auth_header_value` | `string` \| `null` \| `undefined` | | `webhook_status_code` | `number` \| `null` \| `undefined` | | `webhook_url` | `string` \| `null` \| `undefined` | *** ## ISonioxTranslationJob Type contract for SonioxTranslationJob class. **Extends** * [`ISonioxTranscription`](types#isonioxtranscription) **Methods** **delete()** ```ts delete(): Promise; ``` **Returns** `Promise`\<`void`> **Inherited from** [`ISonioxTranscription`](types#isonioxtranscription).[`delete`](types#isonioxtranscription-delete) *** **destroy()** ```ts destroy(): Promise; ``` **Returns** `Promise`\<`void`> **Inherited from** [`ISonioxTranscription`](types#isonioxtranscription).[`destroy`](types#isonioxtranscription-destroy) *** **fetchTranslation()** ```ts fetchTranslation(options?): Promise; ``` **Parameters** | Parameter | Type | | ----------------- | --------------------------------------------------- | | `options?` | \{ `force?`: `boolean`; `signal?`: `AbortSignal`; } | | `options.force?` | `boolean` | | `options.signal?` | `AbortSignal` | **Returns** `Promise`\<[`SonioxTranslation`](types#sonioxtranslation) | `null`> *** **getTranscript()** ```ts getTranscript(options?): Promise; ``` **Parameters** | Parameter | Type | | ----------------- | --------------------------------------------------- | | `options?` | \{ `force?`: `boolean`; `signal?`: `AbortSignal`; } | | `options.force?` | `boolean` | | `options.signal?` | `AbortSignal` | **Returns** `Promise`\<[`ISonioxTranscript`](types#isonioxtranscript) | `null`> **Inherited from** [`ISonioxTranscription`](types#isonioxtranscription).[`getTranscript`](types#isonioxtranscription-gettranscript) *** **getTranslation()** ```ts getTranslation(options?): Promise; ``` **Parameters** | Parameter | Type | | ----------------- | --------------------------------------------------- | | `options?` | \{ `force?`: `boolean`; `signal?`: `AbortSignal`; } | | `options.force?` | `boolean` | | `options.signal?` | `AbortSignal` | **Returns** `Promise`\<[`SonioxTranslation`](types#sonioxtranslation) | `null`> *** **refresh()** ```ts refresh(signal?): Promise; ``` **Parameters** | Parameter | Type | | --------- | ------------- | | `signal?` | `AbortSignal` | **Returns** `Promise`\<`ISonioxTranslationJob`> **Overrides** [`ISonioxTranscription`](types#isonioxtranscription).[`refresh`](types#isonioxtranscription-refresh) *** **toJSON()** ```ts toJSON(): SonioxTranscriptionData; ``` **Returns** [`SonioxTranscriptionData`](types#sonioxtranscriptiondata) **Overrides** [`ISonioxTranscription`](types#isonioxtranscription).[`toJSON`](types#isonioxtranscription-tojson) *** **wait()** ```ts wait(options?): Promise; ``` **Parameters** | Parameter | Type | | ---------- | ---------------------------------- | | `options?` | [`WaitOptions`](types#waitoptions) | **Returns** `Promise`\<`ISonioxTranslationJob`> **Overrides** [`ISonioxTranscription`](types#isonioxtranscription).[`wait`](types#isonioxtranscription-wait) **Properties** | Property | Type | | ------------------------------------------------------------------------------------------------ | -------------------------------------------------------------------------------- | | `audio_duration_ms` | `number` \| `null` \| `undefined` | | `audio_url` | `string` \| `null` \| `undefined` | | `client_reference_id` | `string` \| `null` \| `undefined` | | `context` | \| [`TranscriptionContext`](types#transcriptioncontext) \| `null` \| `undefined` | | `created_at` | `string` | | `enable_language_identification` | `boolean` | | `enable_speaker_diarization` | `boolean` | | `error_message` | `string` \| `null` \| `undefined` | | `error_type` | `string` \| `null` \| `undefined` | | `file_id` | `string` \| `null` \| `undefined` | | `filename` | `string` | | `id` | `string` | | `language_hints` | `string`\[] \| `undefined` | | `model` | `string` | | `status` | [`TranscriptionStatus`](types#transcriptionstatus) | | `transcript` | [`ISonioxTranscript`](types#isonioxtranscript) \| `null` \| `undefined` | | `translation` | [`SonioxTranslation`](types#sonioxtranslation) \| `null` \| `undefined` | | `webhook_auth_header_name` | `string` \| `null` \| `undefined` | | `webhook_auth_header_value` | `string` \| `null` \| `undefined` | | `webhook_status_code` | `number` \| `null` \| `undefined` | | `webhook_url` | `string` \| `null` \| `undefined` | *** ## translateFromTranscript() ```ts function translateFromTranscript(transcript, mode): SonioxTranslation; ``` Reshape a transcript produced by a translation-enabled transcription into a structured [SonioxTranslation](types#sonioxtranslation) result. This is the same logic `SonioxTranslationJob.getTranslation()` applies. Use it directly in webhook handlers or anywhere else you already have a transcript in hand. **Parameters** | Parameter | Type | Description | | ------------ | ------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------- | | `transcript` | `TranscriptLike` | Transcript (or any object with a `tokens` array) emitted for a translation-enabled transcription. | | `mode` | [`TranslateFromTranscriptMode`](types#translatefromtranscriptmode) | Whether to reshape as one-way or two-way; the discriminator tells the helper which result shape to produce. | **Returns** [`SonioxTranslation`](types#sonioxtranslation) A [SonioxTranslation](types#sonioxtranslation) keyed on `mode`. **Example** ```typescript import { translateFromTranscript } from '@soniox/node'; // From a webhook handler that just received the transcript const result = translateFromTranscript(transcript, { type: 'one_way', to: 'es' }); console.log(result.translation_text); ``` # Async transcription with Node SDK URL: /sdk/node-SDK/stt/async-transcription Transcribe audio files asynchronously with the Soniox Node SDK Soniox Node SDK supports asynchronous transcription for audio files. This allows you to transcribe recordings without maintaining a live connection or streaming pipeline. You can either wait for completion or create a job and retrieve the results based on the webhook event. ## Quickstart SDK provides you a convenient method to transcribe audio from a local file, public URL, or previously uploaded file. The **[`transcribe`](/sdk/node-SDK/reference/classes#sonioxsttapi-transcribe) method will:** 1. Upload the file to Soniox if it's not already uploaded (if `file` is provided) 2. Transcribe the audio 3. Await for the transcription to complete (if `wait: true` is provided) 4. Return the transcription object and final transcript (you can disable this by setting `fetch_transcript: false` and fetch transcript later using [`getTranscript`](/sdk/node-SDK/reference/classes#sonioxsttapi-gettranscript) method) 5. Delete the file from Soniox if it was uploaded (configurable using `cleanup` option) Don't forget to remove files and transcriptions from Soniox after you're done with them if `cleanup` option is not set. **Transcribe from a local file and delete everything after transcription is complete** ```ts import { readFile } from 'node:fs/promises'; const audio = await readFile('audio.mp3'); // Buffer | Uint8Array | Blob | ReadableStream const transcription = await client.stt.transcribe({ model: 'stt-async-v5', file: audio, filename: 'audio.mp3', wait: true, cleanup: ['file', 'transcription'], }); ``` **Transcribe from a public URL and fetch the transcript later using [`getTranscript`](/sdk/node-SDK/reference/classes#sonioxsttapi-gettranscript) method** ```ts const transcription = await client.stt.transcribe({ model: 'stt-async-v5', audio_url: 'https://soniox.com/media/examples/coffee_shop.mp3', wait: true, }); const transcript = await transcription.getTranscript(); ``` **Transcribe from a previously uploaded file and setup a [webhook](/sdk/node-SDK/stt/webhooks) to get the transcription when it's complete** ```ts const transcription = await client.stt.transcribe({ model: 'stt-async-v5', file_id: file.id, wait: false, webhook_url: 'https://example.com/webhook', }); ``` Learn more about [testing webhooks locally](/sdk/node-SDK/stt/webhooks#testing-webhooks-locally). ## Retrieve list of transcriptions You can retrieve a list of transcriptions using [`list`](/sdk/node-SDK/reference/classes#sonioxsttapi-list) method. ```ts const transcriptions = await client.stt.list({ limit: 100, // page size — pagination continues automatically when iterated }); ``` The returned result is async iterable - use `for await...of` to iterate through all pages. ```ts for await (const transcription of transcriptions) { console.log(transcription.id, transcription.status); } ``` ## Count transcriptions Use [`count`](/sdk/node-SDK/reference/classes#sonioxsttapi-count) to retrieve the total number of transcriptions. ```ts const counts = await client.stt.count(); console.log(counts.public_api); console.log(counts.playground); console.log(counts.total); ``` ## Get transcription You can get a transcription by ID using [`get`](/sdk/node-SDK/reference/classes#sonioxsttapi-get) method. ```ts const transcription = await client.stt.get('transcription-id'); console.log(transcription.id, transcription.status); ``` ## Get transcription transcript You can get a transcription transcript using [`getTranscript`](/sdk/node-SDK/reference/classes#sonioxsttapi-gettranscript) method. ```ts const transcript = await transcription.getTranscript(); console.log(transcript.text); ``` Or get transcript by transcription ID. ```ts const transcript = await client.stt.getTranscript(transcription.id); console.log(transcript.text); ``` ## Segmenting transcripts Group tokens by speaker and language changes: ```ts const transcript = await transcription.getTranscript(); for (const segment of transcript?.segments() ?? []) { console.log(`[${segment.speaker}][${segment.language}] ${segment.text}`); } ``` ## Delete or destroy transcription You can delete or destroy a transcription using [`delete`](/sdk/node-SDK/reference/classes#sonioxsttapi-delete) or [`destroy`](/sdk/node-SDK/reference/classes#sonioxsttapi-destroy) method. **Delete transcription only** ```ts await client.stt.delete(transcription.id); ``` **Delete transcription and its file if it was uploaded** ```ts await client.stt.destroy(transcription.id); ``` ## Delete all transcriptions and files from your account ### Delete all transcriptions You can delete all transcriptions using [`stt.delete_all`](/sdk/node-SDK/reference/classes#sonioxsttapi-delete_all) method. ```ts await client.stt.delete_all(); ``` ### Delete all transcriptions and their files You can delete all transcriptions and their files using [`stt.destroy_all`](/sdk/node-SDK/reference/classes#sonioxsttapi-destroy_all) method. ```ts await client.stt.destroy_all(); ``` `delete_all` and `destroy_all` operations are irreversible and cannot be undone. # Async translation with Node SDK URL: /sdk/node-SDK/stt/async-translation Translate audio asynchronously with the Soniox Node SDK Soniox Node SDK supports asynchronous speech translation through [`client.stt.translate`](/sdk/node-SDK/reference/classes#sonioxsttapi-translate). It creates an async transcription job with translation enabled, then reshapes the completed transcript into a structured translation result. Use async translation when you want to translate uploaded files, public audio URLs, or previously uploaded files without maintaining a live connection. This method is just a syntactic sugar on top of [`client.stt.transcribe`](/sdk/node-SDK/stt/async-transcription). Use `transcribe` if you need more low level control on the output. ## Quickstart For simple one-way translation, provide `to` and one audio source: ```ts const job = await client.stt.translate({ audio_url: 'https://soniox.com/media/examples/coffee_shop.mp3', to: 'es', wait: true, }); console.log(job.translation?.translation_text); ``` ## Translation modes ### One-way translation Use `to` to translate any detected source language to a target language. ```ts const job = await client.stt.translate({ file: audio, filename: 'meeting.mp3', to: 'es', wait: true, }); const translation = job.translation; console.log(translation?.original_text); console.log(translation?.translation_text); ``` If you know the source language, provide `from`. The SDK forwards it as a language hint for the underlying async transcription job. ```ts const job = await client.stt.translate({ audio_url: 'https://soniox.com/media/examples/coffee_shop.mp3', from: 'en', to: 'es', wait: true, }); ``` ### Two-way translation Use `between` for [bidirectional translation](/translation/stt-translation#two-way-translation) between two languages. This is useful for conversations where either language can appear in the audio. ```ts const job = await client.stt.translate({ file: audio, filename: 'conversation.mp3', between: ['en', 'es'], enable_speaker_diarization: true, wait: true, }); for (const segment of job.translation?.segments ?? []) { console.log(`[${segment.from}] ${segment.original_text}`); if (segment.translation_text) { console.log(` -> [${segment.to}] ${segment.translation_text}`); } } ``` ## Wait later If you do not want to block until the job completes, omit `wait` and call [`wait`](/sdk/node-SDK/reference/classes#sonioxtranslationjob-wait) or [`fetchTranslation`](/sdk/node-SDK/reference/classes#sonioxtranslationjob-fetchtranslation) later. ```ts const job = await client.stt.translate({ audio_url: 'https://soniox.com/media/examples/coffee_shop.mp3', to: 'es', }); const completed = await job.wait(); const translation = await completed.fetchTranslation(); console.log(translation?.translation_text); ``` ## Use with webhooks Async translation jobs are regular async transcription jobs with translation enabled, so you can use the same [webhooks](/sdk/node-SDK/stt/webhooks) flow. ```ts const job = await client.stt.translate({ file_id: file.id, to: 'es', webhook_url: 'https://example.com/webhook', }); ``` In the webhook handler, fetch the transcript and reshape it with [`translateFromTranscript`](/sdk/node-SDK/reference/types#translatefromtranscript). ```ts import { translateFromTranscript } from '@soniox/node'; const transcript = await client.stt.getTranscript(event.id); if (transcript) { const translation = translateFromTranscript(transcript, { type: 'one_way', to: 'es', }); console.log(translation.translation_text); } ``` ## Fetch transcript instead of translation `translate()` also keeps access to the underlying transcript. This is useful when you need both the original token stream and the reshaped translation. ```ts const job = await client.stt.translate({ file: audio, filename: 'meeting.mp3', to: 'es', wait: true, }); const transcript = await job.getTranscript(); const translation = await job.getTranslation(); ``` Translation tokens do not carry timestamps, so `TranscriptToken.start_ms`, `TranscriptToken.end_ms`, and matching segment timestamps can be `undefined` for translation-only tokens. # Handling files with Node SDK URL: /sdk/node-SDK/stt/files Upload audio files and manage them with the Soniox Node SDK Node SDK provides helpers to work with the [Files API](/api-reference/stt/files/upload_file) to upload audio for async transcription or to reuse files across multiple jobs. ## Upload [`upload()`](/sdk/node-SDK/reference/classes#sonioxfilesapi-upload) accepts `Buffer`, `Uint8Array`, `Blob`, `ReadableStream`. ```ts import { readFile } from 'node:fs/promises'; const audio = await readFile('audio.mp3'); const file = await client.files.upload(audio, { filename: 'audio.mp3', client_reference_id: 'meeting-42', }); console.log(file.id, file.filename, file.size); ``` Read more about [Supported audio formats](/stt/async/async-transcription#audio-formats). ## List files [`list()`](/sdk/node-SDK/reference/classes#sonioxfilesapi-list) returns a paginated list of all uploaded files. Use `for await...of` to iterate through all pages. ```ts const result = await client.files.list({ limit: 100 }); // Automatic pagination for await (const file of result) { console.log(file.id, file.filename); } ``` ## Count files [`count()`](/sdk/node-SDK/reference/classes#sonioxfilesapi-count) returns the total number of uploaded files. ```ts const counts = await client.files.count(); console.log(counts.public_api); console.log(counts.playground); console.log(counts.total); ``` ## Get file Get a file by ID using [`get()`](/sdk/node-SDK/reference/classes#sonioxfilesapi-get) method: ```ts const file = await client.files.get('file-id'); ``` ## Delete file Delete file via instance using [`file.delete()`](/sdk/node-SDK/reference/classes#sonioxfilesapi-delete) method: ```ts const file = await client.files.get('file-id'); await file.delete(); ``` Or delete by ID using [`delete()`](/sdk/node-SDK/reference/classes#sonioxfilesapi-delete) method: ```ts await client.files.delete('file-id'); ``` ## Delete all files from your account You can delete all files using `files.delete_all` method. ```ts await client.files.delete_all(); ``` `delete_all` operation is irreversible and cannot be undone. # Real-time transcription with Node SDK URL: /sdk/node-SDK/stt/realtime-transcription Create and manage real-time speech-to-text sessions with the Soniox Node SDK Soniox Node SDK supports real-time streaming transcription over WebSocket. This allows you to transcribe live audio with low latency — ideal for voice agents, live captions, and interactive experiences. You can consume results via events, async iteration, or buffers that group tokens into utterances. SDK provides you helper methods to work both with direct and proxy streaming. ## Direct stream and temporary API keys Read more about [Direct stream](/guides/direct-stream) Node SDK provides you a helper method to issue [temporary API Keys](/api-reference/auth/create_temporary_api_key) to use with [Direct stream](/guides/direct-stream) from the client's browser. `client_reference_id` is optional — when set, every request authenticated with the key is recorded under that identifier in [usage logs](/guides/usage-logs). ```ts const { api_key, expires_at } = await client.auth.createTemporaryKey({ usage_type: "transcribe_websocket", expires_in_seconds: 3600, client_reference_id: "support-call-123", }); console.log(api_key, expires_at); ``` Soniox's [Web Library](/sdk/web-SDK) handles everything client-side — capturing microphone input, managing the WebSocket connection, and authenticating using temporary API keys. ## Proxy stream helpers Read more about [Proxy stream](/guides/proxy-stream) Use the SDK's real-time session for low-latency transcription, live captions, and voice agent experiences. ## Create a real-time session ```ts const session = client.realtime.stt({ model: "stt-rt-v5", audio_format: "pcm_s16le", sample_rate: 16000, num_channels: 1, enable_endpoint_detection: true, enable_speaker_diarization: true, language_hints: ["en"], context: { text: "Support call about billing", terms: ["invoice", "refund"], }, }); ``` ## Connect and stream Use [`sendAudio`](/sdk/node-SDK/reference/classes#realtimesttsession-sendaudio) to send audio chunks to the session. ```ts // audioStream: AsyncIterable // e.g. a Node fs ReadStream, fetch response body, or mic-capture iterable. await session.connect(); session.on("result", (result) => { process.stdout.write(result.tokens.map((t) => t.text).join("")); }); for await (const chunk of audioStream) { session.sendAudio(chunk); } await session.finish(); ``` See the full example with a demo stream in the quickstart: [Create your first real-time session](/sdk/node-SDK#create-your-first-real-time-session) ## Handle session events ```ts session.on("connected", () => console.log("connected")); session.on("disconnected", (reason) => console.log("disconnected:", reason)); session.on("error", (error) => console.error("error:", error)); session.on("result", (result) => console.log(result.tokens.map((t) => t.text).join("")), ); session.on("endpoint", () => console.log("endpoint")); // server-detected end of utterance session.on("finalized", () => console.log("finalized")); // fires after a manual `finalize()` request session.on("finished", () => console.log("finished")); // fires after `finish()` completes ``` ## Session lifecycle ```ts // Connect to the session await session.connect(); // idle -> connected // Send audio chunks to the session for await (const chunk of audioStream) { session.sendAudio(chunk); } // Gracefully end the session (Signal end of audio and wait for remaining results from the server) await session.finish(); // Or cancel immediately without waiting for final results: session.close(); ``` ## Endpoint detection and manual finalization Endpoint detection lets you know when a speaker has finished speaking. This is critical for real-time voice AI assistants, command-and-response systems, and conversational apps where you want to respond immediately without waiting for long silences. Read more about [Endpoint detection](/stt/rt/endpoint-detection) Enable endpoint detection by setting `enable_endpoint_detection: true` in the session configuration. ```ts const session = client.realtime.stt({ model: "stt-rt-v5", enable_endpoint_detection: true, }); ``` Manual finalization gives you precise control over when audio should be finalized — useful for Push-to-talk systems and client-side voice activity detection (VAD). Read more about [Manual finalization](/stt/rt/manual-finalization) ```ts session.finalize(); ``` ## Pause and resume ```ts session.pause(); // keeps connection alive, drops audio while paused session.resume(); // resume sending audio ``` You are billed for the full stream duration even when session is paused. In a typical voice agent loop, you pause the STT session while the agent is responding to avoid transcribing the agent's own audio or processing overlapping speech: ```ts session.on("endpoint", async () => { const utterance = utteranceBuffer.markEndpoint(); // Read more about utterance buffer below if (!utterance) return; // Pause STT while the agent processes and responds session.pause(); const response = await myAgent.respond(utterance.text); // ... send response audio to the client ... // Resume listening for the next utterance session.resume(); }); ``` SDK will `finalize` audio on pause. Make sure to adjust your VAD sensitivity to have enough silence before pause. Learn more about [Manual finalization](/stt/rt/manual-finalization#key-points) ## Keepalive Read more about [Connection keepalive](/stt/rt/connection-keepalive) While the session is paused via [`session.pause()`](/sdk/node-SDK/reference/classes#realtimesttsession-pause), the Node SDK **automatically sends** a keepalive every `keepalive_interval_ms` (default `5000` ms) to hold the WebSocket open. A manual call is only needed if you want to force a keepalive outside of that cadence: ```ts session.keepAlive(); ``` ## Detecting utterance for voice agents When building voice AI agents, you need to know when the user has finished speaking so you can process their input. The SDK provides [`RealtimeUtteranceBuffer`](/sdk/node-SDK/reference/classes#realtimeutterancebuffer) to collect streaming tokens into complete utterances, driven by the server's endpoint detection. ### How it works 1. Set `enable_endpoint_detection: true` in the session config – the server detects when the user stops speaking and emits an endpoint event. 2. Feed every result event into the buffer with [`addResult()`](/sdk/node-SDK/reference/classes#realtimeutterancebuffer-addresult). 3. When an endpoint fires, call [`markEndpoint()`](/sdk/node-SDK/reference/classes#realtimeutterancebuffer-markendpoint) to flush the buffer and get the complete utterance. ### Example ```ts import { SonioxNodeClient, RealtimeUtteranceBuffer } from "@soniox/node"; const client = new SonioxNodeClient(); // Call this for each new user/connection - each session needs its own buffer function createAgentSession(onUtterance: (text: string) => void) { const session = client.realtime.stt({ model: "stt-rt-v5", enable_endpoint_detection: true, }); // Each session gets its own buffer const utteranceBuffer = new RealtimeUtteranceBuffer({ final_only: true, }); session.on("result", (result) => { utteranceBuffer.addResult(result); }); session.on("endpoint", () => { const utterance = utteranceBuffer.markEndpoint(); if (utterance) { onUtterance(utterance.text); } }); return session; } // Usage: create a session per user connection const session = createAgentSession((text) => { console.log("User said:", text); // Pass to your LLM / agent pipeline }); await session.connect(); session.sendAudio(audioChunk); ``` ## Streaming audio from a file Use [`sendStream()`](/sdk/node-SDK/reference/classes#realtimesttsession-sendstream) to pipe audio directly from a file (or any async source) into a real-time session. It accepts any `AsyncIterable` – Node.js file streams, Web `ReadableStream`, Bun file streams, fetch response bodies, or custom async generators. ```ts import { createReadStream } from "node:fs"; await session.connect(); await session.sendStream(createReadStream("audio.mp3"), { pace_ms: 60, // Throttle to simulate real-time pace; omit for live audio finish: true, }); ``` ### Simulating real-time pace When streaming pre-recorded files, you can throttle sending with `pace_ms` to simulate how audio would arrive from a live source (e.g. a microphone). This isn't needed for live audio – it naturally arrives at real-time pace. Use [`sendAudio`](/sdk/node-SDK/stt/realtime-transcription#connect-and-stream) if you need more control. # Handling webhooks with Node SDK URL: /sdk/node-SDK/stt/webhooks Use webhooks to receive transcription results with the Soniox Node SDK SDK provides you a helper method to handle [Webhooks](/stt/async/webhooks) from the Soniox API and transform them into a typed object. ## Configure webhook delivery If webhook is enabled during [transcription creation](/sdk/node-SDK/stt/async-transcription#quickstart), Soniox will send a POST request to your webhook URL with the transcription result. ```ts await client.stt.transcribe({ model: 'stt-async-v5', audio_url: 'https://soniox.com/media/examples/coffee_shop.mp3', webhook_url: 'https://your-server.com/webhooks/soniox', webhook_auth_header_name: 'X-Webhook-Secret', webhook_auth_header_value: process.env.SONIOX_API_WEBHOOK_SECRET, }); ``` You can also append metadata as query parameters: ```ts await client.stt.transcribe({ model: 'stt-async-v5', audio_url: 'https://soniox.com/media/examples/coffee_shop.mp3', webhook_url: 'https://your-server.com/webhooks/soniox', webhook_query: { request_id: 'abc-123' }, }); ``` Learn more about [testing webhooks locally](/sdk/node-SDK/stt/webhooks#testing-webhooks-locally). ## Handling webhooks The SDK provides both framework-agnostic and framework-specific handlers that parse the request body, verify authentication, and return a typed [`WebhookHandlerResultWithFetch`](/sdk/node-SDK/reference/types#webhookhandlerresultwithfetch). All handlers return: * `ok` — whether the webhook was handled successfully * `status` — HTTP status code to return to Soniox * `event` — the parsed [`WebhookEvent`](/sdk/node-SDK/reference/types#webhookevent) (when `ok=true`) * `error` — error message (when `ok=false`) * `fetchTranscript()` — lazily fetch the full transcript (when `event.status === 'completed'`) * `fetchTranscription()` — lazily fetch the transcription object All handler snippets below assume you already have an initialized `client`, for example `const client = new SonioxNodeClient();`. ```ts import express from 'express'; const app = express(); app.use(express.json()); app.post('/webhooks/soniox', async (req, res) => { const result = client.webhooks.handleExpress(req); if (result.ok && result.event.status === 'completed') { const transcript = await result.fetchTranscript(); console.log(transcript?.text); } res.status(result.status).json({ received: true }); }); ``` ```ts import Fastify from 'fastify'; const app = Fastify(); app.post('/webhooks/soniox', async (request, reply) => { const result = client.webhooks.handleFastify(request); if (result.ok && result.event.status === 'completed') { const transcript = await result.fetchTranscript(); console.log(transcript?.text); } return reply.status(result.status).send({ received: true }); }); ``` ```ts import { Hono } from 'hono'; const app = new Hono(); app.post('/webhooks/soniox', async (c) => { const result = await client.webhooks.handleHono(c); if (result.ok && result.event.status === 'completed') { const transcript = await result.fetchTranscript(); console.log(transcript?.text); } return c.json({ received: true }, result.status); }); ``` `handleHono` is async because it reads the request body from the Hono context. ```ts import { Controller, Post, Req, Res } from '@nestjs/common'; import { Request, Response } from 'express'; @Controller('webhooks') export class WebhooksController { @Post('soniox') async handleSoniox(@Req() req: Request, @Res() res: Response) { const result = client.webhooks.handleNestJS(req); if (result.ok && result.event.status === 'completed') { const transcript = await result.fetchTranscript(); console.log(transcript?.text); } res.status(result.status).json({ received: true }); } } ``` Use `handleRequest` with any framework that provides a standard Fetch API `Request` object: ```ts export default { async fetch(request: Request) { if (new URL(request.url).pathname === '/webhooks/soniox') { const result = await client.webhooks.handleRequest(request); if (result.ok && result.event.status === 'completed') { const transcript = await result.fetchTranscript(); console.log(transcript?.text); } return Response.json({ received: true }, { status: result.status }); } return new Response('Not found', { status: 404 }); }, }; ``` `handleRequest` is async because it reads the request body from the `Request` object. The `handle` method is a framework-agnostic handler. You provide the method, headers, and parsed body directly: ```ts const result = client.webhooks.handle({ method: req.method, headers: req.headers, body: req.body, }); if (result.ok && result.event.status === 'completed') { const transcript = await result.fetchTranscript(); console.log(transcript?.text); } if (result.ok && result.event.status === 'error') { const transcription = await result.fetchTranscription(); console.log(transcription?.error_message); } ``` See [`HandleWebhookOptions`](/sdk/node-SDK/reference/types#handlewebhookoptions) for all available options. ## Webhook auth helpers By default, webhook handlers read auth from `SONIOX_API_WEBHOOK_HEADER` and `SONIOX_API_WEBHOOK_SECRET`. You can override auth explicitly: ```ts const result = client.webhooks.handleExpress(req, { name: 'X-Webhook-Secret', value: process.env.SONIOX_API_WEBHOOK_SECRET, }); ``` Learn more info about [Environment Variables](/sdk/node-SDK/reference#environment-variables). But you can also verify the auth manually: ```ts const auth = client.webhooks.getAuthFromEnv(); if (!auth) { throw new Error('Missing webhook auth'); } const isValid = client.webhooks.verifyAuth(req.headers, auth); ``` ## Webhook event helpers ```ts const event = client.webhooks.parseEvent(req.body); const isValid = client.webhooks.isEvent(req.body); ``` ## Testing webhooks locally Since Soniox needs to reach your server over the internet, you'll need a tunnel to expose your local development server. You can use [Cloudflare Tunnel](https://developers.cloudflare.com/pages/how-to/preview-with-cloudflare-tunnel/) or [ngrok](https://ngrok.com/). [Cloudflare Tunnel](https://developers.cloudflare.com/pages/how-to/preview-with-cloudflare-tunnel/) provides a quick way to expose your local server — no account required. Install `cloudflared` and start a tunnel pointing to your local server: ```bash # macOS brew install cloudflared # Start a tunnel to your local server on port 3000 cloudflared tunnel --url http://localhost:3000 ``` The command will output a public URL like `https://random-name.trycloudflare.com`. [ngrok](https://ngrok.com/) creates a secure tunnel to your local server and provides a stable public URL. Install ngrok, authenticate, and start a tunnel: ```bash # macOS brew install ngrok # Authenticate (one-time setup) ngrok config add-authtoken # Start a tunnel to your local server on port 3000 ngrok http 3000 ``` The command will output a public URL like `https://abcd-1234.ngrok-free.app`. Once you have your public tunnel URL, use it as the `webhook_url` when creating a transcription: ```ts import express from 'express'; import { SonioxNodeClient } from '@soniox/node'; const client = new SonioxNodeClient(); const app = express(); app.use(express.json()); // Handle incoming webhook events app.post('/webhooks/soniox', async (req, res) => { const result = client.webhooks.handleExpress(req); // You will receive the webhook event when the transcription is completed if (result.ok && result.event.status === 'completed') { const transcript = await result.fetchTranscript(); // Lazy fetch the transcript console.log(transcript?.text); } res.status(result.status).json({ received: true }); }); app.listen(3000, () => console.log('Listening on port 3000')); // Start a transcription with the tunnel URL as webhook await client.stt.transcribe({ model: 'stt-async-v5', audio_url: 'https://soniox.com/media/examples/coffee_shop.mp3', webhook_url: 'https:///webhooks/soniox', }); ``` # Real-time speech generation with Node SDK URL: /sdk/node-SDK/tts/realtime-speech-generation Stream text to speech with the Soniox Node SDK over WebSocket The Soniox Node SDK supports real-time Text-to-Speech generation over WebSocket. You send text — all at once or incrementally — and receive decoded audio chunks as they are generated. This is the lowest-latency path for voice agents, LLM output narration, and any scenario where text arrives progressively. If you already have the full text up front and don't need chunk-by-chunk streaming, use [REST speech generation](/sdk/node-SDK/tts/rest-speech-generation) instead — it's a single HTTP request. ## Quickstart `client.realtime.tts()` creates a single-stream session: it opens a WebSocket, configures a stream, and returns a [`RealtimeTtsStream`](/sdk/node-SDK/reference/classes#realtimettsstream). Send text, then consume audio by async iteration. ```typescript import { writeFileSync } from "node:fs"; import { SonioxNodeClient } from "@soniox/node"; const client = new SonioxNodeClient(); const stream = await client.realtime.tts({ voice: "Adrian", model: "tts-rt-v2", language: "en", audio_format: "wav", }); // Send all the text and mark the stream as finished. stream.sendText( "Hello from Soniox real-time text-to-speech. This is a single-stream example.", { end: true }, ); // Collect audio as it arrives. const chunks: Uint8Array[] = []; for await (const chunk of stream) { chunks.push(chunk); } const audio = Buffer.concat(chunks); writeFileSync("tts_realtime.wav", audio); console.log(`Wrote ${audio.byteLength} bytes`); ``` The stream closes itself (and the underlying WebSocket) once `terminated` fires. You never have to call `close()` in single-stream mode. ## Send text incrementally Use `sendText(text)` for each chunk as it becomes available, then either set `{ end: true }` on the last call or invoke `finish()` explicitly. This is the pattern for narrating an LLM response token-by-token. ```typescript const stream = await client.realtime.tts({ voice: "Adrian", model: "tts-rt-v2", audio_format: "wav", }); stream.sendText("Hello from Soniox "); stream.sendText("real-time TTS. "); stream.sendText("This is the final chunk.", { end: true }); for await (const chunk of stream) { playback(chunk); // play, buffer, or write the chunk somewhere } ``` Equivalent with an explicit `finish()`: ```typescript stream.sendText("Hello from Soniox real-time TTS."); stream.finish(); for await (const chunk of stream) { playback(chunk); } ``` ## Pipe from an async iterable `stream.sendStream(source)` pipes any `AsyncIterable` into the TTS session and auto-finishes when the iterable completes. This is the idiomatic way to connect an LLM token stream directly to speech output — sending and receiving run concurrently. ```typescript async function* tokensFromLlm(): AsyncIterable { const words = "Hello from Soniox real-time TTS.".split(" "); for (let i = 0; i < words.length; i++) { await new Promise((r) => setTimeout(r, 50)); yield i === 0 ? words[i] : " " + words[i]; } } const stream = await client.realtime.tts({ voice: "Adrian", model: "tts-rt-v2", audio_format: "wav", }); // Start piping text. Audio consumption runs concurrently below. void stream.sendStream(tokensFromLlm()); for await (const chunk of stream) { playback(chunk); } ``` ## Event-based consumption `RealtimeTtsStream` is also a [`TypedEmitter`](/sdk/node-SDK/reference/classes#realtimettsstream). When you prefer an event-driven style over async iteration, listen for [`TtsStreamEvents`](/sdk/node-SDK/reference/types#ttsstreamevents): | Event | Payload | Description | | ------------ | ------------ | ------------------------------------------------------ | | `audio` | `Uint8Array` | Decoded audio chunk. | | `audioEnd` | — | Server marked the final audio payload for this stream. | | `terminated` | — | Stream fully closed by the server. | | `error` | `Error` | Stream-level error. | ```typescript const stream = await client.realtime.tts({ voice: "Adrian", model: "tts-rt-v2", audio_format: "wav", }); stream.on("audio", (chunk) => playback(chunk)); stream.on("audioEnd", () => console.log("last audio payload received")); stream.on("error", (err) => console.error("Stream error:", err)); stream.on("terminated", () => console.log("stream done")); stream.sendText("Hello from event-based TTS.", { end: true }); ``` Choose either async iteration **or** event listeners — not both. The async iterator consumes `audio` events internally. ## Multi-stream connection A single WebSocket connection can carry up to 5 concurrent TTS streams. Use `client.realtime.tts.multiStream()` to open a [`RealtimeTtsConnection`](/sdk/node-SDK/reference/classes#realtimettsconnection), then call `connection.stream()` for each stream. Each stream has its own `streamId` and can have different voice, model, and audio format settings. ```typescript const connection = await client.realtime.tts.multiStream(); const streamA = await connection.stream({ voice: "Adrian", audio_format: "wav", }); const streamB = await connection.stream({ // Enumerate available voices via `client.tts.listModels()`. voice: "", audio_format: "wav", }); streamA.sendText("Hello from stream A.", { end: true }); streamB.sendText("Hello from stream B.", { end: true }); const [audioA, audioB] = await Promise.all([ collect(streamA), collect(streamB), ]); connection.close(); async function collect(stream: AsyncIterable): Promise { const chunks: Uint8Array[] = []; for await (const chunk of stream) { chunks.push(chunk); } return Buffer.concat(chunks); } ``` Call `connection.close()` when you're done — this ends all active streams and closes the WebSocket. ## Cancel, finish, and close | Method | Behavior | | -------------------- | --------------------------------------------------------------------------------------------------------- | | `stream.finish()` | Signals "no more text". The server finishes generating audio and sends `terminated`. | | `stream.cancel()` | Aborts generation immediately. The server stops producing audio and sends `terminated`. | | `stream.close()` | Terminates the stream. In single-stream mode (`client.realtime.tts(...)`) this also closes the WebSocket. | | `connection.close()` | Closes the WebSocket and terminates all streams on a multi-stream connection. | ```typescript // Graceful stop stream.finish(); // User-triggered cancel stream.cancel(); ``` ## Error handling A failed stream does not close the whole WebSocket connection by default. Stream-level errors finalize only that stream (`terminated` fires for the same `streamId`), while other streams on the same connection can continue. Connection-level failures end the whole connection and all active streams. ```typescript import { RealtimeError, SonioxError } from "@soniox/node"; try { const stream = await client.realtime.tts({ voice: "Adrian" }); stream.sendText("Hello!", { end: true }); for await (const _ of stream) { // consume audio } } catch (err) { if (err instanceof RealtimeError) { console.error(`Realtime TTS error (${err.code}):`, err.message); } else if (err instanceof SonioxError) { console.error("Soniox SDK error:", err.message); } else { throw err; } } ``` ## Server-driven defaults Set shared TTS fields once on the client via `tts_defaults` and they'll be merged as the base layer every time you open a stream. Caller-provided fields on `client.realtime.tts(...)` / `connection.stream(...)` override the defaults, so you never need to spread them manually. ```typescript import { SonioxNodeClient } from "@soniox/node"; const client = new SonioxNodeClient({ tts_defaults: { model: "tts-rt-v2", language: "en", voice: "Adrian", audio_format: "wav", }, }); // Inherits everything from tts_defaults. const stream = await client.realtime.tts(); // Overrides the default voice for this stream only. const customVoice = await client.realtime.tts({ voice: "" }); ``` `tts_defaults` is also accepted on [`RealtimeOptions`](/sdk/node-SDK/reference/types#realtimeoptions) if you want to scope defaults to a specific realtime namespace. On the Web and React SDKs, the equivalent is [`SonioxConnectionConfig.tts_defaults`](/sdk/web-SDK/reference/types#sonioxconnectionconfig) — return it from the async `config` resolver alongside the temporary `api_key` so the server owns the defaults. ## See also * [REST speech generation](/sdk/node-SDK/tts/rest-speech-generation) — single-request HTTP TTS. * [`RealtimeTtsStream` reference](/sdk/node-SDK/reference/classes#realtimettsstream) * [`RealtimeTtsConnection` reference](/sdk/node-SDK/reference/classes#realtimettsconnection) * [`TtsStreamInput`](/sdk/node-SDK/reference/types#ttsstreaminput), [`TtsStreamEvents`](/sdk/node-SDK/reference/types#ttsstreamevents) * [TTS WebSocket API](/api-reference/tts/websocket-api) # REST speech generation with Node SDK URL: /sdk/node-SDK/tts/rest-speech-generation Generate speech from text with the Soniox Node SDK over HTTP The Soniox Node SDK supports Text-to-Speech generation over HTTP with `SonioxNodeClient`. Use REST when you have the full text up front and don't need streaming from an LLM — the SDK can return audio bytes, stream them as an async iterable, or write the output directly to a file. Use [real-time speech generation](/sdk/node-SDK/tts/realtime-speech-generation) when you want the lowest latency to first audio or when text arrives incrementally (for example, streamed from an LLM). ## Quickstart The shortest path is `generateToFile` — the SDK calls the REST API, streams the response, and writes the output for you. ```typescript import { SonioxNodeClient } from "@soniox/node"; // The API key is read from the SONIOX_API_KEY environment variable. const client = new SonioxNodeClient(); const bytesWritten = await client.tts.generateToFile("hello.wav", { text: "Hello from the Soniox Node SDK text-to-speech example.", voice: "Adrian", model: "tts-rt-v2", language: "en", audio_format: "wav", }); console.log(`Wrote ${bytesWritten} bytes`); ``` ## Generate to bytes Use `client.tts.generate()` when you want the full audio payload in memory — for custom storage, uploading to another service, or post-processing. ```typescript const audio = await client.tts.generate({ text: "This response is generated in memory.", voice: "Adrian", model: "tts-rt-v2", language: "en", audio_format: "wav", }); console.log(`Received ${audio.byteLength} bytes`); // audio is a Uint8Array ``` ## Stream audio chunks Use `client.tts.generateStream()` to receive the response as an async iterable of `Uint8Array` chunks. This lets you start processing or playing audio before the full payload has arrived. ```typescript import { createWriteStream } from "node:fs"; const output = createWriteStream("streamed.wav"); let totalBytes = 0; for await (const chunk of client.tts.generateStream({ text: "Streaming audio as it arrives from the server.", voice: "Adrian", model: "tts-rt-v2", language: "en", audio_format: "wav", })) { output.write(chunk); totalBytes += chunk.byteLength; } output.end(); console.log(`Received ${totalBytes} bytes`); ``` ## Write directly to a file `client.tts.generateToFile()` accepts either a file path (string) or any `WritableStream` and returns the total number of bytes written. This is the simplest option for common server-side workflows. ```typescript const bytesWritten = await client.tts.generateToFile("hello.pcm", { text: "Hello from Soniox.", voice: "Adrian", model: "tts-rt-v2", language: "en", audio_format: "pcm_s16le", sample_rate: 24000, }); console.log(`Wrote ${bytesWritten} bytes`); ``` ## List available models `client.tts.listModels()` returns the set of available TTS models with their supported voices. ```typescript const models = await client.tts.listModels(); for (const model of models) { const voices = model.voices.map((v) => v.id).join(", "); console.log(`${model.id} (${model.name}): ${voices}`); } ``` ## Generation options All three generator methods accept the same [`GenerateSpeechOptions`](/sdk/node-SDK/reference/types#generatespeechoptions) shape: | Option | Type | Description | | -------------- | ---------------------------------------------------------------- | ------------------------------------------------------- | | `text` | `string` | Input text to synthesize. Required. | | `voice` | `string` | Voice identifier (e.g. `"Adrian"`). Required. | | `model` | `string` | TTS model. Default `"tts-rt-v2"`. | | `language` | `string` | Language code. Default `"en"`. | | `audio_format` | [`TtsAudioFormat`](/sdk/node-SDK/reference/types#ttsaudioformat) | Output audio format. Default `"wav"`. | | `sample_rate` | `number` | Output sample rate in Hz. Required for raw PCM formats. | | `bitrate` | `number` | Codec bitrate in bps (for compressed formats). | | `signal` | `AbortSignal` | Optional signal to cancel the request. | See [Available models](/tts/models) for the full list of TTS models, voices, and supported audio formats. ## Cancel a request Pass an `AbortSignal` to cancel a generation request — useful when the user changes their mind mid-request, or when you need a deadline. ```typescript import { SonioxHttpError } from "@soniox/node"; const controller = new AbortController(); setTimeout(() => controller.abort(), 5000); // abort after 5s try { const audio = await client.tts.generate({ text: "Some long text...", voice: "Adrian", signal: controller.signal, }); console.log(`Received ${audio.byteLength} bytes`); } catch (err) { if (err instanceof SonioxHttpError && err.code === "aborted") { console.log("Generation cancelled"); } else { throw err; } } ``` ## Error handling REST TTS requests can throw the following errors: | Error | When it's thrown | | ----------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `SonioxHttpError` | Covers all HTTP failures: non-2xx responses (`code: 'http_error'`), network failures (`code: 'network_error'`), timeouts (`code: 'timeout'`), aborted requests (`code: 'aborted'`), and parse errors (`code: 'parse_error'`). Inspect `code`, `statusCode`, `message`, and `bodyText`. | | `SonioxError` | Base class for all SDK errors (`SonioxHttpError` extends it). Catch this if you want a single branch for every Soniox-originated failure. | ```typescript import { SonioxHttpError, SonioxError } from "@soniox/node"; try { await client.tts.generateToFile("out.wav", { text: "Hello!", voice: "Adrian", }); } catch (err) { if (err instanceof SonioxHttpError) { console.error( `HTTP ${err.statusCode ?? "n/a"} (${err.code}): ${err.message}` ); } else if (err instanceof SonioxError) { console.error("Soniox SDK error:", err.message); } else { throw err; } } ``` For raw HTTP integration details, see the [TTS REST API reference](/api-reference/tts/generate_tts). ### Error handling limitations Once audio streaming has started, errors cannot be delivered to the client. For guaranteed error delivery, use the [realtime WebSocket TTS](/sdk/node-SDK/tts/realtime-speech-generation) instead. ## See also * [Real-time speech generation](/sdk/node-SDK/tts/realtime-speech-generation) — WebSocket-based streaming TTS for LLM token streaming and low-latency playback. * [`GenerateSpeechOptions` reference](/sdk/node-SDK/reference/types#generatespeechoptions) * [`TtsModel` reference](/sdk/node-SDK/reference/types#ttsmodel) and [`TtsAudioFormat`](/sdk/node-SDK/reference/types#ttsaudioformat) # Async Client URL: /sdk/python-SDK/Full-SDK-reference/async_client Soniox Python SDK - Async Client Reference *** > **Sync mirror:** the synchronous `SonioxClient` exposes the same API as `AsyncSonioxClient` below - drop `await` from each call and treat `AsyncIterator[X]` return types as plain `Iterator[X]`. Only the async surface is documented here to avoid duplicating an otherwise identical reference. Realtime sessions have genuinely different sync/async patterns and are documented in the [Realtime Client](./realtime_client.md) page. *** ## AsyncSonioxClient Asynchronous Soniox REST client exposing HTTP and realtime helpers. ### Constructor ```python AsyncSonioxClient(api_key: str | None = None, api_base_url: str | None = None, websocket_base_url: str | None = None, tts_api_base_url: str | None = None, tts_websocket_base_url: str | None = None, timeout_sec: float | None = None, webhook_secret: str | None = None, webhook_signature_header: str | None = None, **client_kwargs: Any) ``` **Parameters** | Parameter | Type | Description | | -------------------------- | --------------- | --------------------------------------------------------------- | | `api_key` | `str \| None` | API key used for authentication. | | `api_base_url` | `str \| None` | Base URL for Soniox REST API requests. | | `websocket_base_url` | `str \| None` | Base URL for Soniox realtime WebSocket endpoint. | | `tts_api_base_url` | `str \| None` | Base URL for Soniox Text-to-Speech REST API requests. | | `tts_websocket_base_url` | `str \| None` | Base URL for Soniox Text-to-Speech realtime WebSocket endpoint. | | `timeout_sec` | `float \| None` | Maximum wait time in seconds. | | `webhook_secret` | `str \| None` | Webhook secret used for signature verification. | | `webhook_signature_header` | `str \| None` | Webhook signature header name. | | `client_kwargs` | `Any` | Additional HTTP client keyword arguments. | **Returns** `None` ### Properties | Property | Type | Description | | -------------------- | --------------------------- | ----------------------------------------------------------- | | `files` | `AsyncFilesAPI` | List of uploaded files. | | `stt` | `AsyncSttAPI` | Speech-to-text API namespace. | | `tts` | `AsyncTtsAPI` | Text-to-Speech API namespace | | `models` | `AsyncModelsAPI` | Voice readiness status for each available model. | | `tts_models` | `AsyncTtsModelsAPI` | - | | `voices` | `AsyncVoicesAPI` | Voices supported by this model. | | `usage_logs` | `AsyncUsageLogsAPI` | Per-request usage log entries ordered by end\_time. | | `concurrency_limits` | `AsyncConcurrencyLimitsAPI` | - | | `auth` | `AsyncAuthAPI` | Authentication API namespace. | | `webhooks` | `AsyncSonioxWebhooksAPI` | Webhook utilities API namespace. | | `realtime` | `AsyncRealtimeAPI` | Entrypoint for async realtime helpers on AsyncSonioxClient. | ### request() ```python request(method: str, path: str, *, params: Mapping[str, Any] | None = None, json: Any | None = None, data: Mapping[str, Any] | None = None, files: Mapping[str, Any] | None = None) -> httpx.Response ``` Perform a request against the configured Soniox REST endpoint. **Parameters** | Parameter | Type | Description | | --------- | --------------------------- | ----------------------------------- | | `method` | `str` | HTTP method to use for the request. | | `path` | `str` | Relative API path for the request. | | `params` | `Mapping[str, Any] \| None` | Query parameters for the request. | | `json` | `Any \| None` | JSON request payload. | | `data` | `Mapping[str, Any] \| None` | Form-encoded request payload. | | `files` | `Mapping[str, Any] \| None` | Multipart file payload mapping. | **Returns** `httpx.Response` *** ### aclose() ```python aclose() -> None ``` Close any outstanding async HTTP connections. **Returns** `None` *** ## AsyncFilesAPI ### Constructor ```python AsyncFilesAPI(client: AsyncSonioxClient) ``` **Parameters** | Parameter | Type | Description | | --------- | ------------------- | ----------------------- | | `client` | `AsyncSonioxClient` | Soniox client instance. | **Returns** `None` ### list() ```python list(limit: int = 100, cursor: str | None = None) -> GetFilesResponse ``` List uploaded files. Performs a GET request to `/files` with optional pagination. **Parameters** | Parameter | Type | Description | | --------- | ------------- | ----------------------------------------------- | | `limit` | `int` | Maximum number of files to return. | | `cursor` | `str \| None` | Pagination cursor for the next page of results. | **Returns** `GetFilesResponse` **Raises** * `SonioxAPIError` When the API returns an error. *** ### count() ```python count() -> GetFilesCountResponse ``` Return a breakdown of uploaded file counts. Performs a GET request to `/files/count`. **Returns** `GetFilesCountResponse` **Raises** * `SonioxAPIError` When the API returns an error. *** ### list\_all() ```python list_all(limit: int = 100) -> AsyncGenerator[File, None] ``` Iterate through all uploaded files across all pages. **Parameters** | Parameter | Type | Description | | --------- | ----- | ---------------------------------- | | `limit` | `int` | Maximum number of files to return. | **Yields** `AsyncGenerator[File, None]` File: The next file object from the API. **Raises** * `SonioxAPIError` When the API returns an error. *** ### get() ```python get(file_id: str) -> File ``` Retrieve a file by ID. Performs a GET request to `/files/{file_id}`. **Parameters** | Parameter | Type | Description | | --------- | ----- | --------------------------------- | | `file_id` | `str` | ID of a previously uploaded file. | **Returns** `File` **Raises** * `SonioxAPIError` When the API returns an error. *** ### get\_or\_none() ```python get_or_none(file_id: str) -> File | None ``` Retrieve a file by ID. Returns `None` if the file does not exist. **Parameters** | Parameter | Type | Description | | --------- | ----- | --------------------------------- | | `file_id` | `str` | ID of a previously uploaded file. | **Returns** `File | None` **Raises** * `SonioxAPIError` When the API returns an error. *** ### delete() ```python delete(file_id: str) -> None ``` Delete a file by ID. Performs a DELETE request to `/files/{file_id}`. **Parameters** | Parameter | Type | Description | | --------- | ----- | --------------------------------- | | `file_id` | `str` | ID of a previously uploaded file. | **Returns** `None` **Raises** * `SonioxAPIError` When the API returns an error. *** ### delete\_if\_exists() ```python delete_if_exists(file_id: str) -> None ``` Delete a file by ID if it exists. Ignores missing files. **Parameters** | Parameter | Type | Description | | --------- | ----- | --------------------------------- | | `file_id` | `str` | ID of a previously uploaded file. | **Returns** `None` **Raises** * `SonioxAPIError` When the API returns an error. *** ### upload() ```python upload(file: BinaryIO | bytes | Path | str, *, filename: str | None = None, client_reference_id: str | None = None) -> File ``` Upload a file. Performs a multipart POST request to `/files`. **Parameters** | Parameter | Type | Description | | --------------------- | ---------------------------------- | --------------------------------------------------------------- | | `file` | `BinaryIO \| bytes \| Path \| str` | File input to upload or transcribe. | | `filename` | `str \| None` | Filename associated with uploaded file data. | | `client_reference_id` | `str \| None` | Optional tracking identifier string. Does not need to be unique | **Returns** `File` **Raises** * `SonioxAPIError` When the API returns an error. *** ### delete\_all() ```python delete_all(limit: int = 100) -> None ``` Delete all files. Iterates through all pages and deletes each file. Stops and raises on the first failed deletion. **Parameters** | Parameter | Type | Description | | --------- | ----- | ---------------------------------- | | `limit` | `int` | Maximum number of files to return. | **Returns** `None` **Raises** * `SonioxAPIError` When the API returns an error. *** ## AsyncSttAPI ### Constructor ```python AsyncSttAPI(client: AsyncSonioxClient) ``` **Parameters** | Parameter | Type | Description | | --------- | ------------------- | ----------------------- | | `client` | `AsyncSonioxClient` | Soniox client instance. | **Returns** `None` ### list() ```python list(limit: int = 100, cursor: str | None = None) -> GetTranscriptionsResponse ``` List transcriptions. Performs a GET request to `/transcriptions` with optional pagination. **Parameters** | Parameter | Type | Description | | --------- | ------------- | ----------------------------------------------- | | `limit` | `int` | Maximum number of transcriptions to return. | | `cursor` | `str \| None` | Pagination cursor for the next page of results. | **Returns** `GetTranscriptionsResponse` **Raises** * `SonioxAPIError` When the API returns an error. *** ### count() ```python count() -> GetTranscriptionsCountResponse ``` Return a breakdown of transcription counts. Performs a GET request to `/transcriptions/count`. **Returns** `GetTranscriptionsCountResponse` **Raises** * `SonioxAPIError` When the API returns an error. *** ### list\_all() ```python list_all(limit: int = 100) -> AsyncGenerator[Transcription, None] ``` Iterate through all transcriptions across all pages. **Parameters** | Parameter | Type | Description | | --------- | ----- | ------------------------------------------- | | `limit` | `int` | Maximum number of transcriptions to return. | **Yields** `AsyncGenerator[Transcription, None]` File: The next transcription object from the API. **Raises** * `SonioxAPIError` When the API returns an error. *** ### delete\_all() ```python delete_all(limit: int = 100) -> None ``` Delete all transcriptions. Iterates through all pages and deletes each transcription. Stops and raises on the first failed deletion. **Parameters** | Parameter | Type | Description | | --------- | ----- | ------------------------------------------- | | `limit` | `int` | Maximum number of transcriptions to return. | **Returns** `None` **Raises** * `SonioxAPIError` When the API returns an error. *** ### create() ```python create(*, model: str = DEFAULT_MODEL, file_id: str | None = None, audio_url: str | None = None, client_reference_id: str | None = None, config: CreateTranscriptionConfig | None = None) -> Transcription ``` Create a transcription. Performs a POST request to `/transcriptions`. **Parameters** | Parameter | Type | Description | | --------------------- | ----------------------------------- | ----------------------------------------- | | `model` | `str` | Speech-to-text model to use. | | `file_id` | `str \| None` | ID of a previously uploaded file. | | `audio_url` | `str \| None` | Publicly accessible audio URL. | | `client_reference_id` | `str \| None` | Optional tracking identifier. | | `config` | `CreateTranscriptionConfig \| None` | Configuration options for this operation. | **Returns** `Transcription` **Raises** * `SonioxAPIError` When the API returns an error. *** ### get() ```python get(transcription_id: str) -> Transcription ``` Retrieve a transcription by ID. Performs a GET request to `/transcriptions/{transcription_id}`. **Parameters** | Parameter | Type | Description | | ------------------ | ----- | ------------------------- | | `transcription_id` | `str` | Transcription identifier. | **Returns** `Transcription` **Raises** * `SonioxAPIError` When the API returns an error. *** ### get\_or\_none() ```python get_or_none(transcription_id: str) -> Transcription | None ``` Retrieve a transcription by ID. Returns `None` if the transcription does not exist. **Parameters** | Parameter | Type | Description | | ------------------ | ----- | ------------------------- | | `transcription_id` | `str` | Transcription identifier. | **Returns** `Transcription | None` **Raises** * `SonioxAPIError` When the API returns an error. *** ### delete() ```python delete(transcription_id: str) -> None ``` Delete a transcription by ID. Performs a DELETE request to `/transcriptions/{transcription_id}`. **Parameters** | Parameter | Type | Description | | ------------------ | ----- | ------------------------- | | `transcription_id` | `str` | Transcription identifier. | **Returns** `None` **Raises** * `SonioxAPIError` When the API returns an error. *** ### delete\_if\_exists() ```python delete_if_exists(transcription_id: str) -> None ``` Delete a transcription by ID if it exists. Ignores missing transcriptions. **Parameters** | Parameter | Type | Description | | ------------------ | ----- | ------------------------- | | `transcription_id` | `str` | Transcription identifier. | **Returns** `None` **Raises** * `SonioxAPIError` When the API returns an error. *** ### destroy() ```python destroy(transcription_id: str) -> None ``` Delete a transcription and its associated uploaded file. **Parameters** | Parameter | Type | Description | | ------------------ | ----- | ------------------------- | | `transcription_id` | `str` | Transcription identifier. | **Returns** `None` **Raises** * `SonioxAPIError` When the API returns an error. *** ### destroy\_all() ```python destroy_all(limit: int = 100) -> None ``` Delete all transcriptions and their associated files. Stops and raises on the first failed deletion. **Parameters** | Parameter | Type | Description | | --------- | ----- | ------------------------------------------- | | `limit` | `int` | Maximum number of transcriptions to return. | **Returns** `None` **Raises** * `SonioxAPIError` When the API returns an error during listing. *** ### get\_transcript() ```python get_transcript(transcription_id: str) -> TranscriptionTranscript ``` Retrieve the transcript for a transcription. Performs a GET request to `/transcriptions/{transcription_id}/transcript`. **Parameters** | Parameter | Type | Description | | ------------------ | ----- | ------------------------- | | `transcription_id` | `str` | Transcription identifier. | **Returns** `TranscriptionTranscript` **Raises** * `SonioxAPIError` When the API returns an error. *** ### wait() ```python wait(transcription_id: str, *, interval_sec: float = 5.0, timeout_sec: float | None = None) -> Transcription ``` Poll a transcription until it leaves the queued or processing state. **Parameters** | Parameter | Type | Description | | ------------------ | --------------- | ----------------------------- | | `transcription_id` | `str` | Transcription identifier. | | `interval_sec` | `float` | Polling interval in seconds. | | `timeout_sec` | `float \| None` | Maximum wait time in seconds. | **Returns** `Transcription` **Raises** * `SonioxAPIError` When the API returns an error. * `TimeoutError` Waiting for the transcription to finish exceeded `timeout_sec`. *** ### transcribe\_from\_url() ```python transcribe_from_url(*, model: str = DEFAULT_MODEL, audio_url: str, client_reference_id: str | None = None, config: CreateTranscriptionConfig | None = None) -> Transcription ``` Create a transcription from an audio URL. **Parameters** | Parameter | Type | Description | | --------------------- | ----------------------------------- | ----------------------------------------- | | `model` | `str` | Speech-to-text model to use. | | `audio_url` | `str` | Publicly accessible audio URL. | | `client_reference_id` | `str \| None` | Optional tracking identifier. | | `config` | `CreateTranscriptionConfig \| None` | Configuration options for this operation. | **Returns** `Transcription` **Raises** * `SonioxAPIError` When the API returns an error. *** ### transcribe\_from\_file\_id() ```python transcribe_from_file_id(*, model: str = DEFAULT_MODEL, file_id: str, client_reference_id: str | None = None, config: CreateTranscriptionConfig | None = None) -> Transcription ``` Create a transcription from an existing uploaded file. **Parameters** | Parameter | Type | Description | | --------------------- | ----------------------------------- | ----------------------------------------- | | `model` | `str` | Speech-to-text model to use. | | `file_id` | `str` | ID of a previously uploaded file. | | `client_reference_id` | `str \| None` | Optional tracking identifier. | | `config` | `CreateTranscriptionConfig \| None` | Configuration options for this operation. | **Returns** `Transcription` **Raises** * `SonioxAPIError` When the API returns an error. *** ### transcribe\_from\_file() ```python transcribe_from_file(*, model: str = DEFAULT_MODEL, file: BinaryIO | bytes | Path | str, filename: str | None = None, client_reference_id: str | None = None, config: CreateTranscriptionConfig | None = None) -> Transcription ``` Upload a file and create a transcription from it. **Parameters** | Parameter | Type | Description | | --------------------- | ----------------------------------- | -------------------------------------------- | | `model` | `str` | Speech-to-text model to use. | | `file` | `BinaryIO \| bytes \| Path \| str` | File input to upload or transcribe. | | `filename` | `str \| None` | Filename associated with uploaded file data. | | `client_reference_id` | `str \| None` | Optional tracking identifier. | | `config` | `CreateTranscriptionConfig \| None` | Configuration options for this operation. | **Returns** `Transcription` **Raises** * `SonioxAPIError` When the API returns an error. *** ### transcribe() ```python transcribe(*, model: str = DEFAULT_MODEL, audio_url: str | None = None, file_id: str | None = None, file: BinaryIO | bytes | Path | str | None = None, filename: str | None = None, client_reference_id: str | None = None, config: CreateTranscriptionConfig | None = None) -> Transcription ``` Create a transcription from a file, file ID, or audio URL. Validates mutually exclusive inputs before submission. **Parameters** | Parameter | Type | Description | | --------------------- | ------------------------------------------ | -------------------------------------------- | | `model` | `str` | Speech-to-text model to use. | | `audio_url` | `str \| None` | Publicly accessible audio URL. | | `file_id` | `str \| None` | ID of a previously uploaded file. | | `file` | `BinaryIO \| bytes \| Path \| str \| None` | File input to upload or transcribe. | | `filename` | `str \| None` | Filename associated with uploaded file data. | | `client_reference_id` | `str \| None` | Optional tracking identifier. | | `config` | `CreateTranscriptionConfig \| None` | Configuration options for this operation. | **Returns** `Transcription` **Raises** * `SonioxAPIError` When the API returns an error. * `SonioxValidationError` When the payload fails validation. *** ### transcribe\_file\_with\_webhook() ```python transcribe_file_with_webhook(*, model: str = DEFAULT_MODEL, file: BinaryIO | bytes | Path | str, webhook_url: str, filename: str | None = None, client_reference_id: str | None = None, webhook_auth: WebhookAuthConfig | None = None, config: CreateTranscriptionConfig | None = None) -> Transcription ``` Upload a file, configure a webhook, and start transcription. **Parameters** | Parameter | Type | Description | | --------------------- | ----------------------------------- | -------------------------------------------- | | `model` | `str` | Speech-to-text model to use. | | `file` | `BinaryIO \| bytes \| Path \| str` | File input to upload or transcribe. | | `webhook_url` | `str` | URL to receive webhook notifications. | | `filename` | `str \| None` | Filename associated with uploaded file data. | | `client_reference_id` | `str \| None` | Optional tracking identifier. | | `webhook_auth` | `WebhookAuthConfig \| None` | Webhook authentication configuration. | | `config` | `CreateTranscriptionConfig \| None` | Configuration options for this operation. | **Returns** `Transcription` **Raises** * `SonioxAPIError` When the API returns an error. *** ### transcribe\_and\_wait() ```python transcribe_and_wait(*, model: str = DEFAULT_MODEL, audio_url: str | None = None, file_id: str | None = None, file: BinaryIO | bytes | Path | str | None = None, filename: str | None = None, client_reference_id: str | None = None, delete_after: bool = False, wait_interval_sec: float = 5.0, wait_timeout_sec: float | None = None, config: CreateTranscriptionConfig | None = None) -> Transcription ``` Create a transcription and wait for completion. Returns a Transcription object after it is completed. Optionally deletes the transcription and the uploaded file after completion. **Parameters** | Parameter | Type | Description | | --------------------- | ------------------------------------------ | ----------------------------------------------------- | | `model` | `str` | Speech-to-text model to use. | | `audio_url` | `str \| None` | Publicly accessible audio URL. | | `file_id` | `str \| None` | ID of a previously uploaded file. | | `file` | `BinaryIO \| bytes \| Path \| str \| None` | File input to upload or transcribe. | | `filename` | `str \| None` | Filename associated with uploaded file data. | | `client_reference_id` | `str \| None` | Optional tracking identifier. | | `delete_after` | `bool` | Whether to delete created resources after completion. | | `wait_interval_sec` | `float` | Polling interval in seconds while waiting. | | `wait_timeout_sec` | `float \| None` | Maximum wait time in seconds while polling. | | `config` | `CreateTranscriptionConfig \| None` | Configuration options for this operation. | **Returns** `Transcription` **Raises** * `SonioxAPIError` When the API returns an error. * `SonioxValidationError` When the payload fails validation. * `TimeoutError` Waiting for the transcription to finish exceeded `timeout_sec`. *** ### transcribe\_and\_wait\_with\_tokens() ```python transcribe_and_wait_with_tokens(*, model: str = DEFAULT_MODEL, audio_url: str | None = None, file_id: str | None = None, file: BinaryIO | bytes | Path | str | None = None, filename: str | None = None, client_reference_id: str | None = None, delete_after: bool = False, wait_interval_sec: float = 5.0, wait_timeout_sec: float | None = None, config: CreateTranscriptionConfig | None = None) -> TranscriptionTranscript ``` Create a transcription, wait for completion, and return the transcript. Optionally deletes the transcription and uploaded file after completion. **Parameters** | Parameter | Type | Description | | --------------------- | ------------------------------------------ | ----------------------------------------------------- | | `model` | `str` | Speech-to-text model to use. | | `audio_url` | `str \| None` | Publicly accessible audio URL. | | `file_id` | `str \| None` | ID of a previously uploaded file. | | `file` | `BinaryIO \| bytes \| Path \| str \| None` | File input to upload or transcribe. | | `filename` | `str \| None` | Filename associated with uploaded file data. | | `client_reference_id` | `str \| None` | Optional tracking identifier. | | `delete_after` | `bool` | Whether to delete created resources after completion. | | `wait_interval_sec` | `float` | Polling interval in seconds while waiting. | | `wait_timeout_sec` | `float \| None` | Maximum wait time in seconds while polling. | | `config` | `CreateTranscriptionConfig \| None` | Configuration options for this operation. | **Returns** `TranscriptionTranscript` **Raises** * `SonioxAPIError` When the API returns an error. * `SonioxValidationError` When the payload fails validation. * `TimeoutError` Waiting for the transcription to finish exceeded `timeout_sec`. *** ### translate\_from\_url() ```python translate_from_url(*, audio_url: str, to: LanguageCode | None = None, source: LanguageCode | None = None, between: tuple[LanguageCode, LanguageCode] | None = None, model: str = DEFAULT_MODEL, client_reference_id: str | None = None, config: CreateTranscriptionConfig | None = None) -> Transcription ``` Translate audio at a URL. Provide exactly one of `to` (one-way) or `between` (two-way). `source` is an optional language hint and is only valid with `to`. **Parameters** | Parameter | Type | Description | | --------------------- | ------------------------------------------- | ----------------------------------------- | | `audio_url` | `str` | Publicly accessible audio URL. | | `to` | `LanguageCode \| None` | - | | `source` | `LanguageCode \| None` | The source term to translate. | | `between` | `tuple[LanguageCode, LanguageCode] \| None` | - | | `model` | `str` | Speech-to-text model to use. | | `client_reference_id` | `str \| None` | Optional tracking identifier. | | `config` | `CreateTranscriptionConfig \| None` | Configuration options for this operation. | **Returns** `Transcription` **Raises** * `SonioxAPIError` When the API returns an error. * `SonioxValidationError` When the translate kwargs are invalid. *** ### translate\_from\_file\_id() ```python translate_from_file_id(*, file_id: str, to: LanguageCode | None = None, source: LanguageCode | None = None, between: tuple[LanguageCode, LanguageCode] | None = None, model: str = DEFAULT_MODEL, client_reference_id: str | None = None, config: CreateTranscriptionConfig | None = None) -> Transcription ``` Translate an already-uploaded file. Provide exactly one of `to` (one-way) or `between` (two-way). `source` is an optional language hint and is only valid with `to`. **Parameters** | Parameter | Type | Description | | --------------------- | ------------------------------------------- | ----------------------------------------- | | `file_id` | `str` | ID of a previously uploaded file. | | `to` | `LanguageCode \| None` | - | | `source` | `LanguageCode \| None` | The source term to translate. | | `between` | `tuple[LanguageCode, LanguageCode] \| None` | - | | `model` | `str` | Speech-to-text model to use. | | `client_reference_id` | `str \| None` | Optional tracking identifier. | | `config` | `CreateTranscriptionConfig \| None` | Configuration options for this operation. | **Returns** `Transcription` **Raises** * `SonioxAPIError` When the API returns an error. * `SonioxValidationError` When the translate kwargs are invalid. *** ### translate\_from\_file() ```python translate_from_file(*, file: BinaryIO | bytes | Path | str, to: LanguageCode | None = None, source: LanguageCode | None = None, between: tuple[LanguageCode, LanguageCode] | None = None, filename: str | None = None, model: str = DEFAULT_MODEL, client_reference_id: str | None = None, config: CreateTranscriptionConfig | None = None) -> Transcription ``` Upload a file and translate it. Provide exactly one of `to` (one-way) or `between` (two-way). `source` is an optional language hint and is only valid with `to`. **Parameters** | Parameter | Type | Description | | --------------------- | ------------------------------------------- | -------------------------------------------- | | `file` | `BinaryIO \| bytes \| Path \| str` | File input to upload or transcribe. | | `to` | `LanguageCode \| None` | - | | `source` | `LanguageCode \| None` | The source term to translate. | | `between` | `tuple[LanguageCode, LanguageCode] \| None` | - | | `filename` | `str \| None` | Filename associated with uploaded file data. | | `model` | `str` | Speech-to-text model to use. | | `client_reference_id` | `str \| None` | Optional tracking identifier. | | `config` | `CreateTranscriptionConfig \| None` | Configuration options for this operation. | **Returns** `Transcription` **Raises** * `SonioxAPIError` When the API returns an error. * `SonioxValidationError` When the translate kwargs are invalid. *** ### translate() ```python translate(*, to: LanguageCode | None = None, source: LanguageCode | None = None, between: tuple[LanguageCode, LanguageCode] | None = None, audio_url: str | None = None, file_id: str | None = None, file: BinaryIO | bytes | Path | str | None = None, filename: str | None = None, model: str = DEFAULT_MODEL, client_reference_id: str | None = None, config: CreateTranscriptionConfig | None = None) -> Transcription ``` Translate audio from a file, file ID, or URL. Provide exactly one of `to` (one-way) or `between` (two-way). `source` is an optional language hint and is only valid with `to`. Convenience over `transcribe()` that fills in the `translation` config and forces `enable_language_identification=True`. **Parameters** | Parameter | Type | Description | | --------------------- | ------------------------------------------- | -------------------------------------------- | | `to` | `LanguageCode \| None` | - | | `source` | `LanguageCode \| None` | The source term to translate. | | `between` | `tuple[LanguageCode, LanguageCode] \| None` | - | | `audio_url` | `str \| None` | Publicly accessible audio URL. | | `file_id` | `str \| None` | ID of a previously uploaded file. | | `file` | `BinaryIO \| bytes \| Path \| str \| None` | File input to upload or transcribe. | | `filename` | `str \| None` | Filename associated with uploaded file data. | | `model` | `str` | Speech-to-text model to use. | | `client_reference_id` | `str \| None` | Optional tracking identifier. | | `config` | `CreateTranscriptionConfig \| None` | Configuration options for this operation. | **Returns** `Transcription` **Raises** * `SonioxAPIError` When the API returns an error. * `SonioxValidationError` When the payload or translate kwargs are invalid. *** ### translate\_and\_wait() ```python translate_and_wait(*, to: LanguageCode | None = None, source: LanguageCode | None = None, between: tuple[LanguageCode, LanguageCode] | None = None, audio_url: str | None = None, file_id: str | None = None, file: BinaryIO | bytes | Path | str | None = None, filename: str | None = None, model: str = DEFAULT_MODEL, client_reference_id: str | None = None, delete_after: bool = False, wait_interval_sec: float = 5.0, wait_timeout_sec: float | None = None, config: CreateTranscriptionConfig | None = None) -> Transcription ``` Translate and wait for completion. Returns the finished `Transcription`. Provide exactly one of `to` (one-way) or `between` (two-way). `source` is an optional language hint and is only valid with `to`. **Parameters** | Parameter | Type | Description | | --------------------- | ------------------------------------------- | ----------------------------------------------------- | | `to` | `LanguageCode \| None` | - | | `source` | `LanguageCode \| None` | The source term to translate. | | `between` | `tuple[LanguageCode, LanguageCode] \| None` | - | | `audio_url` | `str \| None` | Publicly accessible audio URL. | | `file_id` | `str \| None` | ID of a previously uploaded file. | | `file` | `BinaryIO \| bytes \| Path \| str \| None` | File input to upload or transcribe. | | `filename` | `str \| None` | Filename associated with uploaded file data. | | `model` | `str` | Speech-to-text model to use. | | `client_reference_id` | `str \| None` | Optional tracking identifier. | | `delete_after` | `bool` | Whether to delete created resources after completion. | | `wait_interval_sec` | `float` | Polling interval in seconds while waiting. | | `wait_timeout_sec` | `float \| None` | Maximum wait time in seconds while polling. | | `config` | `CreateTranscriptionConfig \| None` | Configuration options for this operation. | **Returns** `Transcription` **Raises** * `SonioxAPIError` When the API returns an error. * `SonioxValidationError` When the payload or translate kwargs are invalid. * `TimeoutError` Waiting for the transcription to finish exceeded `wait_timeout_sec`. *** ### translate\_and\_wait\_with\_tokens() ```python translate_and_wait_with_tokens(*, to: LanguageCode | None = None, source: LanguageCode | None = None, between: tuple[LanguageCode, LanguageCode] | None = None, audio_url: str | None = None, file_id: str | None = None, file: BinaryIO | bytes | Path | str | None = None, filename: str | None = None, model: str = DEFAULT_MODEL, client_reference_id: str | None = None, delete_after: bool = False, wait_interval_sec: float = 5.0, wait_timeout_sec: float | None = None, config: CreateTranscriptionConfig | None = None) -> TranscriptionTranscript ``` Translate, wait for completion, and return the transcript with tokens. Provide exactly one of `to` (one-way) or `between` (two-way). `source` is an optional language hint and is only valid with `to`. Optionally deletes the transcription and uploaded file after completion. **Parameters** | Parameter | Type | Description | | --------------------- | ------------------------------------------- | ----------------------------------------------------- | | `to` | `LanguageCode \| None` | - | | `source` | `LanguageCode \| None` | The source term to translate. | | `between` | `tuple[LanguageCode, LanguageCode] \| None` | - | | `audio_url` | `str \| None` | Publicly accessible audio URL. | | `file_id` | `str \| None` | ID of a previously uploaded file. | | `file` | `BinaryIO \| bytes \| Path \| str \| None` | File input to upload or transcribe. | | `filename` | `str \| None` | Filename associated with uploaded file data. | | `model` | `str` | Speech-to-text model to use. | | `client_reference_id` | `str \| None` | Optional tracking identifier. | | `delete_after` | `bool` | Whether to delete created resources after completion. | | `wait_interval_sec` | `float` | Polling interval in seconds while waiting. | | `wait_timeout_sec` | `float \| None` | Maximum wait time in seconds while polling. | | `config` | `CreateTranscriptionConfig \| None` | Configuration options for this operation. | **Returns** `TranscriptionTranscript` **Raises** * `SonioxAPIError` When the API returns an error. * `SonioxValidationError` When the payload or translate kwargs are invalid. * `TimeoutError` Waiting for the transcription to finish exceeded `wait_timeout_sec`. *** ## AsyncTtsAPI ### Constructor ```python AsyncTtsAPI(client: AsyncSonioxClient) ``` **Parameters** | Parameter | Type | Description | | --------- | ------------------- | ----------------------- | | `client` | `AsyncSonioxClient` | Soniox client instance. | **Returns** `None` ### generate() ```python generate(*, text: str, voice: str, model: str = DEFAULT_MODEL, config: CreateTtsConfig | None = None, language: str | None = None, audio_format: TtsAudioFormat | None = None, sample_rate: TtsSampleRate | None = None, bitrate: TtsBitrate | None = None) -> bytes ``` Generate speech audio from text and return raw audio bytes. Performs a POST request to the TTS REST endpoint. `audio_format`/`sample_rate`/`bitrate` are deprecated; set them on `CreateTtsConfig` instead. Pass `language` explicitly — relying on the default ("en") is deprecated and `language` will be required in the next major release. **Parameters** | Parameter | Type | Description | | -------------- | ------------------------- | --------------------------------------------------------------------------------------------------- | | `text` | `str` | Longer free-form background text, prior interaction history, reference documents, or meeting notes. | | `voice` | `str` | Voice identifier to generate speech audio with. | | `model` | `str` | Speech-to-text model to use. | | `config` | `CreateTtsConfig \| None` | Configuration options for this operation. | | `language` | `str \| None` | Language code for Text-to-Speech (e.g., "en"). | | `audio_format` | `TtsAudioFormat \| None` | Audio format for realtime transcription. | | `sample_rate` | `TtsSampleRate \| None` | Audio sample rate in Hz. | | `bitrate` | `TtsBitrate \| None` | Output bitrate in bits-per-second for compressed formats. | **Returns** `bytes` **Raises** * `SonioxAPIError` When the API returns an error. *** ### generate\_to\_file() ```python generate_to_file(output: BinaryIO | Path | str, *, text: str, voice: str = DEFAULT_VOICE, model: str = DEFAULT_MODEL, config: CreateTtsConfig | None = None, language: str | None = None, audio_format: TtsAudioFormat | None = None, sample_rate: TtsSampleRate | None = None, bitrate: TtsBitrate | None = None) -> int ``` Generate speech audio from text and write the audio bytes to a file-like output. `audio_format`/`sample_rate`/`bitrate` are deprecated; set them on `CreateTtsConfig` instead. Pass `language` explicitly — relying on the default ("en") is deprecated and `language` will be required in the next major release. **Parameters** | Parameter | Type | Description | | -------------- | ------------------------- | --------------------------------------------------------------------------------------------------- | | `output` | `BinaryIO \| Path \| str` | - | | `text` | `str` | Longer free-form background text, prior interaction history, reference documents, or meeting notes. | | `voice` | `str` | Voice identifier to generate speech audio with. | | `model` | `str` | Speech-to-text model to use. | | `config` | `CreateTtsConfig \| None` | Configuration options for this operation. | | `language` | `str \| None` | Language code for Text-to-Speech (e.g., "en"). | | `audio_format` | `TtsAudioFormat \| None` | Audio format for realtime transcription. | | `sample_rate` | `TtsSampleRate \| None` | Audio sample rate in Hz. | | `bitrate` | `TtsBitrate \| None` | Output bitrate in bits-per-second for compressed formats. | **Returns** `int` Number of bytes written. *** ## AsyncTtsModelsAPI ### Constructor ```python AsyncTtsModelsAPI(client: AsyncSonioxClient) ``` **Parameters** | Parameter | Type | Description | | --------- | ------------------- | ----------------------- | | `client` | `AsyncSonioxClient` | Soniox client instance. | **Returns** `None` ### list() ```python list() -> GetTtsModelsResponse ``` List available Text-to-Speech models. Performs a GET request to `/tts-models`. **Returns** `GetTtsModelsResponse` **Raises** * `SonioxAPIError` When the API returns an error. *** ## AsyncModelsAPI ### Constructor ```python AsyncModelsAPI(client: AsyncSonioxClient) ``` **Parameters** | Parameter | Type | Description | | --------- | ------------------- | ----------------------- | | `client` | `AsyncSonioxClient` | Soniox client instance. | **Returns** `None` ### list() ```python list() -> GetModelsResponse ``` List available models. Performs a GET request to `/models`. **Returns** `GetModelsResponse` **Raises** * `SonioxAPIError` When the API returns an error. *** ## AsyncUsageLogsAPI ### Constructor ```python AsyncUsageLogsAPI(client: AsyncSonioxClient) ``` **Parameters** | Parameter | Type | Description | | --------- | ------------------- | ----------------------- | | `client` | `AsyncSonioxClient` | Soniox client instance. | **Returns** `None` ### list() ```python list(start_time: str, end_time: str, limit: int = 1000, sort: UsageLogsSort = 'end_time_asc', cursor: str | None = None) -> GetUsageLogsResponse ``` List usage-log entries for a time window. Performs a GET request to `/usage-logs`. **Parameters** | Parameter | Type | Description | | ------------ | --------------- | ------------------------------------------------------------- | | `start_time` | `str` | Start of the window (inclusive). Filters by request end time. | | `end_time` | `str` | End of the window (exclusive). Filters by request end time. | | `limit` | `int` | Maximum number of entries to return (1–1000). | | `sort` | `UsageLogsSort` | Sort order by end\_time. | | `cursor` | `str \| None` | Pagination cursor for the next page. | **Returns** `GetUsageLogsResponse` **Raises** * `SonioxAPIError` When the API returns an error. *** ### list\_all() ```python list_all(start_time: str, end_time: str, limit: int = 1000, sort: UsageLogsSort = 'end_time_asc') -> AsyncGenerator[UsageLogEntry, None] ``` Iterate through all usage-log entries across all pages. **Parameters** | Parameter | Type | Description | | ------------ | --------------- | ------------------------------------------------------------- | | `start_time` | `str` | Start of the window (inclusive). Filters by request end time. | | `end_time` | `str` | End of the window (exclusive). Filters by request end time. | | `limit` | `int` | Maximum number of entries to return (1–1000). | | `sort` | `UsageLogsSort` | Sort order by end\_time. | **Returns** `AsyncGenerator[UsageLogEntry, None]` *** ## AsyncConcurrencyLimitsAPI ### Constructor ```python AsyncConcurrencyLimitsAPI(client: AsyncSonioxClient) ``` **Parameters** | Parameter | Type | Description | | --------- | ------------------- | ----------------------- | | `client` | `AsyncSonioxClient` | Soniox client instance. | **Returns** `None` ### get() ```python get() -> GetConcurrencyLimitsResponse ``` Get current concurrent sessions and configured limits. Performs a GET request to `/concurrency-limits`. **Returns** `GetConcurrencyLimitsResponse` Project- and organization-scoped current counts and configured limits for realtime STT and TTS sessions. **Raises** * `SonioxAPIError` When the API returns an error. *** ## AsyncAuthAPI ### Constructor ```python AsyncAuthAPI(client: AsyncSonioxClient) ``` **Parameters** | Parameter | Type | Description | | --------- | ------------------- | ----------------------- | | `client` | `AsyncSonioxClient` | Soniox client instance. | **Returns** `None` ### create\_temporary\_api\_key() ```python create_temporary_api_key(*, usage_type: TemporaryApiKeyUsageType = 'transcribe_websocket', expires_in_seconds: int = 5 * 60, client_reference_id: str | None = None, single_use: bool | None = None, max_session_duration_seconds: int | None = None) -> CreateTemporaryApiKeyResponse ``` Create a temporary API key. Performs a POST request to `/auth/temporary-api-key`. **Parameters** | Parameter | Type | Description | | ------------------------------ | -------------------------- | -------------------------------------------------------------------------------------- | | `usage_type` | `TemporaryApiKeyUsageType` | Intended usage of the temporary API key. | | `expires_in_seconds` | `int` | Duration in seconds until the temporary API key expires | | `client_reference_id` | `str \| None` | Optional tracking identifier string. Does not need to be unique | | `single_use` | `bool \| None` | When true, restricts the temporary API key to a single use. | | `max_session_duration_seconds` | `int \| None` | Maximum connection duration in seconds for WebSocket and TTS HTTP streaming endpoints. | **Returns** `CreateTemporaryApiKeyResponse` **Raises** * `SonioxAPIError` When the API returns an error. *** ## AsyncSonioxWebhooksAPI # Realtime Client URL: /sdk/python-SDK/Full-SDK-reference/realtime_client Soniox Python SDK - Realtime Client Reference *** ## RealtimeAPI Entrypoint for realtime helpers on SonioxClient. ### Constructor ```python RealtimeAPI(client: SonioxClient) ``` **Parameters** | Parameter | Type | Description | | --------- | -------------- | ----------------------- | | `client` | `SonioxClient` | Soniox client instance. | **Returns** `None` ### Properties | Property | Type | Description | | -------- | ------------------- | ----------------------------- | | `stt` | `RealtimeSTTClient` | Speech-to-text API namespace. | | `tts` | `RealtimeTTSClient` | Text-to-Speech API namespace | *** ## AsyncRealtimeAPI Entrypoint for async realtime helpers on AsyncSonioxClient. ### Constructor ```python AsyncRealtimeAPI(client: AsyncSonioxClient) ``` **Parameters** | Parameter | Type | Description | | --------- | ------------------- | ----------------------- | | `client` | `AsyncSonioxClient` | Soniox client instance. | **Returns** `None` ### Properties | Property | Type | Description | | -------- | ------------------------ | ----------------------------- | | `stt` | `AsyncRealtimeSTTClient` | Speech-to-text API namespace. | | `tts` | `AsyncRealtimeTTSClient` | Text-to-Speech API namespace | *** ## RealtimeSTTClient Factory for creating synchronous realtime speech-to-text sessions. This class validates credentials and prepares session configuration, but does not itself manage WebSocket connections. ### Constructor ```python RealtimeSTTClient(client: SonioxClient) ``` Create a realtime STT client bound to an existing API client. **Parameters** | Parameter | Type | Description | | --------- | -------------- | ------------------------------------------------------------- | | `client` | `SonioxClient` | Parent Soniox client providing configuration and credentials. | **Returns** `None` ### connect() ```python connect(*, config: RealtimeSTTConfig, api_key: str | None = None, connect_timeout_sec: float = DEFAULT_CONNECT_TIMEOUT_SEC) -> RealtimeSTTSession ``` Create a new realtime STT session. The returned session is not connected until entered as a context manager. **Parameters** | Parameter | Type | Description | | --------------------- | ------------------- | --------------------------------------------------------------------------------- | | `config` | `RealtimeSTTConfig` | Realtime transcription configuration. | | `api_key` | `str \| None` | Optional API key override. If not provided, the client's default API key is used. | | `connect_timeout_sec` | `float` | Maximum seconds to wait for the WebSocket handshake. Defaults to 10 seconds. | **Returns** `RealtimeSTTSession` A new RealtimeSTTSession instance. **Raises** * `SonioxValidationError` If no API key is available. *** ## AsyncRealtimeSTTClient Factory for creating asynchronous realtime speech-to-text sessions. This class validates credentials and prepares session configuration, but does not itself manage WebSocket connections. ### Constructor ```python AsyncRealtimeSTTClient(client: AsyncSonioxClient) ``` Create a realtime STT client bound to an existing API client. **Parameters** | Parameter | Type | Description | | --------- | ------------------- | ------------------------------------------------------------- | | `client` | `AsyncSonioxClient` | Parent Soniox client providing configuration and credentials. | **Returns** `None` ### connect() ```python connect(*, config: RealtimeSTTConfig, api_key: str | None = None, connect_timeout_sec: float = DEFAULT_CONNECT_TIMEOUT_SEC) -> AsyncRealtimeSTTSession ``` Create a new realtime STT session. The returned session is not connected until entered as an async context manager. **Parameters** | Parameter | Type | Description | | --------------------- | ------------------- | --------------------------------------------------------------------------------- | | `config` | `RealtimeSTTConfig` | Realtime transcription configuration. | | `api_key` | `str \| None` | Optional API key override. If not provided, the client's default API key is used. | | `connect_timeout_sec` | `float` | Maximum seconds to wait for the WebSocket handshake. Defaults to 10 seconds. | **Returns** `AsyncRealtimeSTTSession` A new AsyncRealtimeSTTSession instance. **Raises** * `SonioxValidationError` If no API key is available. *** ## RealtimeSTTSession Synchronous WebSocket session for a single real-time speech-to-text stream. This class manages the full lifecycle of a real-time transcription session: connecting to the WebSocket endpoint, streaming audio data, receiving events, and gracefully closing the stream. A session is stateful and represents exactly one streaming interaction with the Soniox realtime API. Instances are designed to be used as context managers. ### Constructor ```python RealtimeSTTSession(url: str, config: RealtimeSTTConfig, *, connect_timeout_sec: float = DEFAULT_CONNECT_TIMEOUT_SEC) ``` Create a new realtime STT session. This does not open a network connection. The WebSocket connection is established when entering the context manager. **Parameters** | Parameter | Type | Description | | --------------------- | ------------------- | ---------------------------------------------------------------------------------------- | | `url` | `str` | WebSocket URL for the realtime transcription endpoint. | | `config` | `RealtimeSTTConfig` | Configuration describing the audio format and transcription behavior for this session. | | `connect_timeout_sec` | `float` | Maximum seconds to wait for the WebSocket handshake to complete. Defaults to 10 seconds. | **Returns** `None` ### Properties | Property | Type | Description | | -------------- | ----------------------- | --------------------------------------------------------- | | `config` | `RealtimeSTTConfig` | Return the configuration used to initialize this session. | | `paused` | `bool` | Return True if the session is currently paused. | | `last_message` | `RealtimeEvent \| None` | Return the most recently received realtime event, if any. | ### close() ```python close() -> None ``` Close the realtime session and release the WebSocket. Signals end-of-audio to the server and clears the underlying connection. Subsequent calls are no-ops. Called automatically when exiting the context manager. **Returns** `None` *** ### send\_byte\_chunk() ```python send_byte_chunk(chunk: bytes) -> None ``` Send a single chunk of raw audio bytes to the realtime stream. The audio data must match the format declared in the session configuration (sample rate, channels, encoding). **Parameters** | Parameter | Type | Description | | --------- | ------- | ------------------------ | | `chunk` | `bytes` | Raw audio bytes to send. | **Returns** `None` **Raises** * `SonioxRealtimeError` If the session is not connected or the send operation fails. *** ### send\_bytes() ```python send_bytes(chunks: bytes | Iterator[bytes], *, finish: bool = True) -> None ``` Send audio data to the realtime stream. Accepts either a single bytes object or an iterator yielding byte chunks (e.g. from `throttle_audio`). If `finish=True` (the default), an end-of-audio signal is sent after the last chunk; pass `finish=False` when you intend to send more audio later in the same session. **Parameters** | Parameter | Type | Description | | --------- | -------------------------- | ------------------------------------------------------------ | | `chunks` | `bytes \| Iterator[bytes]` | Raw bytes or an iterator of byte chunks. | | `finish` | `bool` | If True (default), signal end-of-audio after the last chunk. | **Returns** `None` *** ### send\_control\_message() ```python send_control_message(control_type: RealtimeControlType) -> None ``` Send a control message to the realtime session. Control messages modify the state of the stream, such as signaling end-of-audio or requesting finalization. **Parameters** | Parameter | Type | Description | | -------------- | --------------------- | ------------------------------------ | | `control_type` | `RealtimeControlType` | The type of control message to send. | **Returns** `None` **Raises** * `SonioxRealtimeError` If the session is not connected or the message cannot be sent. *** ### finish() ```python finish() -> None ``` Signal end-of-audio. The server finalizes any pending tokens and closes the connection. Continue iterating `receive_events()` to consume the remaining tokens. **Returns** `None` *** ### keep\_alive() ```python keep_alive() -> None ``` Send a keep-alive message to prevent the session from timing out. **Returns** `None` *** ### finalize() ```python finalize() -> None ``` Finalize all outstanding non-final tokens while keeping the session open. Subsequent tokens will be delivered with `is_final=True`. **Returns** `None` *** ### recv\_bytes() ```python recv_bytes() -> bytes ``` Receive a raw message from the WebSocket connection. **Returns** `bytes` The received message as bytes. An empty bytes object indicates that the connection has been closed. *** ### parse\_event() ```python parse_event(raw: str | bytes) -> RealtimeEvent ``` Parse a raw WebSocket message into a structured realtime event. **Parameters** | Parameter | Type | Description | | --------- | -------------- | --------------------------------------------- | | `raw` | `str \| bytes` | Raw message payload received from the server. | **Returns** `RealtimeEvent` A validated RealtimeEvent instance. *** ### receive\_event() ```python receive_event() -> RealtimeEvent | None ``` Receive and parse the next realtime event from the server. **Returns** `RealtimeEvent | None` The next RealtimeEvent, or None if the connection has closed. **Raises** * `SonioxRealtimeError` If the session is not connected. *** ### receive\_events() ```python receive_events() -> Iterator[RealtimeEvent] ``` Yield realtime events as they are received from the server. Iteration stops automatically when the connection is closed. **Returns** `Iterator[RealtimeEvent]` *** ### handle\_events() ```python handle_events(handler: Callable[[RealtimeEvent], None]) -> None ``` Receive realtime events and dispatch them to a handler callback. **Parameters** | Parameter | Type | Description | | --------- | --------------------------------- | ------------------------------------------------- | | `handler` | `Callable[[RealtimeEvent], None]` | Callable invoked for each received RealtimeEvent. | **Returns** `None` *** ### pause() ```python pause(*, finalize: bool = True) -> None ``` Pause the session, suppressing outgoing audio and starting a background keepalive thread. While paused, calls to `send_byte_chunk` are silently dropped. A background thread sends a keepalive message every `KEEP_ALIVE_INTERVAL_SEC` seconds to prevent the server from timing out the session. Calling `pause` on an already-paused session is a no-op. **Parameters** | Parameter | Type | Description | | ---------- | ------ | ---------------------------------------------------- | | `finalize` | `bool` | If True (default), call `finalize()` before pausing. | **Returns** `None` **Raises** * `SonioxRealtimeError` If the session is not connected. *** ### resume() ```python resume() -> None ``` Resume a paused session, stopping the keepalive thread and allowing audio to be sent again. Calling `resume` on a session that is not paused is a no-op. **Returns** `None` **Raises** * `SonioxRealtimeError` If the session is not connected. *** ## AsyncRealtimeSTTSession Asynchronous WebSocket session for a single real-time speech-to-text stream. This class manages the full lifecycle of a real-time transcription session: connecting to the WebSocket endpoint, streaming audio data, receiving events, and gracefully closing the stream. A session is stateful and represents exactly one streaming interaction with the Soniox realtime API. Instances are designed to be used as async context managers. ### Constructor ```python AsyncRealtimeSTTSession(url: str, config: RealtimeSTTConfig, *, connect_timeout_sec: float = DEFAULT_CONNECT_TIMEOUT_SEC) ``` Create a new realtime STT session. This does not open a network connection. The WebSocket connection is established when entering the async context manager. **Parameters** | Parameter | Type | Description | | --------------------- | ------------------- | ---------------------------------------------------------------------------------------- | | `url` | `str` | WebSocket URL for the realtime transcription endpoint. | | `config` | `RealtimeSTTConfig` | Configuration describing the audio format and transcription behavior for this session. | | `connect_timeout_sec` | `float` | Maximum seconds to wait for the WebSocket handshake to complete. Defaults to 10 seconds. | **Returns** `None` ### Properties | Property | Type | Description | | -------------- | ----------------------- | --------------------------------------------------------- | | `config` | `RealtimeSTTConfig` | Return the configuration used to initialize this session. | | `paused` | `bool` | Return True if the session is currently paused. | | `last_message` | `RealtimeEvent \| None` | Return the most recently received realtime event, if any. | ### close() ```python close() -> None ``` Close the realtime session and release the WebSocket. Signals end-of-audio to the server and clears the underlying connection. Subsequent calls are no-ops. Called automatically when exiting the async context manager. **Returns** `None` *** ### send\_byte\_chunk() ```python send_byte_chunk(chunk: bytes) -> None ``` Send a single chunk of raw audio bytes to the realtime stream. The audio data must match the format declared in the session configuration (sample rate, channels, encoding). **Parameters** | Parameter | Type | Description | | --------- | ------- | ------------------------ | | `chunk` | `bytes` | Raw audio bytes to send. | **Returns** `None` **Raises** * `SonioxRealtimeError` If the session is not connected or the send operation fails. *** ### send\_bytes() ```python send_bytes(chunks: bytes | AsyncIterator[bytes], *, finish: bool = True) -> None ``` Send audio data to the realtime stream. Accepts either a single bytes object or an async iterator yielding byte chunks (e.g. from `throttle_audio_async`). If `finish=True` (the default), an end-of-audio signal is sent after the last chunk; pass `finish=False` when you intend to send more audio later in the same session. **Parameters** | Parameter | Type | Description | | --------- | ------------------------------- | ------------------------------------------------------------ | | `chunks` | `bytes \| AsyncIterator[bytes]` | Raw bytes or an async iterator of byte chunks. | | `finish` | `bool` | If True (default), signal end-of-audio after the last chunk. | **Returns** `None` *** ### send\_control\_message() ```python send_control_message(control_type: RealtimeControlType) -> None ``` Send a control message to the realtime session. Control messages modify the state of the stream, such as signaling end-of-audio or requesting finalization. **Parameters** | Parameter | Type | Description | | -------------- | --------------------- | ------------------------------------ | | `control_type` | `RealtimeControlType` | The type of control message to send. | **Returns** `None` **Raises** * `SonioxRealtimeError` If the session is not connected or the message cannot be sent. *** ### finish() ```python finish() -> None ``` Signal end-of-audio. The server finalizes any pending tokens and closes the connection. Continue iterating `receive_events()` to consume the remaining tokens. **Returns** `None` *** ### keep\_alive() ```python keep_alive() -> None ``` Send a keep-alive message to prevent the session from timing out. **Returns** `None` *** ### finalize() ```python finalize() -> None ``` Finalize all outstanding non-final tokens while keeping the session open. Subsequent tokens will be delivered with `is_final=True`. **Returns** `None` *** ### recv\_bytes() ```python recv_bytes() -> bytes ``` Receive a raw message from the WebSocket connection. **Returns** `bytes` The received message as bytes. An empty bytes object indicates that the connection has been closed. *** ### parse\_event() ```python parse_event(raw: str | bytes) -> RealtimeEvent ``` Parse a raw WebSocket message into a structured realtime event. **Parameters** | Parameter | Type | Description | | --------- | -------------- | --------------------------------------------- | | `raw` | `str \| bytes` | Raw message payload received from the server. | **Returns** `RealtimeEvent` A validated RealtimeEvent instance. *** ### receive\_event() ```python receive_event() -> RealtimeEvent | None ``` Receive and parse the next realtime event from the server. **Returns** `RealtimeEvent | None` The next RealtimeEvent, or None if the connection has closed. **Raises** * `SonioxRealtimeError` If the session is not connected. *** ### receive\_events() ```python receive_events() -> AsyncIterator[RealtimeEvent] ``` Yield realtime events as they are received from the server. Iteration stops automatically when the connection is closed. **Returns** `AsyncIterator[RealtimeEvent]` *** ### handle\_events() ```python handle_events(handler: Callable[[RealtimeEvent], Awaitable[None]]) -> None ``` Receive realtime events and dispatch them to a handler callback. **Parameters** | Parameter | Type | Description | | --------- | -------------------------------------------- | ------------------------------------------------- | | `handler` | `Callable[[RealtimeEvent], Awaitable[None]]` | Callable invoked for each received RealtimeEvent. | **Returns** `None` *** ### pause() ```python pause(*, finalize: bool = True) -> None ``` Pause the session, suppressing outgoing audio and starting a background keepalive task. While paused, calls to `send_byte_chunk` are silently dropped. A background task sends a keepalive message every `KEEP_ALIVE_INTERVAL_SEC` seconds to prevent the server from timing out the session. Calling `pause` on an already-paused session is a no-op. **Parameters** | Parameter | Type | Description | | ---------- | ------ | ---------------------------------------------------- | | `finalize` | `bool` | If True (default), call `finalize()` before pausing. | **Returns** `None` **Raises** * `SonioxRealtimeError` If the session is not connected. *** ### resume() ```python resume() -> None ``` Resume a paused session, stopping the keepalive task and allowing audio to be sent again. Calling `resume` on a session that is not paused is a no-op. **Returns** `None` **Raises** * `SonioxRealtimeError` If the session is not connected. *** ## RealtimeTTSClient Factory for synchronous realtime Text-to-Speech connections and streams. ### Constructor ```python RealtimeTTSClient(client: SonioxClient) ``` **Parameters** | Parameter | Type | Description | | --------- | -------------- | ----------------------- | | `client` | `SonioxClient` | Soniox client instance. | **Returns** `None` ### connect() ```python connect(*, config: RealtimeTTSConfig, api_key: str | None = None, connect_timeout_sec: float = DEFAULT_CONNECT_TIMEOUT_SEC) -> RealtimeTTSConnection ``` Create a single-stream realtime Text-to-Speech connection. **Parameters** | Parameter | Type | Description | | --------------------- | ------------------- | ----------------------------------------- | | `config` | `RealtimeTTSConfig` | Configuration options for this operation. | | `api_key` | `str \| None` | API key used for authentication. | | `connect_timeout_sec` | `float` | - | **Returns** `RealtimeTTSConnection` *** ### connect\_multi\_stream() ```python connect_multi_stream(*, connect_timeout_sec: float = DEFAULT_CONNECT_TIMEOUT_SEC) -> RealtimeTTSMultiplexedConnection ``` Create a multiplexed realtime Text-to-Speech connection. **Parameters** | Parameter | Type | Description | | --------------------- | ------- | ----------- | | `connect_timeout_sec` | `float` | - | **Returns** `RealtimeTTSMultiplexedConnection` *** ## AsyncRealtimeTTSClient Factory for asynchronous realtime Text-to-Speech connections and streams. ### Constructor ```python AsyncRealtimeTTSClient(client: AsyncSonioxClient) ``` **Parameters** | Parameter | Type | Description | | --------- | ------------------- | ----------------------- | | `client` | `AsyncSonioxClient` | Soniox client instance. | **Returns** `None` ### connect() ```python connect(*, config: RealtimeTTSConfig, api_key: str | None = None, connect_timeout_sec: float = DEFAULT_CONNECT_TIMEOUT_SEC) -> AsyncRealtimeTTSConnection ``` Create a single-stream realtime Text-to-Speech connection. **Parameters** | Parameter | Type | Description | | --------------------- | ------------------- | ----------------------------------------- | | `config` | `RealtimeTTSConfig` | Configuration options for this operation. | | `api_key` | `str \| None` | API key used for authentication. | | `connect_timeout_sec` | `float` | - | **Returns** `AsyncRealtimeTTSConnection` *** ### connect\_multi\_stream() ```python connect_multi_stream(*, connect_timeout_sec: float = DEFAULT_CONNECT_TIMEOUT_SEC) -> AsyncRealtimeTTSMultiplexedConnection ``` Create a multiplexed realtime Text-to-Speech connection. **Parameters** | Parameter | Type | Description | | --------------------- | ------- | ----------- | | `connect_timeout_sec` | `float` | - | **Returns** `AsyncRealtimeTTSMultiplexedConnection` *** ## RealtimeTTSConnection Synchronous WebSocket connection for one realtime Text-to-Speech stream. ### Constructor ```python RealtimeTTSConnection(url: str, config: RealtimeTTSConfig, *, connect_timeout_sec: float = DEFAULT_CONNECT_TIMEOUT_SEC) ``` **Parameters** | Parameter | Type | Description | | --------------------- | ------------------- | ----------------------------------------- | | `url` | `str` | WebSocket URL for realtime transcription. | | `config` | `RealtimeTTSConfig` | Configuration options for this operation. | | `connect_timeout_sec` | `float` | - | **Returns** `None` ### Properties | Property | Type | Description | | -------------- | -------------------------- | -------------------------------------------------- | | `config` | `RealtimeTTSConfig` | Configuration used to initialize this connection. | | `paused` | `bool` | Return True if the connection is currently paused. | | `last_message` | `RealtimeTTSEvent \| None` | Most recently received realtime event, if any. | ### close() ```python close() -> None ``` Close the realtime Text-to-Speech connection. **Returns** `None` *** ### send\_text\_chunk() ```python send_text_chunk(text: str, *, text_end: bool = False) -> None ``` Send one text chunk to the realtime stream. **Parameters** | Parameter | Type | Description | | ---------- | ------ | --------------------------------------------------------------- | | `text` | `str` | Text chunk to generate into speech. | | `text_end` | `bool` | Whether this message marks the final text chunk for the stream. | **Returns** `None` *** ### send\_text\_chunks() ```python send_text_chunks(chunks: str | Iterator[str], *, text_end: bool = True) -> None ``` Send text data to the realtime stream. **Parameters** | Parameter | Type | Description | | ---------- | ---------------------- | --------------------------------------------------------------- | | `chunks` | `str \| Iterator[str]` | Audio chunks to stream to realtime transcription. | | `text_end` | `bool` | Whether this message marks the final text chunk for the stream. | **Returns** `None` *** ### finish() ```python finish() -> None ``` Signal that no more text will be sent for this stream. **Returns** `None` *** ### cancel() ```python cancel() -> None ``` Cancel the realtime Text-to-Speech stream. **Returns** `None` *** ### keep\_alive() ```python keep_alive() -> None ``` Send a keep-alive message to prevent the session from timing out. **Returns** `None` *** ### pause() ```python pause() -> None ``` Pause outgoing text and start periodic keep-alive messages. **Returns** `None` *** ### resume() ```python resume() -> None ``` Resume outgoing text and stop periodic keep-alive messages. **Returns** `None` *** ### recv\_bytes() ```python recv_bytes() -> bytes ``` Receive one raw websocket message payload as bytes. **Returns** `bytes` *** ### parse\_event() ```python parse_event(raw: str | bytes) -> RealtimeTTSEvent ``` Parse a raw websocket message into a realtime event. **Parameters** | Parameter | Type | Description | | --------- | -------------- | ---------------------------------------- | | `raw` | `str \| bytes` | Raw event payload from the realtime API. | **Returns** `RealtimeTTSEvent` *** ### receive\_event() ```python receive_event() -> RealtimeTTSEvent | None ``` Receive and parse the next realtime event. **Returns** `RealtimeTTSEvent | None` *** ### receive\_events() ```python receive_events() -> Iterator[RealtimeTTSEvent] ``` Yield realtime events until the stream ends or closes. **Returns** `Iterator[RealtimeTTSEvent]` *** ### receive\_audio\_chunks() ```python receive_audio_chunks() -> Iterator[bytes] ``` Yield decoded audio chunks from incoming realtime events. **Returns** `Iterator[bytes]` *** ## AsyncRealtimeTTSConnection Asynchronous WebSocket connection for one realtime Text-to-Speech stream. ### Constructor ```python AsyncRealtimeTTSConnection(url: str, config: RealtimeTTSConfig, *, connect_timeout_sec: float = DEFAULT_CONNECT_TIMEOUT_SEC) ``` **Parameters** | Parameter | Type | Description | | --------------------- | ------------------- | ----------------------------------------- | | `url` | `str` | WebSocket URL for realtime transcription. | | `config` | `RealtimeTTSConfig` | Configuration options for this operation. | | `connect_timeout_sec` | `float` | - | **Returns** `None` ### Properties | Property | Type | Description | | -------------- | -------------------------- | -------------------------------------------------- | | `config` | `RealtimeTTSConfig` | Configuration used to initialize this connection. | | `paused` | `bool` | Return True if the connection is currently paused. | | `last_message` | `RealtimeTTSEvent \| None` | Most recently received realtime event, if any. | ### close() ```python close() -> None ``` Close the realtime Text-to-Speech connection. **Returns** `None` *** ### send\_text\_chunk() ```python send_text_chunk(text: str, *, text_end: bool = False) -> None ``` Send one text chunk to the realtime stream. **Parameters** | Parameter | Type | Description | | ---------- | ------ | --------------------------------------------------------------- | | `text` | `str` | Text chunk to generate into speech. | | `text_end` | `bool` | Whether this message marks the final text chunk for the stream. | **Returns** `None` *** ### send\_text\_chunks() ```python send_text_chunks(chunks: str | AsyncIterator[str], *, text_end: bool = True) -> None ``` Send text data to the realtime stream. **Parameters** | Parameter | Type | Description | | ---------- | --------------------------- | --------------------------------------------------------------- | | `chunks` | `str \| AsyncIterator[str]` | Audio chunks to stream to realtime transcription. | | `text_end` | `bool` | Whether this message marks the final text chunk for the stream. | **Returns** `None` *** ### finish() ```python finish() -> None ``` Signal that no more text will be sent for this stream. **Returns** `None` *** ### cancel() ```python cancel() -> None ``` Cancel the realtime Text-to-Speech stream. **Returns** `None` *** ### keep\_alive() ```python keep_alive() -> None ``` Send a keep-alive message to prevent the session from timing out. **Returns** `None` *** ### pause() ```python pause() -> None ``` Pause outgoing text and start periodic keep-alive messages. **Returns** `None` *** ### resume() ```python resume() -> None ``` Resume outgoing text and stop periodic keep-alive messages. **Returns** `None` *** ### recv\_bytes() ```python recv_bytes() -> bytes ``` Receive one raw websocket message payload as bytes. **Returns** `bytes` *** ### parse\_event() ```python parse_event(raw: str | bytes) -> RealtimeTTSEvent ``` Parse a raw websocket message into a realtime event. **Parameters** | Parameter | Type | Description | | --------- | -------------- | ---------------------------------------- | | `raw` | `str \| bytes` | Raw event payload from the realtime API. | **Returns** `RealtimeTTSEvent` *** ### receive\_event() ```python receive_event() -> RealtimeTTSEvent | None ``` Receive and parse the next realtime event. **Returns** `RealtimeTTSEvent | None` *** ### receive\_events() ```python receive_events() -> AsyncIterator[RealtimeTTSEvent] ``` Yield realtime events until the stream ends or closes. **Returns** `AsyncIterator[RealtimeTTSEvent]` *** ### receive\_audio\_chunks() ```python receive_audio_chunks() -> AsyncIterator[bytes] ``` Yield decoded audio chunks from incoming realtime events. **Returns** `AsyncIterator[bytes]` *** ### handle\_events() ```python handle_events(handler: Callable[[RealtimeTTSEvent], Awaitable[None]]) -> None ``` Receive events and pass each one to `handler`. **Parameters** | Parameter | Type | Description | | --------- | ----------------------------------------------- | ------------------------------------------------------------------ | | `handler` | `Callable[[RealtimeTTSEvent], Awaitable[None]]` | Event payload received from the realtime Text-to-Speech websocket. | **Returns** `None` *** ## RealtimeTTSMultiplexedConnection Synchronous websocket connection that can host multiple Text-to-Speech streams. ### Constructor ```python RealtimeTTSMultiplexedConnection(url: str, api_key: str, *, connect_timeout_sec: float = DEFAULT_CONNECT_TIMEOUT_SEC) ``` **Parameters** | Parameter | Type | Description | | --------------------- | ------- | ----------------------------------------- | | `url` | `str` | WebSocket URL for realtime transcription. | | `api_key` | `str` | API key used for authentication. | | `connect_timeout_sec` | `float` | - | **Returns** `None` ### Properties | Property | Type | Description | | -------------- | -------------------------- | -------------------------------------------------- | | `last_message` | `RealtimeTTSEvent \| None` | Most recently received realtime event, if any. | | `paused` | `bool` | Return True if the connection is currently paused. | ### close() ```python close() -> None ``` Close the websocket and clear the stream state. **Returns** `None` *** ### keep\_alive() ```python keep_alive() -> None ``` Send a keep-alive message to prevent the session from timing out. **Returns** `None` *** ### pause() ```python pause() -> None ``` Pause outgoing text and start periodic keep-alive messages. **Returns** `None` *** ### resume() ```python resume() -> None ``` Resume outgoing text and stop periodic keep-alive messages. **Returns** `None` *** ### open\_stream() ```python open_stream(*, config: RealtimeTTSConfig) -> RealtimeTTSStream ``` Register and start a new stream on the shared websocket. **Parameters** | Parameter | Type | Description | | --------- | ------------------- | ----------------------------------------- | | `config` | `RealtimeTTSConfig` | Configuration options for this operation. | **Returns** `RealtimeTTSStream` *** ## AsyncRealtimeTTSMultiplexedConnection Asynchronous websocket connection that can host multiple TTS streams. ### Constructor ```python AsyncRealtimeTTSMultiplexedConnection(url: str, api_key: str, *, connect_timeout_sec: float = DEFAULT_CONNECT_TIMEOUT_SEC) ``` **Parameters** | Parameter | Type | Description | | --------------------- | ------- | ----------------------------------------- | | `url` | `str` | WebSocket URL for realtime transcription. | | `api_key` | `str` | API key used for authentication. | | `connect_timeout_sec` | `float` | - | **Returns** `None` ### Properties | Property | Type | Description | | -------------- | -------------------------- | -------------------------------------------------- | | `last_message` | `RealtimeTTSEvent \| None` | Most recently received realtime event, if any. | | `paused` | `bool` | Return True if the connection is currently paused. | ### close() ```python close() -> None ``` Close the websocket and clear the stream state. **Returns** `None` *** ### keep\_alive() ```python keep_alive() -> None ``` Send a keep-alive message to prevent the session from timing out. **Returns** `None` *** ### pause() ```python pause() -> None ``` Pause outgoing text and start periodic keep-alive messages. **Returns** `None` *** ### resume() ```python resume() -> None ``` Resume outgoing text and stop periodic keep-alive messages. **Returns** `None` *** ### open\_stream() ```python open_stream(*, config: RealtimeTTSConfig) -> AsyncRealtimeTTSStream ``` Register and start a new stream on the shared websocket. **Parameters** | Parameter | Type | Description | | --------- | ------------------- | ----------------------------------------- | | `config` | `RealtimeTTSConfig` | Configuration options for this operation. | **Returns** `AsyncRealtimeTTSStream` *** ## RealtimeTTSStream Handle for one stream on a multiplexed realtime TTS connection. ### Constructor ```python RealtimeTTSStream(connection: RealtimeTTSMultiplexedConnection, config: RealtimeTTSConfig) ``` **Parameters** | Parameter | Type | Description | | ------------ | ---------------------------------- | ------------------------------------------------------------------------------- | | `connection` | `RealtimeTTSMultiplexedConnection` | Synchronous websocket connection that can host multiple Text-to-Speech streams. | | `config` | `RealtimeTTSConfig` | Configuration options for this operation. | **Returns** `None` ### Properties | Property | Type | Description | | -------------- | -------------------------- | ----------------------------------------------------- | | `config` | `RealtimeTTSConfig` | Stream configuration. | | `stream_id` | `str` | Stream identifier. | | `last_message` | `RealtimeTTSEvent \| None` | Most recently received event for this stream, if any. | ### send\_text\_chunk() ```python send_text_chunk(text: str, *, text_end: bool = False) -> None ``` Send one text chunk for this stream. **Parameters** | Parameter | Type | Description | | ---------- | ------ | --------------------------------------------------------------- | | `text` | `str` | Text chunk to generate into speech. | | `text_end` | `bool` | Whether this message marks the final text chunk for the stream. | **Returns** `None` *** ### send\_text\_chunks() ```python send_text_chunks(chunks: str | Iterator[str], *, text_end: bool = True) -> None ``` Send text chunks for this stream. **Parameters** | Parameter | Type | Description | | ---------- | ---------------------- | --------------------------------------------------------------- | | `chunks` | `str \| Iterator[str]` | Audio chunks to stream to realtime transcription. | | `text_end` | `bool` | Whether this message marks the final text chunk for the stream. | **Returns** `None` *** ### finish() ```python finish() -> None ``` Signal that no more text will be sent for this stream. **Returns** `None` *** ### cancel() ```python cancel() -> None ``` Cancel this stream. **Returns** `None` *** ### keep\_alive() ```python keep_alive() -> None ``` Send a keepalive message on the underlying shared connection. **Returns** `None` *** ### pause() ```python pause() -> None ``` Pause the underlying shared connection and start keepalive. **Returns** `None` *** ### resume() ```python resume() -> None ``` Resume the underlying shared connection and stop keepalive. **Returns** `None` *** ### receive\_event() ```python receive_event() -> RealtimeTTSEvent | None ``` Receive the next event for this stream. **Returns** `RealtimeTTSEvent | None` *** ### receive\_events() ```python receive_events() -> Iterator[RealtimeTTSEvent] ``` Yield events for this stream until it ends. **Returns** `Iterator[RealtimeTTSEvent]` *** ### receive\_audio\_chunks() ```python receive_audio_chunks() -> Iterator[bytes] ``` Yield decoded audio chunks for this stream. **Returns** `Iterator[bytes]` *** ## AsyncRealtimeTTSStream Handle for one stream on a multiplexed realtime TTS connection. ### Constructor ```python AsyncRealtimeTTSStream(connection: AsyncRealtimeTTSMultiplexedConnection, config: RealtimeTTSConfig) ``` **Parameters** | Parameter | Type | Description | | ------------ | --------------------------------------- | --------------------------------------------------------------------- | | `connection` | `AsyncRealtimeTTSMultiplexedConnection` | Asynchronous websocket connection that can host multiple TTS streams. | | `config` | `RealtimeTTSConfig` | Configuration options for this operation. | **Returns** `None` ### Properties | Property | Type | Description | | -------------- | -------------------------- | ----------------------------------------------------- | | `config` | `RealtimeTTSConfig` | Stream configuration. | | `stream_id` | `str` | Stream identifier. | | `last_message` | `RealtimeTTSEvent \| None` | Most recently received event for this stream, if any. | ### send\_text\_chunk() ```python send_text_chunk(text: str, *, text_end: bool = False) -> None ``` Send one text chunk for this stream. **Parameters** | Parameter | Type | Description | | ---------- | ------ | --------------------------------------------------------------- | | `text` | `str` | Text chunk to generate into speech. | | `text_end` | `bool` | Whether this message marks the final text chunk for the stream. | **Returns** `None` *** ### send\_text\_chunks() ```python send_text_chunks(chunks: str | AsyncIterator[str], *, text_end: bool = True) -> None ``` Send text chunks for this stream. **Parameters** | Parameter | Type | Description | | ---------- | --------------------------- | --------------------------------------------------------------- | | `chunks` | `str \| AsyncIterator[str]` | Audio chunks to stream to realtime transcription. | | `text_end` | `bool` | Whether this message marks the final text chunk for the stream. | **Returns** `None` *** ### finish() ```python finish() -> None ``` Signal that no more text will be sent for this stream. **Returns** `None` *** ### cancel() ```python cancel() -> None ``` Cancel this stream. **Returns** `None` *** ### keep\_alive() ```python keep_alive() -> None ``` Send a keepalive message on the underlying shared connection. **Returns** `None` *** ### pause() ```python pause() -> None ``` Pause the underlying shared connection and start keepalive. **Returns** `None` *** ### resume() ```python resume() -> None ``` Resume the underlying shared connection and stop keepalive. **Returns** `None` *** ### receive\_event() ```python receive_event() -> RealtimeTTSEvent | None ``` Receive the next event for this stream. **Returns** `RealtimeTTSEvent | None` *** ### receive\_events() ```python receive_events() -> AsyncIterator[RealtimeTTSEvent] ``` Yield events for this stream until it ends. **Returns** `AsyncIterator[RealtimeTTSEvent]` *** ### receive\_audio\_chunks() ```python receive_audio_chunks() -> AsyncIterator[bytes] ``` Yield decoded audio chunks for this stream. **Returns** `AsyncIterator[bytes]` # Types URL: /sdk/python-SDK/Full-SDK-reference/types Soniox Python SDK - Types Reference *** ## Token Token metadata emitted during realtime streaming transcriptions. ### Properties | Property | Type | Description | | -------------------- | --------------- | ------------------------------------------------------------ | | `text` | `str` | The transcribed text. | | `start_ms` | `int \| None` | Start time in milliseconds relative to audio start. | | `end_ms` | `int \| None` | End time in milliseconds relative to audio start. | | `confidence` | `float \| None` | Confidence score (0.0 to 1.0). | | `is_final` | `bool \| None` | Whether this is a finalized token. | | `speaker` | `str \| None` | Speaker identifier (if diarization enabled). | | `translation_status` | `str \| None` | Translation status of this token. | | `language` | `str \| None` | Detected language code (if language identification enabled). | | `source_language` | `str \| None` | Source language for translated tokens. | *** ## ApiError Structured representation of a non-2xx API response payload. ### Properties | Property | Type | Description | | ------------------- | ------------------------------- | ------------------------------------------------------------------------------------------ | | `status_code` | `int` | HTTP status code. | | `error_type` | `str` | High-level error code (e.g., 'bad\_request', 'quota\_exceeded') for programmatic handling. | | `message` | `str` | Detailed error message describing the failure. | | `validation_errors` | `list[ApiErrorValidationError]` | List of specific field validation failures, if applicable. | | `request_id` | `str \| None` | Unique identifier for the request, useful for troubleshooting. | | `more_info` | `str \| None` | Optional URL pointing to documentation for resolving this error. | *** ## ApiErrorValidationError Details a single validation error reported by the Soniox API. ### Properties | Property | Type | Description | | ------------ | ----- | -------------------------------------------------------- | | `error_type` | `str` | The category of validation error. | | `location` | `str` | The location of the error, e.g. \['body', 'audio\_url']. | | `message` | `str` | A human-readable description of the validation failure. | *** ## CreateTemporaryApiKeyPayload Payload for requesting a temporary API key (e.g., websocket). ### Properties | Property | Type | Description | | ------------------------------ | -------------------------- | -------------------------------------------------------------------------------------- | | `usage_type` | `TemporaryApiKeyUsageType` | Intended usage of the temporary API key. | | `expires_in_seconds` | `int` | Duration in seconds until the temporary API key expires | | `client_reference_id` | `str \| None` | Optional tracking identifier string. Does not need to be unique | | `single_use` | `bool \| None` | When true, restricts the temporary API key to a single use. | | `max_session_duration_seconds` | `int \| None` | Maximum connection duration in seconds for WebSocket and TTS HTTP streaming endpoints. | *** ## CreateTemporaryApiKeyResponse Response data for a temp API key request. ### Properties | Property | Type | Description | | ------------ | ---------- | --------------------------------------------------------------------- | | `api_key` | `str` | Created temporary API key. | | `expires_at` | `datetime` | UTC timestamp indicating when generated temporary API key will expire | *** ## CreateTtsPayload Payload sent to generate speech audio from text via REST. ### Properties | Property | Type | Description | | -------------- | ----------------------- | ------------------------------------------------------------------------ | | `model` | `str` | Text-to-Speech model to use. | | `language` | `str` | Language code for Text-to-Speech (e.g., "en"). | | `voice` | `str` | Voice identifier to generate speech audio with. | | `audio_format` | `TtsAudioFormat` | Requested output audio format. | | `sample_rate` | `TtsSampleRate \| None` | Output sample rate in Hz. | | `bitrate` | `TtsBitrate \| None` | Output bitrate in bits-per-second for compressed formats. | | `speed` | `float \| None` | Speaking rate multiplier from 0.7 to 1.3; 1.0 (default) is normal speed. | | `text` | `str` | Input text to generate into speech. | *** ## ConcurrencyCurrentValues Live counts of concurrent sessions. ### Properties | Property | Type | Description | | ----------------------- | ----- | ------------------------------------------------------------ | | `transcribe_concurrent` | `int` | Number of concurrent realtime STT sessions currently active. | | `tts_concurrent` | `int` | Number of concurrent realtime TTS sessions currently active. | *** ## ConcurrencyLimitValues Configured concurrency limits. None means no limit. ### Properties | Property | Type | Description | | ----------------------- | ------------- | --------------------------------------------------------------- | | `transcribe_concurrent` | `int \| None` | Maximum concurrent realtime STT sessions, or None if unlimited. | | `tts_concurrent` | `int \| None` | Maximum concurrent realtime TTS sessions, or None if unlimited. | *** ## ConcurrencyScopeValues Current and limit values for a single scope (project or organization). ### Properties | Property | Type | Description | | --------- | -------------------------- | ------------------------------- | | `current` | `ConcurrencyCurrentValues` | Live counts of active sessions. | | `limits` | `ConcurrencyLimitValues` | Configured limits. | *** ## CreateTtsConfig Helper config used when building Text-to-Speech payloads. ### Properties | Property | Type | Description | | -------------- | ------------------------ | ------------------------------------------------------------------------ | | `model` | `str \| None` | Deprecated: pass `model` to generate()/generate\_to\_file() instead. | | `language` | `str \| None` | Deprecated: pass `language` to generate()/generate\_to\_file() instead. | | `voice` | `str \| None` | Deprecated: pass `voice` to generate()/generate\_to\_file() instead. | | `audio_format` | `TtsAudioFormat \| None` | Requested output audio format. | | `sample_rate` | `TtsSampleRate \| None` | Output sample rate in Hz. | | `bitrate` | `TtsBitrate \| None` | Output bitrate in bits-per-second for compressed formats. | | `speed` | `float \| None` | Speaking rate multiplier from 0.7 to 1.3; 1.0 (default) is normal speed. | *** ## CreateTranscriptionPayload Payload sent to create an asynchronous transcription job. ### Properties | Property | Type | Description | | -------------------------------- | -------------------------------- | ------------------------------------------------------------------------------------------------- | | `model` | `str` | Speech-to-text model to use. | | `audio_url` | `str \| None` | URL of a publicly accessible audio file. | | `file_id` | `str \| None` | ID of a previously uploaded file (UUID). | | `language_hints` | `list[LanguageCode] \| None` | Array of expected ISO language codes to bias recognition. | | `language_hints_strict` | `bool \| None` | When true, model relies more heavily on language hints (best results with one language hint set). | | `enable_speaker_diarization` | `bool \| None` | Enable speaker diarization to identify different speakers. | | `enable_language_identification` | `bool \| None` | Enable automatic language identification. | | `translation` | `TranslationConfigInput \| None` | Translation configuration. | | `context` | `StructuredContextInput \| None` | Additional context to improve transcription accuracy and formatting of specialized terms. | | `webhook_url` | `str \| None` | URL to receive webhook notifications when transcription is completed or fails. | | `webhook_auth_header_name` | `str \| None` | Name of the authentication header sent with webhook notifications | | `webhook_auth_header_value` | `str \| None` | Authentication header value sent with webhook notifications. | | `client_reference_id` | `str \| None` | Optional tracking identifier. | *** ## CreateTranscriptionConfig Helper config used when building transcription payloads. ### Properties | Property | Type | Description | | -------------------------------- | -------------------------------- | ----------------------------------------------------------------------------------------- | | `model` | `str \| None` | Deprecated: pass `model` to the create call instead. | | `language_hints` | `list[LanguageCode] \| None` | Array of expected ISO language codes to bias recognition. | | `language_hints_strict` | `bool \| None` | When true, model relies more heavily on language hints. | | `enable_speaker_diarization` | `bool \| None` | Enable speaker diarization to identify different speakers. | | `enable_language_identification` | `bool \| None` | Enable automatic language identification | | `translation` | `TranslationConfigInput \| None` | Translation configuration | | `context` | `StructuredContextInput \| None` | Additional context to improve transcription accuracy and formatting of specialized terms. | | `webhook_url` | `str \| None` | URL to receive webhook notifications when transcription is completed or fails. | | `webhook_auth_header_name` | `str \| None` | Name of the authentication header sent with webhook notifications | | `webhook_auth_header_value` | `str \| None` | Authentication header value sent with webhook notifications | | `client_reference_id` | `str \| None` | Deprecated: pass `client_reference_id` to the create call instead. | *** ## File Metadata describing an uploaded file in the Soniox API. ### Properties | Property | Type | Description | | --------------------- | ------------- | ---------------------------------------------------- | | `id` | `str` | Unique identifier of the file (UUID). | | `filename` | `str` | Name of the file. | | `size` | `int` | Size of the file in bytes. | | `created_at` | `datetime` | UTC timestamp indicating when the file was uploaded. | | `client_reference_id` | `str \| None` | Optional tracking identifier string. | *** ## GetConcurrencyLimitsResponse Response returned when fetching concurrency limits. ### Properties | Property | Type | Description | | -------------- | ------------------------ | --------------------------------------------------------- | | `project` | `ConcurrencyScopeValues` | Project-scoped current counts and configured limits. | | `organization` | `ConcurrencyScopeValues` | Organization-scoped current counts and configured limits. | *** ## GetFilesCountResponse Breakdown of uploaded file counts by source. ### Properties | Property | Type | Description | | ------------ | ----- | -------------------------------------------- | | `total` | `int` | Total number of files across all sources. | | `public_api` | `int` | Number of files uploaded via Public API. | | `playground` | `int` | Number of files uploaded via the Playground. | *** ## GetFilesPayload Parameters accepted by the file listing endpoint. ### Properties | Property | Type | Description | | -------- | ------------- | ----------------------------------------------- | | `limit` | `int` | Maximum number of files to return. | | `cursor` | `str \| None` | Pagination cursor for the next page of results. | *** ## GetFilesResponse Paginated response returned when listing uploaded files. ### Properties | Property | Type | Description | | ------------------ | ------------- | ------------------------------------------------------------------------------------------------------------ | | `files` | `list[File]` | List of uploaded files. | | `next_page_cursor` | `str \| None` | A pagination token that references the next page of results. When None, no additional results are available. | *** ## GetModelsResponse Response returned when listing available models. ### Properties | Property | Type | Description | | -------- | ------------- | ----------------------------- | | `models` | `list[Model]` | List of all available models. | *** ## GetTTSModelsResponse ```python GetTTSModelsResponse = GetTtsModelsResponse ``` *** ## GetTtsModelsResponse Response returned when listing available Text-to-Speech models. ### Properties | Property | Type | Description | | -------- | ---------------- | ---------------------------------------- | | `models` | `list[TtsModel]` | List of available Text-to-Speech models. | *** ## GetTranscriptionsCountResponse Breakdown of transcription counts by scope. ### Properties | Property | Type | Description | | ------------ | ----- | ---------------------------------------------------- | | `total` | `int` | Total number of transcriptions across all scopes. | | `public_api` | `int` | Number of transcriptions created via Public API. | | `playground` | `int` | Number of transcriptions created via the Playground. | *** ## GetTranscriptionsPayload Parameters for listing transcription jobs. ### Properties | Property | Type | Description | | -------- | ------------- | ----------------------------------------------- | | `limit` | `int` | Maximum number of transcriptions to return. | | `cursor` | `str \| None` | Pagination cursor for the next page of results. | *** ## GetTranscriptionsResponse Paginated response for transcription listings. ### Properties | Property | Type | Description | | ------------------ | --------------------- | ------------------------------------------------------------------------------------------------------------ | | `transcriptions` | `list[Transcription]` | List of transcriptions. | | `next_page_cursor` | `str \| None` | A pagination token that references the next page of results. When None, no additional results are available. | *** ## GetUsageLogsPayload Parameters accepted by the usage logs listing endpoint. ### Properties | Property | Type | Description | | ------------ | --------------- | ---------------------------------------------------------------------------------- | | `start_time` | `str` | Start of the time window (inclusive). Filters by request end time. | | `end_time` | `str` | End of the time window (exclusive). Filters by request end time. | | `limit` | `int` | Maximum number of usage log entries to return. | | `sort` | `UsageLogsSort` | Sort order by end\_time. Use `end_time_desc` to get the most recent entries first. | | `cursor` | `str \| None` | Pagination cursor for the next page of results. | *** ## GetUsageLogsResponse Paginated response for usage-log listings. ### Properties | Property | Type | Description | | ------------------ | --------------------- | ---------------------------------------------------------------------- | | `usage_logs` | `list[UsageLogEntry]` | Per-request usage log entries ordered by end\_time. | | `next_page_cursor` | `str \| None` | Pagination cursor for the next page of results. None if no more pages. | *** ## GetVoicesCountResponse Total number of voices in the project. ### Properties | Property | Type | Description | | -------- | ----- | -------------------------------------- | | `total` | `int` | Total number of voices in the project. | *** ## GetVoicesPayload Parameters for listing voices. ### Properties | Property | Type | Description | | -------- | ------------- | ----------------------------------------------- | | `limit` | `int` | Maximum number of voices to return. | | `cursor` | `str \| None` | Pagination cursor for the next page of results. | *** ## GetVoicesResponse Response returned when listing voices. ### Properties | Property | Type | Description | | ------------------ | ------------- | ---------------------------------------------------------------------------- | | `voices` | `list[Voice]` | List of voices. | | `next_page_cursor` | `str \| None` | Pagination token for the next page of results, or None when no more results. | *** ## Language Deprecated alias for :class:`SupportedLanguage`. *** ## LanguageCode ```python LanguageCode = Annotated[str, Field(min_length=2, max_length=2)] ``` ISO 639-1 two-letter language code (e.g. `"en"`, `"fr"`). *** ## SupportedLanguage Represents a supported language for transcription or translation. ### Properties | Property | Type | Description | | -------- | ----- | ------------------------------------ | | `code` | `str` | 2-letter language code (ISO format). | | `name` | `str` | Language name. | *** ## Model Describes a Soniox transcription model. ### Properties | Property | Type | Description | | --------------------------------------- | ------------------------- | ------------------------------------------------------------------------------------------------------- | | `id` | `str` | Unique identifier of the model. | | `aliased_model_id` | `str \| None` | If this is an alias, the id of the aliased model. None for non-alias models. | | `name` | `str` | Name of the model. | | `context_version` | `int \| None` | Version of context supported. | | `transcription_mode` | `TranscriptionMode` | Transcription mode of the model. | | `languages` | `list[SupportedLanguage]` | List of languages supported by the model. | | `supports_language_hints_strict` | `bool` | If model supports 'language\_hints\_strict' option. | | `supports_max_endpoint_delay` | `bool` | If model supports 'max\_endpoint\_delay\_ms' option. | | `supports_endpoint_sensitivity` | `bool` | If model supports the 'endpoint\_sensitivity' option. | | `supports_endpoint_latency_adjustment` | `bool` | If model supports the 'endpoint\_latency\_adjustment\_level' option. | | `endpoint_latency_adjustment_max_level` | `int` | Maximum endpoint\_latency\_adjustment\_level the model accepts (0 means unsupported). | | `translation_targets` | `list[TranslationTarget]` | List of supported one-way translation targets. If list is empty, check for one\_way\_translation field. | | `two_way_translation_pairs` | `list[str]` | List of supported two-way translation pairs. If list is empty, check for one\_way\_translation field. | | `one_way_translation` | `str \| None` | When contains string 'all\_languages', any language from languages can be used | | `two_way_translation` | `str \| None` | When contains string 'all\_languages',' any language pair from languages can be used | *** ## RealtimeSTTAudioFormat ```python RealtimeSTTAudioFormat = Literal["auto"] | RealtimeSTTHeaderFormat | RealtimeSTTRawFormat ``` Audio formats accepted by the realtime STT websocket. *** ## RealtimeSTTHeaderFormat ```python RealtimeSTTHeaderFormat = Literal[ "aac", "aiff", "amr", "asf", "flac", "mp3", "ogg", "wav", "webm", ] ``` Container formats whose header carries sample rate and channels. *** ## RealtimeSTTRawFormat ```python RealtimeSTTRawFormat = Literal[ "pcm_s8", "pcm_s16le", "pcm_s16be", "pcm_s24le", "pcm_s24be", "pcm_s32le", "pcm_s32be", "pcm_u8", "pcm_u16le", "pcm_u16be", "pcm_u24le", "pcm_u24be", "pcm_u32le", "pcm_u32be", "pcm_f32le", "pcm_f32be", "pcm_f64le", "pcm_f64be", "mulaw", "alaw", ] ``` Raw formats with no header - require `sample_rate` and `num_channels`. *** ## RecomputeVoicePayload Body for preparing a voice for additional models. ### Properties | Property | Type | Description | | -------- | ------------- | ----------------------------------------------------------------------------------- | | `model` | `str \| None` | Model to prepare the voice for. If None, prepares it for every not-yet-ready model. | *** ## StructuredContext Optional structured context provided to the transcription engine. For ergonomics, `general` and `translation_terms` also accept a plain dict in addition to the typed item lists: * `general={"domain": "Healthcare"}` (dict of key -> value) * `translation_terms={"Mr. Smith": "Sr. Smith"}` (dict of source -> target) ### Properties | Property | Type | Description | | ------------------- | ---------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------- | | `general` | `Annotated[StructuredContextGeneralInput, Field(union_mode='left_to_right')] \| None` | Structured key-value pairs describing domain, topic, intent, participant names, etc. | | `text` | `str \| None` | Longer free-form background text, prior interaction history, reference documents, or meeting notes. | | `terms` | `list[str] \| None` | Domain-specific or uncommon words to recognize. | | `translation_terms` | `Annotated[StructuredContextTranslationTermsInput, Field(union_mode='left_to_right')] \| None` | Custom translations for ambiguous terms. | *** ## StructuredContextGeneralInput ```python StructuredContextGeneralInput = list[StructuredContextGeneralItem] | dict[str, str] ``` Accepted input shapes for `StructuredContext.general`. *** ## StructuredContextGeneralItem Single general context key/value pair for transcription context. ### Properties | Property | Type | Description | | -------- | ----- | ------------------------------------------------------------------------ | | `key` | `str` | The key describing the context type (e.g., "domain", "topic", "doctor"). | | `value` | `str` | The value for the context key. | *** ## StructuredContextInput ```python StructuredContextInput = StructuredContext | dict[str, Any] ``` Accepted input for the `context` field - typed object or a plain dict. *** ## StructuredContextTranslationTerm Defines a translation term mapping used in structured context. ### Properties | Property | Type | Description | | -------- | ----- | ------------------------------------ | | `source` | `str` | The source term to translate. | | `target` | `str` | The target translation for the term. | *** ## StructuredContextTranslationTermsInput ```python StructuredContextTranslationTermsInput = list[StructuredContextTranslationTerm] | dict[str, str] ``` Accepted input shapes for `StructuredContext.translation_terms`. *** ## Transcription Represents a transcription job tracked by Soniox. ### Properties | Property | Type | Description | | -------------------------------- | --------------------- | -------------------------------------------------------------------------------------------- | | `id` | `str` | Unique identifier of the transcription (UUID). | | `status` | `TranscriptionStatus` | Current status of the transcription. | | `created_at` | `datetime` | UTC timestamp when the transcription was created. | | `model` | `str` | Speech-to-text model used. | | `audio_url` | `str \| None` | URL of the audio file being transcribed. | | `file_id` | `str \| None` | ID of the uploaded file being transcribed (UUID). | | `filename` | `str` | Name of the file being transcribed. | | `language_hints` | `list[str] \| None` | Expected languages in the audio. If not specified, languages are automatically detected. | | `enable_speaker_diarization` | `bool` | When true, speakers are identified and separated in the transcription output. | | `enable_language_identification` | `bool` | When true, language is detected for each part of the transcription. | | `audio_duration_ms` | `int \| None` | Duration of the audio in milliseconds. Only available after processing begins. | | `error_type` | `str \| None` | Error type if transcription failed. None for successful or in-progress transcriptions. | | `error_message` | `str \| None` | Error message if transcription failed. None for successful or in-progress transcriptions. | | `webhook_url` | `str \| None` | URL to receive webhook notifications when transcription is completed or fails. | | `webhook_auth_header_name` | `str \| None` | Name of the authentication header sent with webhook notifications. | | `webhook_auth_header_value` | `str \| None` | Authentication header value. Always returned masked. | | `webhook_status_code` | `int \| None` | HTTP status code received from your server when webhook was delivered. None if not yet sent. | | `client_reference_id` | `str \| None` | Optional tracking identifier. | *** ## TranscriptionStatus ```python TranscriptionStatus = Literal["queued", "processing", "completed", "error"] ``` Current status of the transcription job. *** ## TranscriptionTranscript Transcript data including the full text and tokens. ### Properties | Property | Type | Description | | -------- | ------------- | ------------------------------------------------------------------------- | | `id` | `str` | Unique identifier of the transcription this transcript belongs to (UUID). | | `text` | `str` | Complete transcribed text content. | | `tokens` | `list[Token]` | List of detailed token information with timestamps and metadata. | *** ## TranslationConfig Configuration describing how translation should be performed. ### Properties | Property | Type | Description | | ----------------- | ---------------------- | ------------------------------------------------------------------------- | | `type` | `TranslationType` | Translation type. | | `target_language` | `LanguageCode \| None` | Target language code for translation (e.g., "fr", "es", "de") (one\_way). | | `language_a` | `LanguageCode \| None` | First language code (two\_way). | | `language_b` | `LanguageCode \| None` | Second language code (two\_way). | ### validate\_logic() ```python validate_logic() -> TranslationConfig ``` **Returns** `TranslationConfig` *** ## TranslationConfigInput ```python TranslationConfigInput = TranslationConfig | dict[str, Any] ``` Accepted input for the `translation` field - typed object or a plain dict. *** ## TranslationTarget Describes translation targets offered by a model. ### Properties | Property | Type | Description | | -------------------------- | ----------- | ------------------------------------------------------------------------- | | `target_language` | `str` | Target language code for translation (e.g., "fr", "es", "de") (one\_way). | | `source_languages` | `list[str]` | List of source language codes. | | `exclude_source_languages` | `list[str]` | Source language codes excluded for this target. | *** ## TranslationType ```python TranslationType = Literal["one_way", "two_way"] ``` Supported translation configuration types. *** ## TTSModel ```python TTSModel = TtsModel ``` *** ## TTSVoice ```python TTSVoice = TtsVoice ``` *** ## TtsAudioFormat ```python TtsAudioFormat = Literal[ "pcm_f32le", "pcm_s16le", "pcm_mulaw", "pcm_alaw", "wav", "aac", "mp3", "opus", "flac", ] ``` Allowed audio formats for Text-to-Speech output. *** ## TtsBitrate ```python TtsBitrate = Literal[32000, 64000, 96000, 128000, 192000, 256000, 320000] ``` Allowed output bitrates in bits-per-second for compressed Text-to-Speech formats. *** ## TtsModel Represents a Text-to-Speech model. ### Properties | Property | Type | Description | | --------------------------- | ------------------------- | ---------------------------------------------------------------------------- | | `id` | `str` | Unique identifier of the model. | | `aliased_model_id` | `str \| None` | If this is an alias, the id of the aliased model. None for non-alias models. | | `name` | `str` | Name of the model. | | `voices` | `list[TtsVoice]` | Voices supported by this model. | | `languages` | `list[SupportedLanguage]` | Languages supported by this model. | | `supports_timestamps` | `bool` | If model supports character-to-audio timestamps ('return\_timestamps'). | | `supports_speed_adjustment` | `bool` | If model supports adjusting the speaking rate via the 'speed' parameter. | | `speed_min` | `float \| None` | Minimum supported speaking rate (None when speed adjustment is unsupported). | | `speed_max` | `float \| None` | Maximum supported speaking rate (None when speed adjustment is unsupported). | *** ## TtsSampleRate ```python TtsSampleRate = Literal[8000, 16000, 24000, 44100, 48000] ``` Allowed output sample rates in Hz for Text-to-Speech. *** ## TtsVoice Represents a Text-to-Speech voice. ### Properties | Property | Type | Description | | ------------- | ---------------- | ------------------------------- | | `id` | `str` | Unique identifier of the voice. | | `description` | `str` | Description of the voice. | | `gender` | `TtsVoiceGender` | Gender of the voice. | *** ## TtsVoiceGender ```python TtsVoiceGender = Literal["male", "female", "neutral"] ``` Reported gender of a Text-to-Speech voice. *** ## TemporaryApiKeyUsageType ```python TemporaryApiKeyUsageType = Literal["transcribe_websocket", "tts_rt"] ``` Intended usage for temporary API keys. *** ## UploadFilePayload Optional metadata supplied at upload time. ### Properties | Property | Type | Description | | --------------------- | ------------- | --------------------------------------------------------------- | | `client_reference_id` | `str \| None` | Optional tracking identifier string. Does not need to be unique | *** ## UsageLogEntry A single usage-log entry describing one API request. ### Properties | Property | Type | Description | | -------------------------- | ---------- | --------------------------------------------------------------------------- | | `uuid` | `str` | Unique identifier of the request. | | `request_scope` | `str` | Scope of the request (api / playground). | | `client_reference_id` | `str` | Client reference ID supplied on the original request. Empty string if none. | | `model` | `str` | Model identifier. | | `start_time` | `datetime` | When the request started. | | `end_time` | `datetime` | When the request ended. | | `input_text_tokens` | `int` | - | | `input_audio_tokens` | `int` | - | | `input_audio_duration_ms` | `int` | - | | `output_text_tokens` | `int` | - | | `output_audio_tokens` | `int` | - | | `output_audio_duration_ms` | `int` | - | | `cost_usd` | `str` | - | | `input_cost_usd` | `str` | - | | `input_text_cost_usd` | `str` | - | | `input_audio_cost_usd` | `str` | - | | `output_cost_usd` | `str` | - | | `output_text_cost_usd` | `str` | - | | `output_audio_cost_usd` | `str` | - | *** ## UsageLogsSort ```python UsageLogsSort = Literal["end_time_asc", "end_time_desc"] ``` Sort order for usage-log entries by end\_time. *** ## Voice A cloned Text-to-Speech voice created from a reference audio clip. ### Properties | Property | Type | Description | | ------------ | ------------------ | ---------------------------------------------------- | | `id` | `str` | Unique identifier of the voice. | | `name` | `str` | Name of the voice, unique within the project. | | `filename` | `str` | Original file name of the uploaded reference clip. | | `created_at` | `datetime` | UTC timestamp indicating when the voice was created. | | `models` | `list[VoiceModel]` | Voice readiness status for each available model. | *** ## VoiceModel Per-model readiness status of a voice. ### Properties | Property | Type | Description | | --------------- | ------------------ | ------------------------------------------------------------------------ | | `model` | `str` | Name of the model. | | `status` | `VoiceModelStatus` | Has to be 'ready' for the voice to be usable with this model. | | `error_type` | `str \| None` | Machine-readable error category when status is 'failed'; None otherwise. | | `error_message` | `str \| None` | Human-readable error message when status is 'failed'; None otherwise. | *** ## VoiceModelStatus ```python VoiceModelStatus = Literal["not_computed", "processing", "ready", "failed"] ``` Readiness of a voice for a given model. Must be 'ready' to use the voice with that model. *** ## RealtimeEvent Event payload received from the realtime STT websocket. ### Properties | Property | Type | Description | | --------------------- | ------------- | -------------------------------------------------- | | `tokens` | `list[Token]` | Tokens in this result. | | `final_audio_proc_ms` | `int \| None` | Milliseconds of audio that have been finalized. | | `total_audio_proc_ms` | `int \| None` | Total milliseconds of audio processed. | | `finished` | `bool` | Whether this is the final result (session ending). | | `error_code` | `int \| None` | Error code if the realtime operation failed. | | `error_message` | `str \| None` | Human-readable description of the error. | ### validate\_event() ```python validate_event(raw: str | bytes) -> RealtimeEvent ``` **Parameters** | Parameter | Type | Description | | --------- | -------------- | ---------------------------------------- | | `raw` | `str \| bytes` | Raw event payload from the realtime API. | **Returns** `RealtimeEvent` *** ## RealtimeSTTConfig Configuration for initiating a realtime transcription session. ### Properties | Property | Type | Description | | ----------------------------------- | -------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `api_key` | `str \| None` | API key for real-time sessions. | | `model` | `str` | Speech-to-text model to use. | | `audio_format` | `RealtimeSTTAudioFormat` | Audio format. Use 'auto' for automatic detection of container formats. | | `num_channels` | `int \| None` | Number of audio channels (required for raw audio formats). | | `sample_rate` | `int \| None` | Sample rate in Hz (required for PCM formats). | | `language_hints` | `list[LanguageCode] \| None` | Expected languages in the audio (ISO language codes). | | `language_hints_strict` | `bool \| None` | When true, recognition is strongly biased toward language hints (best results when using one language in language\_hints). | | `context` | `StructuredContextInput \| None` | Additional context to improve transcription accuracy. | | `enable_speaker_diarization` | `bool \| None` | Enable speaker identification. | | `enable_language_identification` | `bool \| None` | Enable automatic language detection. | | `enable_endpoint_detection` | `bool \| None` | Enable endpoint detection for utterance boundaries. | | `max_endpoint_delay_ms` | `int \| None` | Maximum delay between the end of speech and returned endpoint. Allowed values for maximum delay are between 500ms and 3000ms. The default value is 2000ms | | `endpoint_sensitivity` | `float \| None` | Adjusts how likely the model is to emit an endpoint. Higher values make endpoints more likely (finalizing sooner); lower values make them less likely. Allowed values are between -1.0 and 1.0; the default is 0.0. Introduced in the Soniox v5 model; earlier models reject it. | | `endpoint_latency_adjustment_level` | `int \| None` | Fine-tunes the latency/accuracy trade-off of endpoint detection. Allowed values are integers from 0 to 3. | | `translation` | `TranslationConfigInput \| None` | Translation configuration. | | `client_reference_id` | `str \| None` | Optional tracking identifier (max 256 chars). | ### build\_payload() ```python build_payload(api_key: str) -> RealtimeSTTConfig ``` **Parameters** | Parameter | Type | Description | | --------- | ----- | -------------------------------- | | `api_key` | `str` | API key used for authentication. | **Returns** `RealtimeSTTConfig` *** ## RealtimeTTSConfig Configuration for initiating a realtime Text-to-Speech stream. ### Properties | Property | Type | Description | | ------------------- | ----------------------- | ---------------------------------------------------------------------------- | | `api_key` | `str \| None` | API key for real-time sessions. | | `stream_id` | `str` | Client stream identifier unique among active streams on a connection. | | `model` | `str` | Text-to-Speech model to use. | | `language` | `str` | Language code for Text-to-Speech (e.g., "en"). | | `voice` | `str` | Voice identifier to generate speech audio with. | | `audio_format` | `TtsAudioFormat` | Requested output audio format. | | `sample_rate` | `TtsSampleRate \| None` | Output sample rate in Hz. | | `bitrate` | `TtsBitrate \| None` | Output bitrate in bits-per-second for compressed formats. | | `speed` | `float \| None` | Speaking rate multiplier from 0.7 to 1.3; 1.0 (default) is normal speed. | | `return_timestamps` | `bool \| None` | Request character-to-audio timestamps on response events. Defaults to false. | ### build\_payload() ```python build_payload(api_key: str) -> RealtimeTTSConfig ``` **Parameters** | Parameter | Type | Description | | --------- | ----- | -------------------------------- | | `api_key` | `str` | API key used for authentication. | **Returns** `RealtimeTTSConfig` *** ## RealtimeTTSEvent Event payload received from the realtime Text-to-Speech websocket. ### Properties | Property | Type | Description | | --------------- | ----------------------- | ----------------------------------------------------------------------------- | | `stream_id` | `str \| None` | Stream identifier associated with this event. | | `audio` | `str \| None` | Base64 encoded audio chunk, when present. | | `audio_end` | `bool` | Whether this event contains the last audio payload for the stream. | | `terminated` | `bool` | Whether the stream has been fully terminated. | | `error_code` | `int \| None` | Error code if the Text-to-Speech stream failed. | | `error_message` | `str \| None` | Human-readable error message. | | `timestamps` | `TtsTimestamps \| None` | Character-to-audio alignment for this chunk, when `return_timestamps` is set. | ### validate\_event() ```python validate_event(raw: str | bytes) -> RealtimeTTSEvent ``` **Parameters** | Parameter | Type | Description | | --------- | -------------- | ---------------------------------------- | | `raw` | `str \| bytes` | Raw event payload from the realtime API. | **Returns** `RealtimeTTSEvent` *** ### audio\_bytes() ```python audio_bytes() -> bytes | None ``` Decode and return the audio bytes for this event, if present. **Returns** `bytes | None` *** ## RealtimeTTSTextMessage Text chunk message sent over realtime Text-to-Speech websocket. ### Properties | Property | Type | Description | | ----------- | ------ | --------------------------------------------------------------- | | `text` | `str` | Text chunk to generate into speech. | | `text_end` | `bool` | Whether this message marks the final text chunk for the stream. | | `stream_id` | `str` | Stream identifier the chunk belongs to. | *** ## TtsTimestamps Character-to-audio alignment attached to realtime Text-to-Speech events. The three arrays are parallel and equal-length: each index maps one character of the (preprocessed) spoken text to the audio span that pronounces it. ### Properties | Property | Type | Description | | ------------------------------- | ------------- | --------------------------------------------------------------- | | `characters` | `list[str]` | One entry per character (Unicode codepoint) of the spoken text. | | `character_start_times_seconds` | `list[float]` | Start time of each character, in seconds. | | `character_end_times_seconds` | `list[float]` | End time of each character, in seconds. | *** ## Headers ```python Headers = Mapping[str, str] ``` *** ## WebhookAuthConfig Configuration for webhook authentication headers. ### Properties | Property | Type | Description | | -------- | ----- | --------------------------------------------------- | | `name` | `str` | Expected header name (case-insensitive comparison). | | `value` | `str` | Expected header value (exact match). | *** ## WebhookEvent Basic webhook event metadata. ### Properties | Property | Type | Description | | -------- | ------------------------------- | ---------------------------- | | `id` | `str` | Transcription ID (UUID). | | `status` | `Literal['completed', 'error']` | Transcription result status. | # Helpers URL: /sdk/python-SDK/Full-SDK-reference/utils Soniox Python SDK - Helper Functions Reference *** ## output\_file\_for\_audio\_format() ```python output_file_for_audio_format(audio_format: str, prefix: str) -> Path ``` Build an output file path with the correct extension for the audio format. **Parameters** | Parameter | Type | Description | | -------------- | ----- | ------------------------------------------------------------------------------------------------------------- | | `audio_format` | `str` | Audio format name (e.g. `"wav"`, `"mp3"`, `"pcm_s16le"`). | | `prefix` | `str` | Filename prefix without extension. The chosen extension is appended to this prefix to form the returned path. | **Returns** `Path` *** ## render\_tokens() ```python render_tokens(final_tokens: list[Token], non_final_tokens: list[Token]) -> str ``` Build a human-friendly transcript from token metadata. **Parameters** | Parameter | Type | Description | | ------------------ | ------------- | ---------------------------------------------------------------- | | `final_tokens` | `list[Token]` | Token metadata emitted during realtime streaming transcriptions. | | `non_final_tokens` | `list[Token]` | Token metadata emitted during realtime streaming transcriptions. | **Returns** `str` *** ## start\_audio\_thread() ```python start_audio_thread(session: RealtimeSTTSession, chunks: bytes | Iterator[bytes], *, name: str | None = None, daemon: bool = True) -> threading.Thread ``` Stream audio into the session on a background thread. **Parameters** | Parameter | Type | Description | | --------- | -------------------------- | --------------------------------------------------------------------------- | | `session` | `RealtimeSTTSession` | Synchronous WebSocket session for a single real-time speech-to-text stream. | | `chunks` | `bytes \| Iterator[bytes]` | Audio chunks to stream to realtime transcription. | | `name` | `str \| None` | Name of the model. | | `daemon` | `bool` | - | **Returns** `threading.Thread` *** ## start\_text\_thread() ```python start_text_thread(session: RealtimeTTSConnection | RealtimeTTSStream, chunks: str | Iterator[str], *, text_end: bool = True, name: str | None = None, daemon: bool = True) -> threading.Thread ``` Stream text into a realtime TTS session on a background thread. **Parameters** | Parameter | Type | Description | | ---------- | -------------------------------------------- | ------------------------------------------------------------------------ | | `session` | `RealtimeTTSConnection \| RealtimeTTSStream` | Synchronous WebSocket connection for one realtime Text-to-Speech stream. | | `chunks` | `str \| Iterator[str]` | Audio chunks to stream to realtime transcription. | | `text_end` | `bool` | Whether this message marks the final text chunk for the stream. | | `name` | `str \| None` | Name of the model. | | `daemon` | `bool` | - | **Returns** `threading.Thread` *** ## stream\_audio() ```python stream_audio(file: Path | str | BinaryIO | bytes, *, chunk_size_bytes: int = 4 * 1024) -> Iterator[bytes] ``` Yield fixed-size chunks from an audio source. Supports bytes, file paths, or binary streams and slices them into `chunk_size_bytes` blocks for realtime transmission. **Parameters** | Parameter | Type | Description | | ------------------ | ---------------------------------- | ----------------------------------- | | `file` | `Path \| str \| BinaryIO \| bytes` | File input to upload or transcribe. | | `chunk_size_bytes` | `int` | - | **Returns** `Iterator[bytes]` *** ## stream\_audio\_async() ```python stream_audio_async(file: Path | str | BinaryIO | bytes, *, chunk_size_bytes: int = 4 * 1024) -> AsyncIterator[bytes] ``` Asynchronously yield fixed-size chunks from an audio source. Mirrors `stream_audio` but produces an async iterator for later consumption. **Parameters** | Parameter | Type | Description | | ------------------ | ---------------------------------- | ----------------------------------- | | `file` | `Path \| str \| BinaryIO \| bytes` | File input to upload or transcribe. | | `chunk_size_bytes` | `int` | - | **Returns** `AsyncIterator[bytes]` *** ## throttle\_audio() ```python throttle_audio(file: Path | str | BinaryIO | bytes, *, chunk_size_bytes: int = 4096, delay_seconds: float = 0.0) -> Iterator[bytes] ``` Yield audio chunks at a regulated pace, optionally sleeping between yields. **Parameters** | Parameter | Type | Description | | ------------------ | ---------------------------------- | ----------------------------------- | | `file` | `Path \| str \| BinaryIO \| bytes` | File input to upload or transcribe. | | `chunk_size_bytes` | `int` | - | | `delay_seconds` | `float` | - | **Returns** `Iterator[bytes]` *** ## throttle\_audio\_async() ```python throttle_audio_async(file: Path | str | BinaryIO | bytes, *, chunk_size_bytes: int = 32 * 1024, delay_seconds: float = 0.0) -> AsyncIterator[bytes] ``` Async counterpart of `throttle_audio`, yielding chunks with optional delay. **Parameters** | Parameter | Type | Description | | ------------------ | ---------------------------------- | ----------------------------------- | | `file` | `Path \| str \| BinaryIO \| bytes` | File input to upload or transcribe. | | `chunk_size_bytes` | `int` | - | | `delay_seconds` | `float` | - | **Returns** `AsyncIterator[bytes]` # Async transcription with Python SDK URL: /sdk/python-SDK/stt/async-transcription Transcribe audio files asynchronously with the Soniox Python SDK Soniox Python SDK supports asynchronous transcription for audio files. This allows you to transcribe recordings without maintaining a live connection or streaming pipeline. You can either wait for completion or create a job and retrieve the results based on the webhook event. ## Quickstart The SDK provides a convenient `transcribe` method that accepts a local file, public URL, or previously uploaded file ID. It will upload the file (if provided) and create the transcription job. Don't forget to remove files and transcriptions from Soniox after you're done with them. I am working on SDK and docs now, after finishing it, I will do a few minor updates (as I proceed with docs, probably will find somethin else): 1. service: GET /v1/models should respect model type (stt or tts) 2. asr-lb: POST /tts should return the same structure of error response, ass stt ```python from soniox import SonioxClient client = SonioxClient() # Transcribe from a local file transcription = client.stt.transcribe( model="stt-async-v5", file="audio.mp3", ) # Transcribe from a public URL and fetch the transcript later transcription = client.stt.transcribe( model="stt-async-v5", audio_url="https://soniox.com/media/examples/coffee_shop.mp3", ) # Transcribe from a previously uploaded file transcription = client.stt.transcribe( model="stt-async-v5", file_id="uploaded-file-id", ) ``` After creating the job, you can poll the status with `client.stt.get` or wait for completion with `client.stt.wait`. To get the final transcript, call `client.stt.get_transcript`. ```python # Check status transcription = client.stt.get("transcription-id") print(transcription.status) # Wait for completion client.stt.wait("transcription-id") # Fetch transcript transcript = client.stt.get_transcript("transcription-id") print(transcript.text) ``` Don't forget to delete files and transcriptions from Soniox after you're done with them. ## Get transcription You can get a transcription by ID using `get` method. ```python transcription = client.stt.get("transcription-id") print(transcription.id, transcription.status) ``` Get a transcription or return `None` if it doesn't exist: ```python transcription = client.stt.get_or_none("transcription-id") if transcription is None: print("Transcription not found") ``` ## Get transcription transcript If you want to receive text or tokens from a transcription, fetch the transcription transcript with `get_transcript`. ```python transcript = client.stt.get_transcript("transcription-id") print(transcript.text) ``` ## Retrieve list of transcriptions You can retrieve a list of transcriptions using `list` method. ```python from soniox import SonioxClient client = SonioxClient() response = client.stt.list(limit=100) for transcription in response.transcriptions: print(transcription.id, transcription.status) # Use pagination to list more transcriptions while response.next_page_cursor: response = client.stt.list( limit=100, cursor=response.next_page_cursor, ) for transcription in response.transcriptions: print(transcription.id, transcription.status) ``` ## Delete or destroy transcription You can delete or destroy a transcription using `delete` or `destroy` method. **Delete transcription only:** ```python client.stt.delete("transcription-id") ``` **Delete a transcription only if it exists:** ```python client.stt.delete_if_exists("transcription-id") ``` **Delete transcription and its file if it was uploaded:** ```python client.stt.destroy("transcription-id") ``` ## Delete all transcriptions and files from your account You have limited space for files and transcriptions, see: [Limits and quotas](https://soniox.com/docs/stt/async/limits-and-quotas). These operations are irreversible and cannot be undone. ### Delete all transcriptions You can delete all transcriptions using `transcriptions.delete_all`. ```python client.stt.delete_all() ``` ### Delete all files You can delete all files using `files.delete_all`. ```python client.files.delete_all() ``` ### Delete all transcriptions and its files You can delete all transcriptions and its files (if it exist) using `files.destroy_all`. ```python client.files.destroy_all() ``` # Handling files with Python SDK URL: /sdk/python-SDK/stt/files Upload audio files and manage them with the Soniox Python SDK Use the Files API to upload audio for async transcription or to reuse files across multiple jobs. ## Upload `upload()` accepts `bytes`, file paths (`str` or `Path`), or a file-like object (`BinaryIO`). ```python from soniox import SonioxClient client = SonioxClient() file = client.files.upload("audio.mp3") print(file.id, file.filename, file.size) ``` Read more about [Supported audio formats](/stt/async/async-transcription#audio-formats). ## Get file Get a file by ID, throws `SonioxNotFoundError` if file does not exist: ```python file = client.files.get("file-id") print(file.id, file.filename) ``` Get file or none: ```python file = client.files.get_or_none("file-id") if file is None: print("File not found") ``` ## List files List files returns a paginated response. Use `next_page_cursor` to fetch additional pages until it is `None`. ```python from soniox import SonioxClient client = SonioxClient() response = client.files.list(limit=100) for file in response.files: print(file.id, file.filename) # Pagination while response.next_page_cursor: response = client.files.list(limit=100, cursor=response.next_page_cursor) for file in response.files: print(file.id, file.filename) ``` ## Delete file Delete file by ID, throws `SonioxNotFoundError` if file does not exist: ```python client.files.delete("file-id") ``` Delete file only if exists: ```python client.files.delete_if_exists("file-id") ``` ## Delete all files Delete all files iterates through every page and removes each file. ```python client.files.delete_all() ``` # Real-time transcription with Python SDK URL: /sdk/python-SDK/stt/realtime-transcription Create and connect to Soniox real-time speech-to-text sessions with the Python SDK Soniox Python SDK supports transcribing audio in real-time with **low latency** and **high accuracy**. This makes it ideal for voice assistants, live captions, and conversational AI. ## Connect to a real-time session Example below streams audio from live radio to the Soniox real-time API. If you want to stream from a file instead, see: [Create your first real-time session](/sdk/python-SDK#create-your-first-real-time-speech-to-text-session). ```python from typing import Iterator from soniox import SonioxClient from soniox.types import ( RealtimeSTTConfig, Token, StructuredContext, StructuredContextGeneralItem, ) from soniox.utils import render_tokens, start_audio_thread import httpx AUDIO_URL = "https://npr-ice.streamguys1.com/live.mp3?ck=1742897559135" # Fetch audio from a live radio stream and yield it in chunks. def stream_audio_from_url(audio_url) -> Iterator[bytes]: with httpx.Client() as client: with client.stream("GET", audio_url) as response: response.raise_for_status() for chunk in response.iter_bytes(4096): if chunk: yield chunk client = SonioxClient() # Create config, see below for all parameters config = RealtimeSTTConfig( model="stt-rt-v5", audio_format="mp3", enable_endpoint_detection=True, enable_speaker_diarization=True, language_hints=["en"], context=StructuredContext( general=[StructuredContextGeneralItem(key="domain", value="live radio / news broadcast")], text="Live NPR news and talk radio stream, including interviews, music, and commentary.", terms=["NPR", "news", "interview", "music", "commentary", "report", "broadcast", "anchor"], ), ) final_tokens: list[Token] = [] non_final_tokens: list[Token] = [] def realtime(): # Create new real-time websocket session with client.realtime.stt.connect(config=config) as session: # Stream audio from live radio to websocket start_audio_thread(session, stream_audio_from_url(AUDIO_URL)) # Receive events from Soniox Real-time STT for event in session.receive_events(): for token in event.tokens: if token.is_final: final_tokens.append(token) else: non_final_tokens.append(token) print(render_tokens(final_tokens, non_final_tokens)) non_final_tokens.clear() realtime() ``` For config options see: [WebSocket API](/api-reference/stt/websocket-api#configuration) or [RealtimeSTTConfig reference](/sdk/python-SDK/Full-SDK-reference/types#realtimesttconfig-properties). ## Endpoint detection Endpoint detection lets you know when a speaker has finished speaking. This is critical for real-time voice AI assistants, command-and-response systems, and conversational apps where you want to respond immediately without waiting for long silences. Read more about [Endpoint detection](/stt/rt/endpoint-detection) Enable endpoint detection by setting `enable_endpoint_detection=True` in the session config. You will receive special token `` when speech ends. ```python # Enable endpoint detection config = RealtimeSTTConfig( enable_endpoint_detection=True, ... ) # When receiving events, check for special token for event in session.receive_events(): for token in event.tokens: if token.text == "": print("Endpoint detected") ``` ## Manual finalization Manual finalization gives you precise control over when audio should be finalized. When you know the user stopped talking (push-to-talk or client-side VAD), call `finalize` to mark all outstanding tokens as final. Read more about [Manual finalization](/stt/rt/manual-finalization) ```python # Finalize current buffered audio without closing the session. session.finalize() ``` ## Pause and resume ```python session.pause(); // keeps connection alive, drops audio while paused session.resume(); // resume sending audio ``` You are billed for the full stream duration even when session is paused. ## Keepalive Soniox terminates your session if no audio arrives for \~20 seconds. To keep the connection alive, send a keepalive control message or run a background keepalive loop. Python SDK **automatically sends** keepalive messages when session is paused via `session.pause()`. ```python # Sends keepalive message manually session.keep_alive() ``` Read more about [Connection keepalive](/stt/rt/connection-keepalive) ## Streaming audio from a file Use `stream_audio()` with `start_audio_thread()` to stream from a file while receiving events. If you are streaming live audio (microphone, client stream, etc.), you can feed raw chunks without throttling. If you are streaming a prerecorded file, throttle chunks to simulate real-time delivery. ```python from soniox.utils import stream_audio, start_audio_thread, throttle_audio ... with client.realtime.stt.connect(config=config) as session: # Start streaming audio on a background thread. start_audio_thread(session, stream_audio("audio.wav")) # Or throttle local audio file to simulate streaming (sends chunk every 100 ms) start_audio_thread(session, throttle_audio("audio.wav", delay_seconds=0.1)) ... ``` Use [`send_bytes`](/sdk/python-SDK/Full-SDK-reference/realtime/__init__#send_bytes) if you need more control ## Direct stream and proxy stream Read more about [Direct stream](/guides/direct-stream) and [Proxy stream](/guides/proxy-stream). For direct streaming from a client, issue a temporary API key and pass it to the browser or device that will open the WebSocket connection. `client_reference_id` is optional — when set, every request authenticated with the key is recorded under that identifier in [usage logs](/guides/usage-logs). ```python from soniox import SonioxClient client = SonioxClient() key = client.auth.create_temporary_api_key( expires_in_seconds=3600, client_reference_id="support-call-123", ) print(key.api_key, key.expires_at) ``` For proxy streaming, keep the WebSocket connection on your server and stream audio through your backend. # Handling webhooks with Python SDK URL: /sdk/python-SDK/stt/webhooks Use webhooks to receive transcription results with the Soniox Python SDK Python SDK provides helper functions to work with [Webhooks](/stt/async/webhooks). ## Configure webhook delivery If webhook is enabled during transcription creation, Soniox will send a POST request to your webhook URL with the [transcription status result](/stt/async/webhooks#example). ```python from soniox import SonioxClient from soniox.types import CreateTranscriptionConfig client = SonioxClient() config = CreateTranscriptionConfig( webhook_url="https://your-server.com/webhooks/soniox", webhook_auth_header_name="X-Webhook-Secret", webhook_auth_header_value="your-secret", ) transcription = client.stt.transcribe( audio_url="https://soniox.com/media/examples/coffee_shop.mp3", config=config, ) ``` For `transcribe`, you must pass `webhook_auth_header_name` and `webhook_auth_header_value` explicitly in the config. Environment variables (`SONIOX_API_WEBHOOK_HEADER` and `SONIOX_API_WEBHOOK_SECRET`) are only used by webhook helpers and verification (see below). You can also append additional metadata as query parameters: ```python config = CreateTranscriptionConfig( webhook_url="https://your-server.com/webhooks/soniox?request_id=abc-123", ) ``` If you are uploading a local file, you can also use the convenience helper (reads webhook secret and header automatically from environment if present): ```python from soniox import SonioxClient from soniox.types import WebhookAuthConfig client = SonioxClient() transcription = client.stt.transcribe_file_with_webhook( model="stt-async-v5", file="audio.mp3", webhook_url="https://your-server.com/webhooks/soniox", ) ``` ## Example (FastAPI + ngrok) Expose your local server (for example with [ngrok](https://ngrok.com/)), then create a transcription that points to the public ngrok URL and verify the webhook payload on your FastAPI server: ```python from fastapi import FastAPI, Request from soniox import SonioxClient from soniox.errors import InvalidWebhookSignatureError from soniox.types import CreateTranscriptionConfig, WebhookAuthConfig app = FastAPI() client = SonioxClient() # Replace with your public ngrok URL. NGROK_URL = "https://your-subdomain.ngrok-free.app" WEBHOOK_SECRET_NAME = "X-Webhook-Secret" WEBHOOK_SECRET_VALUE = "your-secret" # When creating transcript you must provide correct webhook secret name and value # config = CreateTranscriptionConfig( # webhook_url=f"{NGROK_URL}/webhooks/soniox", # webhook_auth_header_name=WEBHOOK_SECRET_NAME, # webhook_auth_header_value=WEBHOOK_SECRET_VALUE, # ) # client.stt.transcribe( # model="stt-async-v5", # audio_url="https://soniox.com/media/examples/coffee_shop.mp3", # config=config, # ) @app.post("/webhooks/soniox") async def soniox_webhook(request: Request): payload = await request.body() headers = dict(request.headers) try: event = client.webhooks.unwrap( payload, headers, # This can be omitted if you have set env variables SONOIX_API_WEBHOOK_SECRET and SONIOX_API_WEBHOOK_HEADER auth=WebhookAuthConfig( name=WEBHOOK_SECRET_NAME, value=WEBHOOK_SECRET_VALUE, ), ) except InvalidWebhookSignatureError: print("InvalidWebhookSignatureError") return if event.status == "completed": transcript = client.stt.get_transcript(event.id) print(transcript.text) if __name__ == "__main__": import uvicorn uvicorn.run(app, host="0.0.0.0", port=8080) ``` ## Webhook verification Verify webhook signatures to ensure the request really came from Soniox (and not a third party posting to your endpoint). You can verify signatures manually: ```python from soniox import SonioxClient from soniox.types import WebhookAuthConfig client = SonioxClient() client.webhooks.verify_signature( headers={"X-Webhook-Secret": "your-secret"}, ) ``` Or rely on `unwrap` to validate and parse in one step: ```python from soniox import SonioxClient from soniox.types import WebhookAuthConfig client = SonioxClient() event = client.webhooks.unwrap( payload=request_body, headers={"X-Webhook-Secret": "your-secret"}, ) print(event.id, event.status) ``` If you prefer, you can also use `client.stt.transcribe_file_with_webhook` and `client.webhooks` with `SONIOX_API_WEBHOOK_HEADER` and `SONIOX_API_WEBHOOK_SECRET` set in your environment. # Real-time speech generation with Python SDK URL: /sdk/python-SDK/tts/realtime-speech-generation Convert text to speech with the Python SDK Soniox Python SDK supports real-time Text-to-Speech generation with low latency streaming output. You send text chunks over WebSocket and receive audio chunks as they are generated. ## Connect to a real-time session ```python from uuid import uuid4 from soniox import SonioxClient from soniox.types import RealtimeTTSConfig TEXT_CHUNKS = [ "This is a test of Soniox real-time text to speech. ", "Audio is streamed back as it is generated. ", "Each chunk is sent as soon as it is ready.", ] client = SonioxClient() config = RealtimeTTSConfig( stream_id=f"sync-{uuid4()}", model="tts-rt-v2", language="en", voice="Adrian", audio_format="wav", ) audio_chunks: list[bytes] = [] with client.realtime.tts.connect(config=config) as session: session.send_text_chunks(TEXT_CHUNKS, text_end=True) for chunk in session.receive_audio_chunks(): audio_chunks.append(chunk) audio = b"".join(audio_chunks) with open("tts_realtime_sync_output.wav", "wb") as f: f.write(audio) print(f"Wrote {len(audio)} bytes") print("Captured final message:", session.last_message) ``` For config options see: [TTS WebSocket API](/api-reference/tts/websocket-api) and [RealtimeTTSConfig reference](/sdk/python-SDK/Full-SDK-reference/types#realtimettsconfig-properties). ## Async real-time session ```python import asyncio from collections.abc import AsyncIterator from uuid import uuid4 from soniox import AsyncSonioxClient from soniox.types import RealtimeTTSConfig TEXT_CHUNKS = [ "This is a test of Soniox real-time text to speech. ", "Audio is streamed back as it is generated. ", "Each chunk is sent as soon as it is ready.", ] async def iter_text_chunks(chunks: list[str]) -> AsyncIterator[str]: for chunk in chunks: yield chunk async def main() -> None: client = AsyncSonioxClient() config = RealtimeTTSConfig( stream_id=f"async-{uuid4()}", model="tts-rt-v2", language="en", voice="Adrian", audio_format="wav", ) audio_chunks: list[bytes] = [] async with client.realtime.tts.connect(config=config) as session: await session.send_text_chunks(iter_text_chunks(TEXT_CHUNKS), text_end=True) async for chunk in session.receive_audio_chunks(): audio_chunks.append(chunk) audio = b"".join(audio_chunks) with open("tts_realtime_async_output.wav", "wb") as f: f.write(audio) print(f"Wrote {len(audio)} bytes") print("Captured final message:", session.last_message) await client.aclose() asyncio.run(main()) ``` ## Send text incrementally Use `send_text_chunk` when text arrives dynamically (for example from an LLM stream). Set `text_end=True` on the final chunk, or call `finish()`. ```python with client.realtime.tts.connect(config=config) as session: session.send_text_chunk("Hello ", text_end=False) session.send_text_chunk("from Soniox ", text_end=False) session.send_text_chunk("real-time TTS.", text_end=True) ``` Equivalent explicit finalization: ```python with client.realtime.tts.connect(config=config) as session: session.send_text_chunk("Hello from Soniox real-time TTS.", text_end=False) session.finish() ``` ## Receive events vs audio chunks `receive_audio_chunks()` yields decoded audio bytes directly and stops after finalization. Use `receive_events()` when you want access to raw event metadata like `audio_end`, `terminated`, and errors. ```python with client.realtime.tts.connect(config=config) as session: session.send_text_chunk("Hello!", text_end=True) for event in session.receive_events(): if event.audio is not None: print("Audio chunk received") if event.audio_end: print("Server marked final audio payload") if event.terminated: print("Stream terminated") break ``` ## Multiple streams on one connection A single WebSocket connection can carry up to 5 concurrent streams. Use `connect_multi_stream()` to open a multiplexed connection, then call `open_stream()` for each stream. Each stream has its own `stream_id` and operates independently — you can send text and receive audio on all streams in parallel. ### Async multi-stream ```python import asyncio from collections.abc import AsyncIterator from uuid import uuid4 from soniox import AsyncSonioxClient from soniox.types import RealtimeTTSConfig STREAM_TEXTS = { "a": ["Hello from stream A. ", "Stream A shares a connection with B. ", "Goodbye from A."], "b": ["Hello from stream B. ", "Stream B shares a connection with A. ", "Goodbye from B."], } async def iter_text(chunks: list[str]) -> AsyncIterator[str]: for chunk in chunks: yield chunk async def collect_audio(stream) -> bytes: chunks: list[bytes] = [] async for chunk in stream.receive_audio_chunks(): chunks.append(chunk) return b"".join(chunks) async def main() -> None: client = AsyncSonioxClient() async with client.realtime.tts.connect_multi_stream() as connection: # Open two streams on the same WebSocket. streams = {} for key in STREAM_TEXTS: config = RealtimeTTSConfig( stream_id=f"multi-{key}-{uuid4()}", model="tts-rt-v2", language="en", voice="Adrian", audio_format="wav", ) streams[key] = await connection.open_stream(config=config) # Start receiving audio from each stream concurrently. receiver_tasks = { key: asyncio.create_task(collect_audio(stream)) for key, stream in streams.items() } # Send text to each stream concurrently. sender_tasks = [ asyncio.create_task( stream.send_text_chunks(iter_text(STREAM_TEXTS[key]), text_end=True) ) for key, stream in streams.items() ] await asyncio.gather(*sender_tasks) # Wait for all audio to arrive. for key, task in receiver_tasks.items(): audio = await task with open(f"tts_multi_{key}.wav", "wb") as f: f.write(audio) print(f"Stream {key}: wrote {len(audio)} bytes") await client.aclose() asyncio.run(main()) ``` ### Sync multi-stream In synchronous code, use threads to send text and receive audio from each stream concurrently. ```python import threading from uuid import uuid4 from soniox import SonioxClient from soniox.types import RealtimeTTSConfig STREAM_TEXTS = { "a": ["Hello from stream A. ", "Stream A shares a connection with B. ", "Goodbye from A."], "b": ["Hello from stream B. ", "Stream B shares a connection with A. ", "Goodbye from B."], } def collect_audio(stream, results: dict, key: str) -> None: results[key] = b"".join(stream.receive_audio_chunks()) client = SonioxClient() with client.realtime.tts.connect_multi_stream() as connection: # Open two streams on the same WebSocket. streams = {} for key in STREAM_TEXTS: config = RealtimeTTSConfig( stream_id=f"multi-{key}-{uuid4()}", model="tts-rt-v2", language="en", voice="Adrian", audio_format="wav", ) streams[key] = connection.open_stream(config=config) audio_results: dict[str, bytes] = {} # Start receiving audio from each stream in background threads. receivers = [] for key, stream in streams.items(): t = threading.Thread(target=collect_audio, args=(stream, audio_results, key)) t.start() receivers.append(t) # Send text to each stream (threads handle concurrent receiving). for key, stream in streams.items(): stream.send_text_chunks(STREAM_TEXTS[key], text_end=True) # Wait for all audio to arrive. for t in receivers: t.join() for key, audio in sorted(audio_results.items()): with open(f"tts_multi_{key}.wav", "wb") as f: f.write(audio) print(f"Stream {key}: wrote {len(audio)} bytes") client.close() ``` ## Error handling A failed stream does not close the whole WebSocket connection by default. Stream-level errors finalize only that stream (`terminated=True` for the same `stream_id`), while other streams on the same connection can continue. Connection-level failures end the whole connection and all streams. ```python from soniox.errors import SonioxRealtimeError try: with client.realtime.tts.connect(config=config) as session: session.send_text_chunk("Hello!", text_end=True) for _chunk in session.receive_audio_chunks(): pass except SonioxRealtimeError as exc: print("Realtime TTS error:", exc) ``` # REST speech generation with Python SDK URL: /sdk/python-SDK/tts/rest-speech-generation Convert text to speech with the Soniox Python SDK Soniox Python SDK supports asynchronous Text-to-Speech generation with `AsyncSonioxClient`. You can generate speech directly to audio bytes or write output to a file. ## Quickstart The SDK provides a convenient `generate_to_file` method for writing audio output directly to disk. ```python import asyncio from soniox import AsyncSonioxClient async def main() -> None: client = AsyncSonioxClient() try: written = await client.tts.generate_to_file( "tts_async_output.wav", text="Hello from the Soniox Python SDK async TTS example.", model="tts-rt-v2", language="en", voice="Adrian", audio_format="wav", ) print(f"Wrote {written} bytes") finally: await client.aclose() asyncio.run(main()) ``` ## Generate to bytes Use `generate` when you want audio bytes in memory (for custom storage, streaming, or post-processing). ```python import asyncio from soniox import AsyncSonioxClient async def main() -> None: client = AsyncSonioxClient() try: audio = await client.tts.generate( text="This response is generated in memory.", model="tts-rt-v2", language="en", voice="Adrian", audio_format="wav", ) print(f"Received {len(audio)} bytes") finally: await client.aclose() asyncio.run(main()) ``` ## Generate to file Use `generate_to_file` when you want the SDK to write output for you. ```python import asyncio from soniox import AsyncSonioxClient async def main() -> None: client = AsyncSonioxClient() try: written = await client.tts.generate_to_file( "hello.pcm", text="Hello from Soniox async text to speech.", model="tts-rt-v2", language="en", voice="Adrian", audio_format="pcm_s16le", sample_rate=24000, ) print(f"Wrote {written} bytes") finally: await client.aclose() asyncio.run(main()) ``` ## Use typed config (`CreateTtsConfig`) You can pass generation options through `CreateTtsConfig`. ```python import asyncio from soniox import AsyncSonioxClient from soniox.types import CreateTtsConfig async def main() -> None: client = AsyncSonioxClient() try: written = await client.tts.generate_to_file( "typed-config.wav", text="Typed configuration example.", voice="Adrian", config=CreateTtsConfig( model="tts-rt-v2", language="en", audio_format="wav", sample_rate=24000, ), ) print(f"Wrote {written} bytes") finally: await client.aclose() asyncio.run(main()) ``` ## Error handling Handle `SonioxAPIError` to inspect API-level failures. For raw HTTP integration details, see [TTS REST API error handling](/api-reference/tts/generate_tts). ```python import asyncio from soniox import AsyncSonioxClient from soniox.errors import SonioxAPIError async def main() -> None: client = AsyncSonioxClient() try: await client.tts.generate_to_file( "error-case.wav", text="Hello from Soniox!", voice="Adrian", ) except SonioxAPIError as exc: print("Soniox API error:", exc) if exc.request_id: print("request_id:", exc.request_id) finally: await client.aclose() asyncio.run(main()) ``` # Full React SDK reference URL: /sdk/react-SDK/reference Full SDK reference for the React SDK ## Components | Component | Description | | ----------------------------------------------------------------- | --------------------------------------------------------------------- | | [`SonioxProvider`](/sdk/react-SDK/reference/types#sonioxprovider) | Provider component for the Soniox client | | [`AudioLevel`](/sdk/react-SDK/reference/types#audiolevel) | Component to display the audio level (not available for React Native) | ### `SonioxProvider` props `@soniox/react` exports three related prop types so you can split provider props in wrapper HOCs without re-deriving them: | Type | Description | | --------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------- | | [`SonioxProviderProps`](/sdk/react-SDK/reference/types#sonioxproviderprops) | Discriminated union consumed by `SonioxProvider`. Pass either `config` (config-driven) or `client` (client-driven) — never both. | | [`SonioxProviderConfigProps`](/sdk/react-SDK/reference/types#sonioxproviderprops) | The config-driven leaf of the union. Use when you want the provider to construct a `SonioxClient` internally from `SonioxConnectionConfig`. | | [`SonioxProviderClientProps`](/sdk/react-SDK/reference/types#sonioxproviderprops) | The client-driven leaf of the union. Use when you already own a `SonioxClient` instance and just want to expose it through context. | ## Hooks | Hook | Description | | ----------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------- | | [`useRecording`](/sdk/react-SDK/reference/types#userecording) | Main hook for real-time speech-to-text session | | [`useTts`](/sdk/react-SDK/reference/types#usetts) | Main hook for real-time text-to-speech session | | [`useMicrophonePermission`](/sdk/react-SDK/reference/types#usemicrophonepermission) | Hook for checking microphone permission | | [`useAudioLevel`](/sdk/react-SDK/reference/types#useaudiolevel) | Hook for real-time audio volume and spectrum data (not available for React Native) | | [`useSoniox`](/sdk/react-SDK/reference/types#usesoniox) | Hook toaccess the SonioxClient from context | # Types URL: /sdk/react-SDK/reference/types Soniox React SDK — Types Reference ## SonioxProviderProps ```ts type SonioxProviderProps = { children: ReactNode; } & SonioxProviderConfigProps | SonioxProviderClientProps; ``` Props for SonioxProvider. Supply either a pre-built `client` instance or configuration props **Type Declaration** | Name | Type | | ---------- | ----------- | | `children` | `ReactNode` | *** ## TtsState ```ts type TtsState = "idle" | "connecting" | "speaking" | "stopping" | "error"; ``` Aggregate state for the TTS hook. *** ## UnsupportedReason ```ts type UnsupportedReason = "ssr" | "no-mediadevices" | "no-getusermedia" | "insecure-context"; ``` Reason why the built-in browser `MicrophoneSource` is unavailable: * `'ssr'` — `navigator` is undefined (SSR, React Native, or other non-browser JS runtimes). * `'no-mediadevices'` — `navigator` exists but `navigator.mediaDevices` is missing. * `'no-getusermedia'` — `navigator.mediaDevices` exists but `getUserMedia` is not a function. * `'insecure-context'` — the page is not served over HTTPS. This only reflects whether the **default** `MicrophoneSource` can work. Custom `AudioSource` implementations (e.g. for React Native) bypass this check entirely and can record regardless of this value. *** ## AudioLevelProps **Extends** * [`UseAudioLevelOptions`](types#useaudioleveloptions) **Properties** | Property | Type | Description | | ------------------------------------------------- | ------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `active?` | `boolean` | Whether volume metering is active. When false, resources are released. | | `bands?` | `number` | Number of frequency bands to return. When set, the `bands` array is populated with per-band levels (0-1). Useful for spectrum/equalizer visualizations. | | `children` | (`state`) => `ReactNode` | - | | `fftSize?` | `number` | FFT size for the AnalyserNode. Must be a power of 2. Higher values give more frequency resolution (more bins per band) but update less frequently. **Default** `256` | | `smoothing?` | `number` | Exponential smoothing factor (0-1). Higher = smoother/slower decay. **Default** `0.85` | *** ## AudioSupportResult **Properties** | Property | Type | | ------------------------------------------------------- | ---------------------------------------------- | | `isSupported` | `boolean` | | `reason?` | [`UnsupportedReason`](types#unsupportedreason) | *** ## MicrophonePermissionState **Properties** | Property | Type | Description | | -------------------------------------------------------------- | ------------------------ | ---------------------------------------------------------------------- | | `canRequest` | `boolean` | Whether the permission can be requested (e.g., via a prompt). | | `check` | () => `Promise`\<`void`> | Check (or re-check) the microphone permission. No-op when unsupported. | | `isDenied` | `boolean` | `status === 'denied'`. | | `isGranted` | `boolean` | `status === 'granted'`. | | `isSupported` | `boolean` | Whether permission checking is available. | | `status` | `MicPermissionStatus` | Current permission status. | *** ## RecordingSnapshot Immutable snapshot of the recording state exposed to React. **Extended by** * [`UseRecordingReturn`](types#userecordingreturn) **Properties** | Property | Type | Description | | ---------------------------------------------------------------- | ---------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `error` | `Error` \| `null` | Latest error, if any. | | `finalText` | `string` | Accumulated finalized text. | | `finalTokens` | readonly `RealtimeToken`\[] | All finalized tokens in chronological order. Useful for rendering per-token metadata (language, speaker, etc.) in the order tokens were spoken. Pair with `partialTokens` for the complete ordered stream. | | `groups` | `Readonly`\<`Record`\<`string`, `TokenGroup`>> | Tokens grouped by the active `groupBy` strategy. Auto-populated when `translation` config is provided: - `one_way` → keys: `"original"`, `"translation"` - `two_way` → keys: language codes (e.g. `"en"`, `"es"`) Empty `{}` when no grouping is active. | | `isActive` | `boolean` | `true` when state is not idle/stopped/canceled/error. | | `isPaused` | `boolean` | `true` when `state === 'paused'`. | | `isReconnecting` | `boolean` | `true` when `state === 'reconnecting'`. | | `isRecording` | `boolean` | `true` when `state === 'recording'`. | | `isSourceMuted` | `boolean` | `true` when the audio source is muted externally (e.g. OS-level or hardware mute). | | `partialText` | `string` | Text from current non-final tokens. | | `partialTokens` | readonly `RealtimeToken`\[] | Non-final tokens from the latest result. | | `reconnectAttempt` | `number` | Current reconnection attempt number (0 when not reconnecting). | | `result` | `RealtimeResult` \| `null` | Latest raw result from the server. | | `segments` | readonly `RealtimeSegment`\[] | Accumulated final segments. | | `state` | `RecordingState` | Current recording lifecycle state. | | `text` | `string` | Full transcript: `finalText + partialText`. | | `tokens` | readonly `RealtimeToken`\[] | Tokens from the latest result message. | | `utterances` | readonly `RealtimeUtterance`\[] | Accumulated utterances (one per endpoint). | *** ## TtsSnapshot Immutable snapshot of the TTS state exposed to React. **Extended by** * [`UseTtsReturn`](types#usettsreturn) **Properties** | Property | Type | | -------------------------------------------------- | ---------------------------- | | `error` | `Error` \| `null` | | `isConnecting` | `boolean` | | `isSpeaking` | `boolean` | | `state` | [`TtsState`](types#ttsstate) | *** ## UseAudioLevelOptions **Extended by** * [`AudioLevelProps`](types#audiolevelprops) **Properties** | Property | Type | Description | | ------------------------------------------------------ | --------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `active?` | `boolean` | Whether volume metering is active. When false, resources are released. | | `bands?` | `number` | Number of frequency bands to return. When set, the `bands` array is populated with per-band levels (0-1). Useful for spectrum/equalizer visualizations. | | `fftSize?` | `number` | FFT size for the AnalyserNode. Must be a power of 2. Higher values give more frequency resolution (more bins per band) but update less frequently. **Default** `256` | | `smoothing?` | `number` | Exponential smoothing factor (0-1). Higher = smoother/slower decay. **Default** `0.85` | *** ## UseAudioLevelReturn **Properties** | Property | Type | Description | | ---------------------------------------------- | -------------------- | ------------------------------------------------------------------------------------ | | `bands` | readonly `number`\[] | Per-band frequency levels, each 0-1. Empty array when the `bands` option is not set. | | `volume` | `number` | Current volume level, 0 to 1. Updated every animation frame. | *** ## UseMicrophonePermissionOptions **Properties** | Property | Type | Description | | ---------------------------------------------------------------- | --------- | ---------------------------------------- | | `autoCheck?` | `boolean` | Automatically check permission on mount. | *** ## UseRecordingConfig Configuration for useRecording. Extends the STT session config (model, language\_hints, etc.) with recording-specific and React-specific options. Can be used **with or without** a ``: * **With Provider:** omit `config`/`apiKey` — the client is read from context. * **Without Provider:** pass `config` (or legacy `apiKey`) — a client is created internally. **Extends** * `SttSessionConfig` **Properties** | Property | Type | Description | | ---------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | ~~`apiKey?`~~ | `ApiKeyConfig` | API key — string or async function that fetches a temporary key. Required when not using ``. **Deprecated** Use `config` instead. | | `audio_format?` | `"auto"` \| `AudioFormat` | Audio format. Use 'auto' for automatic detection of container formats. For raw PCM formats, also set sample\_rate and num\_channels. **Default** `'auto'` | | `auto_reconnect?` | `boolean` | Enable automatic reconnection on retriable errors. **Default** `false` | | `buffer_queue_size?` | `number` | Maximum audio chunks to buffer during connection setup. | | `client_reference_id?` | `string` | Optional tracking identifier (max 256 chars). | | `config?` | \| `SonioxConnectionConfig` \| (`context?`) => `Promise`\<`SonioxConnectionConfig`> | Connection configuration — sync object or async function. Required when not using ``. | | `context?` | `TranscriptionContext` | Additional context to improve transcription accuracy. | | `enable_endpoint_detection?` | `boolean` | Enable endpoint detection for utterance boundaries. Useful for voice AI agents. | | `enable_language_identification?` | `boolean` | Enable automatic language detection. | | `enable_speaker_diarization?` | `boolean` | Enable speaker identification. | | `endpoint_latency_adjustment_level?` | `number` | Reduces endpoint latency compared to the default endpointing behavior. Higher values reduce endpoint latency more aggressively, which means endpoints are returned sooner and more endpoints may be emitted. This can split long speech into more segments and may slightly reduce word recognition accuracy because speech is finalized earlier. Allowed values are 0, 1, 2, and 3. The default value is 0 (default semantic endpointing behavior). | | `endpoint_sensitivity?` | `number` | Controls how aggressively endpoints are detected. Adjusts how likely the model is to emit an endpoint. Higher values make endpoints more likely, which can finalize segments sooner. Lower values make endpoints less likely, which can help the system wait longer before finalizing. Allowed values are between -1.0 and 1.0. The default value is 0.0. | | `groupBy?` | `"translation"` \| `"language"` \| `"speaker"` \| (`token`) => `string` | Group tokens by a key for easy splitting (e.g. translation, language, speaker). - `'translation'` — group by `translation_status`: keys `"original"` and `"translation"` - `'language'` — group by token `language` field: keys are language codes - `'speaker'` — group by token `speaker` field: keys are speaker identifiers - `(token) => string` — custom grouping function **Auto-defaults** when `translation` config is provided: - `one_way` → `'translation'` - `two_way` → `'language'` | | `language_hints?` | `string`\[] | Expected languages in the audio (ISO language codes). | | `language_hints_strict?` | `boolean` | When true, recognition is strongly biased toward language hints. Best-effort only, not a hard guarantee. | | `max_endpoint_delay_ms?` | `number` | Maximum delay between the end of speech and returned endpoint. Allowed values for maximum delay are between 500ms and 3000ms. The default value is 2000ms | | `max_reconnect_attempts?` | `number` | Maximum consecutive reconnection attempts before giving up. **Default** `3` | | `model` | `string` | Speech-to-text model to use. | | `num_channels?` | `number` | Number of audio channels (required for raw audio formats). | | `onConnected?` | () => `void` | Called when the WebSocket connects. | | `onEndpoint?` | () => `void` | Called when an endpoint is detected. | | `onError?` | (`error`) => `void` | Called when an error occurs. | | `onFinalized?` | () => `void` | Called when the server acknowledges a finalize request (see [UseRecordingReturn](types#userecordingreturn)). | | `onFinished?` | () => `void` | Called when the recording session finishes. | | `onReconnected?` | (`event`) => `void` | Called after a successful reconnection. | | `onReconnecting?` | (`event`) => `void` | Called before a reconnection attempt. Call `preventDefault()` to cancel. | | `onResult?` | (`result`) => `void` | Called on each result from the server. | | `onSourceMuted?` | () => `void` | Called when the audio source is muted externally (e.g. OS-level or hardware mute). | | `onSourceUnmuted?` | () => `void` | Called when the audio source is unmuted after an external mute. | | `onStateChange?` | (`update`) => `void` | Called on each state transition. | | `onToken?` | (`token`) => `void` | Called for each token received from the server (both final and non-final). | | `permissions?` | `PermissionResolver` \| `null` | Permission resolver override (only used when creating an inline client). Pass `null` to explicitly disable. | | `reconnect_base_delay_ms?` | `number` | Base delay in milliseconds for exponential backoff. **Default** `1000` | | `reset_transcript_on_reconnect?` | `boolean` | Clear accumulated transcript state on reconnect. Window-tracking state is always reset regardless. **Default** `false` | | `resetOnStart?` | `boolean` | Reset transcript state when `start()` is called. **Default** `true` | | `sample_rate?` | `number` | Sample rate in Hz (required for PCM formats). | | `session_options?` | `SttSessionOptions` | SDK-level session options (signal, etc.). | | `sessionConfig?` | (`resolved`) => `SttSessionConfig` | Function that receives the resolved connection config (including `stt_defaults` from the server) and returns session config overrides. When provided, its return value is used as the session config for the recording, and any flat session config fields on this object are ignored. **Example** `const { start } = useRecording({ config: asyncConfigFn, sessionConfig: (resolved) => ({ ...resolved.stt_defaults, enable_endpoint_detection: true, }), });` | | `source?` | `AudioSource` | Custom audio source (bypasses default MicrophoneSource). | | `translation?` | `TranslationConfig` | Translation configuration. | | ~~`wsBaseUrl?`~~ | `string` | WebSocket URL override (only used when `apiKey` is provided). **Deprecated** Use `config.stt_ws_url` or `config.region` instead. | *** ## UseRecordingReturn Immutable snapshot of the recording state exposed to React. **Extends** * [`RecordingSnapshot`](types#recordingsnapshot) **Properties** | Property | Type | Description | | ------------------------------------------------------------------- | ------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `cancel` | () => `void` | Immediately cancel — does not wait for final results. | | `clearTranscript` | () => `void` | Clear transcript state (finalText, partialText, utterances, segments). | | `error` | `Error` \| `null` | Latest error, if any. | | `finalize` | (`options?`) => `void` | Request the server to finalize current non-final tokens. | | `finalText` | `string` | Accumulated finalized text. | | `finalTokens` | readonly `RealtimeToken`\[] | All finalized tokens in chronological order. Useful for rendering per-token metadata (language, speaker, etc.) in the order tokens were spoken. Pair with `partialTokens` for the complete ordered stream. | | `groups` | `Readonly`\<`Record`\<`string`, `TokenGroup`>> | Tokens grouped by the active `groupBy` strategy. Auto-populated when `translation` config is provided: - `one_way` → keys: `"original"`, `"translation"` - `two_way` → keys: language codes (e.g. `"en"`, `"es"`) Empty `{}` when no grouping is active. | | `isActive` | `boolean` | `true` when state is not idle/stopped/canceled/error. | | `isPaused` | `boolean` | `true` when `state === 'paused'`. | | `isReconnecting` | `boolean` | `true` when the WebSocket is reconnecting after a drop. | | `isRecording` | `boolean` | `true` when `state === 'recording'`. | | `isSourceMuted` | `boolean` | `true` when the audio source is muted externally (e.g. OS-level or hardware mute). | | `isSupported` | `boolean` | Whether the built-in browser `MicrophoneSource` is available. Custom `AudioSource` implementations work regardless of this value. | | `partialText` | `string` | Text from current non-final tokens. | | `partialTokens` | readonly `RealtimeToken`\[] | Non-final tokens from the latest result. | | `pause` | () => `void` | Pause recording — pauses audio capture and activates keepalive. | | `reconnect` | () => `void` | Force a reconnection — tears down the current session and audio encoder, then establishes a new session. Requires `auto_reconnect`. Call this from platform lifecycle handlers (e.g. web `visibilitychange`, React Native `AppState`) to recover from stale connections after sleep/wake or backgrounding. | | `reconnectAttempt` | `number` | Current reconnection attempt number (0 when not reconnecting). | | `result` | `RealtimeResult` \| `null` | Latest raw result from the server. | | `resume` | () => `void` | Resume recording after pause. | | `segments` | readonly `RealtimeSegment`\[] | Accumulated final segments. | | `start` | () => `void` | Start a new recording. Aborts any in-flight recording first. | | `state` | `RecordingState` | Current recording lifecycle state. | | `stop` | () => `Promise`\<`void`> | Gracefully stop — waits for final results from the server. | | `text` | `string` | Full transcript: `finalText + partialText`. | | `tokens` | readonly `RealtimeToken`\[] | Tokens from the latest result message. | | `unsupportedReason` | [`UnsupportedReason`](types#unsupportedreason) \| `undefined` | Why the built-in `MicrophoneSource` is unavailable, if applicable. Custom `AudioSource` implementations bypass this check entirely. | | `utterances` | readonly `RealtimeUtterance`\[] | Accumulated utterances (one per endpoint). | *** ## UseTtsConfig Configuration for useTts. Extends TtsStreamInput — flat TTS fields (model, voice, language, audio\_format) are merged on top of server-provided `tts_defaults`. Can be used **with or without** a ``: * **With Provider:** omit `config` — the client is read from context. * **Without Provider:** pass `config` — a client is created internally. In `'rest'` mode, `voice` is required — the REST TTS endpoint (`GenerateSpeechOptions.voice`) has no default. Discover available voices via `client.tts.listModels()`. **Extends** * `TtsStreamInput` **Properties** | Property | Type | Description | | -------------------------------------------------------------- | ----------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `audio_format?` | `TtsAudioFormat` | Output audio format **Example** `'wav'` | | `bitrate?` | `number` | Codec bitrate in bps (for compressed formats). | | `config?` | \| `SonioxConnectionConfig` \| (`context?`) => `Promise`\<`SonioxConnectionConfig`> | Connection configuration — sync object or async function. Required when not using ``. | | `language?` | `string` | Language code for speech generation. **Example** `'en'` | | `mode?` | `"websocket"` \| `"rest"` | Transport mode for TTS generation. - `'websocket'` (default): Real-time streaming via WebSocket. Supports incremental text input (`sendText`/`finish`) and streaming from LLM. - `'rest'`: HTTP request/response via the TTS REST endpoint. Sends full text at once, streams audio back. Simpler but no incremental text input. **Default** `'websocket'` | | `model?` | `string` | Text-to-Speech model to use. **Example** `'tts-rt-v2'` | | `onAudio?` | (`chunk`, `timestamps?`) => `void` | Called when an audio chunk is received. In WebSocket mode with `return_timestamps` enabled, the second argument carries the character-level alignment for that frame (`undefined` for audio-only frames). REST mode never provides timestamps. | | `onAudioEnd?` | () => `void` | Called when the server marks the final audio payload. | | `onError?` | (`error`) => `void` | Called on error. | | `onStateChange?` | (`event`) => `void` | Called on each state transition. | | `onTerminated?` | () => `void` | Called when generation is complete. | | `reduce_silence?` | `boolean` | Shorten pauses between words in the generated speech. `false` (default) keeps the model's natural pacing; `true` tightens delivery by reducing silence between words. | | `return_timestamps?` | `boolean` | Request character-level audio timestamps in the responses. When enabled, audio frames may carry a TtsTimestamps payload aligning each character of the spoken text to its start/end time in the audio. WebSocket (realtime) only — the REST endpoint streams raw audio bytes and ignores this flag. Timestamps map to the model's preprocessed text, not the raw input. Defaults to `false` when omitted. | | `sample_rate?` | `number` | Output sample rate in Hz. Required for raw PCM formats. | | `speed?` | `number` | Speaking rate. `1.0` is the normal rate; values below `1.0` slow speech down and values above `1.0` speed it up. Supported range is `0.7`-`1.3`. Defaults to `1.0` when omitted. | | `stream_id?` | `string` | Client-generated stream identifier. Must be unique among active streams on the same connection. Auto-generated if omitted. | | `voice?` | `string` | Voice identifier. **Example** `'Adrian'` | *** ## UseTtsReturn Immutable snapshot of the TTS state exposed to React. **Extends** * [`TtsSnapshot`](types#ttssnapshot) **Properties** | Property | Type | Description | | --------------------------------------------------- | ---------------------------- | ----------------------------------------------------------------------------------------- | | `cancel` | () => `void` | Cancel the current generation immediately. | | `error` | `Error` \| `null` | - | | `finish` | () => `void` | Signal that no more text will be sent. WebSocket mode only. | | `isConnecting` | `boolean` | - | | `isSpeaking` | `boolean` | - | | `sendText` | (`text`) => `void` | Send one text chunk without finishing. WebSocket mode only. | | `speak` | (`text`) => `void` | Start TTS. Sends text (or pipes an async iterable in WebSocket mode) and generates audio. | | `state` | [`TtsState`](types#ttsstate) | - | | `stop` | () => `Promise`\<`void`> | Gracefully stop — sends finish and waits for completion. | *** ## AudioLevel() ```ts function AudioLevel(__namedParameters): ReactNode; ``` **Parameters** | Parameter | Type | | ------------------- | ------------------------------------------ | | `__namedParameters` | [`AudioLevelProps`](types#audiolevelprops) | **Returns** `ReactNode` *** ## SonioxProvider() ```ts function SonioxProvider(props): ReactNode; ``` **Parameters** | Parameter | Type | | --------- | -------------------------------------------------- | | `props` | [`SonioxProviderProps`](types#sonioxproviderprops) | **Returns** `ReactNode` *** ## checkAudioSupport() ```ts function checkAudioSupport(): AudioSupportResult; ``` Check whether the current environment supports the built-in browser `MicrophoneSource` (which uses `navigator.mediaDevices.getUserMedia`). This does **not** reflect general recording capability — custom `AudioSource` implementations (e.g. for React Native) bypass this check entirely and can record regardless of the result. **Returns** [`AudioSupportResult`](types#audiosupportresult) **Platform** browser *** ## useAudioLevel() ```ts function useAudioLevel(options?): UseAudioLevelReturn; ``` **Parameters** | Parameter | Type | | ---------- | ---------------------------------------------------- | | `options?` | [`UseAudioLevelOptions`](types#useaudioleveloptions) | **Returns** [`UseAudioLevelReturn`](types#useaudiolevelreturn) *** ## useMicrophonePermission() ```ts function useMicrophonePermission(options?): MicrophonePermissionState; ``` **Parameters** | Parameter | Type | | ---------- | ------------------------------------------------------------------------ | | `options?` | [`UseMicrophonePermissionOptions`](types#usemicrophonepermissionoptions) | **Returns** [`MicrophonePermissionState`](types#microphonepermissionstate) *** ## useRecording() ```ts function useRecording(config): UseRecordingReturn; ``` **Parameters** | Parameter | Type | | --------- | ------------------------------------------------ | | `config` | [`UseRecordingConfig`](types#userecordingconfig) | **Returns** [`UseRecordingReturn`](types#userecordingreturn) *** ## useSoniox() ```ts function useSoniox(): SonioxClient; ``` Returns the `SonioxClient` instance provided by the nearest `SonioxProvider` **Returns** `SonioxClient` **Throws** Error if called outside a `SonioxProvider` *** ## useTts() ```ts function useTts(config): UseTtsReturn; ``` **Parameters** | Parameter | Type | | --------- | ------------------------------------ | | `config` | [`UseTtsConfig`](types#usettsconfig) | **Returns** [`UseTtsReturn`](types#usettsreturn) # Real-time transcription with React SDK URL: /sdk/react-SDK/stt/realtime-transcription Create and manage real-time speech-to-text sessions with the Soniox React SDK import { LinkCards } from "@/components/link-card"; Soniox React SDK supports real-time transcription via React hooks, built on top of the [@soniox/client](/sdk/web-SDK) Web SDK. This allows you to transcribe live audio with low latency — ideal for live captions, voice input, and interactive experiences. You can capture audio from the user's microphone, receive transcription results as reactive state, and control sessions with simple `start`/`stop` calls. ## Soniox Provider [`SonioxProvider`](/sdk/react-SDK/reference/types#sonioxprovider) creates and shares a single [`SonioxClient`](/sdk/web-SDK/reference/classes#sonioxclient) instance via React context. Place it near the root of your component tree. ### With configuration props ```tsx import { SonioxProvider } from "@soniox/react"; function App() { return ( { const res = await fetch("/api/get-temporary-key", { method: "POST" }); const { api_key } = await res.json(); return { api_key }; }} > {children} ); } ``` The `config` resolver runs once per recording session and can return any [`SonioxConnectionConfig`](/sdk/web-SDK/reference/types#sonioxclientoptions) fields — for example `region`, `stt_ws_url`, or `tts_ws_url` — alongside the `api_key`. ### With a pre-built client ```tsx import { SonioxClient } from "@soniox/client"; import { SonioxProvider } from "@soniox/react"; const client = new SonioxClient({ config: async () => { const { api_key } = await fetchKey(); return { api_key }; }, }); function App() { return {children}; } ``` ## `useRecording` `useRecording` is the primary hook for real-time speech-to-text. Returns reactive transcript state and control methods. Returns [`UseRecordingReturn`](/sdk/react-SDK/reference/types#userecordingreturn) which contains reactive state and control methods. ```tsx function Transcriber() { const recording = useRecording({ model: "stt-rt-v5", language_hints: ["en", "es"], enable_endpoint_detection: true, }); return (

State: {recording.state}

{recording.text}

); } ``` ### Handle session events | Callback | Signature | Description | | ------------------ | --------------------------------------------------------------- | ---------------------------------------------------------------------- | | `onResult` | `(result: RealtimeResult) => void` | Called on each result from the server. | | `onToken` | `(token: RealtimeToken) => void` | Called once per token, in addition to `onResult`. | | `onEndpoint` | `() => void` | Called when an endpoint is detected. | | `onError` | `(error: Error) => void` | Called when an error occurs. | | `onStateChange` | `(update: { old_state, new_state, reason? }) => void` | Called on each state transition. | | `onFinished` | `() => void` | Called when the recording session finishes. | | `onConnected` | `() => void` | Called when the WebSocket connects. | | `onReconnecting` | `({ attempt, max_attempts, delay_ms, preventDefault }) => void` | Called before each reconnect attempt (requires `auto_reconnect`). | | `onReconnected` | `({ attempt }) => void` | Called after a successful reconnect. | | `onSessionRestart` | `({ reset_transcript }) => void` | Called when a new STT session is started (initial or after reconnect). | | `onSourceMuted` | `() => void` | Called when the audio source is muted externally. | | `onSourceUnmuted` | `() => void` | Called when the audio source is unmuted. | ### Session lifecycle #### Recording state | Field | Type | Description | | --------------- | ---------------- | -------------------------------------------------------------------- | | `state` | `RecordingState` | Current lifecycle state (`'idle'`, `'recording'`, `'paused'`, etc.). | | `isActive` | `boolean` | `true` when state is not `idle`/`stopped`/`canceled`/`error`. | | `isRecording` | `boolean` | `true` when `state === 'recording'`. | | `isPaused` | `boolean` | `true` when `state === 'paused'`. | | `isSourceMuted` | `boolean` | `true` when the audio source is muted externally. | #### Available methods | Method | Signature | Description | | ----------------- | --------------------- | ------------------------------------------------------------------------ | | `start` | `() => void` | Start a new recording. Aborts any in-flight recording first. | | `stop` | `() => Promise` | Gracefully stop — waits for final results from the server. | | `cancel` | `() => void` | Immediately cancel — does not wait for final results. | | `pause` | `() => void` | Pause audio capture (keepalive keeps connection open). | | `resume` | `() => void` | Resume after pause. | | `finalize` | `(options?) => void` | Request the server to finalize current non-final tokens. | | `clearTranscript` | `() => void` | Clear transcript state (`finalText`, `partialText`, `utterances`, etc.). | ### Endpoint detection and manual finalization Endpoint detection lets you know when a speaker has finished speaking. This is critical for real-time voice AI assistants, command-and-response systems, and conversational apps where you want to respond immediately without waiting for long silences. Read more about [Endpoint detection](/stt/rt/endpoint-detection) Enable endpoint detection by setting `enable_endpoint_detection: true` in the hook configuration. Use the `onEndpoint` callback to know when a speaker has finished speaking. ```tsx const { start, stop, text } = useRecording({ model: "stt-rt-v5", enable_endpoint_detection: true, onEndpoint: () => { console.log("--- speaker finished ---"); }, }); ``` Manual finalization gives you precise control over when audio should be finalized — useful for push-to-talk systems and client-side voice activity detection (VAD). Read more about [Manual finalization](/stt/rt/manual-finalization) The finalize function is returned by useRecording and can be called at any time during an active recording: ```tsx const { start, stop, finalize } = useRecording({}); // Later, when you want to force finalization: finalize(); ``` ### Pause, resume and muting audio source The pause and resume functions are returned by `useRecording`. The `isPaused` flag reflects the current pause state reactively. ```tsx const { start, stop, pause, resume, isPaused } = useRecording({}); pause(); // keeps connection alive, drops audio while paused resume(); // resume sending audio ``` SDK will `finalize` audio on pause. Make sure to adjust your VAD sensitivity to have enough silence before pause. Learn more about [Manual finalization](/stt/rt/manual-finalization#key-points) The hook also tracks system-level mute events via `isSourceMuted`. When the audio source is muted externally (e.g. OS-level or hardware mute), keepalive messages are sent automatically to keep the session alive. You can listen for mute state changes with the `onSourceMuted` and `onSourceUnmuted` callbacks. ```tsx const { isSourceMuted } = useRecording({ onSourceMuted: () => { console.log("Microphone muted externally"); }, onSourceUnmuted: () => { console.log("Microphone unmuted"); }, }); ``` You are billed for the full stream duration even when the session is paused. ### Auto-reconnect `useRecording` can transparently recover from transient network drops. Opt in with `auto_reconnect: true` — on a retriable error, the hook tears down the current WebSocket and audio encoder, re-resolves the connection config, and starts a new session with exponential backoff. Audio captured during the reconnect is buffered and flushed on resume. ```tsx const { start, stop, state, isReconnecting, reconnectAttempt } = useRecording({ model: "stt-rt-v5", auto_reconnect: true, max_reconnect_attempts: 3, // default: 3 reconnect_base_delay_ms: 1000, // exponential backoff: 1s, 2s, 4s, ... onReconnecting: ({ attempt, max_attempts, delay_ms, preventDefault }) => { console.log(`Reconnect attempt ${attempt}/${max_attempts} in ${delay_ms}ms`); // preventDefault() cancels this attempt (e.g. for manual backoff control). }, onReconnected: ({ attempt }) => { console.log(`Reconnected after ${attempt} attempt(s)`); }, }); ``` #### Options | Option | Type | Default | Description | | ------------------------------- | --------- | ------- | ----------------------------------------------------------- | | `auto_reconnect` | `boolean` | `false` | Enable automatic reconnection on retriable errors. | | `max_reconnect_attempts` | `number` | `3` | Maximum consecutive attempts before surfacing the error. | | `reconnect_base_delay_ms` | `number` | `1000` | Base delay in ms for exponential backoff (1x, 2x, 4x, ...). | | `reset_transcript_on_reconnect` | `boolean` | `false` | Clear accumulated transcript state on reconnect. | #### Return values | Field | Type | Description | | ------------------ | --------- | -------------------------------------------------------------------- | | `isReconnecting` | `boolean` | `true` while the hook is in the `reconnecting` state. | | `reconnectAttempt` | `number` | Current attempt number (resets to `0` after a successful reconnect). | #### Force a reconnect `useRecording` also returns a `reconnect()` function — call it from platform lifecycle handlers like `visibilitychange` or React Native `AppState` to proactively rebuild the session when you suspect a stale connection. Requires `auto_reconnect: true`. ```tsx const { reconnect } = useRecording({ model: "stt-rt-v5", auto_reconnect: true, }); useEffect(() => { const onVisibility = () => { if (document.visibilityState === "visible") reconnect(); }; document.addEventListener("visibilitychange", onVisibility); return () => document.removeEventListener("visibilitychange", onVisibility); }, [reconnect]); ``` ### Handling translation The React SDK supports one-way and two-way real-time translation. Configure translation in the useRecording hook config. The hook automatically groups tokens by translation status or language via the groups snapshot field, so you can render original and translated text separately without manual filtering. #### One-way translation Translates all spoken audio into a single target language. When translation is provided with type: `one_way`, the hook automatically sets `groupBy: 'translation'`, splitting tokens into `original` and `translation` groups. ```tsx const { groups } = useRecording({ model: "stt-rt-v5", translation: { type: "one_way", target_language: "es", // Translate everything to Spanish }, }); // Render grouped text return (

Original: {groups.original?.text}

Translated: {groups.translation?.text}

); ``` #### Two-way translation Translates between two languages — each speaker's speech is translated into the other language. When translation is provided with type: `two_way`, the hook automatically sets `groupBy: 'language'`, splitting tokens by language code (e.g. `en`, `fr`). ```tsx const { groups } = useRecording({ model: "stt-rt-v5", translation: { type: "two_way", language_a: "en", language_b: "fr", }, }); // Render grouped text by language return (

English: {groups.en?.text}

French: {groups.fr?.text}

); ``` Learn more about [Real-time translation](/translation/stt-translation/rt-translation) ### Utterances When `enable_endpoint_detection` is enabled, the `utterances` array accumulates utterances separated by natural pauses: ```tsx function TranscriptWithUtterances() { const { utterances, partialText, start, stop, isActive } = useRecording({ model: "stt-rt-v5", enable_endpoint_detection: true, }); return (
{utterances.map((utterance, i) => (

{utterance.text}

))} {partialText &&

{partialText}

}
); } ``` Learn more about [Endpoint detection](/stt/rt/endpoint-detection) ### Token grouping The `groupBy` option splits tokens into named groups, accessible via `recording.groups`. This is particularly useful for translation and multi-speaker scenarios. #### `groupBy` strategies | Value | Keys | Description | | ------------------- | ------------------------------------ | -------------------------------- | | `'translation'` | `"original"`, `"translation"` | Group by `translation_status`. | | `'language'` | Language codes (e.g. `"en"`, `"fr"`) | Group by token `language` field. | | `'speaker'` | Speaker IDs (e.g. `"1"`) | Group by token `speaker` field. | | `(token) => string` | Custom keys | Custom grouping function. | Learn more about [Speaker diarization](/stt/concepts/speaker-diarization) #### TokenGroup fields Each group in `recording.groups` contains: | Field | Type | Description | | --------------- | ----------------- | ------------------------------------------------------- | | `text` | `string` | Full text: `finalText + partialText`. | | `finalText` | `string` | Accumulated finalized text in this group. | | `partialText` | `string` | Text from current non-final tokens. | | `partialTokens` | `RealtimeToken[]` | Current non-final tokens (from the latest result only). | #### Automatic grouping for translation When a `translation` config is provided, `groupBy` is set automatically: * `one_way` translation → groups by `'translation'` (keys: `"original"`, `"translation"`) * `two_way` translation → groups by `'language'` (keys: language codes like `"en"`, `"es"`) ```tsx function TranslatedTranscript() { const { groups, start, stop, isActive } = useRecording({ model: "stt-rt-v5", translation: { type: "one_way", target_language: "es" }, }); return (

Original

{groups.original?.text}

Translation

{groups.translation?.text}

); } ``` ## `useSoniox` Returns the `SonioxClient` instance from the nearest `SonioxProvider`. Useful for low-level session access. ```tsx import { useSoniox } from "@soniox/react"; function MyComponent() { const client = useSoniox(); // Low-level session access (non-reactive): // const session = client.realtime.stt({ model: 'stt-rt-v5' }, { api_key: '...' }); // // Permission helpers (if a resolver is configured): // const result = await client.permissions?.check('microphone'); return null; } ``` `client.tts` in the browser only exposes `generate()` and `generateStream()`. To enumerate available TTS models and voices, use the [Node SDK's `client.tts.listModels()`](/sdk/node-SDK/tts/rest-speech-generation#list-available-models) on your server. ## `useMicrophonePermission` Hook for checking and requesting microphone permission before recording. Requires a `SonioxProvider` with a permission resolver configured (default in browsers). ```tsx import { useMicrophonePermission } from "@soniox/react"; function PermissionGate({ children }) { const mic = useMicrophonePermission({ autoCheck: true }); if (!mic.isSupported) { return

Microphone permissions are not available.

; } if (mic.status === "unknown") { return

Checking permission...

; } if (mic.isDenied) { return (

Microphone access denied.

{!mic.canRequest && (

Please enable microphone access in your browser settings.

)}
); } if (mic.status === "prompt") { // `mic.check` re-queries the permission state. To actually show the // browser prompt, start a recording or call `getUserMedia({ audio: true })` // from the click handler. return ( ); } return children; } ``` ### Options | Option | Type | Default | Description | | ----------- | --------- | ------- | ---------------------------------------- | | `autoCheck` | `boolean` | `false` | Automatically check permission on mount. | ### Return value | Field | Type | Description | | ------------- | --------------------- | ------------------------------------------------------------------------------------------------------ | | `status` | `MicPermissionStatus` | Current status: `'granted'`, `'denied'`, `'prompt'`, `'unavailable'`, `'unsupported'`, or `'unknown'`. | | `canRequest` | `boolean` | Whether the user can be prompted again. `false` when permanently denied. | | `isGranted` | `boolean` | `status === 'granted'`. | | `isDenied` | `boolean` | `status === 'denied'`. | | `isSupported` | `boolean` | Whether permission checking is available. | | `check` | `() => Promise` | Check (or re-check) the microphone permission. No-op when unsupported. | ### Status values | Status | Description | | --------------- | --------------------------------------------------- | | `'granted'` | Microphone access is granted. | | `'denied'` | Microphone access is denied. | | `'prompt'` | User hasn't been asked yet. | | `'unavailable'` | Permissions API not available in this browser. | | `'unsupported'` | No `PermissionResolver` configured in the provider. | | `'unknown'` | Initial state before the first `check()` call. | ## `useAudioLevel` Hook for real-time audio volume metering. Useful for building recording indicators and animations. ```tsx import { useAudioLevel } from "@soniox/react"; function VolumeIndicator({ isActive }) { const { volume } = useAudioLevel({ active: isActive }); // float value between 0 and 1 return (
); } ``` ## Next.js (App Router) The package declares `'use client'` at the entry point. All hooks must be used inside Client Components. Server Components cannot use `useRecording` or other hooks directly. # Real-time speech generation with React SDK URL: /sdk/react-SDK/tts/realtime-speech-generation Stream text to speech in React with the useTts hook import { LinkCards } from "@/components/link-card"; The Soniox React SDK exposes a single [`useTts`](/sdk/react-SDK/reference/types#usetts) hook that covers both real-time WebSocket TTS and one-shot REST TTS. It manages the stream lifecycle, surfaces reactive state for rendering, and handles cleanup automatically when your component unmounts. ## Transport modes `useTts` has two modes selected via the `mode` option: | Feature | `'websocket'` (default) | `'rest'` | | -------------------------------------------- | --------------------------------------------------- | ----------------------------------------------------- | | `speak(string)` | Yes | Yes | | `speak(asyncIterable)` — LLM token streaming | Yes | No | | `sendText()` / `finish()` — incremental text | Yes | No-op | | Audio delivery | Streaming chunks | Streaming chunks | | Error detection | Full (in-band WebSocket errors) | Pre-stream only (HTTP status) | | Use when... | Narrating LLM output, lowest latency to first audio | You have the full text and want a simple HTTP request | ## Set up your temporary API key endpoint In a browser environment you don't want to expose your primary API key. Create a temporary key endpoint on your server using the Soniox [Node SDK](/sdk/node-SDK). TTS keys use the `tts_rt` usage type. To attribute browser-side TTS traffic to an end user or session, pass `client_reference_id` to `createTemporaryKey` - every request authenticated with the key is recorded under that identifier in [usage logs](/guides/usage-logs). Clients cannot override it. ```ts import express from 'express'; import { SonioxNodeClient } from '@soniox/node'; const app = express(); const client = new SonioxNodeClient(); // reads SONIOX_API_KEY from env app.get('/tts-tmp-key', async (_req, res) => { try { const { api_key, expires_at } = await client.auth.createTemporaryKey({ usage_type: 'tts_rt', expires_in_seconds: 300, }); res.json({ api_key, expires_at }); } catch (err) { res.status(500).json({ error: err instanceof Error ? err.message : 'Failed to create temporary key' }); } }); app.listen(3000); ``` Because TTS and STT temporary keys have different `usage_type` values, `useTts` always creates its own client from the inline `config` prop — even when a `` is present. Pass the `config` prop on `useTts` with a resolver that fetches a `tts_rt` key. ## Quickstart Pass a `config` resolver and a `voice` — that's all that's required. `speak()` starts generation; `state` reflects the lifecycle; audio plays via the `onAudio` callback or any hook-up of your choice. ```tsx import { useTts } from "@soniox/react"; import { useRef } from "react"; async function fetchTtsKey() { const res = await fetch("/tts-tmp-key"); const { api_key } = await res.json(); return { api_key }; } export function TtsButton() { const audioChunksRef = useRef([]); const { speak, state, isSpeaking, error } = useTts({ config: fetchTtsKey, voice: "Adrian", model: "tts-rt-v2", language: "en", audio_format: "wav", onAudio: (chunk) => { audioChunksRef.current.push(chunk); }, onTerminated: () => { const blob = new Blob(audioChunksRef.current, { type: "audio/wav" }); new Audio(URL.createObjectURL(blob)).play(); audioChunksRef.current = []; }, onError: (err) => console.error(err), }); return (

State: {state}

{error &&

{error.message}

}
); } ``` ## Stream from an LLM WebSocket mode accepts an `AsyncIterable` — pipe LLM tokens straight into `speak()` and audio starts playing as the first tokens arrive. ```tsx async function* llmTokens(prompt: string): AsyncIterable { const res = await fetch("/llm/stream", { method: "POST", body: JSON.stringify({ prompt }), }); const reader = res.body!.getReader(); const decoder = new TextDecoder(); while (true) { const { value, done } = await reader.read(); if (done) return; yield decoder.decode(value); } } function Narrator() { const audioChunksRef = useRef([]); const { speak, isSpeaking, stop } = useTts({ config: fetchTtsKey, voice: "Adrian", audio_format: "wav", onAudio: (chunk) => audioChunksRef.current.push(chunk), }); return ( ); } ``` ## Send text incrementally For finer control over when text arrives, call `sendText` for each chunk and `finish` when done. This is useful when the text source isn't already an async iterable. ```tsx function IncrementalTts() { const audioChunksRef = useRef([]); const { sendText, finish, stop, isSpeaking } = useTts({ config: fetchTtsKey, voice: "Adrian", audio_format: "wav", onAudio: (chunk) => audioChunksRef.current.push(chunk), }); const speakParagraphs = (paragraphs: string[]) => { for (const p of paragraphs) sendText(p + " "); finish(); }; return (
{isSpeaking && }
); } ``` `sendText` and `finish` are no-ops in REST mode — the REST endpoint only accepts the full text in a single request. ## REST mode Set `mode: 'rest'` to run TTS over HTTP. `speak(string)` still works, but `speak(asyncIterable)`, `sendText`, and `finish` are not available (the hook emits an error if you try to stream an async iterable). ```tsx const audioChunksRef = useRef([]); const { speak, state, isSpeaking } = useTts({ config: fetchTtsKey, mode: "rest", voice: "Adrian", audio_format: "wav", onAudio: (chunk) => audioChunksRef.current.push(chunk), onTerminated: () => console.log("done"), }); speak("Hello over REST."); ``` Use REST mode for one-off playback (confirmations, notifications) where the lower latency of WebSocket isn't worth the extra connection. REST mode requires `voice` in the hook config — there is no built-in fallback. Omitting it surfaces an error via `onError` and leaves `state: 'error'`. WebSocket mode also needs a `voice`, but it can come from the hook config **or** from server-returned `tts_defaults` (see [Server-driven defaults](#server-driven-defaults)). To discover available voices, call [`client.tts.listModels()`](/sdk/node-SDK/tts/rest-speech-generation#list-available-models) from the Node SDK on your server — it is not available in the browser client. ## Lifecycle `useTts` exposes a single `state` that transitions through the lifecycle below. `isSpeaking` and `isConnecting` are derived booleans you can use directly in UI. | State | Meaning | | ------------ | ----------------------------------------------------------------------------- | | `idle` | No generation in flight. Initial state, or after a completed / cancelled run. | | `connecting` | Opening the WebSocket (or issuing the REST request). | | `speaking` | Receiving audio chunks. | | `stopping` | `stop()` was called — waiting for `terminated` to flush. | | `error` | Last run failed. Inspect `error` and call `speak()` again to retry. | Transitions, in plain terms: * The hook starts in `idle`. Calling `speak()` (or `sendText()` in WebSocket mode) moves it to `connecting`. * Once the first audio chunk is received, it moves to `speaking`. * When the server finishes (either because `speak()` passed a complete string, you called `finish()`, or the LLM async iterable ended), the hook fires `onTerminated` and returns to `idle`. * Calling `stop()` during `speaking` moves it to `stopping` and then back to `idle` once the server flushes. Calling `cancel()` at any point jumps straight back to `idle`. * Any error (connect failure or in-stream error) moves it to `error`. The next `speak()` resets it back to `connecting`. ## Methods | Method | Signature | Description | | ---------- | ------------------------------------------------- | ---------------------------------------------------------------------------------- | | `speak` | `(text: string \| AsyncIterable) => void` | Start a new run. Cancels any in-flight generation first. | | `sendText` | `(text: string) => void` | WebSocket only. Send one chunk without finishing. Use with `finish()`. | | `finish` | `() => void` | WebSocket only. Signal no more text — server finishes and sends `terminated`. | | `stop` | `() => Promise` | Graceful stop. Sends `finish()` and resolves when the server reaches `terminated`. | | `cancel` | `() => void` | Immediate cancel. Audio stops right away. | ## Callbacks | Callback | Signature | Description | | --------------- | ------------------------------------ | -------------------------------------------------------------- | | `onAudio` | `(chunk: Uint8Array) => void` | Audio chunk received. Fired for both REST and WebSocket modes. | | `onAudioEnd` | `() => void` | Server marked the final audio payload. | | `onTerminated` | `() => void` | Generation is fully complete. | | `onError` | `(error: Error) => void` | Stream or connection error. | | `onStateChange` | `({ old_state, new_state }) => void` | Fired on every state transition. | ## Return values Reactive snapshot exposed by the hook — see [`UseTtsReturn`](/sdk/react-SDK/reference/types#usettsreturn): | Field | Type | Description | | ----------------------------------------------- | ----------------------------------------------------- | ------------------------------ | | `state` | [`TtsState`](/sdk/react-SDK/reference/types#ttsstate) | Current lifecycle state. | | `isSpeaking` | `boolean` | `state === 'speaking'`. | | `isConnecting` | `boolean` | `state === 'connecting'`. | | `error` | `Error \| null` | Last error, or `null` if none. | | `speak`, `sendText`, `finish`, `stop`, `cancel` | — | Control methods (see above). | ## Configuration See [`UseTtsConfig`](/sdk/react-SDK/reference/types#usettsconfig) for the full type. Most-used options: | Option | Type | Description | | -------------- | ------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `config` | `SonioxConnectionConfig \| (() => Promise)` | Connection configuration. Required — `useTts` always uses its own client. | | `mode` | `'websocket' \| 'rest'` | Transport mode. Default `'websocket'`. | | `voice` | `string` | Voice identifier (e.g. `"Adrian"`). **Required** unless provided via server-returned `tts_defaults` (WebSocket mode only). Discover available voices via the Node SDK's `client.tts.listModels()`. | | `model` | `string` | TTS model. Default `"tts-rt-v2"`. | | `language` | `string` | Language code. Default `"en"`. | | `audio_format` | `TtsAudioFormat` | Output audio format. Default `"wav"`. | | `sample_rate` | `number` | Output sample rate in Hz. Required for raw PCM formats. | | `bitrate` | `number` | Codec bitrate in bps (for compressed formats). | | `stream_id` | `string` | Override the auto-generated stream id. WebSocket mode only. | See [Available models](/tts/models) for the full list of TTS models, voices, and supported audio formats. ## Server-driven defaults There's no first-class endpoint for TTS defaults — you own them. Keep them on your server next to the temporary-key endpoint and return them via `SonioxConnectionConfig.tts_defaults`. `useTts` consumes the defaults automatically through its `config` resolver; caller-provided hook options (`voice`, `model`, `audio_format`, ...) still override them. ```ts // server app.get('/tts-tmp-key', async (_req, res) => { const { api_key, expires_at } = await nodeClient.auth.createTemporaryKey({ usage_type: 'tts_rt', expires_in_seconds: 300, }); res.json({ api_key, expires_at, tts_defaults: { model: 'tts-rt-v2', language: 'en', voice: 'Adrian', audio_format: 'wav', }, }); }); ``` ```tsx // client async function fetchTtsKey() { const res = await fetch("/tts-tmp-key"); return await res.json(); // { api_key, tts_defaults, ... } } function TtsButton() { const audioChunksRef = useRef([]); // Inherits model / voice / audio_format from tts_defaults returned by the server. const { speak } = useTts({ config: fetchTtsKey, onAudio: (chunk) => audioChunksRef.current.push(chunk), }); return ; } ``` ## See also * [`useTts` return type](/sdk/react-SDK/reference/types#usettsreturn) and [config](/sdk/react-SDK/reference/types#usettsconfig) * [`TtsState` reference](/sdk/react-SDK/reference/types#ttsstate) * [Web SDK real-time TTS](/sdk/web-SDK/tts/realtime-speech-generation) — the underlying transport used in `'websocket'` mode. * [Web SDK REST TTS](/sdk/web-SDK/tts/rest-speech-generation) — the underlying transport used in `'rest'` mode. # Classes URL: /sdk/web-SDK/reference/classes Soniox Client SDK — Class Reference ## SonioxClient Main entry point for the Soniox client SDK.
### Example ```typescript // Recommended: async config with region const client = new SonioxClient({ config: async () => { const res = await fetch('/api/soniox-config', { method: 'POST' }); return await res.json(); // { api_key, region } }, }); // High-level: record from microphone const recording = client.realtime.record({ model: 'stt-rt-v5' }); recording.on('result', (r) => console.log(r.tokens)); await recording.stop(); // Low-level: direct session access const session = client.realtime.stt({ model: 'stt-rt-v5' }, { api_key: key }); await session.connect(); ``` ### permissions ```ts get permissions(): PermissionResolver | undefined; ``` Permission resolver, if configured. Returns `undefined` if no resolver was provided (SSR-safe). **Example** ```typescript const mic = await client.permissions?.check('microphone'); if (mic?.status === 'denied') { showSettingsMessage(); } ``` **Returns** [`PermissionResolver`](types#permissionresolver) | `undefined` ### Constructor ```ts new SonioxClient(options): SonioxClient; ``` **Parameters** | Parameter | Type | | --------- | -------------------------------------------------- | | `options` | [`SonioxClientOptions`](types#sonioxclientoptions) | **Returns** `SonioxClient` ### Properties | Property | Type | Description | | -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `realtime` | \{ `record`: (`options`) => [`Recording`](classes#recording); `stt`: (`config`, `options`) => [`RealtimeSttSession`](classes#realtimesttsession); `tts`: [`ClientTtsFactory`](types#clientttsfactory); } | Real-time API namespace | | `realtime.record` | (`options`) => [`Recording`](classes#recording) | Start a high-level recording session. Returns synchronously so callers can attach event listeners before any async work (key fetch, mic access, connection) begins. | | `realtime.stt` | (`config`, `options`) => [`RealtimeSttSession`](classes#realtimesttsession) | Create a low-level STT session. The WebSocket URL is derived from the client's `config` (respecting `region` / `base_domain` / `stt_ws_url`) when `config` is a plain object, or from `ws_base_url` on the legacy path. If `config` was passed as an async function, call `client.realtime.record()` instead, or pass `ws_base_url` explicitly to `SonioxClient`. **Throws** [SonioxError](classes#sonioxerror) if the WebSocket URL cannot be resolved synchronously (async-config client without `ws_base_url`). | | `realtime.tts` | [`ClientTtsFactory`](types#clientttsfactory) | TTS factory — callable for single-stream, `.multiStream()` for multi-stream. Uses the client's config resolver to obtain credentials and TTS WebSocket URL. **Examples** `const stream = await client.realtime.tts({ model: 'tts-rt-v2', voice: 'Adrian', language: 'en', audio_format: 'wav', }); stream.sendText("Hello"); stream.finish(); for await (const chunk of stream) { process(chunk); }` `const conn = await client.realtime.tts.multiStream(); const s1 = await conn.stream({ model: 'tts-rt-v2', voice: 'Adrian', language: 'en', audio_format: 'wav', });` | | `tts` | \{ `generate`: `Promise`\<`Uint8Array`\<`ArrayBufferLike`>>; `generateStream`: `AsyncIterable`\<`Uint8Array`\<`ArrayBufferLike`>>; } | REST TTS API namespace. **Example** `const audio = await client.tts.generate({ text: 'Hello', voice: 'Adrian', language: 'en', });` | | `tts.generate` | `Promise`\<`Uint8Array`\<`ArrayBufferLike`>> | - | | `tts.generateStream` | `AsyncIterable`\<`Uint8Array`\<`ArrayBufferLike`>> | - | *** ## Recording ### state ```ts get state(): RecordingState; ``` Current recording state **Returns** [`RecordingState`](types#recordingstate) ### cancel() ```ts cancel(): void; ``` Immediately cancel recording without waiting for final results **Returns** `void` *** ### finalize() ```ts finalize(options?): void; ``` Request the server to finalize current non-final tokens. **Parameters** | Parameter | Type | | ------------------------------ | -------------------------------------- | | `options?` | \{ `trailing_silence_ms?`: `number`; } | | `options.trailing_silence_ms?` | `number` | **Returns** `void` *** ### off() ```ts off(event, handler): this; ``` Remove an event handler **Type Parameters** | Type Parameter | | -------------------------------------------------------------- | | `E` *extends* keyof [`RecordingEvents`](types#recordingevents) | **Parameters** | Parameter | Type | | --------- | ------------------------------------------------ | | `event` | `E` | | `handler` | [`RecordingEvents`](types#recordingevents)\[`E`] | **Returns** `this` *** ### on() ```ts on(event, handler): this; ``` Register an event handler **Type Parameters** | Type Parameter | | -------------------------------------------------------------- | | `E` *extends* keyof [`RecordingEvents`](types#recordingevents) | **Parameters** | Parameter | Type | | --------- | ------------------------------------------------ | | `event` | `E` | | `handler` | [`RecordingEvents`](types#recordingevents)\[`E`] | **Returns** `this` *** ### once() ```ts once(event, handler): this; ``` Register a one-time event handler **Type Parameters** | Type Parameter | | -------------------------------------------------------------- | | `E` *extends* keyof [`RecordingEvents`](types#recordingevents) | **Parameters** | Parameter | Type | | --------- | ------------------------------------------------ | | `event` | `E` | | `handler` | [`RecordingEvents`](types#recordingevents)\[`E`] | **Returns** `this` *** ### pause() ```ts pause(): void; ``` Pause recording. Pauses the audio source (stops microphone capture) and pauses the session (activates automatic keepalive to prevent server disconnect). **Returns** `void` *** ### reconnect() ```ts reconnect(): void; ``` Force a reconnection — tears down the current session and audio encoder, then establishes a new session via the standard reconnect flow (backoff, config re-resolution, buffer drain). Use this to recover from stale connections after platform lifecycle events such as laptop sleep/wake (web `visibilitychange`) or app backgrounding (React Native `AppState`). Requires `auto_reconnect` to be enabled. No-op when the recording is not in `recording` or `paused` state. **Returns** `void` *** ### resume() ```ts resume(): void; ``` Resume recording after pause. Resumes the audio source and session. Audio capture and transmission continue from where they left off. If audio was buffered during a reconnect while paused, the buffer is drained now. **Returns** `void` *** ### stop() ```ts stop(): Promise; ``` Gracefully stop recording Stops the audio source and waits for the server to process all buffered audio and return final results. **Returns** `Promise`\<`void`> Promise that resolves when the server acknowledges completion *** ## MicrophoneSource Browser microphone audio source Uses `navigator.mediaDevices.getUserMedia` to capture audio from the microphone and `MediaRecorder` to encode it into chunks. ### Example ```typescript const source = new MicrophoneSource(); await source.start({ onData: (chunk) => session.sendAudio(chunk), onError: (err) => console.error(err), }); // Later: source.stop(); ``` ### Constructor ```ts new MicrophoneSource(options): MicrophoneSource; ``` **Parameters** | Parameter | Type | | --------- | ---------------------------------------------------------- | | `options` | [`MicrophoneSourceOptions`](types#microphonesourceoptions) | **Returns** `MicrophoneSource` ### pause() ```ts pause(): void; ``` Pause audio capture **Returns** `void` *** ### restart() ```ts restart(): void; ``` Reinitialize the MediaRecorder on the existing stream so the next chunks contain a fresh container header (required after reconnecting to a new server session). **Returns** `void` *** ### resume() ```ts resume(): void; ``` Resume audio capture **Returns** `void` *** ### start() ```ts start(handlers): Promise; ``` Request microphone access and start recording **Parameters** | Parameter | Type | | ---------- | -------------------------------------------------- | | `handlers` | [`AudioSourceHandlers`](types#audiosourcehandlers) | **Returns** `Promise`\<`void`> **Throws** AudioUnavailableError if getUserMedia or MediaRecorder is not supported **Throws** AudioPermissionError if microphone access is denied **Throws** AudioDeviceError if no microphone is found *** ### stop() ```ts stop(): void; ``` Stop recording and release all resources **Returns** `void` *** ## BrowserPermissionResolver Browser permission resolver for checking and requesting microphone access. ### Example ```typescript const resolver = new BrowserPermissionResolver(); const mic = await resolver.check('microphone'); if (mic.status === 'prompt') { const result = await resolver.request('microphone'); if (result.status === 'denied') { showDeniedMessage(); } } ``` ### Constructor ```ts new BrowserPermissionResolver(): BrowserPermissionResolver; ``` **Returns** `BrowserPermissionResolver` ### check() ```ts check(permission): Promise; ``` Check current microphone permission status without prompting the user. **Parameters** | Parameter | Type | | ------------ | -------------- | | `permission` | `"microphone"` | **Returns** `Promise`\<[`PermissionResult`](types#permissionresult)> *** ### request() ```ts request(permission): Promise; ``` Request microphone permission from the user. This may show a browser permission prompt. **Parameters** | Parameter | Type | | ------------ | -------------- | | `permission` | `"microphone"` | **Returns** `Promise`\<[`PermissionResult`](types#permissionresult)> *** ## AudioPermissionError Thrown when microphone access is denied by the user or blocked by the browser. Maps to `getUserMedia` `NotAllowedError` DOMException. ### Extends * [`SonioxError`](classes#sonioxerror) ### toJSON() ```ts toJSON(): Record; ``` Converts to a plain object for logging/serialization **Returns** `Record`\<`string`, `unknown`> **Inherited from** [`SonioxError`](classes#sonioxerror).[`toJSON`](classes#sonioxerror-tojson) *** ### toString() ```ts toString(): string; ``` Creates a human-readable string representation **Returns** `string` **Inherited from** [`SonioxError`](classes#sonioxerror).[`toString`](classes#sonioxerror-tostring) ### Properties | Property | Type | Description | | ------------ | --------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- | | `cause` | `unknown` | The underlying error that caused this error, if any. | | `code` | \| `SonioxErrorCode` \| `string` & \{ } | Error code describing the type of error. Typed as `string` at the base level to allow subclasses (e.g. HTTP errors) to use their own error code unions. | | `statusCode` | `number` \| `undefined` | HTTP status code when applicable (e.g., 401 for auth errors, 500 for server errors). | *** ## AudioDeviceError Thrown when no audio input device is found Maps to `getUserMedia` `NotFoundError` DOMException. ### Extends * [`SonioxError`](classes#sonioxerror) ### toJSON() ```ts toJSON(): Record; ``` Converts to a plain object for logging/serialization **Returns** `Record`\<`string`, `unknown`> **Inherited from** [`SonioxError`](classes#sonioxerror).[`toJSON`](classes#sonioxerror-tojson) *** ### toString() ```ts toString(): string; ``` Creates a human-readable string representation **Returns** `string` **Inherited from** [`SonioxError`](classes#sonioxerror).[`toString`](classes#sonioxerror-tostring) ### Properties | Property | Type | Description | | ------------ | --------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- | | `cause` | `unknown` | The underlying error that caused this error, if any. | | `code` | \| `SonioxErrorCode` \| `string` & \{ } | Error code describing the type of error. Typed as `string` at the base level to allow subclasses (e.g. HTTP errors) to use their own error code unions. | | `statusCode` | `number` \| `undefined` | HTTP status code when applicable (e.g., 401 for auth errors, 500 for server errors). | *** ## AudioUnavailableError Thrown when audio capture is not supported in the current environment For example, when `getUserMedia` or `MediaRecorder` is not available. ### Extends * [`SonioxError`](classes#sonioxerror) ### toJSON() ```ts toJSON(): Record; ``` Converts to a plain object for logging/serialization **Returns** `Record`\<`string`, `unknown`> **Inherited from** [`SonioxError`](classes#sonioxerror).[`toJSON`](classes#sonioxerror-tojson) *** ### toString() ```ts toString(): string; ``` Creates a human-readable string representation **Returns** `string` **Inherited from** [`SonioxError`](classes#sonioxerror).[`toString`](classes#sonioxerror-tostring) ### Properties | Property | Type | Description | | ------------ | --------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- | | `cause` | `unknown` | The underlying error that caused this error, if any. | | `code` | \| `SonioxErrorCode` \| `string` & \{ } | Error code describing the type of error. Typed as `string` at the base level to allow subclasses (e.g. HTTP errors) to use their own error code unions. | | `statusCode` | `number` \| `undefined` | HTTP status code when applicable (e.g., 401 for auth errors, 500 for server errors). | *** ## RealtimeSttSession Real-time Speech-to-Text session Provides WebSocket-based streaming transcription with support for: * Event-based and async iterator consumption * Pause/resume with automatic keepalive while paused * AbortSignal cancellation ### Example ```typescript const session = new RealtimeSttSession(apiKey, wsUrl, { model: 'stt-rt-v5' }); session.on('result', (result) => { console.log(result.tokens.map(t => t.text).join('')); }); await session.connect(); session.sendAudio(audioChunk); await session.finish(); ``` ### paused ```ts get paused(): boolean; ``` Whether the session is currently paused. **Returns** `boolean` *** ### state ```ts get state(): SttSessionState; ``` Current session state. **Returns** `SttSessionState` ### Constructor ```ts new RealtimeSttSession( apiKey, wsBaseUrl, config, options?): RealtimeSttSession; ``` **Parameters** | Parameter | Type | | ----------- | ------------------- | | `apiKey` | `string` | | `wsBaseUrl` | `string` | | `config` | `SttSessionConfig` | | `options?` | `SttSessionOptions` | **Returns** `RealtimeSttSession` ### \[asyncIterator]\() ```ts asyncIterator: AsyncIterator; ``` Async iterator for consuming events. The returned iterator's `return()` resets the internal iterator-attach flag and drops any buffered events, so consumers that exit `for await` early (via `break` etc.) stop accruing memory while the session keeps running. **Returns** `AsyncIterator`\<`RealtimeEvent`> *** ### close() ```ts close(): void; ``` Close (cancel) the session immediately without waiting **Returns** `void` *** ### connect() ```ts connect(): Promise; ``` Connect to the Soniox WebSocket API. **Returns** `Promise`\<`void`> **Throws** AbortError If aborted **Throws** ConnectionError If connection fails **Throws** StateError If already connected *** ### finalize() ```ts finalize(options?): void; ``` Requests the server to finalize current transcription **Parameters** | Parameter | Type | | ------------------------------ | -------------------------------------- | | `options?` | \{ `trailing_silence_ms?`: `number`; } | | `options.trailing_silence_ms?` | `number` | **Returns** `void` *** ### finish() ```ts finish(): Promise; ``` Gracefully finish the session **Returns** `Promise`\<`void`> *** ### keepAlive() ```ts keepAlive(): void; ``` Send a keepalive message **Returns** `void` *** ### off() ```ts off(event, handler): this; ``` Remove an event handler **Type Parameters** | Type Parameter | | -------------------------------------- | | `E` *extends* keyof `SttSessionEvents` | **Parameters** | Parameter | Type | | --------- | ------------------------ | | `event` | `E` | | `handler` | `SttSessionEvents`\[`E`] | **Returns** `this` *** ### on() ```ts on(event, handler): this; ``` Register an event handler **Type Parameters** | Type Parameter | | -------------------------------------- | | `E` *extends* keyof `SttSessionEvents` | **Parameters** | Parameter | Type | | --------- | ------------------------ | | `event` | `E` | | `handler` | `SttSessionEvents`\[`E`] | **Returns** `this` *** ### once() ```ts once(event, handler): this; ``` Register a one-time event handler **Type Parameters** | Type Parameter | | -------------------------------------- | | `E` *extends* keyof `SttSessionEvents` | **Parameters** | Parameter | Type | | --------- | ------------------------ | | `event` | `E` | | `handler` | `SttSessionEvents`\[`E`] | **Returns** `this` *** ### pause() ```ts pause(): void; ``` Pause audio transmission and starts automatic keepalive messages **Returns** `void` *** ### resume() ```ts resume(): void; ``` Resume audio transmission **Returns** `void` *** ### sendAudio() ```ts sendAudio(data): void; ``` Send audio data to the server **Parameters** | Parameter | Type | Description | | --------- | ----------- | --------------------------------------- | | `data` | `AudioData` | Audio data as Uint8Array or ArrayBuffer | **Returns** `void` **Throws** AbortError If aborted **Throws** StateError If not connected *** ### sendStream() ```ts sendStream(stream, options?): Promise; ``` Stream audio data from an async iterable source. **Parameters** | Parameter | Type | Description | | ---------- | ----------------------------- | ---------------------------------------- | | `stream` | `AsyncIterable`\<`AudioData`> | Async iterable yielding audio chunks | | `options?` | `SendStreamOptions` | Optional pacing and auto-finish settings | **Returns** `Promise`\<`void`> **Throws** AbortError If aborted during streaming **Throws** StateError If not connected *** ## RealtimeTtsConnection WebSocket connection for real-time Text-to-Speech. Supports up to 5 concurrent streams multiplexed by `stream_id`. The connection automatically sends keepalive messages while open. ### Example ```typescript const conn = new RealtimeTtsConnection(apiKey, wsUrl, ttsDefaults); await conn.connect(); const s1 = conn.stream({ model, voice, language, audio_format }); s1.sendText("Hello"); s1.finish(); for await (const chunk of s1) { ... } conn.close(); ``` ### Extends * `TypedEmitter`\<[`TtsConnectionEvents`](types#ttsconnectionevents)> ### isConnected ```ts get isConnected(): boolean; ``` Whether the WebSocket is connected. **Returns** `boolean` ### Constructor ```ts new RealtimeTtsConnection( apiKey, wsUrl, ttsDefaults?, options?): RealtimeTtsConnection; ``` **Parameters** | Parameter | Type | | -------------- | ------------------------------------------------------ | | `apiKey` | `string` | | `wsUrl` | `string` | | `ttsDefaults?` | `Partial`\<[`TtsStreamConfig`](types#ttsstreamconfig)> | | `options?` | [`TtsConnectionOptions`](types#ttsconnectionoptions) | **Returns** `RealtimeTtsConnection` **Overrides** ```ts TypedEmitter.constructor ``` ### close() ```ts close(): void; ``` Close the WebSocket connection and terminate all active streams. **Returns** `void` *** ### connect() ```ts connect(): Promise; ``` Open the WebSocket connection and start keepalive. Called automatically by [stream](#stream) if not yet connected. **Returns** `Promise`\<`void`> *** ### emit() ```ts emit(event, ...args): void; ``` Emit an event to all registered handlers. Handler errors do not prevent other handlers from running. Errors are reported to an `error` event if present, otherwise rethrown async. **Type Parameters** | Type Parameter | | ---------------------------------------------------------------------- | | `E` *extends* keyof [`TtsConnectionEvents`](types#ttsconnectionevents) | **Parameters** | Parameter | Type | | --------- | ----------------------------------------------------------------------- | | `event` | `E` | | ...`args` | `Parameters`\<[`TtsConnectionEvents`](types#ttsconnectionevents)\[`E`]> | **Returns** `void` **Inherited from** ```ts TypedEmitter.emit ``` *** ### off() ```ts off(event, handler): this; ``` Remove an event handler. **Type Parameters** | Type Parameter | | ---------------------------------------------------------------------- | | `E` *extends* keyof [`TtsConnectionEvents`](types#ttsconnectionevents) | **Parameters** | Parameter | Type | | --------- | -------------------------------------------------------- | | `event` | `E` | | `handler` | [`TtsConnectionEvents`](types#ttsconnectionevents)\[`E`] | **Returns** `this` **Inherited from** ```ts TypedEmitter.off ``` *** ### on() ```ts on(event, handler): this; ``` Register an event handler. **Type Parameters** | Type Parameter | | ---------------------------------------------------------------------- | | `E` *extends* keyof [`TtsConnectionEvents`](types#ttsconnectionevents) | **Parameters** | Parameter | Type | | --------- | -------------------------------------------------------- | | `event` | `E` | | `handler` | [`TtsConnectionEvents`](types#ttsconnectionevents)\[`E`] | **Returns** `this` **Inherited from** ```ts TypedEmitter.on ``` *** ### once() ```ts once(event, handler): this; ``` Register a one-time event handler. **Type Parameters** | Type Parameter | | ---------------------------------------------------------------------- | | `E` *extends* keyof [`TtsConnectionEvents`](types#ttsconnectionevents) | **Parameters** | Parameter | Type | | --------- | -------------------------------------------------------- | | `event` | `E` | | `handler` | [`TtsConnectionEvents`](types#ttsconnectionevents)\[`E`] | **Returns** `this` **Inherited from** ```ts TypedEmitter.once ``` *** ### removeAllListeners() ```ts removeAllListeners(event?): void; ``` Remove all event handlers. **Parameters** | Parameter | Type | | --------- | ------------------------- | | `event?` | keyof TtsConnectionEvents | **Returns** `void` **Inherited from** ```ts TypedEmitter.removeAllListeners ``` *** ### stream() ```ts stream(input?): Promise; ``` Open a new TTS stream on this connection. Auto-connects if the WebSocket is not yet open. **Parameters** | Parameter | Type | Description | | --------- | ---------------------------------------- | ------------------------------------------------ | | `input?` | [`TtsStreamInput`](types#ttsstreaminput) | Stream configuration (merged with tts\_defaults) | **Returns** `Promise`\<[`RealtimeTtsStream`](classes#realtimettsstream)> A ready-to-use stream handle *** ## RealtimeTtsStream Handle for one TTS stream on a WebSocket connection. Emits typed events and supports async iteration over decoded audio chunks. ### Examples ```typescript stream.on('audio', (chunk) => process(chunk)); stream.on('terminated', () => console.log('done')); stream.sendText("Hello world"); stream.finish(); ``` ```typescript stream.sendText("Hello world"); stream.finish(); for await (const chunk of stream) { process(chunk); } ``` ### Extends * `TypedEmitter`\<[`TtsStreamEvents`](types#ttsstreamevents)> ### state ```ts get state(): TtsStreamState; ``` Current stream lifecycle state. **Returns** [`TtsStreamState`](types#ttsstreamstate) ### \[asyncIterator]\() ```ts asyncIterator: AsyncIterator>; ``` Async iterator that yields decoded audio chunks. The returned iterator's `return()` resets the internal iterator-attach flag and drops any buffered audio, so consumers that exit `for await` early (via `break` etc.) stop accruing memory while the stream keeps receiving server audio. **Returns** `AsyncIterator`\<`Uint8Array`\<`ArrayBufferLike`>> *** ### cancel() ```ts cancel(): void; ``` Cancel this stream. The server will stop generating and send `terminated`. **Returns** `void` *** ### close() ```ts close(): void; ``` Close this stream. For single-stream usage (created via `tts(input)`), also closes the underlying WebSocket connection. **Returns** `void` *** ### emit() ```ts emit(event, ...args): void; ``` Emit an event to all registered handlers. Handler errors do not prevent other handlers from running. Errors are reported to an `error` event if present, otherwise rethrown async. **Type Parameters** | Type Parameter | | -------------------------------------------------------------- | | `E` *extends* keyof [`TtsStreamEvents`](types#ttsstreamevents) | **Parameters** | Parameter | Type | | --------- | --------------------------------------------------------------- | | `event` | `E` | | ...`args` | `Parameters`\<[`TtsStreamEvents`](types#ttsstreamevents)\[`E`]> | **Returns** `void` **Inherited from** ```ts TypedEmitter.emit ``` *** ### finish() ```ts finish(): void; ``` Signal that no more text will be sent for this stream. The server will finish generating audio and send `terminated`. **Returns** `void` *** ### off() ```ts off(event, handler): this; ``` Remove an event handler. **Type Parameters** | Type Parameter | | -------------------------------------------------------------- | | `E` *extends* keyof [`TtsStreamEvents`](types#ttsstreamevents) | **Parameters** | Parameter | Type | | --------- | ------------------------------------------------ | | `event` | `E` | | `handler` | [`TtsStreamEvents`](types#ttsstreamevents)\[`E`] | **Returns** `this` **Inherited from** ```ts TypedEmitter.off ``` *** ### on() ```ts on(event, handler): this; ``` Register an event handler. **Type Parameters** | Type Parameter | | -------------------------------------------------------------- | | `E` *extends* keyof [`TtsStreamEvents`](types#ttsstreamevents) | **Parameters** | Parameter | Type | | --------- | ------------------------------------------------ | | `event` | `E` | | `handler` | [`TtsStreamEvents`](types#ttsstreamevents)\[`E`] | **Returns** `this` **Inherited from** ```ts TypedEmitter.on ``` *** ### once() ```ts once(event, handler): this; ``` Register a one-time event handler. **Type Parameters** | Type Parameter | | -------------------------------------------------------------- | | `E` *extends* keyof [`TtsStreamEvents`](types#ttsstreamevents) | **Parameters** | Parameter | Type | | --------- | ------------------------------------------------ | | `event` | `E` | | `handler` | [`TtsStreamEvents`](types#ttsstreamevents)\[`E`] | **Returns** `this` **Inherited from** ```ts TypedEmitter.once ``` *** ### removeAllListeners() ```ts removeAllListeners(event?): void; ``` Remove all event handlers. **Parameters** | Parameter | Type | | --------- | --------------------- | | `event?` | keyof TtsStreamEvents | **Returns** `void` **Inherited from** ```ts TypedEmitter.removeAllListeners ``` *** ### sendStream() ```ts sendStream(source): Promise; ``` Pipe an async iterable of text chunks into the stream. Automatically calls [finish](#finish) when the iterable completes. Designed for concurrent use: call `sendStream()` and consume audio via `for await` or events simultaneously. **Parameters** | Parameter | Type | | --------- | -------------------------- | | `source` | `AsyncIterable`\<`string`> | **Returns** `Promise`\<`void`> **Example** ```typescript stream.sendStream(llmTokenStream); for await (const audio of stream) { forward(audio); } ``` *** ### sendText() ```ts sendText(text, options?): void; ``` Send one text chunk to the TTS stream. **Parameters** | Parameter | Type | Description | | -------------- | ----------------------- | --------------------------------------------- | | `text` | `string` | Text to synthesize | | `options?` | \{ `end?`: `boolean`; } | - | | `options.end?` | `boolean` | If true, signals this is the final text chunk | **Returns** `void` ### Properties | Property | Type | | ---------- | -------- | | `streamId` | `string` | *** ## SonioxError ### Extends * `Error` ### Extended by * [`AudioPermissionError`](classes#audiopermissionerror) * [`AudioDeviceError`](classes#audiodeviceerror) * [`AudioUnavailableError`](classes#audiounavailableerror) * [`SonioxHttpError`](classes#sonioxhttperror) ### Constructor ```ts new SonioxError( message, code?, statusCode?, cause?): SonioxError; ``` **Parameters** | Parameter | Type | | ------------- | --------------------------------------- | | `message` | `string` | | `code?` | \| `SonioxErrorCode` \| `string` & \{ } | | `statusCode?` | `number` | | `cause?` | `unknown` | **Returns** `SonioxError` **Overrides** ```ts Error.constructor ``` ### toJSON() ```ts toJSON(): Record; ``` Converts to a plain object for logging/serialization **Returns** `Record`\<`string`, `unknown`> *** ### toString() ```ts toString(): string; ``` Creates a human-readable string representation **Returns** `string` ### Properties | Property | Type | Description | | ------------ | --------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- | | `cause` | `unknown` | The underlying error that caused this error, if any. | | `code` | \| `SonioxErrorCode` \| `string` & \{ } | Error code describing the type of error. Typed as `string` at the base level to allow subclasses (e.g. HTTP errors) to use their own error code unions. | | `statusCode` | `number` \| `undefined` | HTTP status code when applicable (e.g., 401 for auth errors, 500 for server errors). | *** ## SonioxHttpError HTTP error class for all HTTP-related failures (REST API). Thrown when HTTP requests fail due to network issues, timeouts, server errors, or response parsing failures. ### Extends * [`SonioxError`](classes#sonioxerror) ### Constructor ```ts new SonioxHttpError(details): SonioxHttpError; ``` **Parameters** | Parameter | Type | | --------- | -------------------------------------------- | | `details` | [`HttpErrorDetails`](types#httperrordetails) | **Returns** `SonioxHttpError` **Overrides** [`SonioxError`](classes#sonioxerror).[`constructor`](classes#sonioxerror-constructor) ### toJSON() ```ts toJSON(): Record; ``` Converts to a plain object for logging/serialization **Returns** `Record`\<`string`, `unknown`> **Overrides** [`SonioxError`](classes#sonioxerror).[`toJSON`](classes#sonioxerror-tojson) *** ### toString() ```ts toString(): string; ``` Creates a human-readable string representation **Returns** `string` **Overrides** [`SonioxError`](classes#sonioxerror).[`toString`](classes#sonioxerror-tostring) ### Properties | Property | Type | Description | | ------------ | -------------------------------------------- | ------------------------------------------------------------------------------------ | | `bodyText` | `string` \| `undefined` | Response body text, capped at 4KB (only for http\_error/parse\_error) | | `cause` | `unknown` | The underlying error that caused this error, if any. | | `code` | [`HttpErrorCode`](types#httperrorcode) | Categorized HTTP error code | | `headers` | `Record`\<`string`, `string`> \| `undefined` | Response headers (only for http\_error) | | `method` | [`HttpMethod`](types#httpmethod) | HTTP method | | `statusCode` | `number` \| `undefined` | HTTP status code when applicable (e.g., 401 for auth errors, 500 for server errors). | | `url` | `string` | Request URL | *** ## TtsRestClient Browser-safe REST client for TTS generation. Provides `generate()` (buffered) and `generateStream()` (streaming) using only `globalThis.fetch`. HTTP failures are surfaced as [SonioxHttpError](classes#sonioxhttperror), matching the rest of the Soniox SDK. Authentication uses the `Authorization: Bearer ` header. ### Example ```typescript const client = new TtsRestClient(apiKey, 'https://tts-rt.soniox.com'); const audio = await client.generate({ text: 'Hello', voice: 'Adrian' }); ``` ### Constructor ```ts new TtsRestClient(apiKey, ttsApiUrl): TtsRestClient; ``` **Parameters** | Parameter | Type | | ----------- | -------- | | `apiKey` | `string` | | `ttsApiUrl` | `string` | **Returns** `TtsRestClient` ### generate() ```ts generate(options): Promise>; ``` Generate speech audio from text. Returns the full audio as a `Uint8Array`. **Parameters** | Parameter | Type | | --------- | ------------------------------------------------------ | | `options` | [`GenerateSpeechOptions`](types#generatespeechoptions) | **Returns** `Promise`\<`Uint8Array`\<`ArrayBufferLike`>> **Throws** [SonioxHttpError](classes#sonioxhttperror) on non-2xx responses, network failures, or aborted requests. *** ### generateStream() ```ts generateStream(options): AsyncIterable>; ``` Generate speech audio from text as a streaming async iterable. Yields `Uint8Array` chunks as they arrive from the server response body. Lower time-to-first-audio than [generate](#generate). **Known limitation:** Mid-stream server errors (reported via HTTP trailers) cannot be detected through the `fetch` API. The iterator may end early without an explicit error. Use WebSocket TTS for reliable error detection. **Parameters** | Parameter | Type | | --------- | ------------------------------------------------------ | | `options` | [`GenerateSpeechOptions`](types#generatespeechoptions) | **Returns** `AsyncIterable`\<`Uint8Array`\<`ArrayBufferLike`>> **Throws** [SonioxHttpError](classes#sonioxhttperror) on non-2xx responses, network failures, or aborted requests (before the stream starts). # Full Web SDK reference URL: /sdk/web-SDK/reference Full SDK reference for the Web SDK ## Client ### Available client methods | Method | Description | | ------------------------------------------------------------------------------------------- | ----------------------------------------------------- | | [`client.realtime.record()`](/sdk/web-SDK/reference/classes#recording) | Create a recording instance | | [`client.realtime.stt()`](/sdk/web-SDK/reference/classes#realtimesttsession) | Direct low-level STT session | | [`client.realtime.tts()`](/sdk/web-SDK/reference/classes#realtimettsstream) | Create a TTS stream instance | | [`client.realtime.tts.multiStream()`](/sdk/web-SDK/reference/classes#realtimettsconnection) | Create a TTS connection instance for multi-stream TTS | ## STT Recording ### Available recording methods | Method | Description | | ----------------------------------------------------------------- | ------------------------------------------------------- | | [`recording.finalize()`](/sdk/web-SDK/reference/classes#finalize) | Request the server to finalize current non-final tokens | | [`recording.on()`](/sdk/web-SDK/reference/classes#on) | Register an event handler | | [`recording.once()`](/sdk/web-SDK/reference/classes#once) | Register a one-time event handler | | [`recording.off()`](/sdk/web-SDK/reference/classes#off) | Remove an event handler | | [`recording.pause()`](/sdk/web-SDK/reference/classes#pause) | Pause recording | | [`recording.resume()`](/sdk/web-SDK/reference/classes#resume) | Resume recording | | [`recording.stop()`](/sdk/web-SDK/reference/classes#stop) | Stop recording | | [`recording.cancel()`](/sdk/web-SDK/reference/classes#cancel) | Cancel recording | ## TTS Stream ### Available TTS stream methods | Method | Description | | ------------------------------------------------------------------------------------ | ----------------------------------------------------- | | [`stream.sendText()`](/sdk/web-SDK/reference/classes#realtimettsstream-sendtext) | Send text to the TTS stream | | [`stream.sendStream()`](/sdk/web-SDK/reference/classes#realtimettsstream-sendstream) | Pipe an async iterable of text chunks into the stream | | [`stream.finish()`](/sdk/web-SDK/reference/classes#realtimettsstream-finish) | Finish the TTS stream | | [`stream.cancel()`](/sdk/web-SDK/reference/classes#realtimettsstream-cancel) | Cancel the TTS stream | | [`stream.close()`](/sdk/web-SDK/reference/classes#realtimettsstream-close) | Close the TTS stream | | [`stream.on()`](/sdk/web-SDK/reference/classes#realtimettsstream-on) | Register an event handler | | [`stream.once()`](/sdk/web-SDK/reference/classes#realtimettsstream-once) | Register a one-time event handler | | [`stream.off()`](/sdk/web-SDK/reference/classes#realtimettsstream-off) | Remove an event handler | ## TTS Connection ### Available TTS connection methods | Method | Description | | -------------------------------------------------------------------------------------- | --------------------------------------------------------------- | | [`connection.connect()`](/sdk/web-SDK/reference/classes#realtimettsconnection-connect) | Open the WebSocket connection | | [`connection.stream()`](/sdk/web-SDK/reference/classes#realtimettsconnection-stream) | Open a new TTS stream on this connection | | [`connection.close()`](/sdk/web-SDK/reference/classes#realtimettsconnection-close) | Close the WebSocket connection and terminate all active streams | | [`connection.on()`](/sdk/web-SDK/reference/classes#realtimettsconnection-on) | Register an event handler | | [`connection.once()`](/sdk/web-SDK/reference/classes#realtimettsconnection-once) | Register a one-time event handler | | [`connection.off()`](/sdk/web-SDK/reference/classes#realtimettsconnection-off) | Remove an event handler | ## AudioSource ### Available audio source methods | Method | Description | | ------------------------------------------------------------- | ---------------------- | | [`source.start()`](/sdk/web-SDK/reference/types#audiosource) | Start capturing audio | | [`source.stop()`](/sdk/web-SDK/reference/types#audiosource) | Stop capturing audio | | [`source.pause()`](/sdk/web-SDK/reference/types#audiosource) | Pause capturing audio | | [`source.resume()`](/sdk/web-SDK/reference/types#audiosource) | Resume capturing audio | ## PermissionResolver ### Available browser permission resolver methods | Method | Description | | ----------------------------------------------------------------------- | -------------------------------- | | [`resolver.check()`](/sdk/web-SDK/reference/types#permissionresolver) | Check current permission status | | [`resolver.request()`](/sdk/web-SDK/reference/types#permissionresolver) | Request permission from the user | # Types URL: /sdk/web-SDK/reference/types Soniox Client SDK — Types Reference ## ApiKeyConfig ```ts type ApiKeyConfig = string | () => Promise; ``` API key configuration. * `string` - A pre-fetched temporary API key (e.g., injected from SSR) * `() => Promise` - An async function that fetches a fresh temporary key from your backend. Called once per recording session. **Deprecated** Use SonioxConnectionConfig with `SonioxClientOptions.config` instead. **Example** ```typescript // Static key (for demos or SSR-injected keys) const client = new SonioxClient({ api_key: 'snx_temp_...' }); // Async function (recommended for production) const client = new SonioxClient({ api_key: async () => { const res = await fetch('/api/get-temporary-key', { method: 'POST' }); const { api_key } = await res.json(); return api_key; }, }); ``` Note: If you use Node.js, you can use the `SonioxNodeClient` to fetch a temporary API key via `client.auth.createTemporaryKey()`. *** ## AudioErrorCode ```ts type AudioErrorCode = "permission_denied" | "device_not_found" | "audio_unavailable"; ``` Error codes for audio-related errors *** ## AudioSourceHandlers ```ts type AudioSourceHandlers = { onData: (chunk) => void; onError: (error) => void; onMuted?: () => void; onUnmuted?: () => void; }; ``` Callbacks for receiving audio data and errors from an AudioSource. **Properties** | Property | Type | Description | | ----------------------------------------------------- | ------------------- | ---------------------------------------------------------------------------------- | | `onData` | (`chunk`) => `void` | Called when an audio chunk is available. | | `onError` | (`error`) => `void` | Called when a runtime error occurs during audio capture (after start). | | `onMuted?` | () => `void` | Called when the audio source is muted externally (e.g. OS-level or hardware mute). | | `onUnmuted?` | () => `void` | Called when the audio source is unmuted after an external mute. | *** ## GenerateSpeechOptions ```ts type GenerateSpeechOptions = { audio_format?: string; bitrate?: number; client_reference_id?: string; language?: string; model?: string; reduce_silence?: boolean; sample_rate?: number; signal?: AbortSignal; speed?: number; text: string; voice: string; }; ``` Options for REST TTS generation (`generate` / `generateStream`). **Properties** | Property | Type | Description | | --------------------------------------------------------------------------- | ------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `audio_format?` | `string` | Output audio format **Default** `'wav'` | | `bitrate?` | `number` | Codec bitrate in bps (for compressed formats). | | `client_reference_id?` | `string` | Optional tracking identifier. Does not need to be unique. Ignored if the request authenticates with a temporary API key. **Max Length** 256 | | `language?` | `string` | Language code. **Default** `'en'` | | `model?` | `string` | Text-to-Speech model to use. **Default** `'tts-rt-v2'` | | `reduce_silence?` | `boolean` | Shorten pauses between words in the generated speech. `false` (default) keeps the model's natural pacing; `true` tightens delivery by reducing silence between words. | | `sample_rate?` | `number` | Output sample rate in Hz. Required for raw PCM formats. | | `signal?` | `AbortSignal` | Optional AbortSignal for cancellation. | | `speed?` | `number` | Speaking rate. `1.0` is the normal rate; values below `1.0` slow speech down and values above `1.0` speed it up. Supported range is `0.7`-`1.3`. Defaults to `1.0` when omitted. | | `text` | `string` | Input text to generate as speech. | | `voice` | `string` | Voice identifier. | *** ## HttpErrorCode ```ts type HttpErrorCode = "network_error" | "timeout" | "aborted" | "http_error" | "parse_error"; ``` Error codes for HTTP client errors *** ## HttpMethod ```ts type HttpMethod = "GET" | "POST" | "PUT" | "PATCH" | "DELETE" | "HEAD"; ``` HTTP methods supported by the client *** ## MicrophoneSourceOptions ```ts type MicrophoneSourceOptions = { constraints?: MediaTrackConstraints; recorderOptions?: MediaRecorderOptions; timesliceMs?: number; }; ``` Options for MicrophoneSource **Properties** | Property | Type | Description | | --------------------------------------------------------------------- | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `constraints?` | `MediaTrackConstraints` | MediaTrackConstraints for the audio track. **Default** `{ echoCancellation: false, noiseSuppression: false, autoGainControl: false, channelCount: 1, sampleRate: 16000 }` | | `recorderOptions?` | `MediaRecorderOptions` | MediaRecorder options. **See** [https://developer.mozilla.org/en-US/docs/Web/API/MediaRecorder/MediaRecorder](https://developer.mozilla.org/en-US/docs/Web/API/MediaRecorder/MediaRecorder) | | `timesliceMs?` | `number` | Time interval in milliseconds between audio data chunks. **Default** `60` | *** ## PermissionResult ```ts type PermissionResult = { can_request: boolean; status: PermissionStatus; }; ``` Result of a permission check or request. **Properties** | Property | Type | Description | | ----------------------------------------------------- | -------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `can_request` | `boolean` | Whether the user can be prompted again. `false` means permanently denied (e.g., browser "Block" or iOS settings). Useful for showing "go to settings" instructions. | | `status` | [`PermissionStatus`](types#permissionstatus) | Current permission status. | *** ## PermissionStatus ```ts type PermissionStatus = "granted" | "denied" | "prompt" | "unavailable"; ``` Unified permission status across all platforms. *** ## PermissionType ```ts type PermissionType = "microphone"; ``` Permission types supported by the resolver. *** ## RecordOptions ```ts type RecordOptions = SttSessionConfig & ReconnectOptions & { buffer_queue_size?: number; session_config?: (resolved) => SttSessionConfig; session_options?: SttSessionOptions; signal?: AbortSignal; source?: AudioSource; }; ``` Options for creating a recording **Type Declaration** | Name | Type | Description | | -------------------- | ---------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `buffer_queue_size?` | `number` | Maximum number of audio chunks to buffer while waiting for key/connection **Default** `1000` | | `session_config()?` | (`resolved`) => `SttSessionConfig` | Function that receives the resolved connection config (including `stt_defaults` from the server) and returns the final session config. When provided, its return value is used as the session config, and any flat session config fields on this object are ignored. **Example** `client.realtime.record({ session_config: (resolved) => ({ ...resolved.stt_defaults, enable_endpoint_detection: true, }), });` | | `session_options?` | `SttSessionOptions` | SDK-level session options (signal, etc.) | | `signal?` | `AbortSignal` | AbortSignal for cancellation | | `source?` | [`AudioSource`](types#audiosource) | Audio source to use. Defaults to MicrophoneSource if not provided. | *** ## RecordingEvents ```ts type RecordingEvents = { connected: () => void; endpoint: () => void; error: (error) => void; finalized: () => void; finished: () => void; reconnected: (event) => void; reconnecting: (event) => void; result: (result) => void; session_restart: (event) => void; source_muted: () => void; source_unmuted: () => void; state_change: (update) => void; token: (token) => void; }; ``` Events emitted by a Recording instance **Properties** | Property | Type | Description | | ------------------------------------------------------------ | -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `connected` | () => `void` | WebSocket connected and ready. | | `endpoint` | () => `void` | Endpoint detected (speaker finished talking). | | `error` | (`error`) => `void` | Error occurred during recording. | | `finalized` | () => `void` | Finalization complete. | | `finished` | () => `void` | Recording finished (server acknowledged end of stream). | | `reconnected` | (`event`) => `void` | Successfully reconnected after a drop. | | `reconnecting` | (`event`) => `void` | About to attempt a reconnection. Call `preventDefault()` to cancel. | | `result` | (`result`) => `void` | Parsed result received from the server. | | `session_restart` | (`event`) => `void` | New STT session started (initial or after reconnect). Consumers should reset any session-local tracking state (e.g. token window comparisons). The `reset_transcript` flag indicates whether accumulated transcript state should also be cleared. | | `source_muted` | () => `void` | Audio source was muted externally (e.g. OS-level or hardware mute). | | `source_unmuted` | () => `void` | Audio source was unmuted after an external mute. | | `state_change` | (`update`) => `void` | Recording state transition. | | `token` | (`token`) => `void` | Individual token received. | *** ## RecordingState ```ts type RecordingState = | "idle" | "starting" | "connecting" | "recording" | "paused" | "reconnecting" | "stopping" | "stopped" | "error" | "canceled"; ``` Unified recording lifecycle states. *** ## SonioxClientOptions ```ts type SonioxClientOptions = { api_key?: ApiKeyConfig; buffer_queue_size?: number; config?: | SonioxConnectionConfig | (context?) => Promise; default_session_options?: SttSessionOptions; permissions?: PermissionResolver; ws_base_url?: string; }; ``` Options for creating a SonioxClient instance. **Properties** | Property | Type | Description | | --------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | ~~`api_key?`~~ | [`ApiKeyConfig`](types#apikeyconfig) | API key configuration. - `string` - A pre-fetched temporary API key (e.g., injected from SSR) - `() => Promise` - Async function that fetches a fresh key from your backend **Deprecated** Use `config` instead. | | `buffer_queue_size?` | `number` | Default maximum number of audio chunks to buffer while waiting for key/connection. Can be overridden per-recording. **Default** `1000` | | `config?` | \| `SonioxConnectionConfig` \| (`context?`) => `Promise`\<`SonioxConnectionConfig`> | Connection configuration — sync object or async function. When provided as a function, it is called once per recording session, allowing you to fetch a fresh temporary API key and connection settings from your backend at runtime. **Example** `// Sync config with region const client = new SonioxClient({ config: { api_key: tempKey, region: 'eu' }, }); // Async config (recommended for production) const client = new SonioxClient({ config: async () => { const res = await fetch('/api/soniox-config', { method: 'POST' }); return await res.json(); // { api_key, region, ... } }, });` | | `default_session_options?` | `SttSessionOptions` | Default session options applied to all sessions. Can be overridden per-recording. | | `permissions?` | [`PermissionResolver`](types#permissionresolver) | Optional permission resolver for pre-flight microphone permission checks. Not set by default (SSR-safe, RN-safe). **Example** `import { BrowserPermissionResolver } from '@soniox/client'; const client = new SonioxClient({ config: { api_key: tempKey }, permissions: new BrowserPermissionResolver(), });` | | ~~`ws_base_url?`~~ | `string` | WebSocket URL for real-time connections. **Default** `'wss://stt-rt.soniox.com/transcribe-websocket'` **Deprecated** Use `config.stt_ws_url` or `config.region` instead. | *** ## SttOptions ```ts type SttOptions = { api_key: string; session_options?: SttSessionOptions; }; ``` Options for creating a low-level STT session. **Properties** | Property | Type | Description | | -------------------------------------------------------- | ------------------- | ---------------------------------------- | | `api_key` | `string` | Resolved API key string (temporary key). | | `session_options?` | `SttSessionOptions` | Session options (signal, etc.). | *** ## TtsAudioFormat ```ts type TtsAudioFormat = | "pcm_f32le" | "pcm_s16le" | "pcm_s16be" | "pcm_mulaw" | "pcm_alaw" | "wav" | "aac" | "mp3" | "opus" | "flac" | string & { }; ``` Supported audio formats for Text-to-Speech output. *** ## TtsConnectionEvents ```ts type TtsConnectionEvents = { close: () => void; error: (error) => void; }; ``` Events emitted by a TTS WebSocket connection. **Properties** | Property | Type | Description | | -------------------------------------------- | ------------------- | ------------------------------------------------------------------------------------------------------------------- | | `close` | () => `void` | The WebSocket connection was closed. | | `error` | (`error`) => `void` | A connection-level error occurred. Always a RealtimeError subclass (e.g. ConnectionError, NetworkError, AuthError). | *** ## TtsConnectionOptions ```ts type TtsConnectionOptions = { connect_timeout_ms?: number; keepalive_interval_ms?: number; }; ``` Options for creating a TTS connection. **Properties** | Property | Type | Description | | ------------------------------------------------------------------------------ | -------- | --------------------------------------------------------------------------------------------- | | `connect_timeout_ms?` | `number` | Maximum time to wait for the WebSocket connection to open (milliseconds). **Default** `20000` | | `keepalive_interval_ms?` | `number` | Interval for sending keepalive messages (milliseconds). **Default** `5000` **Minimum** 1000 | *** ## TtsLanguage ```ts type TtsLanguage = { code: string; name: string; }; ``` A language supported by a Text-to-Speech model. **Properties** | Property | Type | Description | | ---------------------------------- | -------- | ----------------------------- | | `code` | `string` | ISO language code. | | `name` | `string` | Human-readable language name. | *** ## TtsModel ```ts type TtsModel = { aliased_model_id: string | null; id: string; languages: TtsLanguage[]; name: string; speed_max: number; speed_min: number; supports_silence_reduction: boolean; supports_speed_adjustment: boolean; supports_timestamps?: boolean; voices: TtsVoice[]; }; ``` A Text-to-Speech model. **Properties** | Property | Type | Description | | --------------------------------------------------------------------------- | ------------------------------------- | ---------------------------------------------------------------------------------------------- | | `aliased_model_id` | `string` \| `null` | If this is an alias, the id of the aliased model. Null for non-alias models. | | `id` | `string` | Unique identifier of the model. | | `languages` | [`TtsLanguage`](types#ttslanguage)\[] | Languages supported by this model. | | `name` | `string` | Name of the model. | | `speed_max` | `number` | Maximum supported speaking rate. | | `speed_min` | `number` | Minimum supported speaking rate. | | `supports_silence_reduction` | `boolean` | Whether the model supports shortening pauses between words via the `reduce_silence` parameter. | | `supports_speed_adjustment` | `boolean` | Whether the model supports adjusting the speaking rate via the `speed` parameter. | | `supports_timestamps?` | `boolean` | Whether the model can return character-level audio timestamps via `return_timestamps`. | | `voices` | [`TtsVoice`](types#ttsvoice)\[] | Voices supported by this model. | *** ## TtsStreamConfig ```ts type TtsStreamConfig = { audio_format: string; bitrate?: number; language: string; model: string; reduce_silence?: boolean; return_timestamps?: boolean; sample_rate?: number; speed?: number; stream_id: string; voice: string; }; ``` Fully resolved TTS stream config sent over the WebSocket. All required fields are present after merging input with defaults. **Properties** | Property | Type | | ----------------------------------------------------------------- | --------- | | `audio_format` | `string` | | `bitrate?` | `number` | | `language` | `string` | | `model` | `string` | | `reduce_silence?` | `boolean` | | `return_timestamps?` | `boolean` | | `sample_rate?` | `number` | | `speed?` | `number` | | `stream_id` | `string` | | `voice` | `string` | *** ## TtsStreamEvents ```ts type TtsStreamEvents = { audio: (chunk, timestamps?) => void; audioEnd: () => void; error: (error) => void; terminated: () => void; }; ``` Events emitted by a TTS stream. **Properties** | Property | Type | Description | | -------------------------------------------------- | ---------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `audio` | (`chunk`, `timestamps?`) => `void` | Decoded audio chunk received. When `return_timestamps` is enabled, the second argument carries the character-level alignment for this frame (it is `undefined` for audio-only frames). | | `audioEnd` | () => `void` | Server marked the final audio payload for this stream. | | `error` | (`error`) => `void` | A stream-level error occurred. Always a RealtimeError subclass mapped from the server `error_code` / `error_message`. | | `terminated` | () => `void` | Stream has been fully terminated by the server. | *** ## TtsStreamInput ```ts type TtsStreamInput = { audio_format?: TtsAudioFormat; bitrate?: number; language?: string; model?: string; reduce_silence?: boolean; return_timestamps?: boolean; sample_rate?: number; speed?: number; stream_id?: string; voice?: string; }; ``` Input for creating a TTS stream. All fields are optional and are merged with `tts_defaults` from the resolved connection config. After merging, `model`, `language`, `voice`, and `audio_format` must be present. **Properties** | Property | Type | Description | | ---------------------------------------------------------------- | ---------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `audio_format?` | [`TtsAudioFormat`](types#ttsaudioformat) | Output audio format **Example** `'wav'` | | `bitrate?` | `number` | Codec bitrate in bps (for compressed formats). | | `language?` | `string` | Language code for speech generation. **Example** `'en'` | | `model?` | `string` | Text-to-Speech model to use. **Example** `'tts-rt-v2'` | | `reduce_silence?` | `boolean` | Shorten pauses between words in the generated speech. `false` (default) keeps the model's natural pacing; `true` tightens delivery by reducing silence between words. | | `return_timestamps?` | `boolean` | Request character-level audio timestamps in the responses. When enabled, audio frames may carry a TtsTimestamps payload aligning each character of the spoken text to its start/end time in the audio. WebSocket (realtime) only — the REST endpoint streams raw audio bytes and ignores this flag. Timestamps map to the model's preprocessed text, not the raw input. Defaults to `false` when omitted. | | `sample_rate?` | `number` | Output sample rate in Hz. Required for raw PCM formats. | | `speed?` | `number` | Speaking rate. `1.0` is the normal rate; values below `1.0` slow speech down and values above `1.0` speed it up. Supported range is `0.7`-`1.3`. Defaults to `1.0` when omitted. | | `stream_id?` | `string` | Client-generated stream identifier. Must be unique among active streams on the same connection. Auto-generated if omitted. | | `voice?` | `string` | Voice identifier. **Example** `'Adrian'` | *** ## TtsStreamState ```ts type TtsStreamState = "active" | "finishing" | "ended" | "error"; ``` Lifecycle states for a TTS stream. *** ## TtsVoice ```ts type TtsVoice = { description: string; gender: TtsVoiceGender; id: string; }; ``` A Text-to-Speech voice. **Properties** | Property | Type | Description | | --------------------------------------------- | ---------------------------------------- | --------------------------------- | | `description` | `string` | Human-readable voice description. | | `gender` | [`TtsVoiceGender`](types#ttsvoicegender) | Voice gender metadata. | | `id` | `string` | Unique identifier of the voice. | *** ## TtsVoiceGender ```ts type TtsVoiceGender = "male" | "female" | "neutral"; ``` Voice gender metadata returned by the TTS models API. *** ## AudioSource Platform-agnostic audio source interface. Implementations must: * Begin capturing audio in `start()` and deliver chunks via `handlers.onData` * Stop all capture and release resources in `stop()` * Throw typed errors from `start()` if capture cannot begin (e.g., permission denied) **Example** ```typescript // Built-in browser source const source = new MicrophoneSource(); // Custom source (e.g., React Native) class MyAudioSource implements AudioSource { async start(handlers: AudioSourceHandlers) { ... } stop() { ... } } ``` **Methods** **pause()?** ```ts optional pause(): void; ``` Pause audio capture (optional). When paused, no data should be delivered via onData. **Returns** `void` *** **restart()?** ```ts optional restart(): void; ``` Reinitialize the audio encoder without releasing the underlying capture device (optional). Called during reconnection so the new server session receives a fresh audio stream with proper container headers. Implementations that produce a header-less format (e.g. raw PCM) can omit this. **Returns** `void` *** **resume()?** ```ts optional resume(): void; ``` Resume audio capture after pause (optional). **Returns** `void` *** **start()** ```ts start(handlers): Promise; ``` Start capturing audio. **Parameters** | Parameter | Type | Description | | ---------- | -------------------------------------------------- | ----------------------------------- | | `handlers` | [`AudioSourceHandlers`](types#audiosourcehandlers) | Callbacks for audio data and errors | **Returns** `Promise`\<`void`> **Throws** AudioPermissionError if microphone access is denied **Throws** AudioDeviceError if no audio device is found **Throws** AudioUnavailableError if audio capture is not supported *** **stop()** ```ts stop(): void; ``` Stop capturing audio and release all resources. Safe to call multiple times. **Returns** `void` *** ## ClientTtsFactory() Callable TTS factory with `.multiStream()` for multi-stream connections. ```ts ClientTtsFactory(input?): Promise; ``` Callable TTS factory with `.multiStream()` for multi-stream connections. **Parameters** | Parameter | Type | | --------- | ---------------------------------------- | | `input?` | [`TtsStreamInput`](types#ttsstreaminput) | **Returns** `Promise`\<[`RealtimeTtsStream`](classes#realtimettsstream)> **Methods** **multiStream()** ```ts multiStream(): Promise; ``` **Returns** `Promise`\<[`RealtimeTtsConnection`](classes#realtimettsconnection)> *** ## HttpErrorDetails Error details for SonioxHttpError **Properties** | Property | Type | Description | | ---------------------------------------------------- | -------------------------------------- | ---------------------------------- | | `bodyText?` | `string` | Response body text (capped at 4KB) | | `cause?` | `unknown` | - | | `code` | [`HttpErrorCode`](types#httperrorcode) | - | | `headers?` | `Record`\<`string`, `string`> | - | | `message` | `string` | - | | `method` | [`HttpMethod`](types#httpmethod) | - | | `statusCode?` | `number` | - | | `url` | `string` | - | *** ## PermissionResolver Platform-agnostic permission resolver. Implementations handle platform-specific permission APIs: * Browser: `navigator.permissions.query` + `getUserMedia` * React Native: `expo-av` or `react-native-permissions` **Example** ```typescript // Check before recording const mic = await resolver.check('microphone'); if (mic.status === 'denied' && !mic.can_request) { showGoToSettingsMessage(); } ``` **Methods** **check()** ```ts check(permission): Promise; ``` Check current permission status WITHOUT prompting the user. **Parameters** | Parameter | Type | | ------------ | -------------- | | `permission` | `"microphone"` | **Returns** `Promise`\<[`PermissionResult`](types#permissionresult)> *** **request()** ```ts request(permission): Promise; ``` Request permission from the user (may show a system prompt). On platforms where status is already 'granted', this is a no-op. **Parameters** | Parameter | Type | | ------------ | -------------- | | `permission` | `"microphone"` | **Returns** `Promise`\<[`PermissionResult`](types#permissionresult)> *** ## resolveApiKey() ```ts function resolveApiKey(config): Promise; ``` Resolves an ApiKeyConfig to a plain API key string. **Parameters** | Parameter | Type | Description | | --------- | ------------------------------------ | ------------------------- | | `config` | [`ApiKeyConfig`](types#apikeyconfig) | The API key configuration | **Returns** `Promise`\<`string`> The resolved API key string **Throws** If the function rejects or returns a non-string value **Deprecated** Use SonioxConnectionConfig with `SonioxClientOptions.config` instead. # Real-time transcription with Web SDK URL: /sdk/web-SDK/stt/realtime-transcription Create and manage real-time speech-to-text sessions with the Soniox Web SDK Soniox Web SDK supports real-time transcription over WebSocket directly in the browser. This allows you to transcribe live audio with low latency — ideal for live captions, voice input, and interactive experiences. You can capture audio from the user's microphone, consume results via events or buffers that group tokens into utterances, and manage sessions with built-in connection handling. ## Create a real-time recording session `client.realtime.record()` is the high-level API for capturing audio and streaming it to Soniox for real-time transcription. It returns a [`Recording`](/sdk/web-SDK/reference/classes#recording) instance synchronously so you can attach event listeners before any async work (microphone access, API key fetch, WebSocket connection) begins. ```typescript const recording = client.realtime.record({ // speech-to-text model to use model: "stt-rt-v5", // Optional: hint expected languages language_hints: ["en", "es"], // Optional: enable speaker identification enable_speaker_diarization: true, // Optional: detect utterance boundaries (useful for voice agents) enable_endpoint_detection: true, // Optional: provide domain context to improve accuracy context: { terms: ["Soniox", "WebSocket"], general: [{ key: "domain", value: "technology" }], }, // ... other options ... }); ``` ### Listen for results The `result` event fires every time the server returns a transcription update. Each `RealtimeResult` contains an array of `RealtimeToken` objects — both finalized and in-progress tokens. ```typescript recording.on("result", (result) => { const text = result.tokens.map((t) => t.text).join(""); if (text) console.log(text); }); ``` ## Handle session events | Event | Payload | Description | | ----------------- | ----------------------------------------------------- | -------------------------------------------------------------------------- | | `result` | `RealtimeResult` | Transcription result received from the server. | | `token` | `RealtimeToken` | Individual token received (fires once per token, in addition to `result`). | | `error` | `Error` | An error occurred during recording. | | `endpoint` | — | Endpoint detected (speaker finished talking). | | `finalized` | — | Server completed finalization of current tokens. | | `finished` | — | Server acknowledged end of stream. Fires before `stopped` state. | | `connected` | — | WebSocket connected and streaming. | | `state_change` | `{ old_state, new_state, reason? }` | Recording state transition. | | `reconnecting` | `{ attempt, max_attempts, delay_ms, preventDefault }` | About to attempt a reconnection (requires `auto_reconnect`). | | `reconnected` | `{ attempt }` | Successfully reconnected after a drop. | | `session_restart` | `{ reset_transcript }` | New STT session started (initial or after reconnect). | | `source_muted` | — | Audio source was muted externally (e.g. OS-level or hardware mute). | | `source_unmuted` | — | Audio source was unmuted after an external mute. | ## Session lifecycle A `Recording` transitions through a set of states. The lifecycle is fully managed — audio buffering during connection, keepalive during pause, and cleanup on stop or error are all handled automatically. ### States | State | Description | | ------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------- | | `idle` | Initial state before any work begins. | | `starting` | Audio source is starting, API key is being fetched. Audio is buffered. | | `connecting` | WebSocket connection is being established. | | `recording` | Actively capturing and streaming audio. | | `paused` | Audio capture and streaming paused. Keepalive messages maintain the connection.
**You are still charged for the open session even when it is paused.** | | `stopping` | `stop()` called. Waiting for the server to finish processing remaining audio. | | `stopped` | Gracefully stopped. All final results have been received. | | `error` | An error occurred. Resources have been cleaned up. | | `canceled` | Canceled via `cancel()` or `AbortSignal`. | ### Methods #### `stop(): Promise` Gracefully stops the recording. Stops the audio source and waits for the server to process all remaining audio and return final results. ```typescript await recording.stop(); // All final results have been received at this point ``` #### `cancel(): void` Immediately cancels the recording without waiting for final results. Closes the WebSocket connection and releases all resources. ```typescript recording.cancel(); ``` #### `pause(): void` Pauses audio capture and streaming. The WebSocket connection stays open with automatic keepalive messages. ```typescript recording.pause(); console.log(recording.state); // 'paused' ``` You are charged for the full stream duration even when session is paused. #### `resume(): void` Resumes audio capture and streaming after a pause. ```typescript recording.resume(); console.log(recording.state); // 'recording' ``` #### `finalize(options?): void` Requests the server to finalize current non-final tokens. Useful for forcing finalization at a specific point (e.g. before displaying a completed sentence). ```typescript recording.finalize(); // With trailing silence trimming: recording.finalize({ trailing_silence_ms: 500 }); ``` ### Tracking state changes ```typescript recording.on("state_change", ({ old_state, new_state }) => { console.log(`${old_state} → ${new_state}`); }); ``` ## Endpoint detection and manual finalization Endpoint detection lets you know when a speaker has finished speaking. This is critical for real-time voice AI assistants, command-and-response systems, and conversational apps where you want to respond immediately without waiting for long silences. Read more about [Endpoint detection](/stt/rt/endpoint-detection) Enable endpoint detection by setting `enable_endpoint_detection: true` in the session configuration. Listen for the `endpoint` event to know when a speaker has finished speaking. ```typescript recording.on("endpoint", () => { console.log("--- speaker finished ---"); }); ``` Manual finalization gives you precise control over when audio should be finalized — useful for Push-to-talk systems and client-side voice activity detection (VAD). Read more about [Manual finalization](/stt/rt/manual-finalization) ```ts recording.finalize(); ``` ## Pause, resume and muting audio source ```ts recording.pause(); // keeps connection alive, drops audio while paused recording.resume(); // resume sending audio ``` SDK will `finalize` audio on pause. Make sure to adjust your VAD sensitivity to have enough silence before pause. Learn more about [Manual finalization](/stt/rt/manual-finalization#key-points) Recording will also react on system level mute events and will start sending keepalive messages to keep the session alive. You are billed for the full stream duration even when session is paused. ## Auto-reconnect The SDK can transparently recover from transient network drops. Opt in by passing reconnection options to `client.realtime.record()` — on a retriable error, the SDK tears down the current WebSocket and audio encoder, re-resolves the connection config, and starts a new session with exponential backoff. Audio captured during the reconnect is buffered and flushed on resume. ```typescript const recording = client.realtime.record({ model: "stt-rt-v5", auto_reconnect: true, max_reconnect_attempts: 3, // default: 3 reconnect_base_delay_ms: 1000, // exponential backoff: 1s, 2s, 4s, ... }); recording.on("reconnecting", ({ attempt, max_attempts, delay_ms, preventDefault }) => { console.log(`Reconnect attempt ${attempt}/${max_attempts} in ${delay_ms}ms`); // preventDefault() cancels this attempt (useful for manual backoff control). }); recording.on("reconnected", ({ attempt }) => { console.log(`Reconnected after ${attempt} attempt(s)`); }); ``` ### Options | Option | Type | Default | Description | | ------------------------------- | --------- | ------- | ----------------------------------------------------------- | | `auto_reconnect` | `boolean` | `false` | Enable automatic reconnection on retriable errors. | | `max_reconnect_attempts` | `number` | `3` | Maximum consecutive attempts before surfacing the error. | | `reconnect_base_delay_ms` | `number` | `1000` | Base delay in ms for exponential backoff (1x, 2x, 4x, ...). | | `reset_transcript_on_reconnect` | `boolean` | `false` | Clear accumulated segment / utterance state on reconnect. | ### Force a reconnect Call `recording.reconnect()` to proactively rebuild the session — useful from `visibilitychange` or mobile background/foreground handlers when you suspect a stale connection. Requires `auto_reconnect: true`. ```typescript document.addEventListener("visibilitychange", () => { if (document.visibilityState === "visible") { recording.reconnect(); } }); ``` While reconnecting, the recording transitions through the `reconnecting` state — inspect `recording.state` for UI feedback. See the [Handle session events](#handle-session-events) table above for the full `reconnecting` / `reconnected` / `session_restart` payloads. ## Handling translation The SDK supports one-way and two-way real-time translation. Configure translation in the session config, then filter tokens by `translation_status` to separate original and translated text. ### One-way translation Translates all spoken audio into a single target language. ```typescript const recording = client.realtime.record({ model: "stt-rt-v5", translation: { type: "one_way", target_language: "es", // Translate everything to Spanish }, }); recording.on("result", (result) => { for (const token of result.tokens) { if (token.translation_status === "original") { console.log("[Original]", token.text); } else if (token.translation_status === "translation") { console.log("[Translated]", token.text); } } }); ``` ### Two-way translation Translates between two languages — each speaker's speech is translated into the other language. ```typescript const recording = client.realtime.record({ model: "stt-rt-v5", translation: { type: "two_way", language_a: "en", language_b: "fr", }, }); ``` ### Translation token fields When translation is enabled, each `RealtimeToken` includes: | Field | Type | Description | | -------------------- | --------------------------------------- | ------------------------------------------------------- | | `translation_status` | `'none' \| 'original' \| 'translation'` | Whether this token is original speech or a translation. | | `source_language` | `string` | The source language code for translated tokens. | | `language` | `string` | The language of this token's text. | Learn more about [Real-time translation](/translation/stt-translation/rt-translation) You can provide [custom translation terms](/stt/concepts/context#translation-terms) in the context to improve translation accuracy. ## Handle permissions The SDK provides a platform-agnostic permission system for checking and requesting microphone access before starting a recording. This is optional but recommended for a good user experience — you can show appropriate UI based on the permission state rather than waiting for the recording to fail. ### Setup Pass a [`BrowserPermissionResolver`](/sdk/web-SDK/reference/classes#browserpermissionresolver) when creating the client: ```typescript import { SonioxClient, BrowserPermissionResolver } from "@soniox/client"; const client = new SonioxClient({ config: fetchConfig, permissions: new BrowserPermissionResolver(), }); ``` ### Check permission status `check()` queries the current microphone permission without prompting the user: ```typescript const result = await client.permissions?.check("microphone"); switch (result?.status) { case "granted": // Microphone access already granted — safe to record break; case "prompt": // User hasn't been asked yet — show a "start recording" button break; case "denied": if (!result.can_request) { // Permanently denied — show "go to browser settings" instructions } break; case "unavailable": // No microphone or getUserMedia not supported break; } ``` ### Request permission `request()` triggers the browser permission prompt. On platforms where permission is already granted, this is a no-op. ```typescript const result = await client.permissions?.request("microphone"); if (result?.status === "granted") { startRecording(); } else if (result?.status === "denied") { showPermissionDeniedMessage(); } ``` Only create `BrowserPermissionResolver` in browser environments ## Use custom audio source By default, `client.realtime.record()` uses the built-in [`MicrophoneSource`](/sdk/web-SDK/reference/classes#microphonesource) which captures audio via `getUserMedia` and [`MediaRecorder`](https://developer.mozilla.org/en-US/docs/Web/API/MediaRecorder). You can replace it with any object that implements the [`AudioSource`](/sdk/web-SDK/reference/types#audiosource) interface. # Real-time speech generation with Web SDK URL: /sdk/web-SDK/tts/realtime-speech-generation Stream text to speech in the browser with the Soniox Web SDK over WebSocket The Soniox Web SDK supports real-time Text-to-Speech generation over WebSocket directly in the browser. You send text — all at once or incrementally — and receive decoded audio chunks as they arrive, so playback can start before generation is complete. This is the ideal transport for narrating LLM output and building voice agents in the browser. If you already have the full text up front and don't need chunk-by-chunk playback, use [REST speech generation](/sdk/web-SDK/tts/rest-speech-generation) — it's a single HTTP request. ## Set up your temporary API key endpoint Create a temporary key endpoint on your server using the Soniox [Node SDK](/sdk/node-SDK). Real-time TTS keys use the `tts_rt` usage type. To attribute browser-side TTS traffic to an end user or session, pass `client_reference_id` to `createTemporaryKey` - every request authenticated with the key is recorded under that identifier in [usage logs](/guides/usage-logs). Clients cannot override it. ```ts import express from 'express'; import { SonioxNodeClient } from '@soniox/node'; const app = express(); const client = new SonioxNodeClient(); // reads SONIOX_API_KEY from env app.get('/tts-rt-tmp-key', async (_req, res) => { try { const { api_key, expires_at } = await client.auth.createTemporaryKey({ usage_type: 'tts_rt', expires_in_seconds: 300, }); res.json({ api_key, expires_at }); } catch (err) { res.status(500).json({ error: err instanceof Error ? err.message : 'Failed to create temporary key' }); } }); app.listen(3000); ``` ## Quickstart Create a `SonioxClient` with a `config` resolver, then call `client.realtime.tts()` to open a single-stream session. Send text, consume audio by async iteration, and play it back. ```typescript import { SonioxClient } from "@soniox/client"; const client = new SonioxClient({ config: async () => { const res = await fetch("/tts-rt-tmp-key"); const { api_key } = await res.json(); return { api_key }; }, }); const stream = await client.realtime.tts({ voice: "Adrian", model: "tts-rt-v2", language: "en", audio_format: "wav", }); stream.sendText("Hello from Soniox real-time text-to-speech.", { end: true }); const chunks: Uint8Array[] = []; for await (const chunk of stream) { chunks.push(chunk); } const blob = new Blob(chunks, { type: "audio/wav" }); await new Audio(URL.createObjectURL(blob)).play(); ``` The stream closes itself (and the underlying WebSocket) once `terminated` fires. You never have to call `close()` in single-stream mode. ## Play audio as it arrives For the lowest-latency playback, feed chunks into a [`MediaSource`](https://developer.mozilla.org/en-US/docs/Web/API/MediaSource) instead of waiting for the full payload. ```typescript const mediaSource = new MediaSource(); const audioEl = new Audio(URL.createObjectURL(mediaSource)); await audioEl.play(); mediaSource.addEventListener("sourceopen", async () => { const sourceBuffer = mediaSource.addSourceBuffer("audio/wav"); const stream = await client.realtime.tts({ voice: "Adrian", audio_format: "wav", }); stream.sendText("Streaming audio for low-latency playback.", { end: true }); for await (const chunk of stream) { await new Promise((resolve) => { sourceBuffer.addEventListener("updateend", () => resolve(), { once: true }); sourceBuffer.appendBuffer(chunk); }); } mediaSource.endOfStream(); }); ``` ## Send text incrementally Call `sendText(text)` for each chunk as it becomes available, then mark the last chunk with `{ end: true }` or invoke `finish()` explicitly. This is the pattern for narrating an LLM response token-by-token. ```typescript const stream = await client.realtime.tts({ voice: "Adrian", audio_format: "wav" }); stream.sendText("Hello from Soniox "); stream.sendText("real-time TTS. "); stream.sendText("This is the final chunk.", { end: true }); for await (const chunk of stream) { playback(chunk); // your playback function (see "Play audio as it arrives" above) } ``` ## Pipe from an async iterable `stream.sendStream(source)` pipes any `AsyncIterable` into the TTS session and auto-finishes when the iterable completes. Sending and receiving run concurrently. ```typescript async function* llmTokens(prompt: string): AsyncIterable { const res = await fetch("/llm/stream", { method: "POST", body: JSON.stringify({ prompt }), }); const reader = res.body!.getReader(); const decoder = new TextDecoder(); while (true) { const { value, done } = await reader.read(); if (done) return; yield decoder.decode(value); } } const stream = await client.realtime.tts({ voice: "Adrian", audio_format: "wav" }); stream.sendStream(llmTokens("Tell me a story.")); for await (const chunk of stream) { playback(chunk); } ``` ## Event-based consumption `RealtimeTtsStream` is also a typed event emitter. When you prefer an event-driven style over async iteration, listen for [`TtsStreamEvents`](/sdk/web-SDK/reference/types#ttsstreamevents): | Event | Payload | Description | | ------------ | ------------ | ------------------------------------------------------ | | `audio` | `Uint8Array` | Decoded audio chunk. | | `audioEnd` | — | Server marked the final audio payload for this stream. | | `terminated` | — | Stream fully closed by the server. | | `error` | `Error` | Stream-level error. | ```typescript const stream = await client.realtime.tts({ voice: "Adrian", audio_format: "wav" }); stream.on("audio", (chunk) => playback(chunk)); stream.on("audioEnd", () => console.log("last audio payload received")); stream.on("error", (err) => console.error("Stream error:", err)); stream.on("terminated", () => console.log("stream done")); stream.sendText("Hello from event-based TTS.", { end: true }); ``` Choose either async iteration **or** event listeners — not both. The async iterator consumes `audio` events internally. ## Multi-stream connection A single WebSocket connection can carry up to 5 concurrent TTS streams. Use `client.realtime.tts.multiStream()` to open a [`RealtimeTtsConnection`](/sdk/web-SDK/reference/classes#realtimettsconnection), then call `connection.stream()` for each stream — each with its own voice, model, and audio format. ```typescript const connection = await client.realtime.tts.multiStream(); const s1 = await connection.stream({ voice: "Adrian", audio_format: "wav" }); // Enumerate available voices via the Node SDK's `client.tts.listModels()`. const s2 = await connection.stream({ voice: "", audio_format: "wav" }); s1.sendText("Hello from stream 1.", { end: true }); s2.sendText("Hello from stream 2.", { end: true }); // Consume both streams concurrently. `playback(chunk, streamId)` is your // app-specific playback function (e.g. feeding two `MediaSource` instances). await Promise.all([ (async () => { for await (const c of s1) playback(c, "s1"); })(), (async () => { for await (const c of s2) playback(c, "s2"); })(), ]); connection.close(); ``` Call `connection.close()` when you're done — this ends all active streams and closes the WebSocket. ## Cancel, finish, and close | Method | Behavior | | -------------------- | --------------------------------------------------------------------------------------- | | `stream.finish()` | Signals "no more text". The server finishes generating audio and sends `terminated`. | | `stream.cancel()` | Aborts generation immediately. The server stops producing audio and sends `terminated`. | | `stream.close()` | Terminates the stream. In single-stream mode this also closes the WebSocket. | | `connection.close()` | Closes the WebSocket and terminates all streams on a multi-stream connection. | ```typescript stream.finish(); // graceful stop stream.cancel(); // user-triggered cancel ``` ## Error handling A failed stream does not close the whole WebSocket connection by default. Stream-level errors finalize only that stream (`terminated` fires for the same stream id), while other streams on the same connection can continue. Connection-level failures end the whole connection and all active streams. ```typescript import { RealtimeError, SonioxError } from "@soniox/client"; try { const stream = await client.realtime.tts({ voice: "Adrian" }); stream.sendText("Hello!", { end: true }); for await (const _ of stream) { // consume audio } } catch (err) { if (err instanceof RealtimeError) { console.error(`Realtime TTS error (${err.code}):`, err.message); } else if (err instanceof SonioxError) { console.error("Soniox SDK error:", err.message); } else { throw err; } } ``` ## Server-driven defaults There's no first-class endpoint for TTS defaults — you own them. Keep them on your server next to the temporary-key endpoint and return them via `SonioxConnectionConfig.tts_defaults`. The SDK merges them as the base layer when opening TTS streams, and caller-provided fields on `client.realtime.tts(...)` / `connection.stream(...)` override the defaults. ```ts app.get('/tts-rt-tmp-key', async (_req, res) => { const { api_key, expires_at } = await nodeClient.auth.createTemporaryKey({ usage_type: 'tts_rt', expires_in_seconds: 300, }); res.json({ api_key, expires_at, tts_defaults: { model: 'tts-rt-v2', language: 'en', voice: 'Adrian', audio_format: 'wav', }, }); }); ``` The browser client consumes the defaults automatically: ```typescript const client = new SonioxClient({ config: async () => { const res = await fetch("/tts-rt-tmp-key"); return await res.json(); // { api_key, tts_defaults, ... } }, }); const stream = await client.realtime.tts({}); // uses server-provided defaults const override = await client.realtime.tts({ voice: "" }); // overrides voice ``` ## See also * [REST speech generation](/sdk/web-SDK/tts/rest-speech-generation) — single-request HTTP TTS. * [`RealtimeTtsStream` reference](/sdk/web-SDK/reference/classes#realtimettsstream) * [`RealtimeTtsConnection` reference](/sdk/web-SDK/reference/classes#realtimettsconnection) * [`TtsStreamInput`](/sdk/web-SDK/reference/types#ttsstreaminput), [`TtsStreamEvents`](/sdk/web-SDK/reference/types#ttsstreamevents) * [TTS WebSocket API](/api-reference/tts/websocket-api) # REST speech generation with Web SDK URL: /sdk/web-SDK/tts/rest-speech-generation Generate speech from text in the browser with the Soniox Web SDK over HTTP The Soniox Web SDK supports Text-to-Speech generation over HTTP with `SonioxClient`. Use REST when you have the full text up front — the SDK returns audio bytes that you can play, download, or hand to an `