Live transcription supports
"solaria-1" only.- Using our SDKs
- Using the API
The SDK provides methods for starting a live session, sending audio and handling transcript, error and session-ended events. It also manages WebSocket lifecycle functions such as reconnection and buffering. Check the SDK behavior for your integration and handle unrecoverable errors in your application.
Install the SDK
Initiate your real-time session
First, call thePOST /v2/live endpoint and pass your configuration.
It’s important to correctly define the properties encoding, sample_rate, bit_depth and channels as we need them to parse your audio chunks. These values describe the source-audio format you are sending. Each configured value must match the audio being sent. The sample rate in the examples below is one supported configuration, not the only supported value.Why initiate with POST instead of connecting directly to the WebSocket?
Why initiate with POST instead of connecting directly to the WebSocket?
- Security: Generate the WebSocket URL on your backend and keep your API key private. The init call returns a connectable URL and a session
idthat you can safely pass to web, iOS, or Android clients without exposing credentials in the app. - Lower infrastructure load: The secure URL is generated on your backend, the client can connect directly to Gladia’s WebSocket server without a pass-through on your side, saving your own resources.
- Resilient reconnection and session continuity: If the WebSocket disconnects (which can happen on unreliable networks), the session created by the init call lets the client reconnect without losing context. Traditional flows that open a socket first typically force a brand‑new session on disconnect, dropping in‑progress state.
Connect to the WebSocket
Now that you’ve initiated the session and have the URL, you can connect to the WebSocket using your preferred language/framework. Here’s an example in JavaScript:Send audio chunks
You can now start sending us your audio chunks through the WebSocket:A single realtime transcription session cannot exceed 3 hours. For longer events, start a new session before reaching the limit. See Concurrency and rate limits and Supported files & duration for details.
Read messages
During the whole session, we will send various messages through the WebSocket, the callback URL or webhooks. You can specify which kind of messages you want to receive in the initial configuration. Seemessages_config for WebSocket messages and callback_config for callback messages.Partial transcripts are provisional; final transcripts complete an utterance. Events for the same utterance share data.id. A newer partial or final replaces the displayed text for that utterance instead of creating a duplicate. Use is_final to distinguish partial and final events.Here’s an example of how to read a transcript message received through a WebSocket:Need low-latency partial results?Enable partial transcripts by setting
messages_config.receive_partial_transcripts: true.Use the is_final property to distinguish between partial and final transcript messages.Stop the recording
Finalizing an utterance (is_final: true) is not the same as ending the live session. Utterance finals arrive while the session is still open; stopping the session ends the WebSocket flow and triggers final session results. A single live session cannot exceed 3 hours. For concurrency, HTTP 429 behavior, and capacity limits, see concurrency and rate limits.Async accuracy benchmarks and the blind comparison tool evaluate async accuracy within their documented scope; neither establishes streaming latency. See pricing for current charges.Once you’re done, send us the stop_recording message. We will process remaining audio chunks and start the post-processing phase, in which we put together the final audio file and results with the add-ons you requested.You’ll receive a message at every step of the process in the WebSocket, or in the callback if configured. Once the post-processing is done, the WebSocket is closed with a code 1000.Get the final results
If you want to get the complete result, you can call theGET /v2/live/:id endpoint with the id you received from the initial request.