Realtime transcription

Stream live audio over a WebSocket and get text back while the speaker talks: partials about every 2 seconds, finals a few seconds after each pause, English and French in the same session. Your backend creates a session with POST /v1/realtime/transcriptions. The response carries a wss:// URL with a single-use ticket, so the browser never holds your API key.

The session sends partial and final text while the speaker talks. The server keeps no transcript after close. To get a stored, diarized transcript, set finalPass: true.

Pricing: a 5-credit session fee when the first decoder joins, or when a session that no client connected to expires. Then 2 credits per started minute of decoded audio. A session that fails on our side costs nothing.

Read the docs