Realtime transcription

Stream live audio over a WebSocket and get text back while the speaker talks.
View as Markdown

Your backend creates a realtime transcription with its API key. The response carries a wss:// URL with a single-use ticket. A browser, or any other client, opens that socket and sends raw audio. The session sends back partial and final text while the speaker talks.

A partial updates about every 2 seconds of speech. A final arrives about 2 to 4 seconds after the speaker pauses. English and French work in the same session, and the model follows a switch between them.

The API key never leaves your backend. The client holds only the ticket.

The server keeps no transcript after close. Persist every final message yourself, for example POST it to your backend. To get a stored, diarized transcript, set finalPass: true at create.

Routes

Method and pathWhat it does
POST /v1/realtime/transcriptionsCreate a session. Returns 201 and websocket.url.
GET /v1/realtime/transcriptions/{id}Read the session status and counters.
GET /v1/realtime/transcriptionsList sessions. ?limit (1–50, default 20) and &cursor.
POST /v1/realtime/transcriptions/{id}/ticketsMint a new client ticket. Only in created or active.
POST /v1/realtime/transcriptions/{id}/stopStop the session. Returns 202. Idempotent.
GET /v1/realtime/transcriptions/{id}/ws?ticket=…The WebSocket. It is not an API operation, so the API reference does not list it. This page documents it.

The resource is on api.cloudraker.com only. api.paperwork.sh does not serve it.

Send a User-Agent header from server code. Some HTTP clients, such as Python’s urllib, send a default agent that our edge blocks with 403. fetch, httpx, requests, and websockets are fine.

Concepts

  • Session. One live audio stream. The object is realtime_transcription and the id starts with rtt_.
  • Ticket. A single-use capability in the socket URL. A create or re-mint ticket expires after 60 seconds.
  • Resume ticket. Each ready message carries a new single-use resumeTicket. Use it to reconnect without a backend call.
  • Timeline. Timeline seconds count the audio samples that the session received. The start and end fields use timeline seconds, not wall time.
  • Partial and final. A partial shows the current guess for the open utterance. A final commits a span of text and never changes.
  • Final pass. Opt-in. At close, the session saves the audio as a file and runs the batch diarized transcription on it.

States

statusMeaning
createdThe session exists. No client connected yet.
activeA client connected at least once.
closingIntake stopped. The decoder sends the remaining finals. This takes at most about 15 seconds when a decoder is attached.
closedThe session ended. closeReason tells why.
expiredNo client connected within 10 minutes of create.
failedAn internal fault ended the session.

A client disconnect alone does not change the state. The session stays active until 5 minutes pass without audio. That is your reconnect window.

closeReason is one of client_end, stopped, idle_timeout, max_duration, credits_exhausted, worker_unavailable, too_fast, or internal_error.

Create a session

Call create from your backend. Always send an Idempotency-Key. A retry with the same key and the same body returns the same session with a fresh ticket and the header idempotent-replay: true. The same key with a different body returns 409 idempotency_conflict.

curl -s https://api.cloudraker.com/v1/realtime/transcriptions \
-H "Authorization: Bearer $CLOUDRAKER_API_KEY" \
-H "Idempotency-Key: $(uuidgen)" \
-H "Content-Type: application/json" \
-d '{}'

Request body

FieldTypeDefaultMeaning
finalPassbooleanfalseSave the audio as a file and run the batch diarized transcription at close. See Final pass.
maxDurationSecondsinteger14400Hard cap on received audio, 60 to 14400 (4 hours). With finalPass, the default and the maximum is 7200 (2 hours).
metadataobjectnoneAt most 16 string keys (64 characters) and values (512 characters). Kept 30 days after close. Do not put personal data here.

Unknown keys return 400 invalid_request.

Example response

{
"object": "realtime_transcription",
"id": "rtt_01KZ4M2Q8V3N7X5R9T1B6C0D2E",
"status": "created",
"closeReason": null,
"audio": { "encoding": "pcm_s16le", "sampleRate": 16000, "channels": 1 },
"maxDurationSeconds": 14400,
"audioSeconds": 0,
"billedMinutes": 0,
"finalPass": { "enabled": false, "status": "disabled", "fileId": null },
"metadata": { "call": "support-4471" },
"createdAt": "2026-09-27T14:02:11Z",
"connectedAt": null,
"closedAt": null,
"connectBy": "2026-09-27T14:12:11Z",
"websocket": {
"object": "realtime_transcription_ticket",
"url": "wss://api.cloudraker.com/v1/realtime/transcriptions/rtt_01KZ4M2Q8V3N7X5R9T1B6C0D2E/ws?ticket=…",
"expiresAt": "2026-09-27T14:03:11Z"
}
}

Only the create response carries websocket. To get another ticket, call POST /v1/realtime/transcriptions/{id}/tickets with an empty body. It returns a realtime_transcription_ticket object.

The API reference lists every field and response code.

WebSocket protocol

Open the websocket.url from create, from re-mint, or with a resumeTicket. Set binaryType to arraybuffer.

wss://api.cloudraker.com/v1/realtime/transcriptions/{id}/ws?ticket={ticket}[&afterSeq={n}]

Tickets

  • A ticket is 43 characters of base64url. Each ticket works once.
  • A bad, used, or expired ticket closes the socket with 4401.
  • Each new ticket has a higher generation than the previous one. A socket with a newer ticket evicts the current client with 4409. A socket with an older ticket gets 4409 itself.
  • So a stolen ticket shows the finals so far. Your next re-mint evicts the thief for good.

Audio: client to server, binary

  • Send raw PCM16 little-endian, 16 000 Hz, mono. No header and no container.
  • Each binary message holds an even number of bytes, from 640 to 32 000 (20 ms to 1 s). Send 100 ms frames: 3 200 bytes.
  • A shorter frame is allowed only as the last frame before end.
  • Send at most 50 messages per second, binary and text together.
  • Send audio at real time. The session allows 1.25 times real time with a 30-second burst. Faster audio closes the socket with 4429.
  • Send audio only after you receive ready.
  • Frames after end are ignored.

Control: client to server, JSON text

MessageEffect
{"type":"end"}Stop intake. The session sends the remaining finals, then closed, then closes with 1000.
{"type":"ping"}The server answers {"type":"pong"}.

Text messages are at most 1 KiB. Any other text, an odd byte length, or a binary size out of range sends error with code protocol_error and closes with 4400. The session stays active: fix the client, then reconnect.

Results: server to client, JSON text

MessageFieldsMeaning
readyid, status, audioSeconds, lastSeq, worker, resumeTicketFirst message after connect, after any afterSeq replay. audioSeconds is the audio the session holds. worker is queued or attached.
ackaudioSecondsEvery 5 seconds while audio flows. Trim your local audio buffer to it.
workerstateThe decoder slot changed: queued or attached. A failover shows queued, then attached.
partialtext, start, endThe current guess. It replaces the previous partial. Discard it when a final arrives.
finalseq, text, start, endCommitted text for [start, end]. seq starts at 1 and only goes up.
gapstart, endAudio in [start, end] was never decoded. You do not pay for it.
errorcode, messageComes before a close. code is the close-code name.
closedreason, audioSeconds, billedMinutes, finalPassThe last message. finalPass is {status, fileId}. The socket closes next.
pongnoneThe answer to ping.

Live output has no language and no speaker fields. The model detects the language. Speaker labels come only from the final pass.

{"type":"ready","id":"rtt_01KZ4M…","status":"active","audioSeconds":0,"lastSeq":0,"worker":"attached","resumeTicket":"Xb4…"}
{"type":"partial","text":"thanks for calling","start":0.42,"end":1.9}
{"type":"final","seq":1,"text":"Thanks for calling, how can I help?","start":0.42,"end":2.61}
{"type":"closed","reason":"client_end","audioSeconds":61.2,"billedMinutes":2,"finalPass":{"status":"disabled","fileId":null}}

Close codes

CodeName (error.code)Session afterWhat to do
1000client_end, stoppedclosedDone.
4400protocol_errorunchangedFix the client. Do not reconnect in a loop.
4401unauthorizedunchangedGet a new ticket from your backend and reconnect once.
4402credits_exhaustedclosedDone. Top up credits.
4404not_foundnoneFatal. Check the id and the host.
4408idle_timeoutclosedDone. 5 minutes passed without audio.
4409replacedunchangedA newer ticket connected. If you did not start another client, get a new ticket and reconnect once.
4410session_closedterminalDone.
4413max_durationclosedDone.
4429too_fastclosedDone. Send audio at real time.
4503worker_unavailable, internal_errorclosedCreate a new session later.
1006, 1011, 1012, or a close without closedtransport drop or server restartunchangedReconnect. See Reconnect.

These are WebSocket close codes, not HTTP statuses. 4402 is not HTTP 402.

Keepalive and mute

  • To mute, keep the stream open and send silence. In a browser, track.enabled = false sends zeros through the audio pipeline.
  • Silence costs the same as speech. Stop the stream only to end the session.
  • If you pause audio, send ping every 20 seconds. The session still closes after 5 minutes without audio.

Reconnect

Deploys drop every open socket. This is routine: the session state survives, so plan for reconnects.

  1. Reconnect only when this socket received ready and did not receive closed, and the close code is in the reconnect row above.
  2. A failure before the first ready of the first connect is fatal. Show an error and do not retry.
  3. Retry with backoff from 0.5 s to 5 s, with full jitter, for at most 60 seconds.
  4. Connect with the last resumeTicket. On 4401, ask your backend for a new ticket once.
  5. Add &afterSeq=<last final seq>. The session sends the stored finals with a higher seq, then ready.
  6. Resend the audio that the session did not receive (see below), then continue live.

Resend rule. Keep a local ring of captured audio since the last ack, at most 30 seconds. Record R_last, the audioSeconds of the last ready or ack, and L_last, your local capture position at that moment. On the new ready with audioSeconds = R, resend from local position L_last + (R − R_last).

Audio you send twice is transcribed twice and billed twice. Audio you do not resend is absent from the timeline.

Stream audio

Node.js

Node 22 and later have a global WebSocket. ffmpeg -re converts any recording to PCM16 16 kHz mono at real time.

stream.mjs
import { spawn } from "node:child_process";
const res = await fetch("https://api.cloudraker.com/v1/realtime/transcriptions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.CLOUDRAKER_API_KEY}`,
"Idempotency-Key": crypto.randomUUID(),
"Content-Type": "application/json",
},
body: "{}",
});
const session = await res.json();
const ws = new WebSocket(session.websocket.url);
ws.binaryType = "arraybuffer";
ws.onmessage = (event) => {
const msg = JSON.parse(event.data);
if (msg.type === "final") console.log(`[${msg.start.toFixed(1)}] ${msg.text}`);
if (msg.type === "closed") console.log("closed:", msg.reason, msg.billedMinutes, "min");
if (msg.type !== "ready" || ws.started) return;
ws.started = true;
const ffmpeg = spawn("ffmpeg", ["-loglevel", "error", "-re", "-i", process.argv[2],
"-f", "s16le", "-ac", "1", "-ar", "16000", "-"]);
let pending = Buffer.alloc(0);
ffmpeg.stdout.on("data", (chunk) => {
pending = Buffer.concat([pending, chunk]);
while (pending.length >= 3200) {
ws.send(pending.subarray(0, 3200)); // 100 ms frames
pending = pending.subarray(3200);
}
});
ffmpeg.on("close", () => {
const tail = pending.subarray(0, pending.length & ~1); // even byte count
if (tail.length) ws.send(tail);
ws.send(JSON.stringify({ type: "end" }));
});
};
ws.onclose = (event) => console.log("socket closed", event.code);

Run it with node stream.mjs call.m4a.

Python

Use httpx for the create call so that the event loop never blocks. Use websockets for the socket.

stream.py
import asyncio, json, os, sys, uuid
import httpx, websockets
async def main(path: str) -> None:
async with httpx.AsyncClient() as http:
res = await http.post(
"https://api.cloudraker.com/v1/realtime/transcriptions",
headers={
"Authorization": f"Bearer {os.environ['CLOUDRAKER_API_KEY']}",
"Idempotency-Key": str(uuid.uuid4()),
},
json={},
)
session = res.raise_for_status().json()
async with websockets.connect(session["websocket"]["url"]) as ws:
ready = json.loads(await ws.recv())
assert ready["type"] == "ready", ready
async def send_audio() -> None:
ffmpeg = await asyncio.create_subprocess_exec(
"ffmpeg", "-loglevel", "error", "-re", "-i", path,
"-f", "s16le", "-ac", "1", "-ar", "16000", "-",
stdout=asyncio.subprocess.PIPE,
)
while frame := await ffmpeg.stdout.read(3200):
await ws.send(frame[: len(frame) & ~1]) # even byte count
await ws.send(json.dumps({"type": "end"}))
sender = asyncio.create_task(send_audio())
async for raw in ws:
msg = json.loads(raw)
if msg["type"] == "final":
print(f"[{msg['start']:.1f}] {msg['text']}")
elif msg["type"] == "closed":
print("closed:", msg["reason"], msg["billedMinutes"], "min")
await sender
asyncio.run(main(sys.argv[1]))

ffmpeg.stdout.read(3200) can return fewer bytes than asked. Frames under 640 bytes are allowed only as the last frame. For production, buffer to exact 3 200-byte frames, as the Node sample does.

Browser

Your backend creates the session and returns only websocket.url to the page. The page never sees the API key.

Follow these rules:

  1. Create the AudioContext and call ctx.resume() in the click handler, before any await. Do not pass sampleRate.
  2. Call getUserMedia with channelCount: 1 and the three browser processing flags. Do not pass sampleRate. The worklet resamples.
  3. Load the worklet from a Blob URL.
  4. Send audio only after ready. Keep a 30-second ring and reconnect with resumeTicket, as in Reconnect.
  5. To mute, set track.enabled = false. The worklet then sends zeros.
  6. To stop, send end, stop the tracks, and call ctx.close(). Also do this on pagehide.
  7. POST each final to your backend. The server keeps no transcript.
mic.js
const WORKLET = `
class Pcm16 extends AudioWorkletProcessor {
constructor() { super(); this.step = sampleRate / 16000; this.pos = 0; this.acc = 0; this.n = 0
this.out = new Int16Array(1600); this.i = 0 }
process(inputs) {
const ch = inputs[0][0]; if (!ch) return true // mono: first channel only
for (let k = 0; k < ch.length; k++) {
this.acc += ch[k]; this.n++; this.pos += 1 // box low-pass: average the input run
if (this.pos >= this.step) { // fractional phase kept across calls
this.pos -= this.step
const v = Math.max(-1, Math.min(1, this.acc / this.n)); this.acc = 0; this.n = 0
this.out[this.i++] = v < 0 ? v * 0x8000 : v * 0x7fff
if (this.i === 1600) { // 100 ms, exactly-sized buffer
const buf = this.out.buffer; this.port.postMessage(buf, [buf])
this.out = new Int16Array(1600); this.i = 0
}
}
}
return true
}
}
registerProcessor('pcm16', Pcm16)`
startButton.onclick = async () => {
const ctx = new AudioContext(); ctx.resume() // before any await (user activation)
const workletUrl = URL.createObjectURL(new Blob([WORKLET], { type: "text/javascript" }))
const [{ url }, stream] = await Promise.all([
fetch("/my-backend/realtime-ticket", { method: "POST" }).then((r) => r.json()),
navigator.mediaDevices.getUserMedia({ audio: { channelCount: 1, echoCancellation: true,
noiseSuppression: true, autoGainControl: true } }),
])
await ctx.audioWorklet.addModule(workletUrl)
const node = new AudioWorkletNode(ctx, "pcm16")
ctx.createMediaStreamSource(stream).connect(node)
const ws = new WebSocket(url)
ws.binaryType = "arraybuffer"
let ready = false
node.port.onmessage = (e) => { if (ready && ws.readyState === WebSocket.OPEN) ws.send(e.data) }
ws.onmessage = (e) => {
const msg = JSON.parse(e.data)
if (msg.type === "ready") ready = true
if (msg.type === "final") {
fetch("/my-backend/finals", { method: "POST", body: JSON.stringify(msg) })
}
}
const stop = () => {
if (ws.readyState === WebSocket.OPEN) ws.send(JSON.stringify({ type: "end" }))
stream.getTracks().forEach((t) => t.stop())
ctx.close()
}
stopButton.onclick = stop
addEventListener("pagehide", stop)
}

This sample omits the reconnect ring for brevity. Add it before you ship: deploys drop sockets.

Pricing

ItemCreditsWhen
Session fee5 per sessionWhen the first decoder joins the session. Or when the session expires and no client ever connected to it.
Streamed audio2 per started minutePer started minute of decoded audio. The remainder at close rounds up.
  • You pay nothing when the session fails on our side. This covers 503 capacity_unavailable at create and a session where no decoder ever joins (4503).
  • Audio in a gap message is not billed.
  • No minutes accrue before the first decoder joins.
  • Silence costs the same as speech.
  • Examples after a decoder joins: 0 s costs 5 credits, 59 s costs 7, 61 s costs 9.
  • A create without an Idempotency-Key, retried, makes two sessions. The unused one expires and pays its 5-credit fee.

GET returns billedMinutes, the minutes charged so far. The closed message carries the final count.

Final pass

Set finalPass: true at create to get a stored, diarized transcript.

  1. The session saves the received audio as a WAV file in your organization’s API workspace.
  2. At close, the batch diarized transcription runs on that file.
  3. finalPass.fileId names a normal file. Read it with GET /v1/files/{fileId}.
  4. When finalPass.status is ready, read processed.json from the file’s urls.json.

The file carries the standard audio transcript: ordered segments with start, end, text, speaker, and words.

The timestamps use the same timeline as the live finals.

finalPass.status is one of these values:

StatusMeaning
disabledfinalPass is off.
recordingThe session is live and records the audio.
uploadingThe session finishes the file upload.
processingThe batch transcription runs.
readyThe transcript is ready.
failedThe upload or the transcription failed.
skippedThe session closed with no decoded audio, or it closed for credits.

Billing. The existing audio meters bill the final pass, on top of the realtime price. The diarized transcription costs 20 credits per minute, with a minimum of 1 minute. A 30-minute session with finalPass costs 5 + 60 + 600 = 665 credits.

Visibility. The file is an organization file. Every API key and admin in the organization can read it through /v1/files. It stays until you delete it with DELETE /v1/files/{id}.

Limit. With finalPass, a session lasts at most 2 hours.

Limits

LimitValue
Session length4 hours. 2 hours with finalPass. Set a lower cap with maxDurationSeconds.
Idle close5 minutes without audio (4408).
Connect deadline10 minutes after create. Then the session is expired. connectBy gives the time.
Ticket lifetime60 seconds, single use.
Concurrent sessions10 per organization in created, active, or closing. More returns 429 concurrency_limit.
Create rate20 per minute per organization.
Other callsThe general rate limit.
Audio speed1.25 times real time, 30-second burst (4429).
Messages50 per second. Binary frames 640 to 32 000 bytes. Text at most 1 KiB.

Data handling

Location. Montreal decode, North America session state, the final pass may run in the US; no residency guarantee.

Retention without finalPass. Audio exists only in decoder memory and in a short rolling buffer of the session. The finals stay only for reconnect replay. Both are deleted at close. Logs carry counters, codes, and ids, never audio or text.

Retention of session metadata. Ids, status, durations, billed minutes, and metadata values stay 30 days after close. Then the session is deleted.

With finalPass. The audio and the transcript are a normal file in your organization. You control it.

Security notes

  • The API key stays on your backend. Give the client only websocket.url.
  • Tickets appear in URLs, so each ticket works once and expires after 60 seconds.
  • An API key sees only the sessions it created. A session from another key returns 404. An admin session sees every session in the organization.
  • A revoked API key does not end its live sessions. A live session runs until it closes, at most 4 hours. Revocation blocks new creates, tickets, and reads with that key.
  • The MCP server code mode can create a session and mint tickets. It cannot stream audio.

Errors

The HTTP routes return the standard error envelope.

StatuscodeWhenRetry
400invalid_requestInvalid body, unknown key, finalPass with maxDurationSeconds over 7200, or any key in a ticket body.No
402credits_exhaustedThe organization is out of credits. Create only.No
402feature_not_in_planThe plan does not include realtime transcription, or diarized audio with finalPass.No
404not_foundUnknown id, or a session that another API key created.No
409session_closedA ticket for a session in closing, closed, expired, or failed.No
409idempotency_conflictThe same Idempotency-Key with a different body.No
429rate_limitedThe rate limit is exhausted.Yes, after Retry-After
429concurrency_limitThe organization has 10 open sessions.Yes, when one closes
503capacity_unavailableNo decoder capacity is free. Nothing was charged.Yes, after a few seconds
500internal_errorA failure on our side.Yes

stop on a terminal session returns 202, not an error. For socket failures, see Close codes.