Realtime Speech-to-Text API
Real-time speech-to-text streaming over a single WebSocket connection.
This endpoint is a dedicated transcription stream: you push raw audio frames in and receive incremental transcriptions and, optionally, translations back as JSON.
Palabra API client
Consider using the Palabra API Python client and checking the code example with it.
Step 1. Get an API Key
Create an API Key on the Palabra API Keys page. See Authentication for details.
Step 2. Connect
Open a WebSocket to the endpoint below, passing your API Key as the token
query parameter (or in the Authorization header). The server validates the
key and creates a streaming session for the lifetime of the connection
automatically.
All other stream settings are passed as query parameters in the same URL.
wss://stream.palabra.ai/asr/v1/speech-to-text/stream?token=<API_KEY>&language=en&format=pcm_s16le&sample_rate=16000
import websockets
url = (
"wss://api.palabra.ai/asr/v1/speech-to-text/stream"
f"?token={api_key}&language=en&format=pcm_s16le&sample_rate=16000"
)
ws = await websockets.connect(url)
Query parameters
| Parameter | Required | Description |
|---|---|---|
token | yes | Your API Key |
format | yes | Audio format (see Audio formats) |
sample_rate | conditional | Sample rate in Hz. Required for all raw PCM formats; for pcm_s16le required only when the rate is not 16000 |
language | no | Source language code. Defaults to auto |
translate_languages | no | Comma-separated target languages, e.g. es,de,fr |
enable_filler_filter | no | Whether to enable the filler filter. true by default for all languages except ja |
finalization_mode | no | auto (default) or manual — see Finalization modes |
finalization_timeout | no | Manual mode safety-net timeout in seconds (default 60) |
turn_detection and turn_timeout are deprecated aliases of finalization_mode and finalization_timeout.
Supported languages
| Code | Language |
|---|---|
ar | Arabic |
de | German |
en | English |
es | Spanish |
fr | French |
hi | Hindi |
it | Italian |
ja | Japanese |
ko | Korean |
nl | Dutch |
pt | Portuguese |
ru | Russian |
zh | Chinese |
Automatic source language detection (language=auto) is supported in experimental mode.
Step 3. Send audio
Send audio as raw binary WebSocket frames. Chunks of 320 ms are recommended.
await ws.send(data)
The finalize command
Besides binary audio, the client may send a JSON text frame with the finalize command — a manual end-of-segment:
{ "message_type": "finalize" }
It works in both finalization modes: the recognizer finalizes everything received so far and emits the pending segment as final (is_eos: true); recognition then continues as usual. Send it whenever your application knows the speaker's turn is over (push-to-talk release, your own turn detection), or after the last byte of a pre-recorded file (which often ends abruptly, with no trailing silence — so automatic, silence-driven segmentation would never finalize the tail).
The <fin> marker
Every finalize is answered. The final transcript that answers it carries the <fin> marker appended (without a space) to its text — that is how the client knows up to which point the transcription has been finalized:
{ "message_type": "transcription", "is_eos": true, "segment": { "text": "Hello world.<fin>", ... }, ... }
If the finalize produced no text (nothing new was recognized yet), the server still answers with an ordinary transcription message (is_eos: false) whose text is just the marker:
{ "message_type": "transcription", "is_eos": false, "segment": { "text": "<fin>", ... }, ... }
The marker never appears in translated_transcription messages.
Finalization modes
Segmentation ("when is the phrase finished?") is controlled by the finalization_mode query parameter:
auto(default) — segments are finalized automatically. This is the current behavior and requires nothing from the client.manual— automatic finalization is suppressed: the client decides where turns end and sends thefinalizecommand. As a safety net, if nofinalizearrives withinfinalization_timeoutseconds (default 60) since the last segment end, the nearest automatic end-of-segment signal is let through once; everyfinalize(and every delivered final segment) restarts the timer, so a client that keeps finalizing never hits it.
Use manual when your application has better knowledge of turn boundaries than the audio itself — push-to-talk UIs, external diarization/turn-taking logic, or benchmarks that need reproducible segmentation.
Audio formats
format | sample_rate | Notes |
|---|---|---|
pcm_s16le | only if ≠ 16000 | 16-bit signed little-endian PCM. Recommended |
pcm_f32le / pcm_f32be | required | 32-bit float PCM |
pcm_s32le / pcm_s32be | required | 32-bit signed PCM |
mulaw / alaw | required | G.711 |
webm / mp3 / aac / ogg / flac / wav | not used | Container formats; rate is read from the stream |
Step 4. Receive messages
All server-to-client messages are JSON text frames. Switch on message_type.
transcription
Emitted continuously as speech is recognized.
{
"message_type": "transcription",
"transcription_id": "a1b2c3d4",
"language": "en",
"is_eos": false,
"segment": {
"text": "Hello world how are",
"start_time": 0.32,
"end_time": 1.84
},
"delta": {
"text": "how are",
"start_time": 1.20,
"end_time": 1.84
}
}
| Field | Description |
|---|---|
transcription_id | Stable id for the segment. All messages of one segment share the same id. A new id means a new segment has started |
language | Detected (or configured) source language of this segment |
is_eos | false — partial; the segment is still being updated. true — the segment is committed and final |
segment.text | The full text of the segment so far |
segment.start_time / end_time | Segment timing, in seconds relative to session start |
delta | Incremental hint: the text added since the previous partial of the same segment (see below) |
Working with delta
When the filler filter is disabled, delta.text is append-only:
each transcription message carries exactly the text appended
since the previous partial, so you can concatenate deltas directly.
With the filler filter enabled, the recognizer's tail might be rewritten mid-segment, which breaks the append relationship. In that mode treat
segment.text as authoritative and overwrite the current segment on each message; use delta only
as a hint.
translated_transcription
Sent only when translate_languages is set, once per target language, after
each final (is_eos: true) transcription.
{
"message_type": "translated_transcription",
"transcription_id": "a1b2c3d4",
"language": "es",
"is_eos": true,
"segment": {
"text": "Hola mundo, ¿cómo estás?",
"start_time": 0.32,
"end_time": 1.84
}
}
transcription_id matches the id of the source transcription (the
is_eos: true one) this translation was produced from — use it to correlate a
translation back to its original segment. language here is the target
language, and is_eos is always true (translations are produced only for
finalized segments).
Errors
Authentication and routing failures are reported as HTTP status codes during the WebSocket upgrade, before the connection is established:
| HTTP status | Meaning |
|---|---|
401 | Missing or invalid API Key / token |
409 | A session is already active for this identity |
After a successful upgrade, the server does not send application-level error messages over the wire — it closes the connection with a standard WebSocket close frame.
Complete example
Streams microphone audio and prints transcriptions (and translations, if
PALABRA_LANGUAGE targets are configured).
pip install pyaudio websockets
export PALABRA_API_KEY=... # from Step 1
export PALABRA_LANGUAGE=en # source language
import json
import os
import asyncio
import threading
import queue
import pyaudio
import websockets
WS_URL = "wss://api.palabra.ai/asr/v1/speech-to-text/stream"
LANGUAGE = os.environ.get("PALABRA_LANGUAGE", "en")
SAMPLE_RATE = 16000
CHANNELS = 1
CHUNK = 5120 # samples ≈ 320 ms at 16 kHz (recommended chunk size)
def mic_reader(audio_queue: queue.Queue, stop_event: threading.Event):
pa = pyaudio.PyAudio()
stream = pa.open(
format=pyaudio.paInt16,
channels=CHANNELS,
rate=SAMPLE_RATE,
input=True,
frames_per_buffer=CHUNK,
)
print("Microphone open, speak now...")
try:
while not stop_event.is_set():
audio_queue.put(stream.read(CHUNK, exception_on_overflow=False))
finally:
stream.stop_stream()
stream.close()
pa.terminate()
async def stream(token: str):
url = (
f"{WS_URL}?token={token}&language={LANGUAGE}"
f"&format=pcm_s16le&sample_rate={SAMPLE_RATE}"
)
audio_queue: queue.Queue = queue.Queue()
stop_event = threading.Event()
threading.Thread(
target=mic_reader, args=(audio_queue, stop_event), daemon=True
).start()
async with websockets.connect(url) as ws:
print("Connected")
async def send_audio():
loop = asyncio.get_event_loop()
while True:
data = await loop.run_in_executor(None, audio_queue.get)
await ws.send(data) # raw binary frame
async def receive():
async for message in ws:
msg = json.loads(message)
msg_type = msg.get("message_type")
if msg_type == "transcription":
text = msg["segment"]["text"]
tid = msg.get("transcription_id", "")
if msg.get("is_eos"):
print(f"\n[EOS] {text} [{tid}]")
else:
# segment.text is the source of truth — render it whole
print(f"\r {text}", end="", flush=True)
elif msg_type == "translated_transcription":
lang = msg.get("language", "?")
tid = msg.get("transcription_id", "")
print(f"\n[{lang}] {msg['segment']['text']} [{tid}]")
try:
await asyncio.gather(send_audio(), receive())
finally:
stop_event.set()
if __name__ == "__main__":
try:
asyncio.run(stream(os.environ["PALABRA_API_KEY"]))
except KeyboardInterrupt:
print("\nStopped.")