VOICE

Realtime speech infrastructure for your applications.

STT, TTS, turn-taking and tracing as one runtime. Your agent stays yours.

EVERY SEGMENT MEASURED — LATENCIES PER HOP
audio inVAD20ms framesSTT<300ms partialyour agentHTTP + SSETTS96ms first audioaudio out

Target figures — not yet measured.

stt.partial · stt.final
Realtime STT

Partial and final transcripts as the caller speaks. Spanish first, multilingual by design. Open-weight models available for self-host.

tts.first_audio
Streaming TTS

Speech streams per sentence, not per response. Voices are platform resources — switch providers or bring cloned voices without changing code.

VOICE CLONING: COMING
turn.barge_in
Turn-taking & barge-in

The caller interrupts, the agent stops. The runtime detects the turn and cancels TTS mid-sentence — your code never juggles timers.

agent speakingcaller barges in→ tts.cancel · 61ms
latency.turn
Latency tracing

Every conversation leaves a timeline you can query. Per-hop latencies, per-turn totals, exportable.

● 00:03.1 · stt.final142ms
● 00:03.3 · agent.first_token210ms
● 00:03.4 · tts.first_audio96ms
● 00:03.4 · latency.turn448ms

Target figures — not yet measured.

What it's for.

USE CASE
WHAT LETSYLABS DOES
AI agents
Full voice loop — your agent thinks, the runtime listens and speaks.
Call automation
Inbound and outbound calls as programmable sessions with DTMF and transfers.
Live transcription
Partials in under 300 ms, finals with punctuation, per-speaker.
Assistants
Web and in-app voice with barge-in that feels human.
Accessibility
Low-latency speech in and out for interfaces that talk.
POST /v1/sessions create a session
WS   wss://…/audio stream audio + events
GET  /v1/sessions/:id/tracequery the timeline