Give your AI a voice — on the phone and on the web.
Stream audio, transcribe in realtime, speak with low latency and route phone calls — connected to your own agent or LLM. Self-hosted or in our cloud. EU-compliant by design.
curl -X POST https://api.letsylabs.com/v1/sessions \ -H "Authorization: Bearer $LETSY_KEY" \ -d '{ "connector": "my-agent", "language": "es" }' { "session_id": "01JD…", "ws_url": "wss://…/audio" }
Bring your own intelligence.
We handle realtime media — transport, transcription, turn-taking, interruptions, speech and the audit trail. You decide what thinks. Connect your agent over HTTP + SSE, or any OpenAI-compatible endpoint.
Speech infrastructure without the agent lock-in.
Partial transcripts in under 300 ms. Spanish first, multilingual by design.
First audio per sentence, not per response. Speech starts while the model is still thinking.
Interrupt the agent with your voice. The runtime handles turns so your code doesn't.
Every segment measured, every conversation traced. Query the timeline per call.
Turn any phone call into a programmable stream.
Bring your SIP trunk or PBX — Asterisk, 3CX, your carrier — and every call becomes a session your software controls. Inbound, outbound, transfers, DTMF.
The law now requires it. We built it in.
Spain's Ley 10/2025 and the EU AI Act require AI callers to identify themselves, offer a human handoff and keep records. letsylabs ships these as infrastructure: automatic AI disclosure at call start, guaranteed transfer to a human, synthetic-audio marking and a certified, exportable audit log per call.
The caller is told it's an AI at call start. Automatic, logged, verifiable.
A guaranteed transfer path to a person, on request or on rule.
A certified, exportable event log per call. Every disclosure, every transfer.
Recording retention, deletion and rectification — configured, not coded.
Your infrastructure. Your data. One command.
Run the entire platform on your servers with docker compose — including open-weight STT and TTS models, so audio never leaves your infrastructure. The same stack we run in our cloud.
Open where it matters.
The core contracts, VAD and open-weight model adapters are Apache 2.0. Build with the same types we build on — in Rust today, Python soon.
Session, event and audio-frame contracts. The types everything speaks.
Voice activity detection tuned for conversation, not dictation.
Open-weight model adapters — Voxtral and friends.
From first stream to production.
# 1. create a session bound to your connector curl -X POST https://api.letsylabs.com/v1/sessions \ -H "Authorization: Bearer $LETSY_KEY" -d '{ "connector": "my-agent" }' # 2. stream audio over the returned ws_url websocat "wss://api.letsylabs.com/v1/audio?session=01JD…"
let session = letsy::Session::create("my-agent").await?; let mut events = session.connect().await?; while let Some(ev) = events.next().await { if let Event::TranscriptFinal(t) = ev { agent.reply(&t.text).await?; } }
# letsy-py lands in 2027 — same contracts, same events session = letsy.Session.create(connector="my-agent") async for ev in session.events(): if ev.type == "transcript.final": await agent.reply(ev.text)
Voice is the first sense.
letsylabs is built as the interaction layer between intelligent software and the real world. We're shipping it one sense at a time.
Build voice AI that holds up in production — and in an audit.
No online form yet — this opens your email client instead.
{PRIVACY_NOTICE} Privacy policy