REALTIME VOICE INFRASTRUCTURE FOR AI

Give your AI a voice — on the phone and on the web.

Stream audio, transcribe in realtime, speak with low latency and route phone calls — connected to your own agent or LLM. Self-hosted or in our cloud. EU-compliant by design.

bash
curl -X POST https://api.letsylabs.com/v1/sessions \
  -H "Authorization: Bearer $LETSY_KEY" \
  -d '{ "connector": "my-agent", "language": "es" }'

{ "session_id": "01JD…", "ws_url": "wss://…/audio" }
stt.partial · 118ms stt.final · 142ms agent.first_token · 210ms tts.first_audio · 96ms call.transfer.human · ok trace.exported · 1.2s dtmf.received · "3" turn.barge_in · 61ms

Bring your own intelligence.

We handle realtime media — transport, transcription, turn-taking, interruptions, speech and the audit trail. You decide what thinks. Connect your agent over HTTP + SSE, or any OpenAI-compatible endpoint.

CLICK A SOURCE AND A DESTINATION TO REROUTE
PhoneBrowserYour appletsylabs runtimeVAD · STT · turns · TTS · traceYour agentOpenAIOllamaAny LLM
CLICK A SOURCE AND A DESTINATION TO REROUTE
VOICE

Speech infrastructure without the agent lock-in.

stt.partial
Realtime STT

Partial transcripts in under 300 ms. Spanish first, multilingual by design.

tts.first_audio
Streaming TTS

First audio per sentence, not per response. Speech starts while the model is still thinking.

turn.barge_in
Turn-taking & barge-in

Interrupt the agent with your voice. The runtime handles turns so your code doesn't.

latency.turn
Latency tracing

Every segment measured, every conversation traced. Query the timeline per call.

stt.final · es · 142ms
TELEPHONY

Turn any phone call into a programmable stream.

Bring your SIP trunk or PBX — Asterisk, 3CX, your carrier — and every call becomes a session your software controls. Inbound, outbound, transfers, DTMF.

Your trunkAsterisk · 3CX · SIP carrier
letsylabs
AVAILABLE AT GA
Managed numbersnumbers · queues · webphone
letsylabs
COMING SOON
Use your carrier. Or ours, later. Never locked in.
EU-COMPLIANT BY DESIGN

The law now requires it. We built it in.

Spain's Ley 10/2025 and the EU AI Act require AI callers to identify themselves, offer a human handoff and keep records. letsylabs ships these as infrastructure: automatic AI disclosure at call start, guaranteed transfer to a human, synthetic-audio marking and a certified, exportable audit log per call.

AI disclosure

The caller is told it's an AI at call start. Automatic, logged, verifiable.

Human handoff

A guaranteed transfer path to a person, on request or on rule.

Audit trail

A certified, exportable event log per call. Every disclosure, every transfer.

Retention controls

Recording retention, deletion and rectification — configured, not coded.

SELF-HOST & SOVEREIGNTY

Your infrastructure. Your data. One command.

Run the entire platform on your servers with docker compose — including open-weight STT and TTS models, so audio never leaves your infrastructure. The same stack we run in our cloud.

Self-hosted(LICENSED)
Same API, same stack.
Managed cloud
Same API, same stack.
your-server
$ docker compose up -d
✔ letsy-gateway Started 0.4s
✔ letsy-stt Started 1.1s · voxtral, open weights
✔ letsy-tts Started 1.3s · open weights
✔ letsy-trace Started 0.6s
runtime ready wss://0.0.0.0:8443/audio
OPEN SOURCE

Open where it matters.

The core contracts, VAD and open-weight model adapters are Apache 2.0. Build with the same types we build on — in Rust today, Python soon.

letsy-types

Session, event and audio-frame contracts. The types everything speaks.

Apache-2.0Rust
letsy-vad

Voice activity detection tuned for conversation, not dictation.

Apache-2.0Rust
letsy-adapters

Open-weight model adapters — Voxtral and friends.

Apache-2.0Rust
PYTHON BINDINGS: 2027
DEVELOPERS

From first stream to production.

ILLUSTRATIVE · FINAL SNIPPETS AT GA
# 1. create a session bound to your connector
curl -X POST https://api.letsylabs.com/v1/sessions \
  -H "Authorization: Bearer $LETSY_KEY" -d '{ "connector": "my-agent" }'

# 2. stream audio over the returned ws_url
websocat "wss://api.letsylabs.com/v1/audio?session=01JD…"
let session = letsy::Session::create("my-agent").await?;
let mut events = session.connect().await?;
while let Some(ev) = events.next().await {
    if let Event::TranscriptFinal(t) = ev {
        agent.reply(&t.text).await?;
    }
}
# letsy-py lands in 2027 — same contracts, same events
session = letsy.Session.create(connector="my-agent")
async for ev in session.events():
    if ev.type == "transcript.final":
        await agent.reply(ev.text)
<1s
voice-to-voiceTARGET
20ms
audio framesTARGET
100%
events tracedTARGET

Voice is the first sense.

VoiceRealtime mediaVisionAvatarsDevicesCOMING

letsylabs is built as the interaction layer between intelligent software and the real world. We're shipping it one sense at a time.

Build voice AI that holds up in production — and in an audit.

Get early access

No online form yet — this opens your email client instead.

{PRIVACY_NOTICE} Privacy policy

GA lands March 2027 · early-access pilots get founding terms