← Work

Case study · 2026

A telephony voice agent, end to end

A proof of concept for AI agents on real phone calls under HIPAA constraints, from an empty repo to inbound and outbound calls with transfer and DTMF, plus a setup guide that works without me on the call.

Period
2026
Stack
  • Python 3.12
  • LiveKit Agents
  • LiveKit Inference
  • Deepgram
  • OpenAI
  • Cartesia
  • FastAPI
  • Twilio SIP
  • PostgreSQL
  • Docker

Context

A proof of concept for AI agents that handle real phone calls in a HIPAA-conscious environment: LiveKit Agents in Python, a small FastAPI control plane, and a Twilio SIP trunk for the public phone network.

The problem

Get from an empty repo to inbound and outbound calls with transfer and DTMF, with a setup guide someone else can follow without me on the call.

What I did

  • Wrote the ADR mandating LiveKit Inference for STT, LLM and TTS instead of direct provider plugins: one contract, one billing surface, fewer credentials to manage.
  • Set up the SIP trunk with credential-list authentication (LiveKit Cloud has no stable egress IPs to allowlist), TLS transport and required media encryption.
  • Added a “test without Twilio first” path — console mode and the agent console — so nobody burns trunk minutes debugging config.
  • Turned the manual steps into Makefile targets: download the turn-detector model, run the control plane, place a parameterized test call.
  • Debugged the real failures: a config validation error from a missing variable, a guide that contradicted the ADR by wiring provider plugins, a crash from reaching into a private session attribute, and a 403 on transfer that turned out to be a Twilio feature flag, not the international number I suspected.

The trade-off

Inference gives up fine-grained provider control for a simpler compliance and ops story. Console-first testing costs a day of setup and saves every later one.

Outcome

Working inbound and outbound calls with transfers and DTMF, a reproducible guide, and a database connectivity check that fails fast at startup instead of halfway through a call.

What I’d do differently

Encode ADR decisions as checks in CI. A guide and a codebase will drift; a lint rule won’t.