Back to Production Projects

Voice AI · AWS Infrastructure

Voice Agents, Owned End to End

A template-driven voice AI platform, migrated off managed cloud onto self-hosted AWS EC2 — with per-call cost telemetry built into the code and validated on real calls.

$0.068Per-minute cost, validated
2AWS regions
4Production agents
~$30/moFlat infrastructure

The Economics

Why I Took It Off Managed Cloud

Voice AI wrappers — Retell, Vapi, Bland — sell the same stack at $0.10-0.15/min: LiveKit media plus the same OpenAI and Deepgram APIs you can call directly. Managed LiveKit Cloud is fairer at $0.07/min, but it still scales linearly with every minute sold.

Self-hosting flips the model: one ~$30/mo EC2 instance carrying unlimited agents, plus API costs at source rates. The break-even lands around 1,000 minutes a month — after that, every minute sold is margin the wrappers would have kept.

So the platform was migrated: LiveKit Cloud is no longer used. Same agents, same code — owned infrastructure underneath.

Self-hosted vs LiveKit Cloud — validated projections

Monthly volumeSelf-hostedLiveKit Cloud
1,000 min/mo~$74~$70
10,000 min/mo~$474~$700
50,000 min/mo~$2,100~$3,500

At 50,000 minutes a month the owned stack costs $2,100 against $3,500 — and the gap compounds with every client added to the same box.

The Architecture

One Box, Everything On It

The real deployed call path, exactly as it runs — verified against the production repos, not a diagram of aspirations.

Widget — Vercel / Next.js

iframe-embeddable frontend with a server-side token route: Cloudflare Turnstile verification and per-IP rate limiting (5 req/min) before a LiveKit token is ever issued.

nginx → LiveKit Server

WSS terminates at nginx (SSL, port 443) and proxies to a Dockerized LiveKit Server on :7880 — TURN over 3478/5349, WebRTC media on UDP 50000-60000.

Python worker — systemd

One systemd service per agent (Restart=always, Docker-gated): Deepgram nova-3 STT → GPT-4.1 → Deepgram Aura-2 TTS, Silero VAD, streaming throughout.

The world outside

n8n webhooks for lead capture, calendar booking, and call-completed events; Cloudflare R2 for call recordings. AI APIs stay direct — no wrapper markup.

Validated per-minute cost — real test calls

Deepgram STT — nova-3per audio second$0.0077
Deepgram TTS — Aura-2per 1K characters$0.0237
OpenAI GPT-4.1per token in/out$0.0367
AWS EC2 — amortized~$30/mo flat~$0.001
Total, validated$0.068

The LLM dominates — a ~4,000-token system prompt rides every turn, which is exactly the kind of finding you only get from measuring your own calls.

Cost Telemetry

Every Call, Priced by Its Own Code

Provider dashboards can't isolate a single call's cost — so cost tracking was built into the agents themselves. An llm_node() override captures LLM prompt and completion tokens and TTS character counts into per-call state, and the call-completed webhook carries the raw usage out.

An n8n Code node then computes STT, TTS, LLM, and total cost from the raw data — rates declared in one place — and writes every call to Google Sheets and the CRM. No dashboard dependency: code-level tracking is authoritative.

The whole pattern ships as a reusable template doc, so any new agent gets instrumented the day it's cloned.

70%of round-trip is the LLM

Hosting changes the transport, not the model. The self-hosting penalty is 100-300ms — imperceptible next to a 500-1500ms LLM turn.

2regions, chosen for latency

Singapore (ap-southeast-1) with a permanent Elastic IP for Asia; us-east-1 beside OpenAI and Deepgram for US callers.

0serverless shortcuts

Fargate was evaluated and rejected — LiveKit requires host networking. The honest answer was a real box, run like one.

The Fleet

Four Agents, One Template

Every agent is cloned from the same template — customized per client in under 30 minutes: clone, fill the env, edit the system prompt, deploy.

“May”

Reference deployment

The portfolio receptionist — first agent migrated to EC2, and the birthplace of the cost-tracking system. Her repo carries the reusable template doc.

“Sarah”

Commercial cleaning client

Answers questions, qualifies leads (business name included), and books appointments against the client's calendar — timezone-correct through DST.

Bilingual agent

Filipino-English

Speaks Taglish with polite po/nyo markers and a custom Filipino voice persona — running as its own systemd service on the same box.

Agency agent

Marketing agency site

Live receptionist on the agency's contact page — one of the two public deployments you can call yourself, right now.

Production Scars

What Live Traffic Teaches

Every system below earned its keep by failing first — and being fixed with the evidence in hand.

The timezone bug

Bookings used a hardcoded UTC offset — 1 hour off in winter — and the availability check was built in UTC, ~7 hours off. Rebuilt on luxon with America/Los_Angeles zoning: correct through every DST transition.

Token endpoint hardening

The token route is the front door, so it got a door: Turnstile verification plus per-IP rate limiting before any LiveKit credential is minted.

Lead state in every payload

A contextvars pattern carries lead and booking fields into the call-completed webhook, so n8n routes booked calls and missed calls to the right sheet — no post-hoc matching.

Don't take my word for it

Two of these agents are live on public contact pages — call one and hear the stack for yourself.