Building a Voice AI Agent: Twilio + OpenAI TTS + OpenClaw in One Night
What happens when you try to build a fully functional voice AI agent in a single evening? You get 4 hours of debugging, 7 failed calls, 3 root causes, and one very satisfying “it works” moment at 1 AM. Total API cost: $0.03.
The goal: call a phone number, have an AI agent answer, listen to you, understand you, and respond in natural language — in Slovak and English. The stack: Twilio for telephony, OpenAI for text-to-speech and speech-to-text, and OpenClaw as the AI agent platform.
The Architecture
📱 +XXX XXX XXX XXX (your phone)
↕ PSTN (telephone network)
Twilio Cloud (+X XXX XXX XXXX, US number)
↕ HTTPS webhooks + WebSocket media stream
Cloudflare (public domain)
↕ Nginx reverse proxy
OpenClaw Voice Call Plugin (localhost)
↕ /voice/webhook → TwiML response (activates Media Streams)
↕ /voice/stream → WebSocket (PCM μ-law 8kHz audio chunks)
↕ TTS: OpenAI API (tts-1, alloy voice, $0.015/min)
↕ STT: OpenAI Realtime (gpt-4o-transcribe)
↕ AI: GLM-5 LLM (real-time responses)
The key insight is that phone calls require PCM audio — not MP3, not Opus, not whatever your browser uses. The entire pipeline revolves around converting between telephone audio (μ-law 8kHz) and the formats that AI services expect.
The Night: March 29, 2026
21:58 — Plugin Installed
OpenClaw runs on a 1GB RAM VPS via bun, which means npm (required by the plugin installer) wasn’t available. Solution: nvm + Node.js v22. After that, openclaw plugins install @openclaw/voice-call and a gateway restart. Nginx reverse proxy configured for the webhook and WebSocket media stream.
22:43–22:58 — Calls Hang Up After 300ms
7 failed attempts. The call connects but immediately drops. No audio. Nothing.
This is where the debugging began in earnest.
~22:50 — Root Cause #1: Edge TTS Can’t Do PCM
OpenClaw uses Microsoft Edge TTS for Telegram voice messages. But phone calls need PCM audio. Edge TTS is simply ignored for voice calls — the log says “Microsoft speech is ignored for voice calls (telephony audio needs PCM).” No audio buffer means an empty TwiML response, which means Twilio hangs up.
| TTS Provider | PCM Output | Works for Calls |
|---|---|---|
| Microsoft Edge TTS | No | No |
| Google Gemini TTS | No | No |
OpenAI TTS (tts-1) | Yes (24kHz → μ-law 8kHz) | Yes |
| ElevenLabs | Yes | Yes |
~23:05 — Root Cause #2: Cloudflare Breaks Twilio Signatures
Twilio verifies webhook requests using HMAC-SHA1 signatures. When the request goes through Cloudflare, the IP address and headers change, so the signature doesn’t match. The plugin responds with HTTP 401, and Twilio hangs up the call.
Fix: webhookSecurity.trustForwardingHeaders: true in the plugin config.
~23:08 — Root Cause #3: Notify Mode Doesn’t Use Streaming
Notify mode generates inline TwiML with <Say> tags in the webhook response — no Media Streams WebSocket. The TTS provider only initializes when streaming.enabled: true. Result: empty <Response></Response>.
The fix is to use conversation mode instead of notify mode. Conversation mode activates Media Streams and the full TTS pipeline.
23:45 — OpenAI API Key Added
Added OPENAI_API_KEY + OpenAI TTS config for the voice-call plugin. Streaming ON. First test — TTS works (141KB MP3 generated).
00:10 — First Successful Call (11s, English)
Conversation mode + OpenAI TTS + streaming. The agent spoke in English, transitioned to listening mode. After 11 seconds, the Twilio trial announcement cut in and ended the call.
00:43 — Slovak Works!
Weather report played in Slovak through OpenAI TTS in telephony PCM mode. The voice quality was surprisingly good.
01:25 — Full Conversation (STT + AI + TTS)
The agent heard speech, the AI responded, TTS played the response back. 4 full conversation turns in Slovak. This was the moment.
The First Conversation
The actual transcript from the first successful Slovak conversation (23:24–23:25 CEST):
Human: Rozmýšľam, ale rozmýšľam zajtra, aké bude počasie. Vieš mi povedať, či zajtra bude pršať?
AI: Zajtra bude pršať a teplota sa pohybuje okolo 6 stupňov Celzia.
Human: Povedz mi to ešte raz, budem teraz trpezlivo čakať.
AI: Zajtra bude pršať a teplota sa pohybuje okolo 6 stupňov Celzia.
Human: Koľko mám onduliek?
AI: Máš dve andulky, obidve sú zelené.
Human: Super, funguje to, všetko som počul, idem zložiť.
* STT occasionally mistranscribed Slovak (e.g. “on doliek” instead of “onduliek”)
How One Conversation Turn Works
| Step | Process | Latency |
|---|---|---|
| Listening (STT) | You speak → Twilio sends audio via WebSocket → OpenClaw forwards to OpenAI Realtime API → real-time transcription | ~2s |
| AI Response | Transcript sent to LLM (GLM-5) → AI generates response → text sent to TTS | ~2–3s |
| Playback (TTS) | OpenAI TTS generates PCM audio → converts to μ-law 8kHz → sends chunks (160B / 20ms) via WebSocket → Twilio plays | ~2–3s |
| Total per turn | ~5–10s |
5–10 seconds per turn isn’t real-time conversation speed, but it’s perfectly usable for information queries and voice commands.
Configuration
The key pieces of configuration that made it work:
openclaw.json — voice-call plugin:
{
"voice-call": {
"provider": "twilio",
"fromNumber": "+XXXXXXXXXXX",
"toNumber": "+XXXXXXXXX",
"publicUrl": "https://your-domain.com/voice/webhook",
"streaming": { "enabled": true },
"webhookSecurity": {
"allowedHosts": ["your-domain.com"],
"trustForwardingHeaders": true
},
"tts": {
"provider": "openai",
"openai": { "voice": "alloy", "model": "tts-1" }
},
"outbound": { "defaultMode": "conversation" }
}
}
Nginx — voice proxy:
location /voice/webhook {
proxy_pass http://127.0.0.1:PORT/voice/webhook;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}
location /voice/stream {
proxy_pass http://127.0.0.1:PORT/voice/stream;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
}
Lessons Learned
PCM audio is a hard requirement. Regular TTS engines (Edge, Google) don’t work for phone calls. You need a PCM-compatible provider like OpenAI or ElevenLabs.
Cloudflare + Twilio = signature problems. Add
trustForwardingHeadersor use a direct IP without proxy.Streaming must be ON. Without it, the plugin doesn’t initialize the TTS provider and the webhook responds with empty TwiML.
Conversation mode > Notify mode. Notify mode bypasses streaming entirely — it won’t work without PCM TTS.
Slovak works. Both OpenAI TTS and Realtime STT handle Slovak, though occasional transcription errors occur.
Cost is negligible. $0.03 for setup, debugging, 7 failed calls, and the first successful conversation (22 requests, 177K tokens).
What’s Next
| Priority | Item | Description |
|---|---|---|
| High | Upgrade Twilio | Trial account plays announcement and requires dial code. Paid account removes these limits |
| Bug | Webhook server stuck | After each call, the webhook server locks up. Requires gateway restart. Likely a plugin bug (v2026.3.13) |
| Nice | Better voice | Consider shimmer or nova voice, or ElevenLabs for better Slovak quality |
| Config | Stale call reaper | Add staleCallReaperSeconds: 360 to prevent stuck calls from holding the port |
Final Thoughts
Four hours, three root causes, seven failed calls. But by 1 AM, we had a working voice AI agent that speaks Slovak, understands context (it knew about the budgies), and responds in real time over a regular phone call. The entire setup cost $0.03 in API calls and runs on a $5/month VPS. Not bad for a Saturday night.