Overview
How it works
- Telephony: a Twilio voice webhook opens a Media Stream, and call audio arrives over a WebSocket as μ-law frames.
- Speech-to-text: the audio goes straight to Deepgram's streaming API (Nova-2, μ-law encoding) for live transcripts.
- Translation with memory: each transcript segment is translated by OpenAI with the last few English/German exchanges in the prompt. The model keeps terms consistent across the call and fixes obvious transcription errors before translating.
- Text-to-speech: the German text streams to ElevenLabs over WebSockets (eleven_flash_v2_5, tuned chunk scheduling), so audio starts playing before the sentence is finished.
- Observability: REST endpoints expose each call's audio chunks, transcripts and translations, and the React frontend shows them live.
Engineering decisions
- Stream everything: speech-to-text, translation and text-to-speech all run incrementally. Waiting for complete sentences at each stage adds seconds of delay, which ruins a phone conversation.
- A rolling memory instead of the full history: the last 10 exchanges are kept and the last 5 go into the prompt. That's enough to keep translations consistent without growing token cost and latency over a long call.
- Let the translator repair speech-to-text: phone audio is noisy. Asking the model to correct obvious transcription mistakes as it translates fixed errors that would otherwise carry through to the German.
Tech stack
Where it led
Twilio
Deepgram
OpenAI
ElevenLabs
Node.js
React