Asia/Karachi
BlogOctober 1, 2026

Speak English, They Hear German: Building a Live Phone Call Translator

Twilio, Deepgram, OpenAI and ElevenLabs, and the four places latency hides
Abdul Qudoos
Speak English, They Hear German: Building a Live Phone Call Translator — Abdul Qudoos blog cover
Translating a sentence is easy now. Any decent model does it in a second. Translating a sentence while someone is waiting on the other end of a phone line is a different problem. Every stage adds delay, and a phone conversation only tolerates so much before it stops feeling like a conversation. I built a real-time phone call translator to find out where that delay actually comes from. You speak English into a normal phone call, and the other side hears German. Here's the pipeline, stage by stage, with what each stage taught me.
Phone call → Twilio → Deepgram → OpenAI → ElevenLabs → caller hears German
Twilio's Media Streams send live call audio to your server over a WebSocket. The setup is a voice webhook that answers the call with instructions to open a stream to wss://your-server/media-stream. The first surprise is the format. Phone audio arrives as 8 kHz μ-law, not the WAV or MP3 most tutorials assume. Every downstream service has to accept that encoding or you convert it yourself. Get this wrong and you'll spend an afternoon debugging transcripts that are pure noise. The μ-law frames go straight to Deepgram's streaming API (Nova-2, encoding: "mulaw"). Transcripts come back while the person is still talking. The key word is streaming. If you buffer audio until the speaker pauses and then send the whole clip, you've already added seconds of delay before translation even starts. This is where it gets interesting. Phone transcripts are messy: words get misheard, sentences get cut in half, and names come out wrong. Translate each fragment in isolation and you get German that's technically correct but inconsistent from one line to the next. So the translator keeps a rolling memory of the call. It stores the last 10 English/German exchanges and puts the most recent 5 into every prompt:
You are translating a phone conversation from English to German.
Maintain consistency with previous translations and correct any
transcription errors. Here is the recent conversation context:

English: ...
German: ...

Now translate this new English text to German, keeping it consistent
with the conversation flow and fixing any obvious transcription errors:
English: {current segment}
German:
Three choices made this reliable:
  • Temperature 0.1. Translation isn't creative writing. Low temperature means the same English gives the same German.
  • Stop sequences ("English:", "German:", and a blank line) stop the model from inventing the next turn of the conversation.
  • Let the translator fix the transcript. Asking the model to correct obvious mistakes as it translates removes a whole class of errors.
Why only the last 5 exchanges? The whole call would give more context, but every extra line makes the prompt longer, slower and more expensive. Five turned out to be enough to keep names and terms consistent. The German text streams to ElevenLabs over a WebSocket (eleven_flash_v2_5). The setting that matters most is chunk_length_schedule: [50, 120, 160, 250]. It tells ElevenLabs to start generating audio after a short first chunk of text instead of waiting for a long one. The result: the listener hears the start of the German sentence while the rest is still being generated. The difference between "waiting for the robot" and "listening to someone talk" comes down to this one setting. Keep:
  • Streaming at both ends. Speech-to-text and text-to-speech both stream, and translation runs on each transcript segment as soon as it arrives.
  • A rolling context window instead of the full call history.
  • REST endpoints that expose each call's transcripts and translations. Being able to see exactly what was heard and how it was translated made debugging much faster.
Change:
  • Make the prompt context work in both directions. A German-to-English path exists, but only English-to-German uses the rolling memory today.
  • Add proper latency checkpoints on every stage from day one. I added them later, on the voice agent this pipeline grew into.
That last point is the real story. This translator became the audio foundation for a full LangGraph voice agent that takes real actions. Same pipe, more brains.
Need real-time voice in your product, for translation, support or scheduling? Let's talk.
Share this post: