"Hello? Are you still there?"
Measure first, then argue
| Operation | Time |
|---|---|
| Google Calendar: update an appointment | ~2.19 s |
| Google Calendar: fetch appointments | ~617 ms |
Trick #1: say something human while the tool runs
- "Let me pull up your appointments and check your current schedule."
- "I'm updating your appointment in the calendar system right now."
- "Let me update your calendar with the new appointment time."
- The fillers are pre-recorded, not generated live. Running text-to-speech for the filler would itself add latency, which defeats the point. A small FillerAudioService loads audio files from disk at startup and plays the right one in milliseconds.
- The fillers say what's actually happening. A generic "hmm, one moment" on every turn sounds robotic by the third time. A filler that names the task sounds like someone actually doing the work.
Trick #2: fetch before anyone asks
Js
What I'd tell anyone building a voice agent
- Time every step. Put timing checkpoints on every node and tool call from day one. Gut feel about latency is almost always wrong.
- Split your delays into two piles: ones you can remove (preload, cache, parallelise) and ones you can only hide (fillers).
- Pre-record anything you say often. Live text-to-speech is for the answer, not for "one moment".
- Match the filler to the action. It's the difference between "the bot is thinking" and "someone is helping me".
Building a voice agent and fighting latency? Tell me about it. This is exactly the kind of problem I like working on.