How to Test Voice AI Agents: TTS, STT, Audio Metrics and LLM-as-a-Judge Evals in Practice
A voice AI agent talks to customers on the phone by itself: it greets them, checks who is on the line, answers questions, calls tools and transfers the call to a human. In the demo everything sounds smooth. Then a ticket shows up in QA: “check that the agent handles the call correctly”. Very quickly it becomes clear that the usual LLM evals won’t help much here: they check text, but the customer hears sound.