Skip to content

Blog

How to Test Voice AI Agents: TTS, STT, Audio Metrics and LLM-as-a-Judge Evals in Practice

A voice AI agent talks to customers on the phone by itself: it greets them, checks who is on the line, answers questions, calls tools and transfers the call to a human. In the demo everything sounds smooth. Then a ticket shows up in QA: “check that the agent handles the call correctly”. Very quickly it becomes clear that the usual LLM evals won’t help much here: they check text, but the customer hears sound.

DeepEval + Phoenix: Tracking How Agent Metrics Change Between Runs

When an LLM test fails, the log holds more than a number. It holds a verdict — Answer Relevancy 0.71 against a threshold of 0.8, plus a reason from the judge that says, in plain text, what is wrong with the answer. The question “why this score” is answered on the spot. That is the whole point of LLM-as-a-judge — the score comes with an explanation.

Testing LLM Outputs: A Hands-On Guide to DeepEval Metrics

Somebody on the team ships an LLM feature. Then somebody else has to test it. That second person opens the ticket, reads “verify the model does not hallucinate,” and realises there is no assert status == 200 for this. The output is different every run. The expected result is… vibes? Welcome to LLM testing in 2026.

n8n for QA: Automate the Boring Stuff You Keep Doing Manually

Half of what a QA engineer does in a day isn’t testing. It’s the stuff between testing — the notifications, the data preparation, the environment babysitting, the reporting nobody reads until something breaks, the test result summaries, the Jira hygiene, the smoke checks after another team’s deploy to shared staging. It doesn’t require thinking. It just requires someone to do it. And that someone is always QA.