LLM Output Evaluation with DeepEval¶
A pytest-based test suite that evaluates LLM output quality with DeepEval metrics, using an LLM-as-a-judge approach. The judge model is configurable (Cerebras or OpenRouter), and the tests are grouped by metric category, so each area can be run on its own, locally or in Docker.
What's Covered¶
- RAG — AnswerRelevancy, Faithfulness, ContextualPrecision, ContextualRecall, ContextualRelevancy
- LLM quality — Hallucination, Toxicity, Bias, Summarization, GEval, PIILeakage, Misuse, NonAdvice, RoleViolation
- Agents — TaskCompletion, ToolCorrectness, GoalAccuracy, ToolUse, ArgumentCorrectness
- Chatbots — KnowledgeRetention, ConversationCompleteness, RoleAdherence, TopicAdherence, TurnRelevancy, TurnFaithfulness, ConversationalGEval
- Extra metrics — ExactMatch, JsonCorrectness, PatternMatch, PromptAlignment, ArenaGEval, DAGMetric
- Custom metrics — GEval with
evaluation_stepsand an answer correctness pattern - E2E and red teaming — multi-metric conversational scenarios; PII leakage, jailbreak and prompt injection
How It's Built¶
- One module per metric category — RAG, LLM quality, agent, chatbot, custom, extra, E2E and red teaming tests live in separate files, each with a matching
maketarget. - Env-based configuration —
pydantic-settingsloads the judge provider, model and API keys from.env. - Shared fixtures —
conftest.pyprovides the judge model and the inter-test delay. - Rate-limit guard — an
autousefixture pauses between tests (10 seconds by default,TEST_DELAY_SECONDSto override) to avoid API 429 errors. - Structured JSON logs —
JSONLoggerwrites a timestamped log per run with metric name, score, success and reason for each test. - Deterministic subset —
make test-deterministicruns the fast tests that need no LLM.
Tech Stack¶
deepeval— LLM evaluation metricspytest+pytest-html— test runner and HTML reportspydantic+pydantic-settings— env-based configurationuv— Python package managerruff+mypy— linting and type checkingDocker+docker compose— containerised test execution
Running It¶
Install dependencies, create .env with the judge model settings, then run the suite:
make install # install dependencies (uv sync --extra dev)
make env # create .env from .env.example
make test # run all tests
make test-html # run tests with HTML report → reports/test_report.html