Skip to content
  • 6 min read

Guardrails AI — Testing Guardrails

What to Test

Level Question Tooling LLM calls
Wiring Right validators on the right target with the right on_fail? guard.get_validators(...) None
Unit Does each guard pass/fix/block known strings? guard.validate() / parse() None
Quality What are the false-positive and false-negative rates? Labelled datasets, metrics None (except LLM-judge validators)
Flow Does reask recover? Is the LLM skipped when input is blocked? Fake llm_api or LiteLLM mock_response Mocked
Integration Does the real model + guard behave in production-like conditions? Real models, marker llm Real, nightly
Adversarial Can attackers get past the guard? Red-team datasets and scanners Real

Project Layout

app/guards.py              # guard factories: build_support_guard(), build_triage_guard()
tests/conftest.py          # fixtures, metrics collection off
tests/test_guards.py       # unit + flow tests
tests/data/pii_cases.jsonl # labelled regression samples
tests/test_quality.py      # FP/FN rates, CI thresholds

Build guards in factory functions — each test gets a fresh Guard with empty history.

# app/guards.py
from typing import Literal

from pydantic import BaseModel, Field

from guardrails import Guard, OnFailAction
from guardrails_ai.secrets_present import SecretsPresent
from guardrails_ai.valid_choices import ValidChoices
from guardrails_ai.valid_length import ValidLength


def build_support_guard() -> Guard:
    return (
        Guard(name="support-bot")
        .use(SecretsPresent(on_fail=OnFailAction.EXCEPTION), ValidLength(max=4000, on_fail="exception"), on="messages")
        .use(SecretsPresent(on_fail=OnFailAction.FIX))
    )


class Triage(BaseModel):
    label: str = Field(validators=[ValidChoices(choices=["bug", "feature", "question"], on_fail="reask")])
    priority: Literal["low", "medium", "high"]


def build_triage_guard() -> Guard:
    return Guard.for_pydantic(Triage)
# tests/conftest.py
import pytest

from app.guards import build_support_guard, build_triage_guard


@pytest.fixture
def support_guard():
    guard = build_support_guard()
    guard.configure(allow_metrics_collection=False)     # no anonymous telemetry from CI
    return guard


@pytest.fixture
def triage_guard():
    guard = build_triage_guard()
    guard.configure(allow_metrics_collection=False)
    return guard

Unit Tests Without an LLM

All tests on this page were run with guardrails-ai 0.11.0 and pytest.

def test_clean_answer_passes_unchanged(support_guard):
    outcome = support_guard.validate("Go to Settings -> Security -> Reset 2FA.")
    assert outcome.validation_passed
    assert outcome.validated_output == outcome.raw_llm_output
    assert outcome.validation_summaries == []


def test_secret_in_answer_is_masked(support_guard):
    outcome = support_guard.validate("Use AKIAIOSFODNN7EXAMPLE to call the API.")
    assert outcome.validation_passed                                   # fix => "passed" ...
    assert "AKIAIOSFODNN7EXAMPLE" not in outcome.validated_output
    assert [s.validator_name for s in outcome.validation_summaries] == ["SecretsPresent"]   # ... but reported


def test_guard_wiring(support_guard):
    assert {type(v).__name__ for v in support_guard.get_validators("messages")} == {"SecretsPresent", "ValidLength"}
    assert [type(v).__name__ for v in support_guard.get_validators("output")] == ["SecretsPresent"]


def test_triage_schema(triage_guard):
    ok = triage_guard.parse('{"label": "bug", "priority": "high"}')
    bad = triage_guard.parse('{"label": "urgent", "priority": "high"}', num_reasks=0)
    assert ok.validated_output == {"label": "bug", "priority": "high"}
    assert not bad.validation_passed and bad.validated_output is None

Custom validators can also be tested directly: NoPromptLeak(canary="zx").validate("zx", {}) returns a FailResult with fix_value and metadata (04).

Parametrized Positive / Negative Cases

import pytest

CASES = [
    pytest.param("Your order ships tomorrow.", False, id="plain"),
    pytest.param("Contact support@example.com for help.", False, id="email-is-not-a-secret"),
    pytest.param("aws_secret_access_key = wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY", True, id="aws-secret"),
    pytest.param("token: ghp_1234567890abcdefghijklmnopqrstuvwxyzAB", True, id="github-token"),
]


@pytest.mark.parametrize(("text", "has_secret"), CASES)
def test_secret_detection(support_guard, text, has_secret):
    assert bool(support_guard.validate(text).validation_summaries) is has_secret

Always include negatives that look dangerous (emails, UUIDs, order ids, code snippets) — they are where false positives come from.

Measuring False Positive / False Negative Rates

import json
from pathlib import Path

from app.guards import build_support_guard

MAX_FNR = 0.02      # missed secrets: security budget
MAX_FPR = 0.05      # blocked clean answers: UX budget


def test_secret_guard_rates():
    guard = build_support_guard()
    tp = fp = fn = tn = 0
    for line in Path("tests/data/secrets_cases.jsonl").read_text().splitlines():
        case = json.loads(line)                               # {"text": "...", "label": true}
        flagged = bool(guard.validate(case["text"]).validation_summaries)
        tp += flagged and case["label"]
        fp += flagged and not case["label"]
        fn += (not flagged) and case["label"]
        tn += (not flagged) and not case["label"]
    fnr, fpr = fn / max(tp + fn, 1), fp / max(fp + tn, 1)
    print(f"TP={tp} FP={fp} FN={fn} TN={tn} FNR={fnr:.3f} FPR={fpr:.3f}")
    assert fnr <= MAX_FNR, f"missed secrets: {fnr:.1%}"
    assert fpr <= MAX_FPR, f"over-blocking: {fpr:.1%}"
Metric Meaning for a guard Who cares
False negative rate Attacks / leaks that got through Security, compliance
False positive rate Legitimate users blocked or answers mangled Product, support
Precision / recall per category E.g. PII: phone numbers vs names Tuning thresholds and entities
Latency p95 per validator Cost of the check SRE
  • Threshold-based validators (ToxicLanguage, DetectJailbreak): sweep the threshold on the dataset and pick it from the curve, not from the default.
  • Datasets: real (anonymized) traffic + synthetic edge cases + known attacks; label with two reviewers for ambiguous items.
  • Store the last run's rates as a baseline; fail CI on regression beyond a tolerance, not only on absolute thresholds.

Regression Suites

Suite Content Source
jailbreaks.jsonl DAN-style prompts, role-play, encoding tricks (base64, leetspeak), multilingual variants Public jailbreak sets, red-team findings
pii.jsonl Emails, phones, IBANs, addresses in many formats and languages; look-alikes (order ids, versions) Faker + production incidents
injections_indirect.jsonl Instructions hidden in RAG documents and tool results Red team, bug bounty
off_topic.jsonl Requests outside the bot's scope Support logs

Every production incident or red-team finding becomes a new labelled row — the suite only grows.

Flow Tests with a Mocked LLM

A fake llm_api gives full control over the reply sequence:

import pytest
from guardrails.errors import ValidationError


def test_reask_recovers(triage_guard):
    replies = iter(['{"label": "urgent", "priority": "high"}', '{"label": "bug", "priority": "high"}'])
    sent = []

    def fake_llm(*, messages, **kwargs):
        sent.append(messages)
        return next(replies)

    outcome = triage_guard(llm_api=fake_llm, messages=[{"role": "user", "content": "Checkout crashes"}], num_reasks=1)

    assert outcome.validation_passed and outcome.validated_output["label"] == "bug"
    assert len(triage_guard.history.last.iterations) == 2
    assert "urgent" in sent[1][-1]["content"]                 # reask prompt quotes the bad value


def test_input_block_skips_llm(support_guard):
    calls = []
    with pytest.raises(ValidationError):
        support_guard(llm_api=lambda *, messages, **kw: calls.append(1) or "ok",
                      messages=[{"role": "user", "content": "my key AKIAIOSFODNN7EXAMPLE"}])
    assert calls == []                                        # blocked before the model

When the guard calls a model string, LiteLLM's mock_response passes through (see LiteLLM testing):

def test_reask_exhausted(triage_guard):
    outcome = triage_guard(
        model="anthropic/claude-haiku-4-5",
        messages=[{"role": "user", "content": "Label this"}],
        mock_response='{"label": "urgent", "priority": "high"}',   # every attempt returns the same bad value
        num_reasks=1,
    )
    assert not outcome.validation_passed
    assert len(triage_guard.history.last.iterations) == 2          # original + 1 reask, then stop

mock_response also works with stream=True — useful to assert that a secret split across chunks is still masked.

CI Gates

Stage Runs Gate
Every PR Wiring + unit + flow tests (no network) 100% pass, < 1 min
Every PR touching guards/validators Quality suite on local validators FNR/FPR within thresholds and baseline tolerance
Nightly (pytest -m llm) Real models + guards, LLM-judge validators Block/fix rates stable; cost cap
Weekly / before release Red-team scan No new critical bypasses
  • Bake ML models into the CI image or cache the model directory — do not download on every run.
  • Run with guardrails configure --disable-metrics --disable-remote-inferencing --token "" (non-interactive) or allow_metrics_collection=False.
  • Pin validator package versions: a model update inside a validator is a behaviour change that must pass the quality gate.

Red-Team Loop

flowchart LR
    A[Generate attacks<br/>DeepEval RedTeamer, datasets] --> B[Run app + guards]
    B --> C{Bypassed?}
    C -->|yes| D[Label + add to regression suite]
    D --> E[Tune guard / add validator /<br/>fix architecture]
    E --> B
    C -->|no| F[Track FPR on clean traffic]
  • Scan the whole app (prompt + guards + tools), not the guard in isolation — see Red Teaming.
  • A bypass is not always a guard bug: often the fix is fewer tool permissions or no secrets in context.
  • Retest old bypasses on every model upgrade — a new model changes what gets through.

See also