Phoenix — Testing, CI & Production¶
Three ways Phoenix helps a test suite for an LLM system:
- Trace test runs — every test gets its own session in Phoenix; a failing test links to the full trace (prompts, tool calls, tokens).
- Assert on spans — check how the agent reached an answer: tools called, number of LLM calls, token budget, no errors.
- Record tests as experiments — pytest results become experiment runs with scores, comparable across commits.
Phoenix in the Test Environment¶
| Environment | Setup |
|---|---|
| Local | uv run phoenix serve once, keep it running |
| CI | Service container arizephoenix/phoenix:<pinned> or phoenix serve with a temp PHOENIX_WORKING_DIR |
| Shared staging | Central instance with auth, project <app>-ci, PHOENIX_API_KEY from CI secrets |
# tests/conftest.py
import os
import pytest
from openinference.instrumentation import using_attributes
from openinference.instrumentation.anthropic import AnthropicInstrumentor
from phoenix.client import Client
from phoenix.otel import register
PROJECT = os.getenv("PHOENIX_PROJECT_NAME", "triage-ci")
RUN_ID = os.getenv("CI_PIPELINE_ID", "local")
# SimpleSpanProcessor (batch=False): each span is exported when it ends
tracer_provider = register(project_name=PROJECT, protocol="http/protobuf", verbose=False)
AnthropicInstrumentor().instrument(tracer_provider=tracer_provider)
@pytest.fixture(scope="session")
def phoenix() -> Client:
return Client()
@pytest.fixture(autouse=True)
def phoenix_session(request):
"""One Phoenix session per test: session.id = pytest node id."""
session_id = f"{RUN_ID}:{request.node.nodeid}"
with using_attributes(
session_id=session_id,
metadata={"test_id": request.node.nodeid, "run_id": RUN_ID, "git_sha": os.getenv("GIT_SHA", "")},
tags=["pytest", *[m.name for m in request.node.iter_markers()]],
):
yield session_id
request.node.user_properties.append(("phoenix_session", session_id)) # lands in JUnit XML
def pytest_sessionfinish(session, exitstatus):
tracer_provider.force_flush()
In the UI, filter the project by metadata['run_id'] == '1234' to see one CI run, or open Sessions to find a test by its node id.
Asserting on Spans (Trace-Based Testing)¶
Ingestion is asynchronous — poll until spans arrive.
# tests/phoenix_helpers.py
import time
import httpx
from phoenix.client import Client
def wait_for_spans(px: Client, project: str, session_id: str, min_count: int = 1, timeout: float = 15.0) -> list[dict]:
deadline = time.monotonic() + timeout
while True:
try:
spans = px.spans.get_spans(
project_identifier=project,
attributes={"session.id": session_id}, # server >= 14.9
limit=500,
)
except httpx.HTTPStatusError as exc: # 404 until the project's first span lands
if exc.response.status_code != 404:
raise
spans = []
if len(spans) >= min_count or time.monotonic() > deadline:
return spans
time.sleep(0.5)
def by_kind(spans: list[dict], kind: str) -> list[dict]:
return [s for s in spans if s["span_kind"] == kind]
# tests/test_triage_agent.py
from app.agent import triage_agent
from tests.conftest import PROJECT
from tests.phoenix_helpers import by_kind, wait_for_spans
def test_refund_ticket_uses_billing_tool_once(phoenix, phoenix_session):
result = triage_agent("I was charged twice for March, please refund")
assert result["queue"] == "billing"
spans = wait_for_spans(phoenix, PROJECT, phoenix_session, min_count=3)
tools = [s["attributes"]["tool.name"] for s in by_kind(spans, "TOOL")]
llm_calls = by_kind(spans, "LLM")
assert tools.count("lookup_invoices") == 1 # called exactly once
assert "issue_refund" not in tools # agent must not refund on its own
assert len(llm_calls) <= 3 # no runaway loop
assert all(s["status_code"] != "ERROR" for s in spans)
total_tokens = sum(s["attributes"].get("llm.token_count.total", 0) for s in llm_calls)
assert total_tokens < 4_000, f"token budget exceeded: {total_tokens}"
| Span assertion | Catches |
|---|---|
| Tool X called / not called | Wrong routing, unsafe actions (refund, delete) |
Count of LLM spans |
Infinite agent loops, retry storms |
llm.token_count.total budget |
Prompt bloat, cost regressions |
llm.model_name |
Wrong model deployed (e.g. fallback to an expensive one) |
RETRIEVER documents contain an expected doc id |
Retrieval regressions in RAG |
No ERROR status anywhere |
Swallowed tool exceptions behind a "correct" answer |
px.spans.get_spans_dataframe(query=SpanQuery().where("metadata['test_id'] == '...'")) gives the same data as a DataFrame; px.traces.get_traces(project_identifier=, session_id=, include_spans=True) returns whole traces (server ≥ 20.8).
No server needed for unit-level checks
For fast in-process tests add an InMemorySpanExporter to the provider: tracer_provider.add_span_processor(SimpleSpanProcessor(exporter)) and assert on span.attributes["openinference.span.kind"]. See OpenTelemetry — Testing.
The Phoenix pytest Plugin¶
arize-phoenix-client ships a pytest plugin (auto-registered when pytest is installed). Tests marked @pytest.mark.phoenix are recorded as experiment runs: the test file (or dataset=) becomes a dataset, each parametrized case an example, the assertion outcome a pass annotation.
import pytest
from phoenix.client.pytest import evaluate, log_evaluation, log_output
from phoenix.evals import LLM
from phoenix.evals.metrics import CorrectnessEvaluator
from app.rag import answer
judge = LLM(provider="anthropic", model="claude-haiku-4-5")
correctness = CorrectnessEvaluator(llm=judge)
CASES = [
("How long is the refund window?", "30 days"),
("Can I export reports to CSV?", "Settings → Export"),
]
@pytest.mark.phoenix(dataset="docs-rag-regression", repetitions=2)
@pytest.mark.parametrize("question,must_mention", CASES, ids=["refund-window", "csv-export"])
def test_docs_answers(question, must_mention):
reply = answer(question)
log_output({"answer": reply})
log_evaluation(name="length_ok", score=float(len(reply) < 600)) # recorded, does not fail
result = evaluate(correctness, input=question, output=reply) # recorded + returned
assert must_mention.lower() in reply.lower()
assert result[0].label == "correct"
| Control | Effect |
|---|---|
@pytest.mark.phoenix(dataset=, repetitions=, evaluators=[...], experiment_metadata=) |
Per-test settings |
PHOENIX_TEST_DATASET=smoke or ini phoenix_dataset |
Put the whole selection into one dataset |
PHOENIX_TEST_REPETITIONS=3 |
Default repetitions (flakiness measurement) |
PHOENIX_TEST_TRACKING=false |
Run the tests without talking to Phoenix |
The plugin works with pytest-xdist and prints Phoenix: recorded N run(s) across M experiment(s) in the terminal summary. It is a recent addition — pin arize-phoenix-client and check its changelog when upgrading.
Experiments as CI Regression Gates¶
# scripts/eval_gate.py — run in CI after unit tests
import sys
from phoenix.client import Client
from phoenix.client.experiments import run_experiment
from app.triage import route_ticket
THRESHOLDS = {"queue_match": 0.92, "valid_queue": 1.0}
px = Client()
dataset = px.datasets.get_dataset(dataset="triage-golden", splits=["smoke", "adversarial"])
def queue_match(output, expected) -> bool:
return output["queue"] == expected["queue"]
def valid_queue(output) -> bool:
return output["queue"] in {"billing", "bugs", "how-to", "security"}
exp = run_experiment(
dataset=dataset,
task=lambda input: route_ticket(input["ticket"]),
evaluators=[queue_match, valid_queue],
experiment_name=f"ci-{sys.argv[1]}", # commit SHA
experiment_metadata={"sha": sys.argv[1], "branch": sys.argv[2]},
client=px,
)
failed = []
for name, minimum in THRESHOLDS.items():
scores = [r.result["score"] for r in exp["evaluation_runs"] if r.name == name and r.result]
mean = sum(scores) / max(len(scores), 1)
print(f"{name}: {mean:.3f} (min {minimum})")
if mean < minimum:
failed.append(name)
print(px.experiments.get_experiment_url(exp["dataset_id"], exp["experiment_id"]))
sys.exit(1 if failed else 0)
# .github/workflows/llm-eval.yml (fragment)
jobs:
eval-gate:
runs-on: ubuntu-latest
services:
phoenix:
image: arizephoenix/phoenix:20.16.0 # pinned
ports: ["6006:6006", "4317:4317"]
env:
PHOENIX_COLLECTOR_ENDPOINT: http://localhost:6006
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
steps:
- uses: actions/checkout@v4
- uses: astral-sh/setup-uv@v6
- run: uv sync
- run: uv run python scripts/seed_datasets.py # upload golden set from the repo
- run: uv run python scripts/eval_gate.py "$GITHUB_SHA" "$GITHUB_REF_NAME"
- Keep the golden dataset in the repo (CSV/JSONL) and upload it in CI when Phoenix is ephemeral; use a shared instance to keep history across commits.
- Gate on deterministic evaluators first; add LLM-judge thresholds only after calibration, with a tolerance band.
- Run the smoke split on every PR, the full split nightly.
Production: Retention, Sampling, PII¶
| Concern | Setting |
|---|---|
| Retention | PHOENIX_DEFAULT_RETENTION_POLICY_DAYS=30 for new projects; per-project policies in Settings |
| Storage growth | PostgreSQL, monitor disk; separate projects for noisy sources |
| Head sampling | register(..., sampler=ParentBased(TraceIdRatioBased(0.1))) — kwargs go to the OTel TracerProvider |
| Tail sampling | OTel Collector tail_sampling: keep errors, slow traces, negative feedback, test traffic |
| Hide prompts / outputs | OPENINFERENCE_HIDE_INPUTS=true, OPENINFERENCE_HIDE_OUTPUTS=true, OPENINFERENCE_HIDE_INPUT_MESSAGES=true |
| Per-instrumentor masking | AnthropicInstrumentor().instrument(tracer_provider=tp, config=TraceConfig(hide_inputs=True)) |
| Pattern redaction (emails, cards) | Custom span processor or Collector redaction / transform processor before export |
| Access | Auth on, viewer role for most users, system keys per service |
Prompts are user data
LLM spans contain full prompts, retrieved documents and model outputs — often personal data. Decide what is captured before production traffic, not after the first incident.
Production Checklist¶
- Separate projects per environment; test runs never go to the production project
-
batch=Truein services,force_flush()/shutdown()in jobs and test sessions - Session and user ids set for conversational flows; feedback wired to span annotations
- Retention policy set; storage monitored
- Sampling strategy chosen (keep errors and negative feedback at 100%)
- PII masking verified by a test that inspects exported spans
- Production failures regularly curated into the golden dataset
- CI gate: experiment on every PR, thresholds versioned in the repo
- LLM judges calibrated against human labels, judge model pinned
See also¶
- Arize Phoenix — LLM Tracing & Evaluation
- Phoenix — Datasets & Experiments
- Phoenix — Evaluations
- Pytest — Python Testing Framework
- OpenTelemetry — Testing with OpenTelemetry
- Agentic AI — Testing, Evaluation & Observability
- Langfuse — LLM Tracing, Prompts & Evals