Arize Phoenix — LLM Tracing & Evaluation¶
Open-source (Elastic License 2.0) platform for tracing, evaluating and testing LLM applications: agents, RAG pipelines, tool calls. Self-hosted in one container, built on OpenTelemetry.
Where Phoenix Fits¶
flowchart LR
subgraph App["Python app / test suite"]
CODE["LLM code<br/>(OpenAI, Anthropic, LangChain, LiteLLM...)"] --> OI["OpenInference<br/>instrumentors"]
OI --> OTEL["OTel SDK<br/>phoenix.otel.register()"]
CLIENT["phoenix.client<br/>datasets, experiments, spans"]
EVALS["phoenix.evals<br/>LLM-as-judge, code evals"]
end
OTEL -- "OTLP HTTP :6006/v1/traces<br/>OTLP gRPC :4317" --> PX["Phoenix server<br/>UI + REST API :6006"]
CLIENT -- "REST" --> PX
EVALS -. "scores as annotations" .-> CLIENT
PX --> DB[("SQLite / PostgreSQL")]
- OpenTelemetry — the transport. Phoenix is an OTLP backend; any OTel exporter or Collector can send to it.
- OpenInference — Arize's semantic conventions on top of OTel for LLM spans: span kinds (
LLM,CHAIN,TOOL,RETRIEVER,AGENT...),input.value,llm.token_count.*,session.id. - Phoenix — stores spans, renders traces and sessions, holds datasets, experiments, prompts and annotations.
Installation¶
# Server (UI + collector) — for local runs and CI
uv add --dev arize-phoenix
# Lightweight client-side packages — for the app and the tests
uv add arize-phoenix-otel # register(): OTel + OpenInference setup
uv add arize-phoenix-client # REST client: datasets, experiments, spans, prompts, pytest plugin
uv add arize-phoenix-evals # evaluators: LLM-as-judge and code metrics
# Instrumentors — only for the libraries you use
uv add openinference-instrumentation-openai openinference-instrumentation-anthropic
uv add openinference-instrumentation-litellm openinference-instrumentation-langchain
Old API: px.Client()
import phoenix as px; px.Client() is gone in current arize-phoenix releases. Use from phoenix.client import Client (package arize-phoenix-client). The same goes for llm_classify and the old prompt templates in phoenix.evals: replaced by create_classifier and phoenix.evals.metrics in arize-phoenix-evals 3.x.
Section Map¶
| File | Topics |
|---|---|
| 01 Setup & Architecture | phoenix serve, Docker, Compose + Postgres, ports, projects, auth, env vars |
| 02 Tracing & Instrumentation | register(), OpenInference instrumentors, manual spans, sessions, users, Collector, annotations |
| 03 Datasets & Experiments | Datasets from DataFrame / CSV / traces, run_experiment, evaluators, comparisons, prompts |
| 04 Evaluations | phoenix.evals, built-in metrics, create_classifier, code evals, logging results to spans, judge calibration |
| 05 Testing, CI & Production | pytest plugin, asserting on spans, CI regression gates, retention, sampling, PII |
Minimal Setup¶
from openai import OpenAI
from phoenix.otel import register
tracer_provider = register(
project_name="ticket-triage",
endpoint="http://localhost:6006/v1/traces",
auto_instrument=True, # instruments every installed openinference-instrumentation-* package
)
client = OpenAI()
client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Route this ticket: 'I was charged twice'"}],
)
# Open http://localhost:6006 -> project "ticket-triage" -> LLM span with prompt, output, tokens
Quick Commands¶
| Command | Use |
|---|---|
uv run phoenix serve |
Local server, UI on http://localhost:6006, data in ~/.phoenix |
docker run -p 6006:6006 -p 4317:4317 arizephoenix/phoenix:latest |
Server in Docker (pin a version in CI) |
PHOENIX_WORKING_DIR=/tmp/phx uv run phoenix serve |
Throwaway instance with isolated SQLite |
curl -s http://localhost:6006/healthz |
Readiness check before tests start |
PHOENIX_TEST_TRACKING=false uv run pytest |
Run @pytest.mark.phoenix tests without recording |
Key Environment Variables¶
| Variable | Example | Read by |
|---|---|---|
PHOENIX_COLLECTOR_ENDPOINT |
http://localhost:6006 |
register() and Client() |
PHOENIX_PROJECT_NAME |
ticket-triage |
register() — default project for spans |
PHOENIX_API_KEY |
eyJhbGciOi... |
register() and Client() when auth is on |
PHOENIX_CLIENT_HEADERS |
x-team=qa |
Extra headers for OTLP / REST |
PHOENIX_PORT / PHOENIX_GRPC_PORT |
6006 / 4317 |
Server ports |
PHOENIX_SQL_DATABASE_URL |
postgresql://user:pass@db:5432/phoenix |
Server storage (default: SQLite) |
Phoenix vs Langfuse¶
| Aspect | Phoenix | Langfuse |
|---|---|---|
| Licence | Elastic License 2.0, all features self-hostable | MIT core, some enterprise features paid |
| Deployment | Single container, SQLite or Postgres | Web + worker, Postgres + ClickHouse + Redis + S3 |
| Tracing model | Native OTLP + OpenInference conventions | Own SDK (OTel-based in v3), OTLP endpoint |
| Evals | phoenix.evals library, runs in your process |
Managed LLM-as-judge in the UI + SDK scores |
| Experiments | run_experiment on datasets, pytest plugin |
Datasets + experiment runs via SDK |
| Best fit | Local debugging, notebooks, CI test runs, OTel shops | Team-wide prompt management, product analytics, cost dashboards |
Rule of thumb: Phoenix is the lighter choice for a QA engineer running it next to a test suite; Langfuse is heavier but stronger as a shared production platform.
Quick Rules¶
- Instrument with OpenInference, not ad-hoc logging — Phoenix renders LLM, tool and retriever spans only when OpenInference attributes are present.
- One project per app and environment (
triage-dev,triage-ci,triage-prod) — never mix test runs with production traffic. - Pin the server version in Docker and CI; the client packages check server capabilities and fail on old servers.
- Use
batch=Truein services, the defaultSimpleSpanProcessoronly in tests and scripts. - Tag every test run with a session id or metadata (test node id, git SHA) so traces can be found and asserted.
- Turn failures into dataset examples — a production bug becomes a regression test in the next experiment.
- Calibrate LLM judges against human labels before using their scores as CI gates.
- Set a retention policy and mask PII before pointing production traffic at Phoenix.
See also¶
- Digital Garden: Knowledge Base
- Tools — Practical Reference Guides
- Langfuse — LLM Tracing, Prompts & Evals
- MLflow — Experiment Tracking, LLM Tracing & Evaluation
- OpenTelemetry — Python Observability
- DeepEval — LLM Testing Guide
- Agentic AI — Testing, Evaluation & Observability
- Jaeger — Distributed Tracing for OpenTelemetry