Skip to content
  • 5 min read

Phoenix — Tracing & Instrumentation

Three layers, usually combined: register() configures OpenTelemetry, OpenInference instrumentors trace SDK calls automatically, manual spans wrap your own chains, tools and retrievers.

register()

from phoenix.otel import register

tracer_provider = register(
    project_name="ticket-triage-dev",            # or PHOENIX_PROJECT_NAME
    endpoint="http://localhost:6006/v1/traces",  # or PHOENIX_COLLECTOR_ENDPOINT
    batch=True,                                  # BatchSpanProcessor (default: Simple)
    auto_instrument=True,                        # enable all installed openinference instrumentors
)
tracer = tracer_provider.get_tracer(__name__)
Parameter Default Notes
endpoint env / localhost .../v1/traces → HTTP; host:4317 → gRPC
protocol inferred "http/protobuf" or "grpc"; set it when the endpoint is only a base URL
project_name PHOENIX_PROJECT_NAME or default Stored as openinference.project.name resource attribute
batch False True for services; False flushes each span immediately (tests, scripts)
set_global_tracer_provider True False when the app already owns a global provider
headers, api_key env Auth for a secured Phoenix
auto_instrument False Calls .instrument() on every installed openinference-instrumentation-*

register() returns a regular OTel TracerProvider subclass: add extra span processors (e.g. a second exporter) with tracer_provider.add_span_processor(...), and call tracer_provider.shutdown() at the end of short-lived jobs.

OpenInference Instrumentors

Library Package Instrumentor
OpenAI SDK openinference-instrumentation-openai openinference.instrumentation.openai.OpenAIInstrumentor
Anthropic SDK openinference-instrumentation-anthropic openinference.instrumentation.anthropic.AnthropicInstrumentor
LiteLLM openinference-instrumentation-litellm openinference.instrumentation.litellm.LiteLLMInstrumentor
LangChain / LangGraph openinference-instrumentation-langchain openinference.instrumentation.langchain.LangChainInstrumentor
LlamaIndex openinference-instrumentation-llama-index openinference.instrumentation.llama_index.LlamaIndexInstrumentor

Other instrumentors exist for Bedrock, Google GenAI, DSPy, CrewAI, OpenAI Agents SDK, MCP and more — same pattern.

from openinference.instrumentation.anthropic import AnthropicInstrumentor
from openinference.instrumentation.litellm import LiteLLMInstrumentor

# Explicit instead of auto_instrument=True — predictable in tests
AnthropicInstrumentor().instrument(tracer_provider=tracer_provider)
LiteLLMInstrumentor().instrument(tracer_provider=tracer_provider)

import litellm
litellm.completion(
    model="anthropic/claude-sonnet-5",
    messages=[{"role": "user", "content": "Summarise ticket T-812 in one line"}],
)
# -> LLM span: llm.model_name, llm.input_messages.*, llm.output_messages.*, llm.token_count.*

Instrument once

Instrument each library once per process. Double instrumentation (e.g. LiteLLM + the OpenAI SDK it calls) produces nested duplicate LLM spans — pick the outermost layer.

Span Kinds

Kind Represents Key attributes
LLM One model call llm.model_name, llm.provider, llm.input_messages, llm.token_count.total
CHAIN Glue logic, a request handler, a pipeline step input.value, output.value
TOOL A function the model called tool.name, tool.parameters
RETRIEVER Vector / keyword search retrieval.documents.N.document.content, .document.score
AGENT An agent loop (reasoning + tools) input.value, output.value
EMBEDDING Embedding call embedding.model_name, embedding.embeddings
RERANKER, GUARDRAIL, EVALUATOR, PROMPT Reranking, safety checks, eval runs, prompt rendering kind-specific

Manual Spans: Decorators

from phoenix.otel import register

tracer = register(project_name="ticket-triage-dev").get_tracer("triage")


@tracer.tool
def lookup_customer(email: str) -> dict:
    """Fetch the customer plan from CRM."""
    return {"email": email, "plan": "pro", "open_tickets": 2}


@tracer.chain
def classify(ticket: str) -> str:
    customer = lookup_customer("ann@example.com")
    return "billing" if "charged" in ticket.lower() else "bugs"


@tracer.agent
def triage_agent(ticket: str) -> dict:
    queue = classify(ticket)
    return {"queue": queue, "priority": "high"}

Decorators record arguments as input.value, return values as output.value, exceptions as span status ERROR, and tool signatures as tool.parameters (JSON schema). @tracer.llm and @tracer.retriever exist for wrapping custom model and search clients.

Manual Spans: Context Managers

from opentelemetry.trace import Status, StatusCode
from openinference.semconv.trace import DocumentAttributes, SpanAttributes


def retrieve(query: str, k: int = 3) -> list[dict]:
    with tracer.start_as_current_span("search_kb", openinference_span_kind="retriever") as span:
        span.set_input(query)
        docs = kb.search(query, k=k)                     # your vector store
        for i, doc in enumerate(docs):
            prefix = f"{SpanAttributes.RETRIEVAL_DOCUMENTS}.{i}"
            span.set_attribute(f"{prefix}.{DocumentAttributes.DOCUMENT_ID}", doc["id"])
            span.set_attribute(f"{prefix}.{DocumentAttributes.DOCUMENT_CONTENT}", doc["text"])
            span.set_attribute(f"{prefix}.{DocumentAttributes.DOCUMENT_SCORE}", doc["score"])
        span.set_status(Status(StatusCode.OK))
        return docs


def answer(question: str) -> str:
    with tracer.start_as_current_span("rag_answer", openinference_span_kind="chain") as span:
        span.set_input(question)
        docs = retrieve(question)
        reply = llm_answer(question, docs)                # instrumented SDK call -> LLM child span
        span.set_output(reply)
        return reply

Documents on RETRIEVER spans appear in the UI as a ranked list and can be scored with document-level evals (retrieval relevance).

Sessions, Users, Metadata, Tags

from openinference.instrumentation import using_attributes, using_session, using_user


def handle_chat_turn(conversation_id: str, user_id: str, message: str) -> str:
    with using_attributes(
        session_id=conversation_id,              # groups turns into a session (chat view in UI)
        user_id=user_id,
        metadata={"app_version": "2.3.1", "channel": "web"},
        tags=["triage", "beta"],
    ):
        return triage_agent(message)
Context manager Sets on every span inside Use
using_session(session_id) session.id Multi-turn conversations, one pytest test = one session
using_user(user_id) user.id Per-user debugging, feedback
using_metadata(dict) metadata (JSON) App version, experiment arm, test id
using_tags(list) tag.tags Coarse filters: smoke, nightly, canary
using_prompt_template(template=, version=, variables=) llm.prompt_template.* Link LLM spans to a prompt version
using_attributes(...) All of the above One call
suppress_tracing() Nothing is recorded Health checks, warm-up calls, judge calls you do not want in the app project

All of them are also importable from phoenix.otel. They work through OTel context, so they cross async boundaries but not threads started without context propagation.

Sending via an Existing OTel Collector

Phoenix is a plain OTLP backend — route LLM spans through the Collector you already run:

# otel-collector.yaml
exporters:
  otlphttp/phoenix:
    endpoint: http://phoenix:6006          # the exporter appends /v1/traces
    headers:
      authorization: "Bearer ${env:PHOENIX_API_KEY}"
processors:
  filter/llm-only:                          # keep only OpenInference spans
    traces:
      span:
        - 'attributes["openinference.span.kind"] == nil'
service:
  pipelines:
    traces/phoenix:
      receivers: [otlp]
      processors: [filter/llm-only, batch]
      exporters: [otlphttp/phoenix]
  • The app keeps its standard OTLP exporter pointed at the Collector; add OpenInference instrumentors to the same provider.
  • Set the project via resource attribute: OTEL_RESOURCE_ATTRIBUTES=openinference.project.name=triage-prod.
  • Non-OpenInference spans (HTTP, DB) are accepted too, shown as UNKNOWN kind — filter them if they add noise.

Annotations and Feedback

Annotations are labels/scores attached to spans, traces or sessions — from humans (UI or thumbs up/down), code checks, or LLM judges.

from opentelemetry import trace
from phoenix.client import Client

px = Client()


def triage_endpoint(ticket: str) -> dict:
    result = triage_agent(ticket)
    span_id = format(trace.get_current_span().get_span_context().span_id, "016x")
    return {**result, "span_id": span_id}          # return it to the frontend


def on_user_feedback(span_id: str, helpful: bool, comment: str | None) -> None:
    px.spans.add_span_annotation(
        span_id=span_id,
        annotation_name="user_feedback",
        annotator_kind="HUMAN",                    # HUMAN | LLM | CODE
        label="helpful" if helpful else "not_helpful",
        score=1.0 if helpful else 0.0,
        explanation=comment,
    )

Also available: px.traces.add_trace_annotation(...), px.sessions.add_session_annotation(...), px.spans.add_span_note(span_id=..., note=...), and bulk log_span_annotations_dataframe(...) for eval results (04 Evaluations).


See also