Phoenix — Evaluations¶
arize-phoenix-evals is a standalone library: evaluators run in your process (notebook, pytest, CI job), and their scores go to Phoenix as annotations on spans or as experiment evaluations.
| Evaluator type | How it scores | Cost | Use for |
|---|---|---|---|
| Code (heuristic) | Python function | Free, deterministic | Format, exact match, regex, JSON schema, length, latency |
| LLM-as-judge | Classification prompt to a judge model | Tokens, non-deterministic | Correctness, faithfulness, relevance, tone, toxicity |
| Human | Annotation in the UI | People time | Ground truth, judge calibration |
API generations
arize-phoenix-evals 3.x uses LLM, create_classifier, ClassificationEvaluator, phoenix.evals.metrics and evaluate_dataframe. Older guides show llm_classify, OpenAIModel and HALLUCINATION_PROMPT_TEMPLATE — that API is no longer part of the package.
Judge Model: LLM¶
from phoenix.evals import LLM
judge = LLM(provider="anthropic", model="claude-haiku-4-5") # cheap, fast judge
strict_judge = LLM(provider="openai", model="gpt-4o-mini")
litellm_judge = LLM(provider="litellm", model="anthropic/claude-sonnet-5") # needs litellm installed
- Native providers:
openai,azure,anthropic,google; with LiteLLM installed alsolitellm,bedrock,vertex. - API keys come from the usual env vars (
ANTHROPIC_API_KEY,OPENAI_API_KEY). - The judge must support tool calling or structured output — the label is extracted from a structured response, not parsed from free text.
- Extra kwargs on evaluators go to the model call:
CorrectnessEvaluator(llm=judge, temperature=0.0).
Built-in Metrics¶
from phoenix.evals.metrics import CorrectnessEvaluator, FaithfulnessEvaluator
faithfulness = FaithfulnessEvaluator(llm=judge)
scores = faithfulness.evaluate({
"input": "How long is the refund window?",
"context": "Refunds are accepted within 30 days of delivery. Electronics: 14 days.",
"output": "You can get a refund within 90 days.",
})
print(scores[0].label, scores[0].score, scores[0].explanation)
# unfaithful 0.0 The context states 30 days, not 90.
| Evaluator | Inputs | Labels (score) |
|---|---|---|
CorrectnessEvaluator |
input, output |
correct (1) / incorrect (0) |
FaithfulnessEvaluator |
input, output, context |
faithful (1) / unfaithful (0) — RAG groundedness |
HallucinationEvaluator |
input (conversation), output |
hallucinated (1) / grounded (0) — direction: minimize |
DocumentRelevanceEvaluator |
input, document_text |
relevant (1) / unrelated (0) — per retrieved doc |
RetrievalRelevanceEvaluator |
input, context |
Retrieved context vs request |
ToxicityEvaluator |
text |
toxic (1) / non-toxic (0) — minimize |
RefusalEvaluator |
input, output |
refused / answered |
ConcisenessEvaluator, CompletenessEvaluator |
input, output / conversation |
Style and coverage |
ToolSelectionEvaluator, ToolInvocationEvaluator |
input, available_tools, tool_selection |
Agent picked / called the right tool correctly |
ToolResponseHandlingEvaluator |
input, tool_call, tool_result, output |
Agent used the tool result correctly |
PiiDetectionEvaluator, UserFrictionEvaluator |
conversation (+ user_message) |
PII leaks, frustrated users |
exact_match, MatchesRegex(pattern=), PrecisionRecallFScore() |
code metrics | Deterministic checks |
Each call returns a list of Score(name, score, label, explanation, metadata, kind, direction).
Custom LLM Judge: create_classifier¶
from phoenix.evals import create_classifier
TRIAGE_JUDGE = """You are reviewing a support-ticket router.
Ticket: {ticket}
Chosen queue: {queue}
Queues: billing (payments, invoices), bugs (crashes, errors), how-to (usage questions),
security (account takeover, prompt injection, fraud).
Is the chosen queue the best fit for the ticket?"""
routing_quality = create_classifier(
name="routing_quality",
prompt_template=TRIAGE_JUDGE,
llm=judge,
choices={"correct": 1.0, "acceptable": 0.5, "wrong": 0.0},
)
result = routing_quality.evaluate({"ticket": "Someone changed my password", "queue": "how-to"})
assert result[0].label == "wrong"
- Keep choices few and mutually exclusive; describe each label in the prompt.
choicesmay be a list (labels only), a dict label→score, or label→(score, description).- Explanations are on by default — read them when calibrating.
Code Evaluators: create_evaluator¶
import json
from phoenix.evals import Score, create_evaluator
@create_evaluator(name="valid_json", kind="code")
def valid_json(output: str) -> Score:
try:
payload = json.loads(output)
except json.JSONDecodeError as exc:
return Score(score=0.0, label="invalid", explanation=str(exc))
missing = {"queue", "priority"} - payload.keys()
return Score(score=float(not missing), label="valid" if not missing else "incomplete")
@create_evaluator(name="within_budget", kind="code", direction="maximize")
def within_budget(output: str) -> bool:
return len(output) <= 400
Return a Score, bool, number or string. Code evaluators run in the same pipelines as LLM judges.
Mapping Columns: bind_evaluator¶
Evaluators expect fixed input names; map your fields with bind_evaluator (JSONPath-like strings or lambdas):
from phoenix.evals import bind_evaluator
correctness = bind_evaluator(
evaluator=CorrectnessEvaluator(llm=judge),
input_mapping={
"input": "attributes.input.value", # column in a spans DataFrame
"output": "attributes.output.value",
},
)
# Inside experiments: input / output / expected are dicts
exp_correctness = bind_evaluator(
evaluator=CorrectnessEvaluator(llm=judge),
input_mapping={"input": lambda x: x["input"]["question"], "output": lambda x: x["output"]["answer"]},
)
Evaluating Traces and Logging Results Back¶
from phoenix.client import Client
from phoenix.client.types.spans import SpanQuery
from phoenix.evals import evaluate_dataframe
from phoenix.evals.utils import to_annotation_dataframe
px = Client()
spans = px.spans.get_spans_dataframe(
query=SpanQuery().where("span_kind == 'CHAIN' and name == 'rag_answer'"),
project_identifier="docs-rag-staging",
limit=500,
) # index: context.span_id
results = evaluate_dataframe(
dataframe=spans,
evaluators=[correctness, valid_json],
exit_on_error=False,
max_retries=3,
) # adds <name>_score and <name>_execution_details
px.spans.log_span_annotations_dataframe(dataframe=to_annotation_dataframe(dataframe=results))
Scores appear on each span in the UI (filterable: annotations['correctness'].label == 'incorrect') and in project-level aggregates. Use async_evaluate_dataframe for large batches — it runs judge calls concurrently with rate limiting.
For RAG, evaluate documents too: explode attributes.retrieval.documents from RETRIEVER spans, run DocumentRelevanceEvaluator, log with log_document_annotations_dataframe — the UI then shows relevance per retrieved chunk.
Choosing Judge Models¶
| Guideline | Why |
|---|---|
| Judge ≠ model under test (or at least a different size) | Self-preference bias inflates scores |
Cheap judge (claude-haiku-4-5, gpt-4o-mini) for high-volume binary labels |
Cost scales with dataset × evaluators × repetitions |
Stronger judge (anthropic/claude-sonnet-5) for nuanced rubrics |
Better agreement with humans on subtle cases |
temperature=0.0, pinned model version |
Reproducible gates |
| Binary or 3-level labels, not 1–10 scales | Judges are far more consistent on classification |
Judge Calibration¶
A judge is a classifier — measure it like one before trusting it in CI.
- Take 50–200 real outputs; have humans label them (Phoenix annotation UI,
annotator_kind="HUMAN"). - Run the judge on the same rows.
- Compare judge vs human labels: precision, recall, Cohen's kappa for the failure label.
- Read disagreements and their explanations; fix the rubric, add label descriptions or few-shot examples.
- Re-run; freeze the prompt and judge model once agreement is acceptable (e.g. recall ≥ 0.9 on failures).
- Re-calibrate on every judge model or prompt change, and periodically on fresh production data.
from sklearn.metrics import cohen_kappa_score, classification_report
human = labels_df["human_label"] # from UI annotations export
judge_labels = labels_df["judge_label"]
print(classification_report(human, judge_labels))
print("kappa:", cohen_kappa_score(human, judge_labels))
download_benchmark_dataset(...) fetches Arize's public labelled benchmarks for sanity-checking built-in evaluators with your judge model.
See also¶
- Arize Phoenix — LLM Tracing & Evaluation
- Phoenix — Datasets & Experiments
- Phoenix — Testing, CI & Production
- DeepEval — LLM Testing Guide
- Agentic AI — Testing, Evaluation & Observability