Langfuse — Datasets & Evaluations¶
Datasets are the LLM equivalent of a regression suite: fixed inputs with expected outputs. Experiments run your application over a dataset, evaluators turn outputs into scores, and Langfuse stores every run so versions can be compared.
Evaluation Loop¶
flowchart LR
PROD["Production traces"] -- "bad answer / low score" --> DS["Dataset<br/>input + expected_output"]
MAN["Hand-written cases"] --> DS
DS --> EXP["Experiment run<br/>task = your app"]
EXP --> EV["Evaluators<br/>code / LLM judge"]
EV --> SC["Scores per item + run"]
SC --> CMP["Compare runs in UI<br/>gate in CI"]
PROD --> ON["Online evals<br/>LLM judge, feedback, annotation"]
ON --> SC
| Mode | Data | Evaluated by | When |
|---|---|---|---|
| Offline (experiments) | Dataset items with ground truth | SDK evaluators, LLM-as-a-judge on experiments | Before merge / release |
| Online | Live production observations | LLM-as-a-judge rules, user feedback, annotation queues | Continuously |
Datasets¶
from langfuse import get_client
langfuse = get_client()
langfuse.create_dataset(
name="triage-regression",
description="Support ticket triage: priority + routing",
metadata={"owner": "qa", "component": "support-bot"},
)
CASES = [
("Checkout returns 500 for all EU users", {"priority": "P1", "team": "payments"}),
("Typo on the About page", {"priority": "P4", "team": "content"}),
("Password reset email never arrives", {"priority": "P2", "team": "identity"}),
]
for ticket, expected in CASES:
langfuse.create_dataset_item(
dataset_name="triage-regression",
input={"ticket": ticket},
expected_output=expected,
metadata={"source": "manual", "category": "smoke"},
)
| Field | Use |
|---|---|
input |
Whatever your task function needs (dict recommended) |
expected_output |
Ground truth for evaluators; optional for judge-only checks |
metadata |
Category, difficulty, source — group results by it |
source_trace_id / source_observation_id |
Link item back to the production trace it came from |
id |
Pass your own stable id to upsert instead of duplicating |
status |
ARCHIVED hides an item from new runs without deleting history |
create_dataset(..., input_schema=..., expected_output_schema=...)enforces a JSON schema on new items.- Datasets are versioned by time:
get_dataset(name, version=<datetime>)returns the items as they were — reproducible runs.
From production traces¶
In the UI: open a trace or observation → Add to dataset (input/output are prefilled, edit the expected output). In code, e.g. every trace with negative user feedback:
from datetime import datetime, timedelta, timezone
since = datetime.now(timezone.utc) - timedelta(days=1)
bad = langfuse.api.scores.get_many(name="user_feedback", operator="=", value=0, from_timestamp=since, limit=50)
for score in bad.data:
trace = langfuse.api.trace.get(score.trace_id)
langfuse.create_dataset_item(
dataset_name="triage-regression",
id=f"prod-{score.trace_id}", # idempotent re-runs
input=trace.input,
expected_output=None, # fill in during review
source_trace_id=score.trace_id,
metadata={"source": "negative-feedback"},
)
Experiments via SDK¶
import json
from langfuse import Evaluation, get_client
from langfuse.openai import openai
langfuse = get_client()
def triage_task(*, item, **kwargs) -> dict:
prompt = langfuse.get_prompt("ticket-triage", type="chat", label="staging", cache_ttl_seconds=0)
resp = openai.chat.completions.create(
model=prompt.config["model"],
messages=prompt.compile(product="WebShop", ticket=item.input["ticket"], history=[]),
response_format={"type": "json_object"},
langfuse_prompt=prompt,
)
return json.loads(resp.choices[0].message.content)
def priority_match(*, input, output, expected_output, metadata, **kwargs) -> Evaluation:
ok = output.get("priority") == expected_output["priority"]
return Evaluation(name="priority_match", value=1.0 if ok else 0.0,
comment=f"got {output.get('priority')}, want {expected_output['priority']}")
def team_match(*, input, output, expected_output, **kwargs) -> Evaluation:
return Evaluation(name="team_match", value=output.get("team") == expected_output["team"], data_type="BOOLEAN")
def accuracy(*, item_results, **kwargs) -> Evaluation:
vals = [e.value for r in item_results for e in r.evaluations if e.name == "priority_match"]
return Evaluation(name="priority_accuracy", value=sum(vals) / max(len(vals), 1))
dataset = langfuse.get_dataset("triage-regression")
result = dataset.run_experiment(
name="triage-prompt-v4",
run_name="v4-gpt-4o-mini-2026-09-27",
description="New routing schema",
task=triage_task,
evaluators=[priority_match, team_match],
run_evaluators=[accuracy],
max_concurrency=5, # respect provider rate limits
metadata={"prompt_label": "staging", "model": "gpt-4o-mini"},
)
print(result.format())
print(result.dataset_run_url)
| Piece | Signature / notes |
|---|---|
| Task | task(*, item, **kwargs) — sync or async; item is a DatasetItem (.input, .expected_output) or a dict for local data |
| Item evaluator | (*, input, output, expected_output, metadata, **kwargs) -> Evaluation | list[Evaluation] |
| Run evaluator | (*, item_results, **kwargs) -> Evaluation — aggregates (accuracy, pass rate) |
Evaluation |
name, value (number, string, bool), comment, metadata, data_type |
| Result | item_results (output, evaluations, trace_id), run_evaluations, dataset_run_url, format() |
- Each item runs in its own trace; the dataset run links items ↔ traces ↔ scores.
langfuse.run_experiment(name=..., data=[{"input": ..., "expected_output": ...}], task=...)runs over local data — handy in unit tests; results are traced but not stored as a dataset run.- Task exceptions are logged and the item is skipped — check
len(result.item_results)against the dataset size. - v3's
for item in dataset.items: with item.run(...)loop is gone in v4.
Scores¶
| Data type | value |
Example |
|---|---|---|
NUMERIC |
float | faithfulness=0.83, latency_budget_ratio=1.2 |
CATEGORICAL |
string | priority_error="over-escalated" |
BOOLEAN |
1 / 0 |
json_valid=1, user_feedback=0 |
TEXT |
string | Free-text reviewer note |
# Inside an active trace
langfuse.score_current_trace(name="json_valid", value=1, data_type="BOOLEAN")
langfuse.score_current_span(name="retrieval_hit", value=1, data_type="BOOLEAN")
# Later, by id (async job, feedback endpoint, CI)
langfuse.create_score(
trace_id=trace_id,
observation_id=generation_id, # optional: score one step
name="priority_error",
value="over-escalated",
data_type="CATEGORICAL",
comment="P1 for a cosmetic typo",
config_id="cfg-priority-error", # optional: validate against a score config
)
- Score configs (Project settings → Scores) define name, type, range or categories;
config_idrejects out-of-range values — use them so humans, judges and code write comparable scores. - Scores can target a trace, observation, session (
session_id=) or dataset run (dataset_run_id=). - Score
sourceisAPI,EVAL(LLM judge) orANNOTATION(human) — filterable.
LLM-as-a-Judge (UI)¶
- Add an LLM connection (OpenAI, Anthropic, Azure, Bedrock, …) in project settings.
- Evaluators → New: start from a template (hallucination, relevance, toxicity, correctness…) or write your own judge prompt with
{{variables}}. - Choose score type (numeric 0–1, categorical, boolean) and the judge model, e.g.
claude-haiku-4-5. - Choose the target:
- Observations — recommended for production: filter by observation type/name plus trace filters (tags, user, version), set a sampling rate.
- Experiments — runs automatically on new dataset runs; can map
expected_output. - Traces — legacy target, being phased out.
- Map variables to
input,output,metadataof the matched observation; preview on real data before saving.
Tip
Judges cost money and drift too. Sample production (5–10%), pin the judge model, and keep a small human-labelled dataset to measure judge agreement.
Annotation Queues¶
Human review workflow for traces, observations or sessions:
- Create a queue with one or more score configs (e.g.
correctness,priority_error). - Add items from table filters in the UI, or via API.
- Reviewers work through the queue; scores get
source=ANNOTATION.
queue = langfuse.api.annotation_queues.create_queue(name="triage-weekly-review", score_config_ids=["cfg-correctness"])
langfuse.api.annotation_queues.create_queue_item(queue.id, object_id=trace_id, object_type="TRACE")
User Feedback¶
# FastAPI endpoint receiving thumbs up/down from the chat UI
@app.post("/feedback")
def feedback(body: FeedbackIn):
langfuse.create_score(
trace_id=body.trace_id, # returned to the frontend with the answer
name="user_feedback",
value=1 if body.thumbs_up else 0,
data_type="BOOLEAN",
comment=body.comment,
)
return {"ok": True}
Return langfuse.get_current_trace_id() with every answer so the UI can send feedback for the right trace.
See also¶
- Langfuse — LLM Tracing, Prompts & Evals
- Langfuse — Prompt Management
- Langfuse — Testing, CI & Production
- DeepEval — LLM Testing Guide
- Arize Phoenix