Langfuse — Testing, CI & Production¶
Three ways Langfuse shows up in a QA workflow:
- Tests as traced traffic — every pytest case produces a trace, tagged with the run id, so a failure links straight to the full LLM call tree.
- Experiments as regression gates — a dataset run with evaluators decides whether a prompt/model change may merge.
- Production monitoring — cost, latency and quality scores watched over time, with checks on the Metrics API.
pytest: One Trace per Test¶
# tests/conftest.py
import os
import uuid
import pytest
from langfuse import get_client, propagate_attributes
langfuse = get_client()
RUN_ID = os.getenv("GITHUB_RUN_ID", f"local-{uuid.uuid4().hex[:8]}")
@pytest.hookimpl(hookwrapper=True)
def pytest_runtest_makereport(item, call):
outcome = yield
rep = outcome.get_result()
setattr(item, f"rep_{rep.when}", rep) # item.rep_call.passed / .failed
@pytest.fixture(scope="session", autouse=True)
def langfuse_session():
yield langfuse
langfuse.flush() # tests exit fast — send everything
@pytest.fixture(autouse=True)
def lf_trace(request):
trace_id = langfuse.create_trace_id(seed=f"{RUN_ID}:{request.node.nodeid}") # deterministic
with langfuse.start_as_current_observation(
name=request.node.name,
trace_context={"trace_id": trace_id},
input={"test": request.node.nodeid},
):
with propagate_attributes(
session_id=RUN_ID, # whole run = one session in the UI
tags=["pytest", f"run:{RUN_ID}"],
metadata={"nodeid": request.node.nodeid[:200], "branch": os.getenv("GITHUB_REF_NAME", "local")},
):
yield trace_id
rep = getattr(request.node, "rep_call", None)
if rep is not None:
langfuse.create_score(trace_id=trace_id, name="test_passed", value=1 if rep.passed else 0, data_type="BOOLEAN")
if rep.failed:
print(f"\nLangfuse trace: {langfuse.get_trace_url(trace_id=trace_id)}")
- Everything the code under test traces (
@observe, OpenAI wrapper, LangChain handler) nests under the test's root span. session_id=RUN_IDgroups a CI run; filter by tagrun:<id>to compare two runs.- Run with
LANGFUSE_TRACING_ENVIRONMENT=ci(or a dedicated project) so test traffic stays out of production views. - Offline or fork PRs without secrets:
LANGFUSE_TRACING_ENABLED=false— the fixtures still work, nothing is sent.
Asserting on Quality¶
Compute scores in the test, assert on them, and also push them to Langfuse for trend analysis:
import json
import pytest
from langfuse import get_client
from support_bot import triage
langfuse = get_client()
CASES = [
("Checkout returns 500 for all EU users", "P1"),
("Typo on the About page", "P4"),
]
@pytest.mark.parametrize("ticket,expected_priority", CASES)
def test_triage_priority(ticket, expected_priority, lf_trace):
raw = triage(ticket)
parsed = json.loads(raw) # contract: valid JSON
langfuse.score_current_trace(name="json_valid", value=1, data_type="BOOLEAN")
match = parsed["priority"] == expected_priority
langfuse.score_current_trace(name="priority_match", value=1.0 if match else 0.0)
assert match, f"got {parsed['priority']}, want {expected_priority}"
Don't read scores back synchronously
Ingestion is asynchronous (queue → worker → ClickHouse). A score or trace written a moment ago may not be queryable yet. Assert on values you computed locally; read from the API only in dedicated checks with polling.
Trace contract test (polling the API)¶
Verifies that the application itself emits correctly attributed traces — useful after SDK upgrades.
import time
from langfuse import get_client
langfuse = get_client()
def wait_for_trace(trace_id: str, timeout: float = 30.0):
deadline = time.monotonic() + timeout
while time.monotonic() < deadline:
try:
return langfuse.api.trace.get(trace_id)
except Exception:
time.sleep(1)
raise AssertionError(f"trace {trace_id} not ingested within {timeout}s")
def test_chat_turn_trace_contract(lf_trace):
handle_chat_turn(user_id="qa-user-1", thread_id="thread-42", message="How do I reset my password?")
langfuse.flush()
trace = wait_for_trace(lf_trace)
names = {o.name for o in trace.observations}
assert trace.user_id == "qa-user-1" and trace.session_id == "thread-42"
assert {"search_kb", "answer-llm"} <= names
gen = next(o for o in trace.observations if o.name == "answer-llm")
assert gen.type == "GENERATION" and gen.model
Experiments as a CI Gate¶
Option A — plain pytest¶
# tests/llm/test_triage_regression.py
import pytest
from langfuse import get_client
from evals.triage import triage_task, priority_match, team_match, accuracy
THRESHOLD = 0.9
@pytest.mark.llm
def test_triage_dataset_regression():
langfuse = get_client()
dataset = langfuse.get_dataset("triage-regression")
result = dataset.run_experiment(
name="triage-ci",
task=triage_task,
evaluators=[priority_match, team_match],
run_evaluators=[accuracy],
max_concurrency=5,
)
langfuse.flush()
assert len(result.item_results) == len(dataset.items), "some items crashed — see logs"
acc = next(e.value for e in result.run_evaluations if e.name == "priority_accuracy")
assert acc >= THRESHOLD, f"priority_accuracy {acc:.2f} < {THRESHOLD} — {result.dataset_run_url}"
Option B — langfuse/experiment-action¶
# experiments/triage_experiment.py
from langfuse import RegressionError, RunnerContext
from evals.triage import triage_task, priority_match, accuracy
def experiment(context: RunnerContext):
result = context.run_experiment( # dataset + GitHub metadata injected by the action
name="triage-ci",
task=triage_task,
evaluators=[priority_match],
run_evaluators=[accuracy],
)
acc = next(e.value for e in result.run_evaluations if e.name == "priority_accuracy")
if acc < 0.9:
raise RegressionError(result=result, metric="priority_accuracy", value=acc, threshold=0.9)
return result
# .github/workflows/llm-regression.yml
on: pull_request
permissions: { contents: read, pull-requests: write, actions: read }
jobs:
experiment:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v6
- uses: actions/setup-python@v6
with: { python-version: "3.13" }
- uses: langfuse/experiment-action@v1.0.10 # pin to a release SHA in real pipelines
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
with:
langfuse_public_key: ${{ secrets.LANGFUSE_PUBLIC_KEY }}
langfuse_secret_key: ${{ secrets.LANGFUSE_SECRET_KEY }}
langfuse_base_url: https://cloud.langfuse.com
experiment_path: experiments/triage_experiment.py
dataset_name: triage-regression
github_token: ${{ github.token }}
The action runs the script, comments the run summary on the PR and fails the job on RegressionError (should_fail_on_regression, default true).
Gate design: threshold the aggregate score (one flaky item shouldn't block), keep a small exact-check smoke set separate from the quality set, compare with the last main run to catch slow drift, and use temperature=0, pinned models and max_concurrency limits to cut noise. Mark with @pytest.mark.llm and run only on PRs touching prompts or LLM code.
Dashboards and Metrics API¶
Built-in dashboards show cost, token usage, latency (incl. time to first token) and scores, filterable by environment, tags, model, prompt name/version, release. Custom dashboards can be built from the same dimensions.
import json
from datetime import datetime, timedelta, timezone
from langfuse import get_client
langfuse = get_client()
now = datetime.now(timezone.utc)
query = {
"view": "observations",
"dimensions": [{"field": "promptName"}, {"field": "promptVersion"}],
"metrics": [
{"measure": "totalCost", "aggregation": "sum"},
{"measure": "latency", "aggregation": "p95"},
{"measure": "count", "aggregation": "count"},
],
"filters": [
{"column": "type", "operator": "=", "value": "GENERATION", "type": "string"},
{"column": "environment", "operator": "=", "value": "production", "type": "string"},
],
"fromTimestamp": (now - timedelta(days=1)).isoformat(),
"toTimestamp": now.isoformat(),
}
for row in langfuse.api.metrics.metrics(query=json.dumps(query)).data:
print(row) # one dict per promptName/promptVersion with the requested aggregates
- Views:
observations,scores-numeric,scores-boolean,scores-categorical; aggregationssum,avg,count,min,max,p50–p99,histogram. - High-cardinality fields (
traceId,userId,sessionId) work as filters, not dimensions. langfuse.api.metricstargets the v2 endpoint of a v4 server; older self-hosted servers needlangfuse.api.legacy.metrics_v1.- A nightly job over this query (daily cost per prompt version, p95 latency, avg judge score) is a cheap production SLO check.
Retention and PII¶
| Topic | Practice |
|---|---|
| Retention | Per project, minimum 3 days; deletes traces, observations, scores, media; keeps dataset items and audit logs. Cloud Pro/Enterprise or self-hosted EE. Self-hosted OSS keeps data forever unless you clean up |
| Test projects | Short retention (e.g. 14–30 days) |
| PII in prompts/outputs | mask= / mask_otel_spans= in the SDK — masked before data leaves the process |
| Large or sensitive payloads | @observe(capture_input=False, capture_output=False) on specific functions |
| Right to erasure | Delete traces via langfuse.api.trace.delete(trace_id) / delete_multiple(trace_ids=[...]); find them by user_id |
| Secrets | Secret key only on the server and in CI secrets; rotate per consumer |
Production Checklist¶
-
LANGFUSE_BASE_URL, keys andLANGFUSE_TRACING_ENVIRONMENTset per environment -
flush()/shutdown()wired into workers, jobs and serverless handlers -
user_id,session_id,tags,version/release propagated on every request - Prompts fetched by label with
fallback; generations linked to prompts - PII masking in place and unit-tested
- Sampling rate chosen for high-volume endpoints (100% in CI)
- Online LLM-as-a-judge on a sampled subset; user feedback captured as scores
- Regression dataset grows from production failures; CI gate on experiment scores
- Cost/latency alerts or nightly Metrics API checks
- Retention configured; test traffic isolated in its own project or environment
- Self-hosted: backups for Postgres + ClickHouse, S3 lifecycle rules,
CHANGEMEsecrets replaced
See also¶
- Langfuse — LLM Tracing, Prompts & Evals
- Langfuse — Datasets & Evaluations
- Pytest — Python Testing Framework
- OpenTelemetry — Testing with OpenTelemetry
- DeepEval — LLM Testing Guide
- Arize Phoenix