Langfuse — Tracing with the Python SDK¶
The current Python SDK (langfuse v4) is built on OpenTelemetry: every Langfuse observation is an OTel span, exported over OTLP/HTTP to Langfuse. Context propagates automatically through nested calls, threads with context copying, and async code.
SDK Generations¶
| SDK | API style | Status |
|---|---|---|
| v2 | langfuse.trace(), trace.span(), trace.generation() — manual object tree |
Legacy; ignore old tutorials |
| v3 | OTel-based, start_as_current_span(), update_current_trace() |
Superseded |
| v4 | OTel-based, start_as_current_observation(as_type=...), propagate_attributes() |
Current |
v3 → v4 changes worth knowing: start_span/start_generation became start_observation(as_type=...); update_current_trace() was split into propagate_attributes() (user, session, tags, metadata), set_current_trace_io() (deprecated) and set_current_trace_as_public(); only LLM-related OTel spans are exported by default.
Client¶
from langfuse import get_client
langfuse = get_client() # singleton configured from LANGFUSE_* env vars
langfuse.auth_check() # True if keys + base URL are valid (network call — not per request)
get_client()returns the same instance everywhere — no need to pass it around.- Create
Langfuse(...)explicitly only for non-default settings (masking, sampling, custom exporter);get_client()returns it afterwards.
@observe Decorator¶
from langfuse import observe, get_client
langfuse = get_client()
@observe(as_type="retriever")
def search_kb(query: str) -> list[str]:
return ["Reset password: Settings → Security", "SSO users: contact IT"]
@observe(as_type="generation", name="answer-llm")
def answer(query: str, docs: list[str]) -> str:
text = call_model(query, docs) # your LLM client
langfuse.update_current_generation(
model="gpt-4o-mini",
usage_details={"input": 412, "output": 58},
metadata={"docs": len(docs)},
)
return text
@observe() # root span → becomes the trace
def rag_pipeline(query: str) -> str:
docs = search_kb(query)
return answer(query, docs)
| Argument | Effect |
|---|---|
name |
Observation name (default: function name) |
as_type |
span (default), generation, embedding, agent, tool, chain, retriever, evaluator, guardrail |
capture_input / capture_output |
Turn off automatic argument/return capture (large or sensitive payloads) |
transform_to_string |
Join streamed generator chunks into one output string |
- Works on sync and
asyncfunctions and generators. - Exceptions are recorded with level
ERRORand re-raised. - Global switch for I/O capture:
LANGFUSE_OBSERVE_DECORATOR_IO_CAPTURE_ENABLED=false.
Context Managers¶
Use them when the unit of work is not a function, or when you need the span object.
from langfuse import get_client
langfuse = get_client()
with langfuse.start_as_current_observation(as_type="span", name="triage-ticket", input={"ticket_id": "T-981"}) as root:
with langfuse.start_as_current_observation(
as_type="generation",
name="classify-priority",
model="gpt-4o-mini",
model_parameters={"temperature": 0},
input=[{"role": "user", "content": "Checkout is down for all EU users"}],
) as gen:
label = "P1"
gen.update(output=label, usage_details={"input": 35, "output": 2})
langfuse.create_event(name="routing-rule-applied", metadata={"queue": "oncall"})
root.update(output={"priority": label, "queue": "oncall"})
| Method | Behavior |
|---|---|
start_as_current_observation(...) |
Context manager; becomes the active parent; ends on exit |
start_observation(...) |
Manual object; not set as current; call .end() yourself |
span.start_as_current_observation(...) |
Child of a specific span |
obs.update(...) |
Set output, metadata, level, status_message, model, usage_details, cost_details, prompt |
update_current_span() / update_current_generation() |
Update the active observation without holding a reference |
create_event(name=...) |
Zero-duration observation |
get_current_trace_id() / get_trace_url() |
IDs and deep links for logs and test reports |
Trace Attributes: propagate_attributes¶
User, session, tags and metadata are propagated to every span created inside the context — so filters and aggregations work on all observations.
from langfuse import get_client, observe, propagate_attributes
langfuse = get_client()
@observe()
def handle_chat_turn(user_id: str, thread_id: str, message: str) -> str:
with propagate_attributes(
user_id=user_id,
session_id=thread_id,
tags=["support-bot", "eu"],
metadata={"tenant": "acme", "flow": "refund"}, # dict[str, str], values ≤ 200 chars
version="bot-2.3.1",
trace_name="support-chat-turn",
):
return rag_pipeline(message)
- Enter it as early as possible — spans created before the context don't get the attributes.
user_id/session_idlonger than 200 chars are dropped with a warning.environment=can be set per request; process-wide useLANGFUSE_TRACING_ENVIRONMENT.as_baggage=Truealso puts attributes into OTel baggage so they cross service boundaries — they then travel in HTTP headers, so never use it for sensitive values.
Integrations¶
| Integration | How | Result |
|---|---|---|
| OpenAI SDK | from langfuse.openai import openai |
Each call = generation with model, usage, cost; streaming supported |
| LangChain / LangGraph | from langfuse.langchain import CallbackHandler |
Chain/agent/tool/LLM tree |
| LiteLLM | litellm.callbacks = ["langfuse_otel"] |
Generations for every provider behind LiteLLM |
| Any OTel instrumentation | OTLP/HTTP to /api/public/otel |
Spans with gen_ai.* attributes mapped to generations |
# OpenAI drop-in: extra kwargs name/metadata/langfuse_prompt are consumed by Langfuse
from langfuse.openai import openai
resp = openai.chat.completions.create(
name="summarize-bug-report",
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Summarize: login fails after password reset"}],
metadata={"suite": "smoke"},
)
# LangChain: pass the handler in config
from langchain_openai import ChatOpenAI
from langfuse.langchain import CallbackHandler
handler = CallbackHandler()
llm = ChatOpenAI(model="gpt-4o-mini")
llm.invoke("Write 3 negative test cases for a login form", config={"callbacks": [handler]})
# LiteLLM: route to any provider, trace in Langfuse via OTel callback
import litellm
litellm.callbacks = ["langfuse_otel"]
litellm.completion(
model="anthropic/claude-sonnet-5",
messages=[{"role": "user", "content": "Is this stack trace a flaky test or a real bug?"}],
metadata={"generation_name": "flaky-classifier", "session_id": "run-2026-09-27", "tags": ["ci"]},
)
# Raw OTel exporters (other languages, collectors): OTLP over HTTP only, no gRPC
AUTH=$(echo -n "$LANGFUSE_PUBLIC_KEY:$LANGFUSE_SECRET_KEY" | base64 -w 0)
export OTEL_EXPORTER_OTLP_ENDPOINT="$LANGFUSE_BASE_URL/api/public/otel"
export OTEL_EXPORTER_OTLP_HEADERS="Authorization=Basic $AUTH,x-langfuse-ingestion-version=4"
v4 exports only LLM spans by default
Spans from HTTP/DB auto-instrumentation are filtered out unless they come from the Langfuse SDK, carry gen_ai.* attributes, or come from a known LLM instrumentor. Restore "export everything" with Langfuse(should_export_span=lambda span: True) — or compose with langfuse.span_filter.is_default_export_span.
Flushing in Short-Lived Processes¶
Spans are batched in a background thread. A script, pytest session, Celery task or Lambda can exit before the batch is sent.
| Situation | Call |
|---|---|
| End of a script / CLI command | langfuse.flush() |
| Serverless handler | langfuse.flush() before returning |
| Process shutdown (long-running service) | langfuse.shutdown() — flushes and stops threads (also registered atexit) |
| pytest | flush() in a session-scoped fixture teardown |
Sampling¶
Langfuse(sample_rate=0.2) # or LANGFUSE_SAMPLE_RATE=0.2
- Sampling is per trace: a trace is kept or dropped with all its observations.
- Keep
1.0in tests and CI; sample only high-volume production traffic. - Scores for dropped traces have nothing to attach to — sample after deciding what you evaluate online.
Masking¶
import re
from typing import Any
from langfuse import Langfuse
EMAIL = re.compile(r"[\w.+-]+@[\w-]+\.[\w.]+")
CARD = re.compile(r"\b(?:\d[ -]?){13,19}\b")
def mask_pii(*, data: Any, **kwargs: Any) -> Any:
if isinstance(data, str):
return CARD.sub("[CARD]", EMAIL.sub("[EMAIL]", data))
if isinstance(data, dict):
return {k: mask_pii(data=v) for k, v in data.items()}
if isinstance(data, list):
return [mask_pii(data=v) for v in data]
return data
langfuse = Langfuse(mask=mask_pii)
maskruns on inputs, outputs and metadata set through the Langfuse SDK and its integrations.- It does not see raw attributes of third-party OTel instrumentations — for those, use
mask_otel_spans, which patches spans at export time (MaskOtelSpansParams→MaskOtelSpansResultwithOtelSpanPatch). - Unit-test the mask function itself: it is a data-leak control, not a nice-to-have.
See also¶
- Langfuse — LLM Tracing, Prompts & Evals
- Langfuse — Prompt Management
- OpenTelemetry — Tracing
- LiteLLM — One API for 100+ LLM Providers
- LangChain — LLM Application Framework
- Arize Phoenix
- LangGraph — Observability & Deployment