Open-source LLM engineering platform: trace every LLM call, version prompts outside the code, build datasets from real traffic and score outputs with code, humans or LLM judges.
uvaddlangfuse# Python SDK v4 (OpenTelemetry-based)
uvaddopenailitellmlangchain# optional: integrations you actually useexportLANGFUSE_PUBLIC_KEY="pk-lf-..."exportLANGFUSE_SECRET_KEY="sk-lf-..."exportLANGFUSE_BASE_URL="https://cloud.langfuse.com"# or http://localhost:3000 when self-hosted
fromlangfuseimportget_client,observe,propagate_attributesfromlangfuse.openaiimportopenai# drop-in: every call becomes a generationlangfuse=get_client()@observe()deftriage(ticket:str)->str:withpropagate_attributes(user_id="qa-bot",tags=["triage","smoke"]):resp=openai.chat.completions.create(model="gpt-4o-mini",messages=[{"role":"user","content":f"Classify priority P1-P4: {ticket}"}],)returnresp.choices[0].message.contentprint(triage("Checkout returns 500 for all EU users"))langfuse.flush()# short-lived script: send buffered spans before exit
Production LLM platform: tracing, prompts, datasets, evals, dashboards
Tracing + evaluation workbench, strong in notebooks and local debugging
Instrumentation
Own SDK on top of OpenTelemetry; accepts any OTLP/HTTP spans
OpenInference conventions on OpenTelemetry
Prompt management
Built in: versions, labels, caching, playground
Available, less central
Storage
Postgres + ClickHouse + Redis + S3
Lightweight; SQLite or Postgres
Self-hosting effort
Several services (Compose / Helm)
Single container or pip install
License
MIT core, some features in Enterprise Edition
Elastic License 2.0
Rule of thumb: Phoenix for quick local debugging and eval notebooks; Langfuse when a team needs a shared, long-lived place for traces, prompts, datasets and scores.