September 29, 2026
6 min read
MLflow — Experiment Tracking, LLM Tracing & Evaluation
Open-source (Apache 2.0, Linux Foundation) platform for the ML and GenAI lifecycle: experiment tracking, a model registry, and — since MLflow 3 — tracing, evaluation and a prompt registry for LLM apps and agents. One server with a UI and REST API, backed by a SQL database and an artifact store.
What MLflow Gives a QA Engineer
A history of every test and eval run — params, metrics, tags and report files per CI run, searchable and comparable in the UI.
LLM traces — autologging for OpenAI, Anthropic, LangChain / LangGraph, LiteLLM and others; @mlflow.trace for your own code.
Evaluation — mlflow.genai.evaluate() with built-in LLM judges, custom scorers, DeepEval metrics as scorers, versioned datasets.
Release gates — prompt and model versions with aliases (@candidate, @production), promoted only when the eval run passes.
Where MLflow Fits
flowchart LR
subgraph Client["Test suite / LLM app"]
SDK["mlflow SDK<br/>runs, metrics, artifacts"]
TR["mlflow tracing<br/>autolog, @mlflow.trace"]
EV["mlflow.genai.evaluate<br/>scorers, LLM judges"]
OT["Any OTel SDK"]
end
SDK -- "REST" --> S["MLflow server :5000<br/>UI + REST API"]
TR -- "REST" --> S
EV -- "REST" --> S
OT -- "OTLP HTTP /v1/traces" --> S
S --> DB[("Backend store<br/>SQLite / PostgreSQL / MySQL")]
S --> AR[("Artifact store<br/>local dir / S3 / GCS / Azure")]
TR -. "optional OTLP export" .-> J["Jaeger / OTel Collector"]
Area
Entities
Used for in QA
Experiment tracking
Experiment → Run → params, metrics, tags, artifacts
One run per CI job or eval, compared over time
Tracing
Trace → spans, tags, metadata, assessments
Debugging LLM calls, asserting on tool calls
Evaluation
Dataset, scorer, evaluation run
Quality scores on a golden dataset per build
Prompt registry
Prompt → versions → aliases
Testing a prompt version before promoting it
Model registry
Registered model → versions → aliases, tags
Gating which model / app version goes to production
Section Map
File
Topics
01 Setup & Architecture
mlflow server, backend and artifact stores, Docker Compose + Postgres, allowed hosts, auth, env vars
02 Experiment Tracking
Experiments, runs, params, metrics, tags, artifacts, autolog, search_runs, comparing runs
03 GenAI Tracing
mlflow.<flavor>.autolog(), @mlflow.trace, span types, sessions, feedback, search, OpenTelemetry in and out
04 Evaluation & Prompts
mlflow.genai.evaluate, built-in judges, @scorer, make_judge, DeepEval scorers, datasets, prompt registry
05 Testing, CI & Model Registry
pytest + DeepEval results per CI run, trace assertions, regression gate, GitHub Actions, aliases, pitfalls
Installation
uv add mlflow # full package: SDK, server, UI, evaluation
uv add mlflow-tracing # lightweight tracing-only SDK for production services
mlflow-tracing contains only the tracing SDK with a minimal set of dependencies — enough for an app that sends traces to a remote MLflow server. The server, UI and evaluation need the full mlflow package.
Minimal Setup
uv run mlflow server --port 5000 # SQLite ./mlflow.db + ./mlartifacts, UI on http://localhost:5000
# uv add mlflow openai
import mlflow
from openai import OpenAI
mlflow . set_tracking_uri ( "http://localhost:5000" )
mlflow . set_experiment ( "support-bot" )
mlflow . openai . autolog () # every OpenAI call becomes a trace
with mlflow . start_run ( run_name = "smoke" ):
mlflow . log_param ( "model" , "gpt-4.1-mini" )
reply = OpenAI () . chat . completions . create (
model = "gpt-4.1-mini" ,
messages = [{ "role" : "user" , "content" : "What is the capital of France?" }],
)
mlflow . log_metric ( "answer_chars" , len ( reply . choices [ 0 ] . message . content ))
# UI -> Experiments -> support-bot: the run on the Runs tab, the LLM call on the Traces tab
Quick Commands
Command
Use
uv run mlflow server --port 5000
Local server with SQLite and local artifacts
uv run mlflow server --backend-store-uri postgresql://... --artifacts-destination s3://bucket --host 0.0.0.0
Shared team server
curl -sf http://localhost:5000/health
Readiness check before tests start
uv run mlflow experiments search
List experiments
uv run mlflow runs list --experiment-id 1
List runs of one experiment
uv run mlflow traces search --experiment-id 1 --max-results 10
Latest traces from the CLI
uv run mlflow scorers list -b
Built-in scorers and the data they need
uv run mlflow db upgrade <backend-store-uri>
Migrate the database schema after an MLflow upgrade
uv run mlflow doctor
Versions and environment info for bug reports
Key Environment Variables
Variable
Example
Purpose
MLFLOW_TRACKING_URI
http://mlflow:5000
Where the SDK sends runs and traces (default: local sqlite:///mlflow.db)
MLFLOW_EXPERIMENT_NAME
llm-regression
Default experiment when set_experiment() is not called
MLFLOW_EXPERIMENT_ID
5
Same, by id; if both are set, they must point to the same experiment
MLFLOW_TRACKING_USERNAME / MLFLOW_TRACKING_PASSWORD
ci-bot / ***
Basic auth against a server with --app-name basic-auth
MLFLOW_TRACKING_TOKEN
***
Bearer token when a proxy in front of MLflow expects one
MLFLOW_TRACE_SAMPLING_RATIO
0.1
Keep 10% of traces
MLFLOW_GENAI_EVAL_MAX_WORKERS
4
Parallelism of mlflow.genai.evaluate (default 10)
OTEL_EXPORTER_OTLP_TRACES_ENDPOINT
http://jaeger:4318/v1/traces
Send MLflow traces to an OTLP backend instead of MLflow
MLflow vs Phoenix, Langfuse and Jaeger
Aspect
MLflow
Phoenix
Langfuse
Jaeger
Focus
Whole ML / GenAI lifecycle
LLM tracing and evals
LLM tracing, prompts, analytics
Distributed tracing of services
Unique strength
Runs with params and metrics, model registry
Light, OTel-native, notebooks
Prompt management, team workflows
Service map, trace compare
Tracing
Own SDK + autolog, accepts and exports OTLP
OTLP + OpenInference
Own SDK (OTel-based), OTLP endpoint
Plain OTLP
Evaluation
mlflow.genai.evaluate, judges, DeepEval / RAGAS scorers
phoenix.evals, experiments
Experiments, managed LLM judges
—
Deployment
One server + SQL DB + artifact store
One container, SQLite / Postgres
Web + worker, Postgres + ClickHouse + Redis + S3
One binary, several storages
Best fit
Teams that already track ML experiments, or need runs + registry + LLM evals in one place
Local LLM debugging and CI evals
Shared LLM product platform
Microservice and API debugging
Rule of thumb: pick MLflow when you want test and eval runs as first-class, comparable records and a registry to gate releases; pick Phoenix or Langfuse when LLM observability is the only need; keep Jaeger for non-LLM service traces.
Quick Rules
Always set MLFLOW_TRACKING_URI in CI — without it every job writes to a throwaway local mlflow.db.
One experiment per suite , one run per CI job; the run name carries the pipeline id.
Tag every run with git_sha, git_branch, ci_pipeline_id — regressions are found by these tags.
Params for inputs, metrics for results — params are immutable strings; metrics are numbers with history.
Autolog once per framework (mlflow.openai.autolog()) and add @mlflow.trace only for your own code — not on top of autologged calls.
Prefer code scorers , then built-in judges, then custom judges — and pin the judge model.
Use aliases, not version numbers , for prompts and models that tests and apps load.
Pin the server version and run mlflow db upgrade on purpose when upgrading.
See also
evaluation
llm
mlflow
observability
tools