MLflow — Experiment Tracking¶
Experiment tracking is the original core of MLflow: every execution is a run with its inputs (params), results (metrics), labels (tags) and files (artifacts). For a QA engineer a run is a test or eval execution — one CI job, one benchmark, one prompt variant — and the experiment is the history of all of them.
Data Model¶
flowchart TD
E["Experiment<br/>llm-regression"] --> R1["Run ci-1234<br/>status FINISHED"]
E --> R2["Run ci-1235"]
R1 --> P["Params<br/>model=gpt-4.1-mini<br/>prompt_version=v3"]
R1 --> M["Metrics<br/>pass_rate=0.97<br/>faithfulness_mean=0.91"]
R1 --> T["Tags<br/>git_sha, git_branch,<br/>ci_pipeline_id"]
R1 --> A["Artifacts<br/>report.xml, eval_results.json"]
| Entity | Type | Mutability | Use for |
|---|---|---|---|
| Experiment | Named container | Name can be changed | One test suite or one app under evaluation |
| Run | One execution | Ends as FINISHED, FAILED or KILLED |
One CI job, one eval, one benchmark |
| Param | str key → str value |
Write-once: a different value for the same key raises INVALID_PARAMETER_VALUE |
Inputs: model, prompt version, dataset, config |
| Metric | str key → float, with step and timestamp |
Append-only history | Results: pass rate, latency, scores |
| Tag | str key → str value |
Overwritable | Search keys: git SHA, branch, suite, owner |
| Artifact | File or directory | Upload-only | Reports, logs, tables, models |
Logging a Run¶
import mlflow
mlflow.set_tracking_uri("http://localhost:5000") # or MLFLOW_TRACKING_URI
mlflow.set_experiment("checkout-api-tests") # created on first use
with mlflow.start_run(run_name="ci-1234") as run:
mlflow.set_tags({"git_sha": "a1b2c3d", "ci_pipeline_id": "1234", "suite": "smoke"})
mlflow.log_params({"base_url": "https://staging.example.com", "workers": 4})
mlflow.log_metrics({"tests_total": 120, "tests_failed": 3, "pass_rate": 0.975})
for step, latency in enumerate([120, 135, 128]): # a series, plotted as a chart
mlflow.log_metric("p95_latency_ms", latency, step=step)
mlflow.log_dict({"failed": ["test_refund", "test_cancel"]}, "failures.json")
mlflow.log_text("staging, feature flags: new-checkout", "notes/env.txt")
mlflow.log_artifact("report.xml", artifact_path="reports") # an existing local file
print(run.info.run_id, run.info.experiment_id)
- The
withblock ends the run asFINISHED, or asFAILEDif an exception escapes it. - The run is active only in this process and thread;
mlflow.log_*outside a run starts a new one implicitly — always usestart_run()explicitly. log_paramsvalues are converted to strings:4becomes"4", so filter asparams.workers = '4'.
Artifact Helpers¶
| Function | Writes | Good for |
|---|---|---|
mlflow.log_artifact(path, artifact_path=...) |
One local file | JUnit XML, Allure archive, HAR, screenshots |
mlflow.log_artifacts(dir, artifact_path=...) |
A whole directory | allure-results/, playwright-report/ |
mlflow.log_text(text, "file.txt") |
A string | Environment description, diffs |
mlflow.log_dict(obj, "file.json") |
JSON / YAML from a dict | Failure lists, configs |
mlflow.log_table(data=df, artifact_file="t.json") |
A pandas.DataFrame (or dict of lists) as a table |
Per-case eval results, viewable and comparable in the UI |
mlflow.log_figure(fig, "chart.png") |
A matplotlib / plotly figure | Score distributions |
System Tags¶
MLflow sets some tags itself: mlflow.runName, mlflow.user, mlflow.source.name, mlflow.source.type and — when the entry-point script lives in a git repository — mlflow.source.git.commit and mlflow.source.git.branch.
No git tags under pytest
Under pytest the entry point is .venv/.../pytest, outside the repository, so mlflow.source.git.commit is not set. Set your own git_sha / git_branch tags from CI variables (GITHUB_SHA, CI_COMMIT_SHA) — see Testing, CI & Model Registry.
Autologging¶
mlflow.autolog() patches supported libraries so params, metrics and models are logged without explicit calls.
import mlflow
from sklearn.datasets import load_iris
from sklearn.linear_model import LogisticRegression
mlflow.set_experiment("iris")
mlflow.autolog() # or only one library: mlflow.sklearn.autolog()
X, y = load_iris(return_X_y=True)
with mlflow.start_run():
LogisticRegression(max_iter=200).fit(X, y)
# Params: C, max_iter, ...; metrics: training_accuracy_score, training_f1_score, ...; the model as an artifact
| Call | Captures |
|---|---|
mlflow.autolog() |
All installed and supported libraries |
mlflow.sklearn.autolog(), mlflow.pytorch.autolog(), mlflow.xgboost.autolog() ... |
Classic ML: params, training metrics, model |
mlflow.openai.autolog(), mlflow.langchain.autolog(), mlflow.anthropic.autolog() ... |
GenAI: traces of every call — see GenAI Tracing |
mlflow.autolog(disable=True) |
Turns autologging off again |
Searching Runs¶
import mlflow
runs = mlflow.search_runs(
experiment_names=["checkout-api-tests"],
filter_string="tags.suite = 'smoke' and metrics.pass_rate < 0.99",
order_by=["attributes.start_time DESC"],
max_results=20,
) # pandas.DataFrame: run_id, metrics.*, params.*, tags.*
print(runs[["run_id", "metrics.pass_rate", "tags.git_sha"]])
latest = mlflow.search_runs(
experiment_names=["checkout-api-tests"],
filter_string="attributes.status = 'FINISHED'",
order_by=["attributes.start_time DESC"],
max_results=1,
output_format="list", # list[Run] instead of a DataFrame
)[0]
print(latest.data.metrics["pass_rate"], latest.data.tags["git_sha"])
Filter syntax essentials:
| Filter | Meaning |
|---|---|
metrics.pass_rate < 0.95 |
Numeric comparison on the last logged value |
params.model = 'gpt-4.1-mini' |
Params and tags are strings — quote the value |
tags.git_branch = 'main' |
Exact match on a tag |
tags.git_branch LIKE 'release/%' |
Pattern match (ILIKE for case-insensitive) |
tags.`git.sha` = 'a1b2c3d' |
Backticks for keys with dots, dashes or spaces |
metrics.`correctness/mean` > 0.8 |
Backticks for eval metrics named <scorer>/mean |
attributes.status = 'FINISHED' |
Run attributes: status, start_time, run_name, ... |
a = 1 and b = 2 |
Only and — there is no or; run two searches instead |
For metric history (all steps) use MlflowClient().get_metric_history(run_id, "p95_latency_ms").
Comparing Runs in the UI¶
- Experiments → your experiment → Runs — the table shows params, metrics and tags as columns; the search box takes the same filter syntax.
- Select two or more runs → Compare — side-by-side params and metrics, parallel-coordinates and scatter plots, metric charts over steps.
- Chart view — one chart per metric across runs, good for "pass rate over the last 30 builds".
- Artifacts tab of a run — download reports, preview JSON, text, images and tables.
Naming Conventions for Test Runs¶
| Thing | Convention | Example |
|---|---|---|
| Experiment | <app>-<suite> |
support-bot-regression, checkout-api-smoke |
| Run name | ci-<pipeline id> or local-<user> |
ci-18231 |
| Tags | Lower-case, underscores, no dots | git_sha, git_branch, ci_pipeline_id, trigger |
| Params | Everything that defines the run's inputs | model, prompt_version, dataset_version, judge_model |
| Metrics | Aggregates, one value per run | pass_rate, tests_failed, faithfulness_mean, p95_latency_ms |
Tracking Checklist¶
- One experiment per suite; experiment name comes from
MLFLOW_EXPERIMENT_NAMEin CI - Every run has
git_sha,git_branchandci_pipeline_idtags - Everything that changes the result is a param (model, prompt version, dataset version, judge)
- Metrics are aggregates with stable names; per-case detail goes to a table artifact
- Reports (JUnit XML, Allure, HTML) are attached as artifacts
- Runs are always started explicitly with
start_run()inside awithblock
See also¶
- MLflow — Experiment Tracking, LLM Tracing & Evaluation
- MLflow — Setup & Architecture
- MLflow — GenAI Tracing
- MLflow — Testing, CI & Model Registry
- Pytest — Python Testing Framework