OpenTelemetry — Testing with OpenTelemetry¶
Two different goals:
- Test the instrumentation — unit tests assert that code emits the right spans, attributes and metrics.
- Use telemetry in tests — link test runs to backend traces, debug failures, assert on system behavior (trace-based testing).
In-Memory Span Exporter¶
# tests/conftest.py
import pytest
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
from opentelemetry.sdk.trace.export.in_memory_span_exporter import InMemorySpanExporter
_exporter = InMemorySpanExporter()
_provider = TracerProvider()
_provider.add_span_processor(SimpleSpanProcessor(_exporter))
trace.set_tracer_provider(_provider) # global, can be set only once per process
@pytest.fixture
def spans() -> InMemorySpanExporter:
_exporter.clear()
yield _exporter
_exporter.clear()
SimpleSpanProcessorexports onspan.end()— spans are available right after the code under test returns.- The global provider can be set once, so configure it at import time of
conftest.pyand clear the exporter per test. - Alternatively, inject a tracer into the code under test (
TracerProvider().get_tracer(...)) and avoid globals entirely.
Asserting Spans¶
import pytest
from opentelemetry.trace import SpanKind, StatusCode
from shop.checkout import checkout
def test_checkout_creates_span_with_order_attributes(spans):
checkout(order_id="ord-42", items=3)
finished = spans.get_finished_spans()
span = next(s for s in finished if s.name == "checkout")
assert span.kind is SpanKind.INTERNAL
assert span.attributes["shop.order.id"] == "ord-42"
assert span.attributes["shop.order.items"] == 3
assert span.status.status_code is StatusCode.UNSET
def test_payment_failure_marks_span_as_error(spans, failing_gateway):
with pytest.raises(PaymentDeclined):
checkout(order_id="ord-43", items=1)
span = next(s for s in spans.get_finished_spans() if s.name == "charge_card")
assert span.status.status_code is StatusCode.ERROR
assert any(e.name == "exception" for e in span.events)
assert span.events[-1].attributes["exception.type"].endswith("PaymentDeclined")
Asserting the span tree¶
def by_name(spans):
return {s.name: s for s in spans}
def test_charge_is_child_of_checkout(spans):
checkout(order_id="ord-44", items=2)
s = by_name(spans.get_finished_spans())
assert s["charge_card"].parent.span_id == s["checkout"].context.span_id
assert s["charge_card"].context.trace_id == s["checkout"].context.trace_id
Spans are finished child first, so get_finished_spans() lists children before parents. Look up spans by name instead of relying on indices.
Testing Auto-Instrumented FastAPI¶
from fastapi.testclient import TestClient
from opentelemetry.instrumentation.fastapi import FastAPIInstrumentor
from app.main import create_app
@pytest.fixture
def client():
app = create_app()
FastAPIInstrumentor.instrument_app(app)
yield TestClient(app)
FastAPIInstrumentor.uninstrument_app(app)
def test_route_span_uses_template(client, spans):
client.get("/orders/9812")
server = next(s for s in spans.get_finished_spans() if s.kind is SpanKind.SERVER)
assert server.name == "GET /orders/{order_id}" # template, not the raw ID
assert server.attributes["http.route"] == "/orders/{order_id}"
This catches a common regression: span names containing IDs, which explodes cardinality in the backend.
Testing Metrics¶
from opentelemetry.sdk.metrics import MeterProvider
from opentelemetry.sdk.metrics.export import InMemoryMetricReader
def test_orders_counter():
reader = InMemoryMetricReader()
meter = MeterProvider(metric_readers=[reader]).get_meter("test")
service = OrderService(meter=meter) # inject the meter
service.create(payment_method="card")
service.create(payment_method="card")
data = reader.get_metrics_data()
metric = next(
m for rm in data.resource_metrics for sm in rm.scope_metrics for m in sm.metrics
if m.name == "shop.orders.created"
)
point = metric.data.data_points[0]
assert point.value == 2
assert point.attributes == {"payment.method": "card"}
Metrics use an injected MeterProvider here, because the global one — like the tracer provider — can be set only once.
Linking Test Runs to Backend Traces¶
For end-to-end and API tests against a deployed system, give every test its own trace and send traceparent with each request. A failing test then points straight to the backend trace.
# tests/conftest.py
import pytest
from opentelemetry import baggage, context, trace
from opentelemetry.propagate import inject
tracer = trace.get_tracer("tests")
@pytest.fixture(autouse=True)
def test_trace(request):
ctx = baggage.set_baggage("test.run.id", request.config.getoption("--run-id", "local"))
token = context.attach(ctx)
with tracer.start_as_current_span(request.node.nodeid, attributes={"test.name": request.node.name}) as span:
yield span
trace_id = format(span.get_span_context().trace_id, "032x")
request.node.user_properties.append(("trace_id", trace_id)) # lands in JUnit XML
context.detach(token)
@pytest.fixture
def trace_headers() -> dict[str, str]:
headers: dict[str, str] = {}
inject(headers) # traceparent + baggage for the current test span
return headers
# API test (Playwright APIRequestContext)
def test_create_order(playwright, trace_headers):
api = playwright.request.new_context(base_url=BASE_URL, extra_http_headers=trace_headers)
response = api.post("/orders", data={"sku": "book-1", "qty": 1})
assert response.status == 201
# UI test: every browser request of the page carries the test's trace context
def test_checkout_ui(page, trace_headers):
page.set_extra_http_headers(trace_headers)
page.goto("/checkout")
Combined with the keep-test-traffic tail sampling policy (05 Collector & Backends), all test traffic is kept, and the trace_id in the test report opens the full backend trace.
Attach the trace link to the report
Add a link like https://grafana.example.com/explore?...traceId=<trace_id> to Allure (allure.dynamic.link) or the pytest HTML report on failure.
Trace-Based Testing¶
Instead of asserting only on the HTTP response, assert on what happened inside the system:
| Assertion | Example |
|---|---|
| A call happened | POST /charge span exists in payments-api |
| A call did not happen | No span to the fraud service for trusted customers |
| Order of operations | reserve_stock ends before charge_card starts |
| Performance budget | SELECT orders span < 50 ms |
| Number of calls | Exactly one DB query per request (N+1 detection) |
| Errors are recorded | Retry spans have error.type set |
Implementation options:
- Query the trace backend after the test (Tempo / Jaeger HTTP API by
trace_id) and assert on the returned spans. Wait until spans arrive — export is asynchronous. - Test in-process with
InMemorySpanExporterwhen the whole flow runs in one Python process (FastAPITestClient+ instrumented SQLAlchemy). - Dedicated tools such as Tracetest, which define trace assertions declaratively and run them in CI.
def test_order_listing_has_no_n_plus_one(client, spans):
client.get("/orders?limit=20")
db_spans = [s for s in spans.get_finished_spans() if s.attributes.get("db.system.name")]
assert len(db_spans) <= 2, [s.attributes.get("db.query.text") for s in db_spans]
Checklist¶
- Critical business operations have spans with meaningful names and attributes
- Error paths set
ERRORstatus and record exceptions - Span names use route templates, not raw IDs
- Metric attributes are bounded (no IDs)
-
traceparentis propagated through queues and background jobs - Test runs carry their own trace and
test.run.idbaggage - Health checks and test infrastructure noise are excluded
See also¶
- OpenTelemetry — Python Observability
- OpenTelemetry — Tracing
- Pytest
- Playwright — API Testing
- FastAPI — Testing
- Jaeger — Testing, CI & Troubleshooting