OpenTelemetry — Collector & Backends¶
Why a Collector¶
| Without Collector | With Collector |
|---|---|
| Every service holds vendor endpoints and API keys | Services only know collector:4317 |
| Changing backend = redeploying every service | Changing backend = editing one config |
| Sampling decided per service, blind to outcome | Tail sampling on complete traces |
| Retries and buffering eat app memory | Buffering and retries offloaded |
| Sensitive attributes leave the app as-is | Redaction in one place |
Pipeline Model¶
flowchart LR
R["Receivers<br/>otlp, prometheus, filelog"] --> P["Processors<br/>memory_limiter → attributes → tail_sampling → batch"]
P --> E["Exporters<br/>otlp, prometheus, debug"]
C["Connectors<br/>spanmetrics"] -.-> E
- Receivers — accept data (push or pull).
- Processors — transform, filter, sample, batch. Order matters: they run top to bottom.
- Exporters — send data to backends.
- Connectors — act as an exporter of one pipeline and a receiver of another (e.g. derive metrics from spans).
- Extensions — health check, pprof, auth.
Distributions¶
| Image | Contents | Use |
|---|---|---|
otel/opentelemetry-collector |
Core components only | Minimal, OTLP in / OTLP out |
otel/opentelemetry-collector-contrib |
Core + community components (tail_sampling, filelog, k8sattributes, vendor exporters) |
Most real setups |
Custom build (ocb) |
Only the components you list | Smaller attack surface in production |
Baseline Configuration¶
# otel-collector.yaml
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
memory_limiter: # first: protects the collector from OOM
check_interval: 1s
limit_percentage: 80
spike_limit_percentage: 20
resource:
attributes:
- key: deployment.environment.name
value: staging
action: upsert
attributes/redact:
actions:
- key: http.request.header.authorization
action: delete
- key: user.email
action: hash
batch: # last: fewer, larger export requests
send_batch_size: 8192
timeout: 5s
exporters:
otlp/tempo:
endpoint: tempo:4317
tls:
insecure: true
prometheus:
endpoint: 0.0.0.0:8889
otlphttp/loki:
endpoint: http://loki:3100/otlp
debug:
verbosity: basic # detailed = full span dump to collector stdout
extensions:
health_check:
endpoint: 0.0.0.0:13133
service:
extensions: [health_check]
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, resource, attributes/redact, batch]
exporters: [otlp/tempo, debug]
metrics:
receivers: [otlp]
processors: [memory_limiter, resource, batch]
exporters: [prometheus]
logs:
receivers: [otlp]
processors: [memory_limiter, resource, attributes/redact, batch]
exporters: [otlphttp/loki]
A component declared but not listed in service.pipelines is ignored. type/name (e.g. otlp/tempo) lets you have several instances of one component.
Tail Sampling¶
Head sampling in the SDK drops traces before knowing the outcome. Tail sampling waits for a whole trace and then decides.
processors:
tail_sampling:
decision_wait: 10s # how long to buffer spans of one trace
num_traces: 50000
policies:
- name: keep-errors
type: status_code
status_code: { status_codes: [ERROR] }
- name: keep-slow
type: latency
latency: { threshold_ms: 1000 }
- name: keep-test-traffic
type: string_attribute
string_attribute: { key: test.run.id, values: [".+"], enabled_regex_matching: true }
- name: sample-rest
type: probabilistic
probabilistic: { sampling_percentage: 5 }
A trace is kept if any policy matches. All spans of a trace must reach the same collector instance — with several replicas put a loadbalancing exporter (routing by trace_id) in front.
Span Metrics¶
Derive RED metrics from traces, so every service gets rate, errors and latency even without metric instrumentation:
connectors:
spanmetrics:
histogram:
explicit:
buckets: [50ms, 100ms, 250ms, 500ms, 1s, 2.5s]
dimensions:
- name: http.request.method
- name: http.response.status_code
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [otlp/tempo, spanmetrics]
metrics/spanmetrics:
receivers: [spanmetrics]
exporters: [prometheus]
Put spanmetrics before tail sampling (in a separate pipeline), otherwise metrics only count the sampled traces.
Local Stack with Docker Compose¶
Option 1 — all-in-one (fastest)¶
# compose.yaml
services:
lgtm:
image: grafana/otel-lgtm
ports:
- "3000:3000" # Grafana (admin / admin)
- "4317:4317" # OTLP gRPC
- "4318:4318" # OTLP HTTP
Contains a Collector, Tempo (traces), Loki (logs), Prometheus (metrics) and Grafana with data sources preconfigured. For development and demos only.
Option 2 — Collector + Jaeger¶
# compose.yaml
services:
collector:
image: otel/opentelemetry-collector-contrib
command: ["--config=/etc/otelcol/config.yaml"]
volumes:
- ./otel-collector.yaml:/etc/otelcol/config.yaml:ro
ports:
- "4317:4317"
- "4318:4318"
- "13133:13133" # health check
depends_on: [jaeger]
jaeger:
image: jaegertracing/jaeger
ports:
- "16686:16686" # Jaeger UI
app:
build: .
environment:
OTEL_SERVICE_NAME: orders-api
OTEL_EXPORTER_OTLP_ENDPOINT: http://collector:4317
OTEL_EXPORTER_OTLP_PROTOCOL: grpc
ports:
- "8000:8000"
depends_on: [collector]
With exporter otlp/jaeger: { endpoint: jaeger:4317, tls: { insecure: true } } in the collector config. Open http://localhost:16686, pick orders-api, find traces. Jaeger storage, UI and test usage: Jaeger guide.
Deployment Patterns¶
| Pattern | Where the collector runs | Good for |
|---|---|---|
| Agent | Sidecar or DaemonSet next to the app | Host/k8s metadata, local buffering, low latency |
| Gateway | Central deployment behind a load balancer | Tail sampling, redaction, auth to vendors |
| Agent + Gateway | Both | Production at scale |
Backend Choices¶
| Signal | Open source | Notes |
|---|---|---|
| Traces | Jaeger, Grafana Tempo, Zipkin | Tempo stores traces in object storage, queried by TraceQL |
| Metrics | Prometheus, Grafana Mimir, VictoriaMetrics | Prometheus accepts OTLP natively (--web.enable-otlp-receiver) |
| Logs | Grafana Loki, OpenSearch / Elasticsearch | Loki has a native OTLP endpoint (/otlp) |
| All-in-one | SigNoz, Uptrace, Grafana LGTM | One UI for all signals |
Commercial platforms (Datadog, New Relic, Honeycomb, Dynatrace, Grafana Cloud, Elastic) accept OTLP — the app code stays the same.
See also¶
- OpenTelemetry — Python Observability
- OpenTelemetry — Testing with OpenTelemetry
- Client–Server: Observability
- uv — Workspaces & Docker
- Jaeger — Setup & Architecture