Distributed Tracing
Properties
created
14.04.2026, 17:00
modified
26.07.2026, 13:46
published
Empty
sources
OTel Distributed Tracing · W3C Trace Context
topics
Distributed Tracing, Observability
authors
Empty
ai-assisted
Yes
- Technique for tracking a single request as it flows through multiple services in a distributed system
- Answers: which service is slow? Where did this error originate? What is the call graph?
- Essential for microservices architectures where a single user action touches many services
# Topics
- Spans and Traces
- Trace and span data model, span kinds, attributes, sampling strategies
- Context Propagation
- W3C traceparent header, B3/Jaeger formats, baggage, cross-service linking
# The Problem It Solves
- In a monolith, a stack trace shows the full call path
- In microservices, a request might hit API → Auth → Inventory → Payment → Notification
- Each service has its own logs — correlating them manually is nearly impossible
- Distributed tracing stitches them into a single visual timeline
# Core Data Model
- A trace is a directed acyclic graph (tree) of spans
- Each service creates spans; they’re linked via
traceIdandparentSpanId - The full trace is assembled by the backend (Tempo, Jaeger) by joining on
traceId
# Instrumentation Requirements
- Each service must
- Extract incoming trace context from request headers
- Create a child span for its own work
- Inject trace context into all outbound requests
- OTel SDKs handle this automatically via auto-instrumentation
# Key Backends
- Grafana Tempo — object storage backed, TraceQL, Grafana-native
- Jaeger — CNCF project, older but widely deployed
- Zipkin — one of the original tracing systems, simpler feature set
- Datadog APM, AWS X-Ray, Honeycomb — commercial options