Distributed Tracing Platform
Unified observability with metrics, logs, and traces in one place.
observability-platform — grafana
SPANS / SEC
1M
RETENTION
30 days
P99 LATENCY
45ms
MONTHLY COST
< $5K
TRACE: abc123def456 — POST /api/orders — 234ms
Simulated trace waterfall — total: 234ms across 5 services
Consolidation Plan
Datadog
Metrics + APM — $8K/month
Splunk
Log aggregation — $6K/month
New Relic
Browser RUM — $3K/month
Proposed Stack
OpenTelemetry
Vendor-neutral instrumentation
ClickHouse
Columnar DB for traces & logs
Grafana + Tempo
Unified visualization layer
This project is in the planning phase
Currently evaluating vendors, running PoCs, and gathering requirements from engineering teams. Expected rollout begins next quarter.
Key Design Decisions
Sampling Strategy
Head-based probabilistic sampling at 10% for production traffic, 100% for staging.
Storage Tiers
Hot storage (ClickHouse SSD) for 7 days, warm (S3-compatible) for 30 days retention.
Correlation
Trace IDs injected into structured logs for seamless trace-to-log navigation in Grafana.
Alerting
Grafana Alertmanager with SLO-based burn rate alerts, integrated with PagerDuty.