Back to project details
planned

Distributed Tracing Platform

Unified observability with metrics, logs, and traces in one place.

observability-platform — grafana

SPANS / SEC

1M

RETENTION

30 days

P99 LATENCY

45ms

MONTHLY COST

< $5K

TRACE: abc123def456 — POST /api/orders — 234ms

gateway
2ms
auth
3ms
orders
200ms
inventory
18ms
payment
11ms

Simulated trace waterfall — total: 234ms across 5 services

Consolidation Plan

DC

Datadog

Metrics + APM — $8K/month

SP

Splunk

Log aggregation — $6K/month

NR

New Relic

Browser RUM — $3K/month

Total current spend$17K/month
Target spend< $5K/month

Proposed Stack

OT

OpenTelemetry

Vendor-neutral instrumentation

CH

ClickHouse

Columnar DB for traces & logs

GF

Grafana + Tempo

Unified visualization layer

This project is in the planning phase

Currently evaluating vendors, running PoCs, and gathering requirements from engineering teams. Expected rollout begins next quarter.

Key Design Decisions

🔒

Sampling Strategy

Head-based probabilistic sampling at 10% for production traffic, 100% for staging.

📦

Storage Tiers

Hot storage (ClickHouse SSD) for 7 days, warm (S3-compatible) for 30 days retention.

🔗

Correlation

Trace IDs injected into structured logs for seamless trace-to-log navigation in Grafana.

⚡

Alerting

Grafana Alertmanager with SLO-based burn rate alerts, integrated with PagerDuty.