Unified Observability Stack cover
All work

2023

Unified Observability Stack

A cost-aware metrics, logs, and traces platform that gave every team one place to answer "is it my service?"

Overview

Consolidated a sprawl of disconnected monitoring tools into a single OpenTelemetry-based stack: Prometheus and Thanos for long-term metrics, Loki for logs, and Tempo for distributed traces, all fronted by Grafana. Instrumented the shared libraries so new services emit correlated telemetry for free, defined SLOs and burn-rate alerts, and tuned retention and sampling to keep the bill flat as traffic grew.

Role & impact

My role

SRE — observability architecture and on-call tooling.

Impact

Reduced mean time to resolution by 55% and cut monitoring spend by 38% despite 3x traffic growth.

Stack

OpenTelemetryPrometheusGrafanaLokiTempo