← All Case Studies

Amazon2022 – Present · Platform Architect & Standards Owner

Standardized Infrastructure Monitoring Across Thousands of Services

Defined the monitoring standard across thousands of services and drove cross-team adoption.

The Problem

Infrastructure teams operated with fragmented, inconsistent monitoring across thousands of services. Different teams used different tooling, alert standards, and dashboard conventions, creating blind spots, duplicated effort, and unreliable signal during incidents.

There was no unified "paved road" for infrastructure monitoring, meaning each team had to reinvent their observability stack independently.

The Approach

Architected and drove adoption of a unified monitoring platform standardizing infrastructure observability across thousands of services. Defined common monitoring standards: alert policies, SLO baselines, dashboard conventions, and deployment playbooks.

Led cross-team alignment to migrate from fragmented legacy monitoring stacks to the unified platform. Established the platform as the "paved road" for all infrastructure monitoring, onboarding 2,750+ application stages and 3,000,000 monitors at scale.

The Impact

  • Brought thousands of services onto one monitoring paved road, replacing the per-team tooling and alert conventions each team had reinvented on its own
  • Onboarded 2,750+ application stages and brought 3,000,000 monitors under standardized management
  • Replaced inconsistent per-team alerting with shared alert policies and SLO baselines, so incident signal reads the same across services
  • Retired the fragmented legacy monitoring stacks teams had maintained in parallel
  • Became the paved road new infrastructure monitoring is built on org-wide
ObservabilityPlatform EngineeringSREStandardization

Related