Case 02
Unified Observability for Carrier-Grade Cloud Calling
Situation
A globally distributed cloud calling platform was experiencing a recurring operational problem: when something went wrong, figuring out where and why took too long. Engineers and carrier partners were working across separate, disconnected toolchain, each with their own dashboards, alerting logic, and incident workflows. The result was fragmented triage, delayed resolution, and a growing mismatch between the platform's scale and its ability to respond to it.
Complication
The challenge wasn't just technical, it was organizational. Consolidating observability across internal SRE teams and 100+ external carrier partners meant aligning groups with different systems, different incentives, and different definitions of what "resolved" meant. A platform change alone wouldn't work without a change management program alongside it.
Approach
Drove the adoption of a Unified Observability Platform that consolidated telemetry, alerting, and triage workflows into a single operational layer shared across internal teams and carrier partners. Designed the cross-partner operational model, defined the data sharing agreements, and led the change management program that brought carriers onto the platform. Built the governance structure to sustain it.
Outcome
Incident analysis time reduced by 45%. Partner triage coordination moved from fragmented email and ticket chains to a shared operational surface, improving both the speed of resolution and the quality of post-incident analysis. The platform shifted from reactive firefighting to structured, measurable incident management.