Back to portfolio
RS
Observability

Infrastructure Monitoring

24×7 monitoring stack across servers, network devices and cloud with escalation to on-call — for problems that surface *before* users notice.

Observability Engineer 2 months
-70%
MTTD reduction
-45%
MTTR reduction
-80%
False positives
300+
Monitored endpoints
Challenges
  • Incidents discovered by users, not systems.
  • Alert fatigue from noisy, undifferentiated notifications.
  • No historical trend data for capacity planning.
Solutions
  • Deployed Zabbix templates for OS, hypervisor, network and app tiers.
  • Tuned alert thresholds and severity; wired escalation into on-call.
  • Long-term metric retention for capacity and trend analysis.
Business Impact

Mean-time-to-detect reduced by 70%, mean-time-to-recover by 45%.

Technology Stack
ZabbixNagiosCloudWatchSNMPSMTP alerts