Azure Site Reliability Engineer

Gsshrsolutions — Philippines · Posted ~1 day ago

Mid Full-time Hybrid

Skills

Azure Monitor Azure Log Analytics KQL ServiceNow ITOM Grafana Prometheus SRE Observability APM AppDynamics ThousandEyes

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary

Help design and improve enterprise observability and monitoring standards, implement reliability practices, reduce incidents, and build scalable monitoring solutions for cloud platforms.

Highlights

Permanent hybrid role focused on cloud observability, enterprise monitoring, automation, and reliability improvements with cross-functional collaboration.

Description

Experience: 3-8 years Location: Quezon - Hybrid (1-2 days WFO in a week) Shift: Night shift Job type: Direct hire with MNC (Permanent) Key Responsibilities: Design and define standards, patterns, and automations opportunities that elevate monitoring and reliability across platforms and applications, with a strong focus on Azure Monitor, ServiceNow ITOM Event Management, Grafana, and APM/Synthetics toolingPartner with product teams to implement SLO/SLI‑driven operations, reduce alert noise, accelerate incident response, and embed self‑healing where it matters most.Engineer enterprise monitoring & event patterns by authoring and maintaining reference architectures, runbooks, and event management models (alert → event → incident) with actionable alerts and incidents routing.Contribute to Monitoring and Observability & Event Management Strategy and tooling intake/governance checkpoints and coach product teamsExcellent communication skills to drive continuous improvement by reducing alert noise, shorten MTTR, and improve change success by embedding postmortem learnings into patterns, rules, and pipelines. Technologies and Tools: - SRE Practices: Observability and Monitoring - Cloud Observability: Azure Monitor/App Insights/Log Analytics (KQL) - Grafana/Prometheus for metrics visualization where applicable - ServiceNow ITOM Event Management - Azure Fundamentals, Azure Monitor - DevOps and Automation Tools - Grafana, Prometheus, App Dynamics, ThousandEyes - Application Performance Monitoring and Digital User Experience tools Required Qualifications: - 3+ years in monitoring/observability/SRE roles with hands‑on experience in Azure Monitor/App Insights (KQL) and ServiceNow Event Management. - Strong knowledge in Azure Log Analytics, KQL, Telemetry, APM implementations - Demonstrated ability to collaborate across IT Operations team, platform, cyber, network, and product teams, strong written verbal communication for standards and enablement.