Summary
✨ AI‑Generated
A senior engineering role responsible for designing enterprise monitoring strategies, improving reliability practices, integrating observability tools, and guiding teams toward automated and intelligent operations.
Highlights
Leadership role focused on building modern observability strategies, driving platform improvements, and enabling reliable cloud and application operations.
Description
Position Summary
We are seeking a Lead Site Reliability & Environment Monitoring Engineer to establish and evolve our enterprise observability and monitoring strategy across cloud and application platforms.
This is a full-time leadership role responsible for owning monitoring design, driving platform decisions, and guiding engineering teams toward modern SRE practices.
This individual will act as the technical authority for monitoring and alerting, shaping how signals from Dynatrace flow into ServiceNow and enterprise messaging/paging platforms, and enabling a shift toward automated, intelligent, and self-healing operations.
Location: This role is based in Charlotte, NC.
Key Responsibilities
Strategic Leadership & Decision-Making
Define and own the enterprise monitoring and SRE observability strategyServe as the subject matter expert for Dynatrace, ServiceNow integration, and alerting architectureEvaluate and recommend tooling, integration patterns, and platform directionDrive decisions on alerting philosophy, noise reduction, and signal quality improvement
Platform Ownership & Architecture
Architect and standardize end-to-end monitoring and SRE pipelines:Dynatrace → ServiceNow incident lifecycleAlert correlation, deduplication, and prioritizationIntegration with paging systems (PagerDuty, SMS, voice, Teams)Establish best practices for:Event ingestion and enrichmentIncident routing and automated assignmentIntegration with CMDB and service mapping
Site Reliability Engineering (SRE) Leadership
Lead adoption of SRE principles, including:SLIs, SLOs, and error budgetsReliability engineering practices across servicesProactive monitoring and resilience designChampion a shift from reactive operations to proactive reliability engineeringInfluence application and platform teams to build observable, resilient systems by design
Automation & Self-Healing Enablement
Drive development of automated remediation and self-healing capabilitiesLeverage Dynatrace workflows, Azure services, and automation frameworks to:Reduce manual incident handlingEliminate repeatable operational tasksMinimize unnecessary paging
ServiceNow & Observability Integration Leadership
Own integration between Dynatrace and ServiceNow ITSM/ITOM, including:Incident, Event Management, and CMDB alignmentService mapping and dependency visibilityGovernance for application/service taggingDefine standards for:Automated incident creation and resolutionPriority assignment and routing logicMonitoring-to-ITSM data synchronization
Team Leadership & Cross-Functional Influence
Provide technical leadership and mentorship across SRE, platform, and application teamsAct as a central point of coordination between engineering, cloud, and ITSM teamsLead workshops and working sessions to:Drive monitoring standardizationAlign teams on reliability practicesInfluence upstream architectural decisions
Operational Excellence
Establish KPIs and drive improvement in:Incident response and resolution timesAlert quality and paging effectivenessMonitoring coverage across critical servicesProvide leadership with clear visibility into service health and reliability trends
Required Qualifications
7+ years in Site Reliability Engineering, monitoring, or production engineeringProven experience in a technical leadership or lead engineer roleDeep hands-on experience with:Dynatrace (or equivalent observability platforms)Microsoft Azure (IaaS, PaaS, networking, identity)ServiceNow ITSM / ITOM (incident, event management, CMDB)Demonstrated ability to:Design and lead enterprise monitoring/SRE architecturesDrive platform and tooling decisionsIntegrate observability, ITSM, and paging solutions
Preferred Qualifications
Experience leading SRE or observability transformation initiativesStrong expertise with Dynatrace–ServiceNow integrationsExperience modernizing or consolidating paging/on-call toolingFamiliarity with:Azure-based SRE tooling or AI-assisted operationsAutomation frameworks (GitHub Actions, Runbooks, etc.)Infrastructure as Code (Terraform, ARM, Bicep)
Success Metrics
Reduction in alert noise and unnecessary pagingImproved incident routing accuracy and MTTRIncreased adoption of self-healing and automated workflowsStrong alignment between monitoring, CMDB, and service ownershipEnterprise-wide adoption of SRE and monitoring standards