Azure SRE Lead – Azure DevOps / Kubernete - Only W2
Pipsfortune — United States · Posted ~3 hours ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
Job DescriptionWe are seeking an experienced Azure SRE Lead with strong expertise in Site Reliability Engineering, Azure DevOps, Kubernetes, observability, and cloud infrastructure.
The ideal candidate will have 5+ years of hands-on SRE experience and a proven track record supporting large-scale distributed applications and production environments.
The candidate will be responsible for improving system reliability, monitoring and observability, automation, incident management, performance, and production operations across cloud-based environments.
Key ResponsibilitiesLead Site Reliability Engineering initiatives for large-scale distributed applications.Design and implement reliable, scalable, and highly available cloud infrastructure.Build and maintain Azure DevOps CI/CD pipelines and deployment automation.Manage and support Kubernetes and Docker containerized environments.Implement observability solutions covering logs, metrics, traces, and application telemetry.Develop dashboards, alerts, and monitoring solutions using tools such as Splunk, Datadog, Dynatrace, Grafana, Prometheus, and OpenTelemetry.Collect, analyze, and troubleshoot telemetry data to identify system performance and reliability issues.Lead production incident response, root cause analysis, and remediation activities.Perform performance tuning and capacity planning for distributed applications.Implement Infrastructure as Code and automation for cloud environments.Develop automation scripts using Python, Bash, or Go.Establish SRE best practices around reliability, availability, scalability, and operational efficiency.Collaborate with development, DevOps, cloud, and infrastructure teams to improve application reliability.Support continuous improvement of production environments and operational processes.Must-Have Skills5+ years of hands-on experience in Site Reliability Engineering (SRE).Strong experience with Azure DevOps.Strong hands-on experience with Kubernetes.Experience with Docker and containerized applications.Strong knowledge of observability and monitoring.Experience with Splunk, Datadog, Dynatrace, Grafana, Prometheus, or OpenTelemetry.Experience with logging, metrics, tracing, telemetry collection, and dashboard development.Strong experience with cloud platforms, preferably Microsoft Azure.Hands-on experience with CI/CD pipelines and automation.Experience with Infrastructure as Code.Strong troubleshooting, performance tuning, and incident management skills.Proficiency in Python, Bash, or Go.Experience supporting large-scale distributed applications in production.Nice-to-Have SkillsMicrosoft FabricExperience with advanced Azure monitoring and observability services.Experience with multi-cloud environments such as AWS or GCP.Experience implementing SRE frameworks and reliability metrics/SLIs/SLOs.Experience with automated remediation and self-healing infrastructure.QualificationsBachelor’s degree in Computer Science, Information Technology, Engineering, or a related field preferred.8–10 years of overall IT experience with strong SRE/DevOps experience.Strong communication, leadership, troubleshooting, and problem-solving skills.Ability to work in a hybrid environment in Atlanta, GA (2 days per week onsite)
We have 142,037 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 142,037 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume