Summary
✨ AI‑Generated
A lead-level DevOps position responsible for designing and maintaining scalable cloud infrastructure. The role involves improving developer productivity, operational reliability, automation, and modern engineering practices.
Highlights
Leadership opportunity focused on scalable cloud platforms, automation, reliability improvements, and collaboration with engineering teams.
Description
Lead DevOps Engineer
About the Company
This organization is focused on modernizing the mortgage insurance industry through a technology-first approach.
Rather than relying on legacy systems and manual processes, the company leverages software, automation, artificial intelligence, analytics, and scalable operating models to deliver better customer experiences and drive business efficiency.
Its mission is to build a more modern, data-driven insurance platform designed for long-term growth and innovation.
About the Role
The company is seeking a highly skilled and motivated Lead DevOps Engineer to play a key role in designing, implementing, and maintaining reliable, scalable, and high-performing cloud infrastructure within AWS.
This individual will work closely with software engineering, operations, and cross-functional teams to improve platform reliability, enhance developer productivity, and drive operational excellence through automation, monitoring, and incident response practices.
Key Responsibilities
Lead and mentor a team of Site Reliability Engineers, including both full-time employees and contractors.Prioritize, assign, and review technical work while providing guidance and feedback on code and infrastructure changes.Design, implement, and maintain scalable, secure, and highly available cloud infrastructure in AWS.Build and support monitoring, alerting, and observability solutions to ensure platform health and uptime.Automate infrastructure provisioning and configuration management using Infrastructure-as-Code tools.Develop and enhance CI/CD pipelines to improve deployment efficiency and software delivery.Lead incident response efforts, conduct root cause analysis, and implement long-term solutions.Partner with engineering teams to optimize performance, reliability, scalability, and cloud costs.Promote operational best practices across infrastructure and application environments.Develop and maintain disaster recovery and business continuity capabilities.
Qualifications
Bachelor's degree in Computer Science or a related field, or equivalent professional experience.Advanced degree in Computer Science or a related discipline is preferred.
Required Technical Skills
AWS cloud infrastructure and servicesKubernetes and container orchestration platformsInfrastructure as Code (Terraform or similar tools)CI/CD and deployment automationGit and modern version control practicesContainerization technologies (Docker)Monitoring, logging, and observability platformsScripting and programming experienceDatabase administration and managementIncident and problem managementSecurity, compliance, and cloud governance
Preferred Experience
Experience with enterprise monitoring and observability platformsCloud networking and security best practicesDisaster recovery and resilience planningWorkflow automation and orchestration toolsExperience in insurance, financial services, or other regulated industriesAWS certifications or equivalent cloud certifications preferred
Desired Skills and Experience
Infrastructure as a Code
AWS
Git
SQL
Python
Kubernetes
Datadog
Scripting
DevOps
Site Reliability Engineer
Cloud