Summary
✨ AI‑Generated
A lead-level DevOps role responsible for building reliable cloud platforms, automation pipelines, monitoring systems, and security practices. The position requires deep cloud expertise and leadership across infrastructure operations.
Highlights
Leadership opportunity to build reliability practices, own cloud infrastructure improvements, and drive automation in a modern engineering environment.
Description
This organization is seeking a Lead DevOps Engineer to help build and lead its Site Reliability Engineering function as it scales toward production.
This is a full-time hybrid role based in Wilmington, NC, with onsite expectations Tuesday through Thursday.
The infrastructure team supports engineering across all product teams and is building an operating model centered on software, automation, and AI.
The environment includes AWS, Kubernetes, Terraform/Terragrunt, Argo CD, GitHub Actions, Datadog, and cloud security tooling, with significant ownership over reliability, monitoring, disaster recovery, and cost optimization.
Required Skills & Experience
Deep hands-on AWS experience, including EKS, Fargate, and Aurora Terraform and Terragrunt Lead-level Kubernetes expertise with strong understanding of core fundamentals Argo CD and Argo Workflows GitHub Actions for CI/CD pipeline development and hardening Datadog monitoring, alerting, dashboarding, and incident response AWS security services including GuardDuty and Security Hub IAM roles and permissions best practices Networking fundamentals including VPCs, subnets, and routing Experience mentoring engineers, assigning work, and reviewing pull requests Hands-on infrastructure design and implementation experience
Desired Skills & Experience
Cloudflare Cilium Familiarity with Java build tooling
What You Will Be Doing
Build and lead the Site Reliability Engineering function Mentor full-time SREs and contractors Assign work and review pull requests Design and build cloud infrastructure Develop and maintain Kubernetes-based platforms Manage CI/CD pipelines with GitHub Actions Own monitoring, alerting, dashboarding, and incident response processes in Datadog Conduct root cause analysis for production incidents Lead disaster recovery planning initiatives Drive cloud cost optimization efforts Implement and maintain cloud security best practices
Tech Breakdown [Technology percentages not provided]
AWS (EKS, Fargate, Aurora) Terraform Terragrunt Kubernetes Argo CD Argo Workflows GitHub Actions Datadog GuardDuty Security Hub IAM VPCs, Subnets, Routing Cloudflare (preferred) Cilium (preferred)
Daily Responsibilities [Responsibility percentages not provided]
Lead and mentor SRE team members and contractors Design and implement cloud infrastructure Manage Kubernetes environments and related tooling Develop and harden CI/CD pipelines Monitor platform health and respond to incidents Perform root cause analysis and reliability improvements Plan and maintain disaster recovery strategies Optimize cloud spend and resource utilization Support security initiatives and access management practices
The Offer
Full-time employment Strong benefits package PTO offered Hybrid work arrangement (Tuesday through Thursday onsite) Location: Wilmington, NC
Posted By: John Gordon