DevOps Engineer III

Enable Dental — United States · Posted ~5 hours ago

Senior

Skills

DevOps engineering Amazon Web Services (AWS) Docker Kubernetes Amazon EKS Terraform Infrastructure as Code GitHub Actions CI/CD pipeline management Cloud infrastructure management Container orchestration Cloud security Infrastructure automation High availability engineering Incident prevention Deployment automation AWS CI/CD

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An engineering organization is seeking a skilled DevOps professional to build and maintain secure, reliable cloud infrastructure and automated delivery pipelines. You will manage containerized workloads, implement infrastructure as code, improve deployment workflows, and support application teams. The role offers meaningful ownership of operational reliability, security, and engineering efficiency in a data-sensitive industry.

Highlights

Take ownership of scalable cloud infrastructure and modern deployment pipelines. Work independently with engineering teams, improve reliability and security, automate repetitive tasks, and help deliver resilient services supporting important healthcare operations.

Description

Role Overview We are looking for a skilled DevOps Engineer to help build, maintain, and scale our cloud infrastructure and deployment pipelines. In this role, you will be the backbone of our engineering delivery, ensuring that our applications are highly available, secure, and deployable at a moment's notice. You will take ownership of our containerized environments using Docker and Kubernetes (EKS), manage our cloud resources in AWS as code with Terraform, and streamline our CI/CD workflows using GitHub Actions. Our platform supports healthcare operations, so protecting patient data is part of everything we build. You will work independently while partnering closely with our engineers, supporting both the applications we build and a growing set of self-hosted open-source tools. If you are passionate about automating manual processes, eliminating downtime, and building resilient systems that never depend on a single person, this is the role for you. Key Responsibilities Cloud Infrastructure & ArchitectureAWS Management: Provision, configure, and maintain scalable cloud infrastructure across various AWS services (e.g., EC2, RDS, S3, VPC, IAM)Infrastructure as Code: Own our Terraform and Terragrunt codebase across AWS, Cloudflare and GitHub. Keep environments reproducible, secure, version-controlled and changed only through reviewed pull requestsAccount and environment structure: Plan and carry out changes to our AWS account structure, including moving selected workloads and data between accounts with no data loss and minimal downtimeSecurity & Compliance: Enforce security best practices across our cloud environments, including access control, network security and encryption. Help keep our infrastructure HIPAA compliant and support SOC 2 work Containerization & OrchestrationKubernetes Administration: Deploy, manage, upgrade and scale our EKS clusters. Ensure high availability, proper resource allocation and good performance for our servicesDockerization: Work closely with software engineers to containerize applications, optimizing images for size, security and build speedSelf-hosted open-source applications: Run, upgrade and secure the open-source tools we host ourselves (for example, CRM, analytics, workflow automation and internal databases). Help engineering decide when to self-host and when to use a managed service CI/CD & AutomationPipeline Development: Design, build, and maintain robust Continuous Integration and Continuous Deployment (CI/CD) pipelines using GitHub Actions, including self-hosted runnersRelease Engineering: Automate testing, staging, and production deployments to ensure smooth, zero-downtime releasesProcess Automation: Identify bottlenecks in the development lifecycle and write scripts to automate repetitive operational tasks Monitoring, Logging & ReliabilityObservability: Run our monitoring, alerting, logging and tracing so we can see system health and performance, while keeping PHI out of telemetryBackups and recovery: Own database backups and disaster recovery, and regularly test that restores actually workIncident Response: Participate in an on-call rotation to troubleshoot and resolve production issues, conducting blameless root cause analyses (RCAs) to prevent recurrence Collaboration and shared ownershipPartner with engineering: Work independently day to day, in close collaboration with our engineers on architecture, reviews and prioritiesAvoid single points of failure: Keep runbooks and documentation current, train at least one backup for each critical system, and make sure there are break-glass access proceduresEnable self-service: Give engineers safe, scoped ways to deploy and operate their own services without waiting on you Requirements Qualifications Experience: 5+ years of hands-on experience in a DevOps, Site Reliability (SRE) or Cloud Engineering role, including owning production infrastructureInfrastructure as Code: Production experience managing infrastructure with Terraform (Terragrunt, Pulumi or CloudFormation experience also counts). This includes state management, imports, refactoring without downtime, and reviewing plans as part of pull requestsCloud expertise: Strong production experience running infrastructure in AWS, including IAM, networking and managed databasesContainer orchestration: Deep understanding of Docker, and production experience running Kubernetes (EKS preferred), including Helm, networking, upgrades and scalingCI/CD: A track record of building complex, automated pipelines in GitHub Actions, including secure cloud authentication (OIDC)Networking fundamentals: Solid understanding of cloud networking: DNS, load balancing, VPCs, subnets, security groups and private connectivityScripting: Strong Python or Bash skills for automating operational workIndependent ownership: A history of being the primary owner of infrastructure while working closely with application engineers. You document your work, share knowledge and design things so no system depends on you alone Preferred Skills Healthcare and HIPAA: Experience running infrastructure that handles PHI in a HIPAA-regulated environment. That includes encryption, access logging, retention and working with vendors under BAAsSelf-hosted open source: Experience running third-party open-source apps in production, including version pinning, upgrades with database migrations, and rollbacksMulti-account AWS: Experience with AWS Organizations, or with moving workloads and data between AWS accounts Bonus Skills SOC 2: Experience preparing for or supporting a SOC 2 audit, especially automating evidence collectionObservability tools: Experience with the Grafana stack (Loki, Tempo, Mimir), Prometheus, OpenTelemetry or similar