Summary
✨ AI‑Generated
Design and operate the cloud infrastructure and deployment systems that keep production services reliable at scale. You will build AWS infrastructure with Terraform, operate Kubernetes environments, improve CI/CD automation, strengthen observability, and participate in incident response. The role combines infrastructure engineering with close partnership across software teams to improve delivery speed and reliability.
Highlights
Broad DevOps ownership across AWS infrastructure, Kubernetes, infrastructure as code, CI/CD, observability, and incident response. The role directly influences deployment velocity, recovery time, system performance, availability, and engineering efficiency.
Description
About The Role
The DevOps Engineer designs and operates the infrastructure, deployment systems, and observability platform that keep production services reliable at scale.
The role spans AWS cloud infrastructure, Kubernetes, infrastructure as code, CI/CD automation, and incident response across distributed systems.
This engineer will partner with software teams to improve release velocity without compromising availability or security.
The work directly impacts deployment frequency, recovery time, system performance, and the day-to-day operating efficiency of the engineering organization.
Key Responsibilities
Build and maintain AWS infrastructure using Terraform, including VPCs, IAM, ECS or EKS, RDS, S3, CloudFront, and managed messaging servicesOperate Kubernetes-based production environments, improving cluster reliability, resource utilization, deployment safety, and horizontal scalingDevelop and optimize CI/CD pipelines with GitHub Actions, GitLab CI, or Jenkins for automated testing, container builds, progressive delivery, and rollbacksImplement observability using Prometheus, Grafana, CloudWatch, OpenTelemetry, or equivalent tools to monitor service health, latency, capacity, and error ratesAutomate operational workflows with Python, Go, or Bash, reducing manual intervention in provisioning, incident response, and routine maintenanceLead or participate in incident response, including troubleshooting production failures, coordinating remediation, and documenting blameless post-incident reviewsPartner with engineering and security teams to establish platform standards for secrets management, access control, vulnerability remediation, disaster recovery, and compliance
What We Are Looking For
3–8 years of experience in DevOps, site reliability engineering, platform engineering, or a closely related infrastructure roleStrong hands-on experience with AWS and production infrastructure spanning networking, IAM, compute, storage, databases, and security controlsProficiency with Terraform or an equivalent infrastructure-as-code tool, including reusable modules, state management, and code review practicesProduction experience operating Kubernetes and Docker, including deployments, services, ingress, autoscaling, troubleshooting, and resource managementDemonstrated ability to build CI/CD pipelines and release automation using GitHub Actions, GitLab CI, Jenkins, Argo CD, or comparable toolsWorking knowledge of Linux systems, networking fundamentals, observability practices, and scripting with Python, Go, or BashBonus: Experience with service meshes, GitOps, Helm, OpenTelemetry, Kafka, multi-region architecture, SOC 2 controls, or formal computer science and engineering education