Summary
✨ AI‑Generated
A DevOps Engineer will build and operate reliable, secure production infrastructure across cloud platforms and containerized environments. The role covers Terraform-managed infrastructure, Kubernetes, automated CI/CD, observability, and incident response, with close collaboration with software engineering and reliability teams.
Highlights
Own production infrastructure across cloud, containers, deployment automation, observability, and reliability, with strong opportunities to turn operational challenges into durable platform improvements.
Description
About The Role
The DevOps Engineer builds and operates the infrastructure that keeps production services reliable, secure, and responsive as traffic and system complexity grow.
The role spans cloud platforms, container orchestration, deployment automation, observability, and incident response across development and production environments.
You will partner with software engineers and SREs to improve release velocity without sacrificing stability.
The team needs an engineer who can turn recurring operational problems into durable platform improvements through infrastructure as code, automated delivery pipelines, and measurable reliability practices.
Key Responsibilities
Build and maintain AWS infrastructure using Terraform, including VPCs, IAM, EKS, RDS, and supporting managed servicesDesign and improve CI/CD pipelines with GitHub Actions, Argo CD, or similar tools to automate testing, deployment, rollback, and environment provisioningOperate Kubernetes workloads across production environments, improving resource management, scaling, security, and service availabilityImplement observability with Prometheus, Grafana, OpenTelemetry, and centralized logging to provide actionable metrics, traces, alerts, and service dashboardsStrengthen incident response by troubleshooting production failures, leading root-cause analysis, and delivering follow-up remediation workDevelop reusable automation and platform tooling in Python, Go, or Bash to eliminate manual operational tasks and reduce deployment riskPartner with engineering teams on capacity planning, disaster recovery, secrets management, and practical SLOs for critical services
What We Are Looking For
3–8 years of experience in DevOps, SRE, platform engineering, or a closely related infrastructure roleHands-on experience operating production workloads in AWS or another major cloud provider, with strong knowledge of networking, IAM, compute, storage, and managed databasesProficiency with Kubernetes and Docker, including deployments, services, ingress, Helm, troubleshooting, and resource configurationExperience building infrastructure as code with Terraform or an equivalent tool and managing changes through version control and code reviewStrong understanding of CI/CD principles and practical experience with tools such as GitHub Actions, GitLab CI, Jenkins, Argo CD, or SpinnakerWorking knowledge of Linux administration, TCP/IP networking, observability, incident management, and scripting in Python, Go, or BashBachelor’s degree in computer science, engineering, or a related technical field, or equivalent professional experience; Bonus: experience with SRE practices, service mesh technologies, FinOps, compliance automation, and multi-region systems