Summary
✨ AI‑Generated
A DevOps Engineer is needed to build and operate production infrastructure at scale, spanning cloud infrastructure, Kubernetes, CI/CD, observability and incident response. You’ll use infrastructure as code to eliminate manual work, create reusable platforms, improve deployment safety, and strengthen production resilience while working closely with engineering and reliability teams.
Highlights
Build production infrastructure at scale with a strong focus on automation, reliability and deployment safety, while partnering closely with software engineers and SREs to improve delivery and resilience.
Description
About The Role
The DevOps Engineer builds and operates the infrastructure that runs production services at scale, with a focus on automation, availability, deployment safety, and measurable system performance.
The role spans cloud infrastructure, Kubernetes, CI/CD, observability, and incident response across development and production environments.
You will partner with software engineers and SREs to improve delivery velocity without compromising reliability.
The team needs an engineer who can turn operational requirements into reusable platforms, eliminate manual work through infrastructure as code, and lead practical improvements to production resilience.
Key Responsibilities
Design, provision, and maintain AWS infrastructure using Terraform, including VPCs, IAM, EKS, RDS, S3, and related production servicesBuild and improve CI/CD pipelines with GitHub Actions, GitLab CI, or equivalent tools for automated testing, security checks, deployments, and rollbackOperate Kubernetes workloads across development and production environments, including deployments, Helm charts, ingress, autoscaling, secrets, and resource managementImplement observability with Prometheus, Grafana, OpenTelemetry, and centralized logging to track service health, latency, capacity, and error ratesAutomate operational workflows with Python, Go, or Bash, reducing toil and improving consistency across infrastructure and application teamsParticipate in incident response, troubleshoot complex production failures, and produce clear post-incident actions that improve system reliabilityDefine and enforce infrastructure, security, and operational standards through code reviews, documentation, runbooks, and platform tooling
What We Are Looking For
3–8 years of experience in DevOps, site reliability engineering, platform engineering, or production infrastructureHands-on experience operating workloads in AWS or another major cloud provider, with a strong understanding of networking, IAM, compute, storage, and managed databasesProduction experience with Kubernetes and container technologies, including Docker, Helm, service discovery, ingress, and autoscalingProficiency with Terraform or another infrastructure-as-code tool and experience managing infrastructure through version-controlled workflowsPractical experience building CI/CD pipelines and implementing automated testing, deployment strategies, secrets management, and rollback proceduresStrong troubleshooting and communication skills, with experience participating in on-call rotations and resolving production incidentsBachelor’s degree in computer science, engineering, information technology, or a related field, or equivalent professional experience; Bonus: experience with Go, Argo CD, Istio, PostgreSQL, security automation, SLOs, and cost optimization