Summary
✨ AI‑Generated
A DevOps engineering role responsible for infrastructure, deployment systems, and observability supporting reliable production services at scale. You will work across AWS architecture, Kubernetes operations, infrastructure as code, CI/CD, and incident response. The position involves improving deployment safety, system performance, and operational automation while troubleshooting distributed systems in high-availability environments.
Highlights
Opportunity to build and operate highly available cloud infrastructure, improve deployment safety and system performance, automate operational work, and solve challenging distributed-system reliability problems.
Description
About The Role
The DevOps Engineer will build and operate the infrastructure, deployment systems, and observability platform that keep production services reliable at scale.
The role spans cloud architecture, Kubernetes operations, infrastructure as code, CI/CD, and incident response across a high-availability environment.
You will partner with application engineers and SREs to reduce deployment risk, improve system performance, and automate repetitive operational work.
The team values engineers who can troubleshoot distributed systems under pressure and turn recurring production issues into durable platform improvements.
Key Responsibilities
Design and maintain highly available AWS infrastructure using Terraform, Kubernetes, Docker, and managed cloud servicesBuild and improve CI/CD pipelines with GitHub Actions, GitLab CI, or Jenkins to support safe, repeatable application releasesOperate production Kubernetes clusters, including capacity planning, workload scheduling, upgrades, access controls, and disaster recovery proceduresImplement observability standards using Prometheus, Grafana, OpenTelemetry, and centralized logging to track service health and performanceAutomate operational workflows with Python, Go, or Bash, eliminating manual provisioning, configuration, and remediation tasksParticipate in on-call rotations and lead technical incident response, including root-cause analysis, postmortems, and corrective action trackingCollaborate with engineering teams to define reliability targets, improve deployment practices, and strengthen security across the software delivery lifecycle
What We Are Looking For
3–8 years of experience in DevOps, SRE, platform engineering, or a closely related infrastructure roleStrong hands-on experience with AWS, Kubernetes, Docker, and Linux in production environmentsProficiency with infrastructure as code, especially Terraform, and configuration or policy automation tools such as Helm, Ansible, or Argo CDExperience designing and operating CI/CD pipelines with automated testing, deployment approvals, rollback strategies, and release observabilityWorking knowledge of networking, DNS, TLS, IAM, containers, distributed systems, and cloud security fundamentalsProficiency in at least one scripting or systems programming language, such as Python, Go, or Bash, with a focus on maintainable automationBachelor’s degree in computer science, engineering, or a related technical field, or equivalent practical experience; Bonus: experience with multi-region systems, service mesh technologies, GitOps, incident management, or compliance-focused infrastructure