Summary
✨ AI‑Generated
Build and operate reliable infrastructure for production services at scale. You will design highly available cloud environments, manage Kubernetes workloads, improve CI/CD pipelines, implement infrastructure as code, strengthen observability, and participate in incident response. The role emphasizes repeatable deployments, measurable reliability, and safe operational practices.
Highlights
Infrastructure-focused role spanning AWS, Kubernetes, CI/CD, infrastructure as code, observability, and reliability engineering, with strong collaboration across software and security teams.
Description
About The Role
The DevOps Engineer will build and operate the infrastructure, deployment systems, and observability platforms that keep production services reliable at scale.
The role spans AWS cloud infrastructure, Kubernetes, CI/CD, infrastructure as code, and incident response across distributed systems.
The engineer will partner with software development and security teams to improve release velocity without compromising availability or operational safety.
Success means repeatable deployments, clear service ownership, measurable reliability, and infrastructure that scales with product demand.
Key Responsibilities
Design and maintain highly available AWS infrastructure using Terraform, including VPCs, IAM, ECS or EKS, RDS, S3, and CloudFrontBuild and improve CI/CD pipelines with GitHub Actions, GitLab CI, or Jenkins, including automated testing, security scanning, progressive delivery, and rollback proceduresOperate Kubernetes-based workloads, managing deployments, ingress, autoscaling, secrets, resource limits, and cluster upgradesImplement observability using Prometheus, Grafana, CloudWatch, OpenTelemetry, and centralized logging to improve detection and resolution of production issuesAutomate operational workflows with Python, Go, or Bash, reducing manual intervention in provisioning, deployments, incident response, and routine maintenanceDefine and track reliability practices including SLIs, SLOs, error budgets, capacity planning, and disaster recovery testingParticipate in on-call rotations, lead incident response, document root-cause analyses, and deliver follow-up improvements that prevent recurrence
What We Are Looking For
3–8 years of experience in DevOps, site reliability engineering, platform engineering, or a closely related infrastructure roleHands-on experience operating production workloads in AWS, with strong knowledge of networking, IAM, compute, storage, databases, and security controlsProficiency with Terraform or an equivalent infrastructure-as-code tool and practical experience managing infrastructure through version-controlled workflowsExperience with Docker and Kubernetes, including workload scheduling, service discovery, ingress, autoscaling, and troubleshootingStrong understanding of CI/CD design, Git-based development workflows, automated testing, release strategies, and deployment rollback patternsBachelor’s degree in computer science, engineering, information systems, or equivalent practical experienceBonus: Experience with Go, Argo CD, Helm, service mesh technologies, OpenTelemetry, PostgreSQL operations, compliance automation, or multi-region disaster recovery