Summary
A DevOps engineering role responsible for reliable, scalable, and secure production infrastructure serving high request volumes across multiple regions. You will automate infrastructure provisioning, strengthen CI/CD, build observability systems, enforce cloud security practices, and participate in on-call incident response and post-mortems.
Highlights
Own reliability, scalability, and security for high-throughput production infrastructure across multiple cloud regions. The role offers substantial work in infrastructure automation, CI/CD, observability, security, and incident management.
Description
About The Role
The role owns the reliability, scalability, and security of production infrastructure across multi-region cloud environments, serving millions of high-throughput requests daily.
The team works alongside software engineering squads to build and maintain robust CI/CD pipelines, automated infrastructure provisioning, and proactive observability platforms.
Key Responsibilities
Design, build, and maintain production infrastructure using Terraform and Ansible within AWS and Kubernetes ecosystemsOptimize and scale CI/CD pipelines using GitHub Actions or GitLab CI to ensure rapid, safe, and automated deploymentsImplement comprehensive observability and monitoring frameworks using Prometheus, Grafana, and Datadog for proactive alertingEnforce security best practices, IAM policies, secret management, and compliance standards across all cloud resourcesParticipate in an on-call rotation to troubleshoot and resolve production incidents, conducting rigorous post-mortem analysesWrite clean, maintainable infrastructure-as-code and automate repetitive operational workflows using Python or Go
What We Are Looking For
3โ6 years of experience in DevOps, Site Reliability Engineering, or platform engineering roles within high-growth tech environmentsDeep expertise in Kubernetes, Docker, and container orchestration at scaleHands-on experience with Infrastructure as Code (IaC) tools, specifically Terraform and CloudFormationStrong proficiency in scripting languages such as Python, Bash, or Go for automation tasksSolid understanding of networking fundamentals, TCP/IP, DNS, TLS, VPC architecture, and load balancingBonus: Experience with service mesh technologies (Istio, Linkerd), eBPF, or holding relevant AWS/Kubernetes certifications