Summary
✨ AI‑Generated
A DevOps Engineer is sought to own the infrastructure powering high-traffic production systems. You’ll build CI/CD pipelines, operate and scale Kubernetes clusters in major cloud environments, manage infrastructure as code, and develop observability systems with actionable monitoring and alerting. The role offers significant ownership across deployment, scaling, automation, and incident response.
Highlights
Own production infrastructure in a high-traffic environment, work with modern cloud and Kubernetes technologies, automate deployments and incident response, and shape observability and infrastructure practices at scale.
Description
About The Role
The role owns the infrastructure that keeps production systems online - CI/CD pipelines, Kubernetes clusters, cloud networking, and the observability stack that detects problems before customers do.
The engineer will work closely with backend and platform teams to automate everything: deployments, scaling, incident response, and infrastructure provisioning across a multi-service, high-traffic environment.
Key Responsibilities
Build and maintain CI/CD pipelines (GitHub Actions, GitLab CI, or CircleCI) enabling safe, frequent deployments across dozens of microservicesOperate and scale Kubernetes clusters on AWS or GCP, including cluster upgrades, autoscaling policies, and workload optimizationDefine infrastructure as code using Terraform and Terragrunt, with peer-reviewed modules and automated plan/apply workflowsDesign and maintain observability tooling - Prometheus, Grafana, and distributed tracing - with actionable SLOs and alerting that minimizes noiseLead incident response and blameless postmortems; drive reduction of MTTR through runbooks, automation, and chaos testingHarden production environments: IAM policies, network segmentation, secrets management, and security patching pipelinesPartner with development teams to improve service reliability, defining error budgets and capacity plans as the platform scales
What We Are Looking For
3+ years of experience in DevOps, SRE, or infrastructure engineering, with production ownership of systems serving significant trafficDeep hands-on experience with Kubernetes: deployment, networking, Helm, and troubleshooting in production environmentsStrong Terraform/IaC skills and proficiency with at least one major cloud provider (AWS, GCP, or Azure)Solid scripting and automation skills in Python, Bash, or GoExperience building observability from the ground up: metrics, logs, traces, and alert design (Prometheus, Datadog, or similar)Bachelor's degree in Computer Science or equivalent practical experienceBonus: Experience with service mesh (Istio/Linkerd), GitOps tooling (ArgoCD/Flux), multi-region failover design, or compliance environments (SOC 2, PCI)