DevOps Engineer

Evlo Ai — United States · Posted ~13 hours ago

Senior

Skills

CI/CD Kubernetes AWS GCP Terraform Terragrunt Infrastructure as Code Observability Prometheus Grafana Distributed tracing Incident response GitHub Actions GitLab CI CircleCI

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A DevOps Engineer is sought to own the infrastructure powering high-traffic production systems. You’ll build CI/CD pipelines, operate and scale Kubernetes clusters in major cloud environments, manage infrastructure as code, and develop observability systems with actionable monitoring and alerting. The role offers significant ownership across deployment, scaling, automation, and incident response.

Highlights

Own production infrastructure in a high-traffic environment, work with modern cloud and Kubernetes technologies, automate deployments and incident response, and shape observability and infrastructure practices at scale.

Description

About The Role The role owns the infrastructure that keeps production systems online - CI/CD pipelines, Kubernetes clusters, cloud networking, and the observability stack that detects problems before customers do. The engineer will work closely with backend and platform teams to automate everything: deployments, scaling, incident response, and infrastructure provisioning across a multi-service, high-traffic environment. Key Responsibilities Build and maintain CI/CD pipelines (GitHub Actions, GitLab CI, or CircleCI) enabling safe, frequent deployments across dozens of microservicesOperate and scale Kubernetes clusters on AWS or GCP, including cluster upgrades, autoscaling policies, and workload optimizationDefine infrastructure as code using Terraform and Terragrunt, with peer-reviewed modules and automated plan/apply workflowsDesign and maintain observability tooling - Prometheus, Grafana, and distributed tracing - with actionable SLOs and alerting that minimizes noiseLead incident response and blameless postmortems; drive reduction of MTTR through runbooks, automation, and chaos testingHarden production environments: IAM policies, network segmentation, secrets management, and security patching pipelinesPartner with development teams to improve service reliability, defining error budgets and capacity plans as the platform scales What We Are Looking For 3+ years of experience in DevOps, SRE, or infrastructure engineering, with production ownership of systems serving significant trafficDeep hands-on experience with Kubernetes: deployment, networking, Helm, and troubleshooting in production environmentsStrong Terraform/IaC skills and proficiency with at least one major cloud provider (AWS, GCP, or Azure)Solid scripting and automation skills in Python, Bash, or GoExperience building observability from the ground up: metrics, logs, traces, and alert design (Prometheus, Datadog, or similar)Bachelor's degree in Computer Science or equivalent practical experienceBonus: Experience with service mesh (Istio/Linkerd), GitOps tooling (ArgoCD/Flux), multi-region failover design, or compliance environments (SOC 2, PCI)