DevOps Engineer

Evlo Ai — United States · Posted ~1 day ago

Mid Full-time

Skills

DevOps Kubernetes CI/CD Cloud Networking Infrastructure as Code Terraform Helm Observability Infrastructure Scaling Production Reliability GitHub Actions GitLab CI

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A hands-on DevOps engineering role on a small platform team responsible for infrastructure reliability at scale. You will manage Kubernetes clusters, improve CI/CD pipelines, own Infrastructure-as-Code workflows using Terraform and Helm, and develop observability and cloud networking capabilities. Your work will support production uptime, safer deployments, and infrastructure cost efficiency.

Highlights

Take meaningful ownership of production infrastructure and reliability in a small platform team. Build tools used across engineering, improve deployment speed and safety, and directly influence system uptime and infrastructure costs while working closely with product engineers.

Description

About The Role The role owns the infrastructure that keeps production systems running reliably at scale - CI/CD pipelines, Kubernetes clusters, cloud networking, and the observability stack that surfaces problems before customers notice them. This is a hands-on position on a small platform team working directly with product engineers. Expect real ownership: the tooling built here is used by every engineering team, and reliability decisions made in this role directly affect uptime, deploy velocity, and infrastructure cost. Key Responsibilities Build, maintain, and scale Kubernetes-based infrastructure across staging and production environments, including cluster upgrades, autoscaling policies, and node optimizationDesign and improve CI/CD pipelines using GitHub Actions or GitLab CI, with a focus on build speed, deployment safety, and automated rollbackOwn Infrastructure-as-Code using Terraform and Helm; manage state, module reuse, and drift detection across cloud accountsBuild and maintain observability tooling with Prometheus, Grafana, and OpenTelemetry; define actionable SLOs and on-call alerting that minimizes noiseLead incident response and blameless postmortems; drive root-cause fixes and resilience patterns (circuit breakers, graceful degradation, multi-AZ failover)Optimize cloud spend on AWS or GCP - right-sizing instances, storage tiering, and reserved capacity planning - without sacrificing performance or reliabilityPartner with application teams on service design reviews, load testing, and production readiness for high-traffic launches What We Are Looking For 3–7 years of experience in DevOps, SRE, or infrastructure engineering, including running production systems serving meaningful trafficDeep hands-on expertise with Kubernetes and containers in production - not just cluster setup, but day-2 operations: upgrades, debugging, and incident responseStrong Infrastructure-as-Code skills with Terraform (or Pulumi); experience managing multi-environment cloud infrastructure on AWS or GCPProficiency with Linux systems, networking fundamentals (DNS, TLS, load balancing, VPC design), and bash/Python scriptingExperience building or maintaining CI/CD pipelines and GitOps workflows (ArgoCD, Flux, or similar)Track record of improving measurable reliability outcomes: MTTR, error budgets, on-call burden reduction, or cost optimizationBachelor's degree in Computer Science, Engineering, or equivalent practical experience. Bonus: experience with service meshes (Istio, Linkerd), secrets management (Vault), multi-region architectures, or SOC 2 / compliance-related infrastructure work