DevOps Engineer

Evlo Ai — United States · Posted ~2 hours ago

Mid

Skills

AWS Kubernetes CI/CD Terraform Infrastructure as code Incident response Cloud infrastructure Observability GitHub Actions GitLab CI Jenkins ECS EKS RDS S3 CloudFront

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Build and operate reliable infrastructure for production services at scale. You will design highly available cloud environments, manage Kubernetes workloads, improve CI/CD pipelines, implement infrastructure as code, strengthen observability, and participate in incident response. The role emphasizes repeatable deployments, measurable reliability, and safe operational practices.

Highlights

Infrastructure-focused role spanning AWS, Kubernetes, CI/CD, infrastructure as code, observability, and reliability engineering, with strong collaboration across software and security teams.

Description

About The Role The DevOps Engineer will build and operate the infrastructure, deployment systems, and observability platforms that keep production services reliable at scale. The role spans AWS cloud infrastructure, Kubernetes, CI/CD, infrastructure as code, and incident response across distributed systems. The engineer will partner with software development and security teams to improve release velocity without compromising availability or operational safety. Success means repeatable deployments, clear service ownership, measurable reliability, and infrastructure that scales with product demand. Key Responsibilities Design and maintain highly available AWS infrastructure using Terraform, including VPCs, IAM, ECS or EKS, RDS, S3, and CloudFrontBuild and improve CI/CD pipelines with GitHub Actions, GitLab CI, or Jenkins, including automated testing, security scanning, progressive delivery, and rollback proceduresOperate Kubernetes-based workloads, managing deployments, ingress, autoscaling, secrets, resource limits, and cluster upgradesImplement observability using Prometheus, Grafana, CloudWatch, OpenTelemetry, and centralized logging to improve detection and resolution of production issuesAutomate operational workflows with Python, Go, or Bash, reducing manual intervention in provisioning, deployments, incident response, and routine maintenanceDefine and track reliability practices including SLIs, SLOs, error budgets, capacity planning, and disaster recovery testingParticipate in on-call rotations, lead incident response, document root-cause analyses, and deliver follow-up improvements that prevent recurrence What We Are Looking For 3–8 years of experience in DevOps, site reliability engineering, platform engineering, or a closely related infrastructure roleHands-on experience operating production workloads in AWS, with strong knowledge of networking, IAM, compute, storage, databases, and security controlsProficiency with Terraform or an equivalent infrastructure-as-code tool and practical experience managing infrastructure through version-controlled workflowsExperience with Docker and Kubernetes, including workload scheduling, service discovery, ingress, autoscaling, and troubleshootingStrong understanding of CI/CD design, Git-based development workflows, automated testing, release strategies, and deployment rollback patternsBachelor’s degree in computer science, engineering, information systems, or equivalent practical experienceBonus: Experience with Go, Argo CD, Helm, service mesh technologies, OpenTelemetry, PostgreSQL operations, compliance automation, or multi-region disaster recovery