DevOps Engineer

Evlo Ai — United States · Posted ~3 hours ago

Skills

AWS Kubernetes Terraform GitHub Actions CI/CD Infrastructure automation Cloud infrastructure Observability Production troubleshooting Site reliability Infrastructure security IAM

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A DevOps Engineer role responsible for building and operating production infrastructure at scale. You will manage AWS resources with Terraform, operate Kubernetes across multiple environments, automate CI/CD using GitHub Actions, and improve observability, security, reliability, and deployment processes. The role suits an engineer comfortable troubleshooting complex production issues and reducing operational work through automation.

Highlights

Infrastructure engineering role focused on cloud platforms, deployment automation, observability, and resilient production systems. Offers hands-on ownership of AWS infrastructure, Kubernetes environments, CI/CD automation, reliability improvements, and collaboration with engineering and security teams.

Description

About The Role The DevOps Engineer builds and operates the infrastructure that runs production services at scale. The role focuses on cloud platforms, deployment automation, observability, and resilient systems across environments using tools such as AWS, Kubernetes, Terraform, and GitHub Actions. This role partners with software engineers and security teams to make releases safer, reduce operational toil, and improve system reliability. The team needs an engineer who can troubleshoot complex production issues, define practical infrastructure standards, and strengthen the platform through automation and measurable service-level objectives. Key Responsibilities Build and maintain AWS infrastructure using Terraform, including networking, IAM, compute, storage, and managed database servicesOperate Kubernetes workloads across development, staging, and production environments; manage deployments, scaling, ingress, secrets, and resource policiesAutomate CI/CD pipelines with GitHub Actions, Argo CD, or equivalent tools, including testing gates, environment promotion, and rollback proceduresImplement observability with Prometheus, Grafana, CloudWatch, and centralized logging to track service health, latency, capacity, and error ratesDefine and improve reliability practices, including SLOs, alerting standards, incident response procedures, and post-incident corrective actionsPartner with application teams to improve containerization, release workflows, configuration management, and production readinessTroubleshoot infrastructure and distributed-system failures, document root causes, and deliver durable fixes rather than manual workarounds What We Are Looking For 3–8 years of experience in DevOps, site reliability engineering, platform engineering, or infrastructure engineeringHands-on experience managing AWS production environments and writing maintainable Terraform modulesStrong Kubernetes and Docker experience, including debugging workloads, networking issues, deployments, and resource constraintsProficiency with Linux, Bash or Python, Git, and CI/CD systems such as GitHub Actions, GitLab CI, Jenkins, or Argo CDWorking knowledge of observability practices and tools, including metrics, logs, tracing, alert design, and incident responseBachelor’s degree in computer science, engineering, information systems, or a related technical field; equivalent practical experience is also consideredBonus: Experience with Helm, Argo CD, service meshes, PostgreSQL or Redis operations, compliance controls, and building internal developer platforms