Summary
✨ AI‑Generated
A DevOps Engineer role responsible for building and operating production infrastructure at scale. You will manage AWS resources with Terraform, operate Kubernetes across multiple environments, automate CI/CD using GitHub Actions, and improve observability, security, reliability, and deployment processes. The role suits an engineer comfortable troubleshooting complex production issues and reducing operational work through automation.
Highlights
Infrastructure engineering role focused on cloud platforms, deployment automation, observability, and resilient production systems. Offers hands-on ownership of AWS infrastructure, Kubernetes environments, CI/CD automation, reliability improvements, and collaboration with engineering and security teams.
Description
About The Role
The DevOps Engineer builds and operates the infrastructure that runs production services at scale.
The role focuses on cloud platforms, deployment automation, observability, and resilient systems across environments using tools such as AWS, Kubernetes, Terraform, and GitHub Actions.
This role partners with software engineers and security teams to make releases safer, reduce operational toil, and improve system reliability.
The team needs an engineer who can troubleshoot complex production issues, define practical infrastructure standards, and strengthen the platform through automation and measurable service-level objectives.
Key Responsibilities
Build and maintain AWS infrastructure using Terraform, including networking, IAM, compute, storage, and managed database servicesOperate Kubernetes workloads across development, staging, and production environments; manage deployments, scaling, ingress, secrets, and resource policiesAutomate CI/CD pipelines with GitHub Actions, Argo CD, or equivalent tools, including testing gates, environment promotion, and rollback proceduresImplement observability with Prometheus, Grafana, CloudWatch, and centralized logging to track service health, latency, capacity, and error ratesDefine and improve reliability practices, including SLOs, alerting standards, incident response procedures, and post-incident corrective actionsPartner with application teams to improve containerization, release workflows, configuration management, and production readinessTroubleshoot infrastructure and distributed-system failures, document root causes, and deliver durable fixes rather than manual workarounds
What We Are Looking For
3–8 years of experience in DevOps, site reliability engineering, platform engineering, or infrastructure engineeringHands-on experience managing AWS production environments and writing maintainable Terraform modulesStrong Kubernetes and Docker experience, including debugging workloads, networking issues, deployments, and resource constraintsProficiency with Linux, Bash or Python, Git, and CI/CD systems such as GitHub Actions, GitLab CI, Jenkins, or Argo CDWorking knowledge of observability practices and tools, including metrics, logs, tracing, alert design, and incident responseBachelor’s degree in computer science, engineering, information systems, or a related technical field; equivalent practical experience is also consideredBonus: Experience with Helm, Argo CD, service meshes, PostgreSQL or Redis operations, compliance controls, and building internal developer platforms