Summary
✨ AI‑Generated
Join a dynamic team seeking a Senior Platform DevOps Engineer to manage day-to-day operations, deployments, and infrastructure automation in an AWS-first Kubernetes environment. This remote role is available to qualified candidates in Colombia or Costa Rica, offering the opportunity to work with complex, highly available workloads and collaborate closely with engineering teams.
Highlights
Remote position in Colombia or Costa Rica, full-time role with ownership of platform operations, deployments, and infrastructure automation in an AWS-first Kubernetes environment.
Description
- This position is open to candidates located in Colombia or Costa Rica only -
Gorilla Logic is looking for a Senior Platform / DevOps Engineer with strong hands-on experience in Kubernetes, Terraform, and Python to join our team and support a production platform running complex, highly available workloads.
In this role, you will take ownership of day-to-day platform operations, deployments, infrastructure automation, reliability, and production support.
You will work in an AWS-first Kubernetes environment and collaborate closely with engineering teams to ensure systems are scalable, reliable, secure, and maintainable.
We are looking for someone who combines strong technical expertise with ownership, sound engineering judgment, and a quality-first mindset.
The ideal candidate is comfortable challenging decisions when necessary, protecting engineering standards, and making thoughtful trade-offs rather than sacrificing long-term quality for short-term speed.
What you'll do
Manage, operate, and troubleshoot production Kubernetes environments and workloads.Build, maintain, and improve infrastructure using Terraform and Infrastructure as Code practices.Manage application deployments, upgrades, configuration changes, and complex deployment lifecycles.Work with GitOps-based deployment processes and tools such as Argo CD.Develop and maintain Python scripts and tooling to support platform automation and operational workflows.Troubleshoot and resolve production incidents across infrastructure, applications, and platform services.Improve platform reliability, scalability, observability, and operational efficiency.Support Kubernetes scaling and autoscaling strategies for production workloads.Collaborate closely with development and platform teams to identify and resolve infrastructure and deployment challenges.Participate in technical decisions and proactively identify risks, reliability concerns, and opportunities for improvement.Maintain high engineering and quality standards, providing technical pushback when necessary to ensure reliable and maintainable solutions.Take ownership of platform initiatives and drive issues through resolution with minimal supervision.
Required Qualifications
Strong hands-on experience managing Kubernetes in production environments.Experience with Kubernetes deployments, scaling, troubleshooting, and operational management.Hands-on experience with Terraform for provisioning and managing cloud infrastructure.Ability to read and write Python for scripting, automation, troubleshooting, and platform tooling.Experience with AWS cloud infrastructure, ideally including EKS or similar managed Kubernetes environments.Experience with GitOps practices and deployment tools such as Argo CD.Experience managing complex application deployment and upgrade lifecycles.Proven experience troubleshooting, triaging, and supporting production incidents.Strong understanding of infrastructure reliability, scalability, and operational best practices.Strong problem-solving skills and the ability to independently investigate complex production issues.Strong sense of ownership and accountability, with the ability to operate effectively with limited supervision.Quality-first mindset with the judgment to balance delivery speed, reliability, and long-term maintainability.Strong communication and collaboration skills.
Preferred Qualifications
Experience with Helm and Kubernetes package/deployment management.Familiarity with PyTorch and Hugging Face Transformers.Experience supporting GPU-based workloads or ML inference platforms.Familiarity with NVIDIA Triton Inference Server.Experience with Chainguard, distroless container images, Trivy, or container vulnerability reduction.Experience implementing or improving Kubernetes autoscaling solutions.Familiarity with streaming or messaging platforms such as Apache Kafka or similar technologies.Experience with Elasticsearch or ArangoDB.Experience troubleshooting complex service-to-service networking.Exposure to OpenShift, IL5, FedRAMP, or similarly constrained environments.Familiarity with AI/ML or agentic AI development environments.