Summary
✨ AI‑Generated
A senior infrastructure specialist is sought to design, operate, and continuously improve highly available cloud platforms used by engineering teams across multiple regions. The role combines AWS, Linux, networking, security, CI/CD, and operational troubleshooting.
Highlights
Design and operate highly available cloud infrastructure supporting teams across multiple regions, with deep exposure to AWS, Linux, networking, security, CI/CD, and scalable infrastructure engineering.
Description
Position Overview
We are looking for an experienced Senior CloudOps Engineer to join our infrastructure team.
You will be responsible for designing, operating, and continuously improving a highly available cloud platform that supports engineering teams across multiple regions and continents.
This role requires deep technical expertise, a strong operational mindset, and the ability to collaborate closely with software engineers to deliver scalable, secure, and reliable infrastructure.
Technical Skills & Qualifications
• Strong expertise in networking fundamentals, including:
TCP/IP, DNS, HTTP/HTTPSRouting and switchingLoad balancingVPNsFirewalls and security groupsNetwork troubleshooting• Extensive hands-on experience administering and troubleshooting Linux systems in production environments.
• Strong experience with Amazon Web Services (AWS), including designing, deploying, and operating cloud-native infrastructure.
• Experience building, maintaining, and optimizing CI/CD pipelines for automated software delivery.
• Strong knowledge of Kubernetes, preferably Amazon EKS, including cluster administration, deployments, scaling, and troubleshooting.
• Experience using GitHub, including Git workflows, repository management, GitHub Actions, and collaborative development practices.
• Ability to automate operational tasks using scripting languages such as Bash or Python.
• Experience with infrastructure monitoring, alerting, and observability tools.
• Strong troubleshooting skills with the ability to quickly identify root causes across infrastructure, networking, Kubernetes, and application layers.
• Excellent understanding of infrastructure reliability, high availability, disaster recovery, and scalability principles.
• Strong documentation and communication skills.
Skills, that are a plus:
• Experience with Terraform and Infrastructure as Code (IaC).
• Understanding of cloud security best practices, including:
IAMLeast privilege accessSecrets managementEncryptionNetwork securityCompliance concepts• Experience with container security and Kubernetes security best practices.
• Familiarity with observability platforms such as Prometheus, Grafana, CloudWatch, Datadog, or similar tools.
• Experience supporting multi-account AWS environments.
• Experience with incident response, postmortems, and operational excellence practices.
• Experience leveraging AI tools, particularly the Claude API, to improve operational efficiency.
This includes understanding how to write effective prompts, provide appropriate context, filter and validate responses, and use AI to accelerate troubleshooting, documentation, automation, and engineering workflows rather than treating it as a simple chatbot.
Key Responsibilities
Operate and maintain a highly available, globally distributed cloud infrastructure serving users across multiple continents.Ensure maximum uptime, reliability, scalability, and performance of production systems.Monitor infrastructure health, proactively identify potential issues, and implement preventive improvements.Respond to incidents, troubleshoot complex production issues, and drive root cause analysis.Build, maintain, and improve CI/CD pipelines to enable safe, reliable, and frequent deployments.Manage and optimize Kubernetes (EKS) clusters running production workloads.Collaborate closely with software engineers to improve application deployment, performance, reliability, and operational excellence.Support daily infrastructure requests from engineering teams and provide technical guidance on cloud best practices.Design and implement infrastructure automation to reduce manual operational work.Continuously improve infrastructure through automation, standardization, and Infrastructure as Code.Participate in infrastructure planning, capacity management, and cost optimization initiatives.Implement and maintain monitoring, logging, and alerting solutions across the platform.Contribute to security improvements by implementing cloud security best practices and participating in security reviews.Create and maintain technical documentation, operational runbooks, and disaster recovery procedures.Participate in an on-call rotation when required and assist in resolving critical production incidents.Stay current with emerging cloud technologies and recommend improvements that enhance platform reliability, developer productivity, and operational efficiency.
What We're Looking For
A proactive engineer who enjoys solving complex infrastructure challenges.Strong analytical and problem-solving skills.Ability to work independently while collaborating effectively across engineering teams.Excellent communication skills and the ability to explain technical concepts clearly.A mindset focused on automation, continuous improvement, and operational excellence.Passion for building resilient, secure, and scalable cloud platforms that enable engineering teams to move quickly and confidently.