Summary
✨ AI‑Generated
A technology organization is seeking a Cloud Infrastructure Engineer with strong AWS, containerization, and Azure DevOps expertise. You will design and manage highly available infrastructure, build and optimize CI/CD pipelines, improve observability and resilience, automate operational processes, and support platform modernization in a cloud-native environment.
Highlights
Reliability-focused engineering role centered on automation, observability, scalability, and resilience. You will improve highly available cloud infrastructure, optimize CI/CD pipelines, reduce operational toil, and help engineering teams deliver reliable services efficiently.
Description
Role Overview
We are seeking a Cloud Infrastructure/Engineer with strong expertise in AWS cloud infrastructure, containerised platforms, and Azure DevOps CI/CD pipelines.
The successful candidate will improve system reliability, availability, performance, and scalability while enabling engineering teams to deliver high-quality services efficiently.
This role combines engineering and operational excellence, focusing on automation, observability, scalability, and resilience across cloud-native environments.
As a Cloud Engineer, you will drive engineering-led solutions to reduce operational toil, enhance system reliability, and promote DevOps and SRE best practices.
Note: This is a reliability-focused engineering role with on-call responsibilities and involvement in platform modernisation initiatives.
Key Responsibilities
Design, implement, and manage highly available and scalable infrastructure on AWS.Build, maintain, and optimise DevOps Pipelines (CI/CD) for automated build, test, and deployment processes.Implement end-to-end CI/CD workflows, including multi-stage pipelines, approvals, and release strategies.Manage and support Windows (IIS, .NET) and Linux-based production systems.Deploy, manage, and optimise containerised applications using Docker and Kubernetes (EKS/AKS).Implement Infrastructure as Code (IaC) using Terraform, CloudFormation, or ARMDevelop and maintain automation scripts using PowerShell, Bash, or Python.Define and monitor SLIs, SLOs, and SLAs to ensure system reliability.Implement robust monitoring, logging, and alerting solutions (CloudWatch, Prometheus, Grafana, Azure Monitor).Lead incident management, troubleshooting, and root cause analysis (RCA) for production issues.Drive performance tuning and capacity planning for applications and infrastructure.Collaborate with development teams to improve deployment strategies (blue-green, canary releases).Ensure security, compliance, and best practices across CI/CD pipelines and infrastructure.
Qualifications
Required Skills & Experience
5-7 years of experience in Site Reliability Engineering / DevOps / Infrastructure EngineeringStrong hands-on experience with AWS services (EC2, S3, RDS, VPC, IAM, ELB, Auto Scaling, CloudWatch)Deep expertise in Azure DevOps Pipelines (CI/CD), including YAML pipelines and release automationExperience designing multi-stage pipelines and deployment strategiesExpertise in Windows Server administration, including IIS and .NET application supportStrong experience with Linux system administrationHands-on experience with Docker and Kubernetes (EKS/AKS)Experience with Infrastructure as Code (Terraform, CloudFormation, or ARM templates)Strong scripting skills in PowerShell (mandatory) and Bash/PythonExperience with monitoring and logging tools (Prometheus, Grafana, ELK, CloudWatch)Solid understanding of networking, security, and cloud architecture principles
Preferred Qualifications
Experience with hybrid cloud or multi-cloud environmentsKnowledge of Active Directory, Group Policy, and enterprise Windows environmentsFamiliarity with Helm, GitOps practices, or service mesh technologiesExperience with performance testing and tuningRelevant certifications (AWS)Preferred (not mandatory):Azure DevOpsNice-to-have:KubernetesAdvanced SRE practices (SLIs/SLOs, observability)
Key Competencies / Characteristics
Reliability-driven: Focused on uptime, performance, and system resilienceAutomation-first mindset: Continuously reduces manual effort and operational toilOwnership mentality: Takes end-to-end responsibility from design through productionStrong communicator: Clearly articulates incidents, RCA outcomes, and technical conceptsCollaborative: Works effectively with platform, security, and application teamsMentorship mindset: Actively supports and develops junior team membersContinuous learner: Keeps up with evolving SRE practices and cloud-native technologies