Senior AWS Site Reliability Engineer

Spectrum It Recruitment โ€” United Kingdom ยท Posted ~1 day ago

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Description

The company deliver cutting-edge enterprise software solutions across both cloud and on-premises environments, empowering organisations to enhance customer experiences, maintain regulatory compliance, and proactively fight fraud. The company are trusted by businesses worldwide to drive seamless, intelligent customer interactions. In this role, you'll oversee the production environment by ensuring system availability and maintaining a comprehensive perspective on overall health. You'll develop tools and software to support and streamline the management of platform infrastructure and key applications. A major focus will be enhancing the dependability, performance, and delivery speed of our software products. You'll also be responsible for analysing and fine-tuning system performance to anticipate user demands and drive innovation. Additionally, you'll take the lead in providing operational support and technical oversight for several large-scale distributed applications. How You'll Contribute: Monitor and interpret system and application metrics to fine-tune performance and troubleshoot issues effectivelyCollaborate closely with developers to enhance service quality through thorough testing and structured release practicesEngage in architectural discussions, manage platform operations, and contribute to capacity forecastingDesign and implement automated solutions to build resilient, scalable systemsMaintain a strong focus on delivering new features while ensuring stability and adherence to service level goalsYou'll Stand Out If You Have: Practical experience managing large-scale Kubernetes clusters; certifications in Kubernetes are a strong bonusHands-on familiarity with the Grafana Observability Suite, including tools like Loki, Mimir, and TempoBackground in administering or developing with popular monitoring and automation tools such as Splunk, Datadog, PagerDuty, or RundeckExperience using configuration management platforms like Ansible, Puppet, or ChefProfessional certifications in cloud DevOps, such as AWS Certified DevOps Engineer or Google Cloud Professional DevOps Engineer, or similar credentialsDo You Have What It Takes? 3-6 years of hands-on experience in a similar role, with a strong emphasis on systems engineering, automation, and service reliabilityProficient in at least one programming language such as Python, Go, Java, or C#, along with scripting skills in Bash or PowerShellSolid grasp of cloud platforms like AWS, including an understanding of how core services like EC2, ECS, Lambda, and DynamoDB operate under reliability constraintsPractical experience using infrastructure-as-code tools like CloudFormation or TerraformIn-depth knowledge of CI/CD principles and hands-on experience with tools such as Jenkins, GitLab CI/CD, or CircleCIStrong understanding of containerization (e.g., Docker, Kubernetes) and microservices architectureSkilled in using observability and monitoring tools such as Prometheus, Grafana, ELK stack, or AWS CloudWatchExcellent analytical and troubleshooting abilities, especially within complex distributed systemsProven experience handling incident management and conducting blameless postmortems, including leading cross-functional teams through resolution and communication during critical outagesBenefits Life Insurance - 4 x Annual SalaryPrivate Medical InsuranceBonus Scheme Employee Assistance ProgrammeHybrid Working - 3 Days from HomeGP Online Assistance Portal.+ Much MorePlease click the "Apply" button to state your interest in this position. Spectrum IT Recruitment (South) Limited is acting as an Employment Agency in relation to this vacancy.