Site Reliability Engineer

Wearehaystack โ€” United Kingdom ยท Posted ~1 day ago

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Description

We're working with a global leader in AI-powered customer experience and cloud technology, recently awarded a major government programme. This organisation is expanding its engineering teams to build and support highly secure, cloud-native platforms that deliver sensitive communication services, investing heavily in modern cloud engineering, automation, and reliability. The Role Monitor and maintain highly available production platforms in AWS Respond to and manage production incidents across a 24/7 service Investigate complex technical issues and restore services quickly Develop automation to reduce manual operational tasks and improve platform resilience Build and improve monitoring, alerting, and observability across cloud environments Support containerised workloads using Kubernetes and Docker What You'll Need Experience in Site Reliability Engineering, Production Engineering, Cloud Operations, or NOC environments Linux systems administration and AWS cloud infrastructure experience Proficiency with Kubernetes and Docker Strong production support and incident management skills Scripting abilities in Python, Bash, or Go Familiarity with monitoring platforms like Grafana, Prometheus, Datadog, Splunk, or CloudWatch What's On Offer Join an engineering-led organisation with a focus on reliability and automation Opportunity to build and support highly secure, cloud-native platforms Contribute to continuous service improvements and shape resilient cloud services Competitive compensation package with bonus and excellent benefits Apply via Haystack today!