Site Reliability Engineer

Wearehaystack — United Kingdom · Posted ~21 hours ago

Mid Visa History ✓

Skills

Site Reliability Engineering AWS Linux system administration Cloud operations Production incident management Automation Monitoring and observability Linux Cloud Monitoring Observability

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join an engineering team responsible for highly available, secure cloud-native platforms supporting sensitive communication services. You’ll monitor production systems, handle complex incidents, automate operational work, strengthen observability, and collaborate with software, platform, cloud, and security specialists to improve reliability.

Highlights

Work on highly secure, cloud-native platforms with strong emphasis on automation, reliability, observability, and operational excellence. The role offers cross-functional collaboration and exposure to modern cloud engineering practices.

Description

We're working with a global leader in AI-powered customer experience and cloud technology, expanding their engineering teams to build and support highly secure, cloud-native platforms that deliver sensitive communication services. This organization is investing heavily in modern cloud engineering, automation, and reliability. The Role Monitor and maintain highly available production platforms running in AWS Respond to and manage production incidents across a 24/7 service Investigate complex technical issues and restore services quickly and effectively Develop automation to reduce manual operational tasks and improve platform resilience Build and improve monitoring, alerting, and observability across cloud environments Work alongside Software, Platform, Cloud, and Security Engineers to improve reliability and operational excellence What You'll Need Experience in Site Reliability Engineering, Production Engineering, Cloud Operations or NOC environment Linux systems administration and AWS cloud infrastructure Kubernetes and Docker experience Production support and incident management expertise Python, Bash or Go scripting skills Experience with monitoring and observability platforms such as Grafana, Prometheus, Datadog, Splunk or CloudWatch What's On Offer Competitive salary with bonus opportunities Excellent benefits package Opportunity to build resilient cloud platforms for critical national services A role focused on preventing incidents through system improvement and automation Apply via Haystack today!