Summary
✨ AI‑Generated
A technology organization is hiring an SRE to manage cloud-native infrastructure and improve reliability of large-scale systems. The role involves Kubernetes, automation, and deployment practices.
Highlights
Cloud-native infrastructure role focused on reliability, automation, scalability, and modern DevOps practices.
Description
Job Title: Site Reliability Engineer (SRE) – Cloud Kubernetes Platform
Experience: 4–8 years
Location: Edinburgh, Scotland, UK
About the Role:
We are looking for a passionate and experienced Site Reliability Engineer (SRE) to join our Cloud Platform team.
The ideal candidate will have hands-on experience managing large-scale Kubernetes clusters on public cloud environments (AKS, EKS, or GKE) and a strong understanding of modern SRE and DevOps practices.
You will be responsible for ensuring high availability, reliability, scalability, and performance of our cloud-native infrastructure and CI/CD systems.
Key Responsibilities:
• Manage, monitor, and optimize large-scale Kubernetes clusters hosted on public cloud platforms (Azure AKS, AWS EKS, or Google GKE).
• Implement and maintain infrastructure as code using tools such as Terraform.
• Collaborate with development and operations teams to improve system reliability and deployment automation.
• Build and maintain CI/CD pipelines using Jenkins or similar tools.
• Troubleshoot production issues, conduct root cause analysis, and implement preventive measures.
• Automate operational tasks using Python or other scripting languages.
• Contribute to observability and monitoring improvements using modern tools and best practices.
• Participate in on-call rotations and incident response processes.
Required Skills and Experience:
• 4–8 years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles.
• Strong hands-on experience managing Kubernetes clusters in production (AKS/EKS/GKE).
• Proficiency with Terraform and cloud infrastructure automation.
• Practical experience with Jenkins and CI/CD pipeline management.
• Sound understanding of SRE principles (incident management, blameless postmortems, capacity planning, error budgets, etc.).
• Good programming or scripting skills in Python (preferred) or similar languages.
• Strong analytical, troubleshooting, and problem-solving abilities.
• Excellent written and verbal communication skills.
• Experience with Prometheus, Grafana, or OpenTelemetry for observability.
• Exposure to GitOps practices and tools (e.g.
Flux).