Senior Site Reliability Engineer

Q S R โ€” United Kingdom ยท Posted ~21 hours ago

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Description

Senior Site Reliability Engineer (SRE) ๐Ÿ’ฐ ยฃ400/day (Inside IR35) ๐Ÿ“ London City, UK (3 days onsite) ๐Ÿ“… 12-Month Contract Our client is looking for an experienced Senior Site Reliability Engineer (SRE) to join a high-performing engineering team delivering enterprise-scale cloud infrastructure and observability solutions. As an SRE, you'll take ownership of observability, cloud infrastructure, automation, and operational resilience across distributed systems. Working closely with engineering teams, you'll help build highly available platforms, optimise monitoring strategies, automate deployments, and improve service reliability through modern DevOps and GitOps practices. Key Responsibilities Design, implement, and maintain enterprise observability platforms using Datadog and Geneos.Build and manage scalable cloud infrastructure using Terraform and Infrastructure as Code.Champion GitOps methodologies using GitLab CI/CD and deployment automation.Collaborate with software engineering teams to improve application reliability, scalability, and performance.Develop monitoring dashboards, metrics, tracing, and intelligent alerting strategies.Lead incident response, root cause analysis, and continuous service improvement initiatives.Optimise cloud infrastructure costs using FinOps principles.Mentor engineers and promote Site Reliability Engineering best practices.Maintain secure, scalable, and highly available AWS infrastructure.Contribute to long-term cloud strategy and platform engineering initiatives.Essential Skills & Experience 7+ years' experience in Site Reliability Engineering, DevOps, or Platform Engineering.Strong hands-on AWS experience including EC2, S3, RDS, Lambda, IAM, VPC, CloudWatch, ECS, and EKS.Strong Infrastructure as Code experience using Terraform.Deep experience with Datadog monitoring, dashboards, logging, tracing, and observability.Experience with Geneos monitoring platforms.Strong GitOps and CI/CD experience using GitLab (or similar).Scripting experience using Python and/or Bash.Strong Linux, networking, and distributed systems knowledge.Experience with Docker and Kubernetes.Excellent troubleshooting, incident management, and problem-solving skills.Desirable Skills AWS Professional certifications.Experience with SLOs, SLIs, and Error Budgets.Knowledge of multi-cloud or hybrid cloud environments.ITIL or SRE methodologies.Experience working within financial services, fintech, or other highly regulated environments.Knowledge of Equity, Fixed Income, Benchmarks, or Indices is advantageous. Please apply with your Cv and we'll be in touch. Thanks!