AWS Site Reliability Engineer - Datadog

Vallum Associates Limited โ€” United Kingdom ยท Posted ~2 days ago

Senior Contract Hybrid

Skills

AWS Datadog Geneos Terraform Infrastructure as Code GitLab GitOps CI/CD Python Bash Linux Networking Distributed Systems Docker Kubernetes Incident Response Root Cause Analysis Cloud FinOps GitLab CI Jenkins EKS ECS EC2 S3 RDS Lambda VPC IAM CloudWatch

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary

A contract SRE opportunity for an experienced cloud and observability engineer to build and operate scalable monitoring infrastructure. You will work extensively with AWS, Terraform, GitOps, CI/CD, containers, scripting, and distributed systems while improving reliability, incident response, performance, and cloud cost efficiency. The role also involves mentoring engineers and partnering with cross-functional stakeholders on long-term infrastructure strategy.

Highlights

High-impact SRE role focused on modern observability, cloud infrastructure, automation, reliability, and incident response. Offers opportunities to shape infrastructure strategy, mentor engineers, collaborate across international teams, and drive continuous improvement.

Description

The Role: AWS SRE - Datadog Location: London, UK Position Type: Contract Inside IR35 Remote work option Available: 3 Days Onsite Job Description: We are looking for someone with a strong background in Datadog, Geneos, Gitlab and Infra-as-Code (Terraform). You will play a crucial role in our technology team, contributing to the development, deployment, and maintenance of our monitoring and alerting infrastructure. Key Responsibilities Architect, implement, and maintain observability platforms using Datadog and Geneos to ensure comprehensive monitoring and alerting.Design and manage scalable infrastructure using Terraform and Infrastructure as Code (IaC) principles.Champion GitOps methodologies using GitLab for CI/CD, configuration management, and deployment automation.Collaborate with engineering teams to embed reliability, scalability, and performance into the software development lifecycle.Collaborate with software developers across multiple geographies and cross-functional teams to understand project objectives, gather requirements, and deliver systems and software within agreed upon timelines.Optimize alerting strategies to reduce noise and improve actionable insights.Mentor junior engineers and contribute to the evolution of SRE best practices.Oversee and continuously optimize cloud cost management strategies for our Observability infrastructure in line with Cloud FinOps principles.Collaborate with executive leadership and cross-functional stakeholders to align infrastructure strategy with long-term business objectives.Lead incident response, perform root cause analysis, and drive continuous improvement through blameless post-mortems.Ensure code is up-to-date, maintainable, scalable, and secure.Stay updated on industry best practices and emerging technologies in Observability, DevOps and Cloud. Must Have Skills Strong hands-on experience with AWS cloud services: EC2, S3, RDS, Lambda, VPC, IAM, CloudWatch, EKS, ECSStrong hands-on expertise in Infrastructure as Code (IaC) using Terraform.Deep knowledge and hands-on experience with CI/CD tools (e.g., GitLab CI, Jenkins, etc.) for Java and Python based Microservices-style architectures. Deep expertise in Datadog for metrics, logs, traces, and dashboards.Proficiency with Geneos for real-time monitoring and alerting.Solid understanding of GitLab and GitOps workflows.Strong scripting and automation skills (e.g., Python, Bash).Solid understanding of Linux systems, networking, and distributed systems.Experience with containerization and orchestration (e.g., Docker, Kubernetes).Excellent problem-solving skills and ability to work collaboratively.