Senior Site Reliability Engineer

Webologix — United Kingdom · Posted ~4 days ago

Senior

Skills

SRE DevOps Kubernetes AWS Azure GCP Datadog Python Go

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A technology organization is seeking an experienced Site Reliability Engineer to manage cloud environments, observability platforms, and large-scale distributed systems. The role requires strong troubleshooting and automation expertise.

Highlights

Senior reliability engineering role focused on cloud platforms, observability, and distributed systems troubleshooting. Offers complex technical challenges and ownership.

Description

Your responsibilities: (Up to 10, Avoid repetition) Experience: 10+ years of hands-on experience in Site Reliability Engineering (SRE), DevOps, or Systems Architecture, with at least 3+ years specializing deeply in Datadog administration and configuration.Cloud & Container Expertise: Deep professional experience working with AWS, Azure, or GCP, pairedwith heavy production experience managing Kubernetes clusters.Instrumentation & Coding: Proficiency in systems or scripting languages (e.g., Python, Go, Bash, orJavaScript) and experience instrumenting applications for APM.Datadog Mastery: Deep understanding of Datadog’s core pillars—Infrastructure, APM, Logs, Metrics,Synthetics, and Security Monitoring. Datadog Certifications are a strong plus.Problem-Solving Mindset: Demonstrated ability to debug complex, distributed microservicesarchitectures under high-pressure incident response scenarios.Communication Skills: Excellent interpersonal and stakeholder management skills, with the ability to translate technical telemetry data into actionable business and engineering insights. Your Profile Essential skills/knowledge/experience: (Up to 10, Avoid repetition) Experience: 10+ years of hands-on experience in Site Reliability Engineering (SRE), DevOps, or Systems Architecture, with at least 3+ years specializing deeply in Datadog administration and configuration.Cloud & Container Expertise: Deep professional experience working with AWS, Azure, or GCP, pairedwith heavy production experience managing Kubernetes clusters.Instrumentation & Coding: Proficiency in systems or scripting languages (e.g., Python, Go, Bash, orJavaScript) and experience instrumenting applications for APM.Datadog Mastery: Deep understanding of Datadog’s core pillars—Infrastructure, APM, Logs, Metrics,Synthetics, and Security Monitoring. Datadog Certifications are a strong plus.Problem-Solving Mindset: Demonstrated ability to debug complex, distributed microservicesarchitectures under high-pressure incident response scenarios.Communication Skills: Excellent interpersonal and stakeholder management skills, with the ability to translate technical telemetry data into actionable business and engineering insights. Desirable skills/knowledge/experience: (As applicable) Architecture & ImplementationPlatform Ownership: Design, deploy, and manage Datadog agents, integrations, and custom metricsacross multi-cloud (AWS/Azure/GCP) and containerized (Kubernetes, Docker) environments.Observability Pipelines: Architect and scale high-throughput log processing, routing, and transformationsystems using Datadog.APM & Infrastructure Monitoring: Configure and optimize Application Performance Monitoring (APM),Distributed Tracing, Real User Monitoring (RUM), and Infrastructure metrics.Governance & Best PracticesStandardization: Establish company-wide standards for dashboards, monitors, SLOs/SLIs, and alertrouting (integrating with PagerDuty, Jira, Opsgenie, etc.).Cost & Performance Optimization: Audit and optimize Datadog usage, index management, logretention policies, and custom metric volume to maximize ROI and control licensing costs.Security & Compliance: Leverage Datadog Security products (CSPM, CWPP, Cloud SIEM, ContainerSecurity) to maintain compliance postures and mitigate runtime threats.Collaboration & EnablementCross-Functional Mentorship: Act as the go-to escalation point and technical mentor for DevOps, SRE,and Software Engineering teams regarding troubleshooting and instrumentation.Training & Documentation: Create internal documentation, runbooks, and training modules to elevate organizational proficiency in observability.Vendor Management: Act as the primary technical point of contact for Datadog account teams,