Summary
✨ AI‑Generated
A lead reliability engineering role focused on assessing monitoring capabilities, designing observability roadmaps, and collaborating with technical teams to improve service visibility.
Highlights
Contract leadership role focused on defining monitoring strategies, improving reliability, and working across complex enterprise environments.
Description
Location: London, onsite 3 days per week (Sheffield as an alternative)
Rate: £tbd/day inside IR35
Duration: 6 months+
Are you a Senior Observability Engineer / SRE Lead, with demonstrable experience of assessing and defining observability and monitoring roadmaps within enterprise scale environments? If so, apply now for this new contract opportunity.
The Lead Observability Engineer / SRE Lead will be required to assess a complex hybrid estate, understand how services, platforms, infrastructure and networks should be monitored, and work across multiple internal teams, partners and suppliers to build a consolidated view of existing telemetry, monitoring and alerting capabilities.
As well as short term tactical objectives, the role will also be focussed on longer term strategic ones.
Responsibilities of the Lead Observability Engineer / SRE Lead will be to:
Discover and document existing telemetry sources, monitoring tools, dashboards and ownershipWork with technical teams and suppliers to gain access to telemetryDeliver meaningful dashboards and service health viewsIdentify gaps in telemetry, monitoring and alerting coverage and implement pragmatic improvementsDefine health indicators for critical business journeysIntroduce modern observability practices where practical, including SLIs, SLOs and a roadmap towards burn-rate alertingDevelop a roadmap for OpenTelemetry adoptionAssess options for a centralised telemetry platform, including Grafana CloudEvaluate tooling rationalisation opportunities, operating costs and telemetry economicsDefine an observability target architecture, standards and implementation roadmap
The successful Lead Observability Engineer / SRE Lead will demonstrate the following:
Proven experience leading enterprise-scale observability initiativesStrong hands-on expertise with Grafana, Grafana Cloud, OpenTelemetry and modern telemetry pipelinesDeep understanding of metrics, logs, traces, distributed tracing, alerting, SLIs, SLOs and error-budget conceptsExperience designing observability solutions across cloud PaaS, IaaS, on-premises, legacy and third-party hosted platformsStrong knowledge of Azure observability tooling, including Azure Monitor, Log Analytics and Application InsightsIf this sounds like you, please apply to find out more.
Lead Observability Engineer / SRE Lead / Lead Site Reliability Engineer