Site Reliability Engineering Lead

Relx Group — United States · Posted ~1 day ago

Lead Full-time

Skills

site reliability engineering cloud platforms automation resilience engineering operational excellence engineering leadership team leadership Kubernetes Terraform cloud computing SRE

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Lead a Site Reliability Engineering team responsible for reliable, scalable applications and infrastructure in modern cloud environments. You will drive automation and resilience, improve operational practices, and help engineering teams deliver dependable platforms at scale.

Highlights

Lead high-performing SRE teams while improving platform reliability, scalability, automation, resilience, and operational excellence across modern cloud environments.

Description

Are you passionate about building reliable, scalable platforms and helping engineering teams thrive? Do you enjoy leading high-performing teams while driving automation, resilience, and operational excellence across modern cloud environments? About The Business LexisNexis Risk Solutions is the essential partner in the assessment of risk. Within our Business Services vertical, we offer a multitude of solutions focused on helping businesses of all sizes drive higher revenue growth, maximize operational efficiencies, and improve customer experience. Our solutions help our customers solve difficult problems in the areas of Anti-Money Laundering/Counter Terrorist Financing, Identity Authentication & Verification, Fraud and Credit Risk mitigation and Customer Data Management. You can learn more about LexisNexis Risk at https://risk.lexisnexis.com/ About Our Team You will be joining the Core SRE Team in Business Services, a team that oversees all the applications and infrastructure in the biggest business unit in LexisNexis Risk Solutions. We build cloud environments, migrate on prem applications to the cloud, work with self-hosted and 3rd party solutions. The successful candidate is a self-starter who assesses the situation, collaboratively develops a solution, and takes the initiative to improve performance, cost, and reliability at each opportunity. About The Role This is a professional management level role. Individuals are required to provide line management to a small to medium sized team of engineers, including authority over performance management, pay, and recruitment. They will ensure that tasks and projects are prioritized appropriately and provide support to team members when tasks are blocked. They will support engineers in their personal development and ensure they are working within the SRE framework. They will lead the post-mortem reviews and the timely production of RCAs. They address issues with impact beyond their own team based on knowledge of related disciplines. Responsibilities Manage, mentor, and grow a team of SREs; conduct 1:1s, performance reviews, and career development planningOwn hiring, onboarding, and team capacity/resourcing decisionsSet team goals, prioritize backlog, and drive planningFoster a blameless post-incident culture and cross-team collaboration with Dev, Security, and ProductLead reliability initiatives across infrastructure and servicesDrive incident response activities and continuous service improvementChampion automation and operational excellence across the platformSupport the development of scalable, secure, and resilient cloud-native environments Requirements Expert knowledge of Kubernetes, including cluster architecture, upgrades, autoscaling, security hardening, and troubleshooting at scaleExpert experience with Terraform, including modular IaC design, state management, multi-environment provisioning, and policy-as-codeDeep knowledge of Azure Cloud, including compute, networking, identity (AAD), storage, and cost optimizationExperience designing and scaling CI/CD pipelines using GitHub Actions, release strategies, and rollback automationExperience with observability platforms including Prometheus, Grafana, OpenTelemetry, and SLO/SLA/error-budget managementStrong automation skills, focused on eliminating toil through self-healing systems and infrastructure automationAdvanced proficiency in Python, Bash, and/or PowerShell for tooling and automationDeep understanding of networking concepts including TCP/IP, DNS, load balancing, VPNs, and cloud-native networkingExperience in SRE, DevOps, or Infrastructure roles, including experience leading engineering teamsProven track record leading incident response and driving reliability improvements U.S. National Base Pay Range: $118,300 - $219,800. Geographic differentials may apply in some locations to better reflect local market rates. This job is eligible for an annual incentive bonus. We know your well-being and happiness are key to a long and successful career. We are delighted to offer country specific benefits. Click here to access benefits specific to your location.