Senior Site Reliability Engineer
Evridelivery — United Kingdom · Posted ~23 hours ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
Senior Site Reliability Engineer (Team Lead)
Remote | £66,000 + Bonus + Benefits
Build Reliability at Scale.
Lead Engineers.
Drive Modern Cloud Architecture.
Want to shape the reliability, performance and scalability of platforms that support one of the UK's largest parcel delivery networks?
At Evri, Site Reliability Engineering plays a critical role in how we build, deploy and operate technology.
We're looking for a Senior Site Reliability Engineer (Team Lead) to lead and develop a team of six engineers while remaining hands-on across AWS, platform engineering, reliability, automation and cloud-native architecture.
This is a highly influential role requiring both strong leadership and deep technical expertise.
We're looking for someone who can drive architectural decisions, champion modern cloud practices, lead complex technical initiatives and help shape the future of reliability engineering across the organisation.
You'll work closely with engineering teams to deliver resilient, scalable and observable platforms that support business-critical services used by millions of customers every year.
What You'll Be Doing
Lead, mentor and develop a team of Site Reliability EngineersProvide technical leadership across cloud, platform and operational engineering initiativesDefine and drive engineering standards, reliability practices and operational excellenceLead technical design and architecture discussions to ensure solutions are secure, scalable and maintainableImprove platform reliability, resiliency, scalability and operational efficiencyDrive platform modernisation initiatives including container adoption, automation and observability improvementsBuild and maintain monitoring, logging, alerting and observability solutions that provide meaningful operational insightDefine, measure and improve SLIs, SLOs and key service performance metricsDrive automation initiatives that reduce manual effort and improve software deliverySupport engineering teams in solving complex availability, performance and scalability challengesAct as a technical escalation point during major production incidents, leading recovery efforts and driving post-incident improvementsChampion security, resilience and engineering best practices across cloud platforms and servicesManage and evolve shared tooling including CI/CD platforms and automation frameworksIdentify technical debt, challenge existing approaches and drive continuous improvement initiativesDeliver cost optimisation, platform stability and operational efficiency improvements across the estate
What You'll Bring
Essential Experience
Experience leading, mentoring or managing engineers within an SRE, DevOps, Platform Engineering or Cloud Engineering environmentDeep AWS architecture experience designing, building and operating cloud-native platforms at scaleStrong hands-on experience developing and managing cloud infrastructure using AWS CDK and TypeScriptStrong knowledge of AWS networking and security including VPC design, routing, Security Groups, NACLs, IAM and highly available architecturesExperience designing and supporting modern container platforms including Amazon ECS/Fargate and/or Kubernetes (EKS)Proven experience owning and leading major production incidents, driving root cause analysis and implementing long-term resiliency improvementsExperience implementing and operating observability platforms including monitoring, logging, tracing and alerting solutionsStrong Infrastructure as Code experience using tools such as CloudFormation, AWS CDK and AnsibleExperience building, supporting and optimising CI/CD pipelines using Jenkins, GitLab, Concourse CI or similar technologiesStrong Linux and/or Windows administration experienceScripting and automation experience using Python, Bash, Go or similar languagesExcellent troubleshooting and problem-solving skills within complex production environmentsStrong stakeholder management and communication skills with the ability to influence technical and non-technical audiencesDemonstrable experience guiding architectural decisions, identifying technical debt and driving platform modernisation initiatives
Desirable Experience
Experience operating distributed systems at enterprise scaleExperience with service mesh, platform engineering or internal developer platformsExperience working within high-availability, customer-facing environmentsFinOps and cloud cost optimisation experience
Why Evri?
At Evri, reliability matters.
You'll join a business where engineering is central to our success and where technology powers millions of parcel deliveries every year.
This role offers the opportunity to influence engineering strategy, define reliability standards, lead talented engineers and shape the future of our cloud platforms.
We're looking for someone who doesn't just keep systems running, but actively improves them.
Someone who can challenge, innovate and lead from the front.
In return, you'll have the opportunity to make a genuine impact across a large-scale cloud-first environment while continuing to grow your technical and leadership capability.
Let's Deliver It Together
We are Evri.
Where everyone is welcome.
Where great engineering thrives.
Where reliability is built by design.
We have 154,647 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 154,647 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume