Site Reliability Engineer

Venquis Ltd — United Kingdom · Posted ~21 hours ago

Mid Full-time

Skills

Site Reliability Engineering Production monitoring Incident response Scalability Availability Security DevOps Service management Automation SRE Cloud infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A Site Reliability Engineering role focused on keeping critical services scalable, available, secure, and responsive to incidents. You will monitor production systems, automate operational work, remove service bottlenecks, support application releases, and help design and scale a global technology platform.

Highlights

Opportunity to work on business-critical services operating continuously, improve scalability and reliability, automate repetitive work, collaborate closely with development and DevOps teams, and grow into Site Reliability Engineering from adjacent backgrounds.

Description

As a Site Reliability Engineer, your overarching responsibility is to ensure we meet our customers’ Service Level Agreements, and that we respond to incidents in a timely and professional manner. You will proactively monitor production environments to ensure scalability, availability, and security, improving the alignment of service performance to customers and co-workers by reducing or eliminating manual and repetitive tasks, and removing bottlenecks and inefficiencies from services. You will create, deliver, and manage business critical services that are used 24/7 by customers and co-workers. You will work closely with Development and DevOps teams to give them the tools they need and support the application release process, and you will be involved in designing, building, and scaling our global product platform. We welcome engineers from development, DevOps, SRE, or similar backgrounds who want to grow their career in Site Reliability Engineering. Responsibilities • Spend an equal amount of time building software to automate manual work and providing operational support to the products you cover, balancing feature development speed and reliability against service-level objectives • Lead incident response, diagnosis, and follow-up on system outages or alerts • Perform and assist in root cause analysis and blameless post-mortems, enabling incidents to be understood and avoided in future • Propose improvements to infrastructure and product • Improve the reliability, quality, and time-to-market of our software solutions • Provide out-of-hours support based on an on-call rota Skills and Experience • Experience with Azure • Experience with Kubernetes • Proficiency in a programming language such as Python or Go • A track record of writing code you care about, including unit tests, integration tests, static analysis, and resilience tests • Experience with database technologies such as MySQL • Experience with Infrastructure as Code tools such as Terraform • A strong drive to engineer solutions using best practices • A security-first design philosophy