Summary
Lead a DevOps and reliability engineering function, improve cloud infrastructure, mentor engineers, strengthen observability, and build scalable delivery pipelines for a growing SaaS platform.
Highlights
Technical leadership role with ownership of cloud reliability, automation, mentoring, and modern engineering platform improvements.
Description
About Render
Render's purpose is to continuously innovate and evolve technology to build networks better and connect communities.
With $60B+ in annual U.S.
network infrastructure investment forecast to continue for the next decade, the need for technology-enabled efficiency has never been greater.
Backed by IFM Investors, we're accelerating product innovation and growing fast.
We're looking for exceptional people who share our passion for technology's role in connecting communities and who lead with bold innovation, integrity, and a genuine commitment to customers.
About the Role
Reporting to the VP Engineering, the Lead DevOps Engineer is responsible for the reliability, security and operational excellence of Render's cloud platform.
This is a hands-on technical leadership role where you'll lead our DevOps & SRE function while remaining close to the technology.
You'll own production operations, cloud infrastructure, platform reliability, observability and CI/CD while mentoring a team of DevOps engineers.
Working closely with Product Engineering, you'll help shape the future of our engineering platform, improve developer productivity, and ensure our SaaS platform continues to meet the highest standards of availability, resilience and security.
What You'll Do
Lead the operational health, reliability and resilience of Render's production platform.Lead major incident response and service restoration activities, driving effective triage, communication, root cause analysis and continuous improvement.Own the technical investigation, triage and resolution of customer-impacting platform issues and escalations.Operate and maintain Render's AWS serverless platform, ensuring secure, scalable and reliable cloud operations.Build and evolve our AWS cloud platform using Infrastructure as Code and modern cloud engineering practices.Drive operational excellence through automation, reliability engineering practices, and the adoption of SLIs, SLOs and service improvement initiatives.Own our observability strategy using Datadog and improve platform visibility.Maintain and improve Buildkite and associated CI/CD pipelines, enabling fast, reliable and secure software delivery.Support the operational reliability of our event streaming and data platforms, including Confluent Cloud and Databricks.Partner with Product Engineering teams to improve platform scalability, resilience and developer experience.Maintain operational controls supporting ISO 27001 and organisational security requirements.Lead, mentor and grow a high-performing DevOps/SRE team.
What You'll Bring
8+ years of experience in DevOps, Site Reliability Engineering, or Platform Engineering roles, with proven experience operating and improving production SaaS platforms.Proven experience leading DevOps, Platform Engineering or Site Reliability Engineering teams, with the ability to provide technical direction, mentorship and operational leadership.Strong hands-on experience operating production workloads on AWS using modern cloud architectures, with a focus on scalability, security and reliability.Deep understanding of Infrastructure as Code practices, preferably with Terraform.Experience designing and operating CI/CD pipelines, observability platforms and cloud operations practices.Experience supporting distributed systems, event-driven architectures or cloud data platforms.Strong knowledge of incident management, operational resilience, production support and continuous service improvement practices.Experience implementing reliability engineering practices, including SLIs, SLOs, monitoring and automation.Excellent leadership, communication and stakeholder management skills, with the ability to collaborate effectively across engineering and product teams.A passion for automation, continuous improvement and building engineering platforms that enable teams to deliver software faster and more reliably.Experience applying FinOps principles to cloud operations, including cloud cost visibility, optimisation, and balancing cost, performance and reliability trade-offs.Preferred: Bachelor's degree in Information Technology, Computer Science or a related discipline.Preferred: Programming experience with Python or similar languages to support automation, tooling and platform engineering initiatives.