Senior Site Reliability Engineer
Italenters — Spain · Posted ~20 hours ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
Build the reliability culture, not just the incident response.
🚀
Join a profitable, growing product company transforming how vacation-rental hosts manage their businesses.
With ~160 people and a mature engineering organisation, the company is focused on product stability, collaboration and sustainable growth.
As a Senior Site Reliability Engineer, you’ll help evolve Platform from reactive operations to a proactive, data-driven SRE culture.
You’ll code, automate and enable teams to build and run reliable services autonomously.
Working with ~75 microservices, 800 concurrent containers and 3B monthly requests, you’ll strengthen observability, define reliability targets and help build an internal developer platform that makes infrastructure and deployments easier.
Join a small, senior team and make a measurable impact on a product used by customers worldwide.
💼 Responsibilities
Drive the transition from reactive DevOps to a preventive, scalable SRE model, embedding SRE practices and increasing engineering teams’ operational autonomy.Define and evolve SLIs, SLOs and error budgets to support better engineering and product decisions, while improving production reliability through automation, resilient design and sustainable operations.Design and enhance observability, monitoring and alerting across a complex microservices environment.Build and evolve an internal developer platform that enables product teams to deploy and manage infrastructure safely and independently.Automate repetitive operational work through hands-on coding, primarily with Python or similar languages, and contribute to infrastructure-as-code and the continuous evolution of the cloud platform.Solve complex reliability challenges across Kubernetes, cloud infrastructure, messaging systems and databases.
🔎 Profile Requirements
5+ years of experience in Site Reliability Engineering, with a clear preventive and reliability-first mindset.Hands-on experience building automation and writing production-quality code in Python or a similar language.Strong practical knowledge of Kubernetes and cloud environments; experience with Google Cloud Platform is highly valuable.Solid experience with infrastructure as code, ideally using Pulumi or comparable tooling.Proven ability to design and improve observability, monitoring and actionable alerting.Practical knowledge of SLIs, SLOs and error budgets—and the judgment to use them to drive meaningful improvements.A systems-thinking approach: you look beyond immediate fixes and design solutions that last.Experience collaborating with software engineering teams and helping them adopt reliable, autonomous ways of working.Knowledge of Kafka, RabbitMQ, PostgreSQL, SQL Server or similar distributed systems is a plus.Fluent English at C1 level, as you will work in an international environment.
💰 Benefits
100% remote work model, with offices in Barcelona available for workshops and occasional meetings.25 vacation days per year.Private health insurance, with travel, dental, and psychological coverage.Flexible compensation through meal and transport vouchers.Full equipment and setup to work from home.Flexible working hours, with a start time between 8:00 and 10:00 and no strict time tracking.Reduced working hours in August: 35 hours per week.Referral program; for recommending candidates.Possibility of relocation and visa support for professionals joining from outside Spain.
If you want to build reliability into the way an engineering organisation operates—not merely keep the lights on—this is the challenge for you.
Apply and let’s talk.
We have 143,359 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 143,359 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume