Senior Site Reliability Engineer
Roche — Poland · Posted ~1 day ago
🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.
Log in to add to target listDescription
At Roche you can show up as yourself, embraced for the unique qualities you bring.
Our culture encourages personal expression, open dialogue, and genuine connections, where you are valued, accepted and respected for who you are, allowing you to thrive both personally and professionally.
This is how we aim to prevent, stop and cure diseases and ensure everyone has access to healthcare today and for generations to come.
Join Roche, where every voice matters.
The Position
The Position
We are building a global Site Reliability Engineering (SRE) team to support critical commercial and internal platforms and applications.
As an SRE, you will help design, build, and scale reliable distributed systems that power healthcare innovation worldwide.
This role is focused on reliability, scalability, automation and operational excellence.
You will influence system design, define reliability standards and reduce operational toil through engineering solutions.
This role includes participation in a structured on-call rotation.
Who We Are
At Roche, we are passionate about transforming patients’ lives, and we are bold in both decision and action - we believe that good business means a better world.
That is why we come to work every single day.
We commit ourselves to scientific rigor, unassailable ethics and access to medical innovations for all.
We do this today to build a better tomorrow.
Roche is strongly committed to a diverse and inclusive workplace.
We strive to build teams that represent a range of backgrounds, perspectives and skills.
Embracing diversity enables us to create a great place to work and to innovate for patients.
Step into the Future of IT with Roche!
As a seasoned Site Reliability Engineer (SRE) at Roche, you will leverage your deep software engineering expertise to propel our products to new heights of robustness, scalability and reliability.
This isn't just a role—it's an invitation to shape the backbone of technological innovations forward.
Your Mission
Design and maintain cutting-edge tools, scripts and frameworks that automate repetitive tasks, streamline software deployment and manage expansive systems with unparalleled efficiency.
Partner closely with forward-thinking development teams to architect and implement high-performance solutions that elevate system efficiency, optimize resource utilization and enhance deployment processes for superior uptime and user satisfaction.
Your Impact
Lead the charge in incident management and response.
Detect system anomalies, troubleshoot swiftly and conduct thorough root cause analyses to prevent recurring issues.
Champion continuous improvement by refining monitoring and alerting mechanisms, conducting insightful post-incident reviews and embedding best practices in software lifecycle management.
Your strategic foresight and meticulous planning will ensure our systems are not only reliable but also superlatively performant.
By joining our elite team, you will play a pivotal role in delivering seamless experiences to our end-users, exceeding business and customer demands, and solidifying Roche's reputation as a leader in IT innovation.
Your Core Responsibilities
Reliability Engineering & Architecture
Define and implement SLIs, SLOs, and error budgets with product and engineering teamsConduct reliability reviews for new and existing servicesDesign scalable, fault-tolerant architectures in AWS and Azure environmentsLead capacity planning, performance and cost optimization initiativesImprove system resilience through automation and self-healing patternsDrive organizational observability maturity (metrics, logs, traces, alert quality)
Incident Management & Continuous Improvement
Perform complex root cause analysis and drive rapid mitigationParticipate in blameless postmortems and follow-throughImprove MTTR, reduce incident frequency, and elevate production standardsCollaborate seamlessly with engineering teams to enable timely and effective resolutionsHandle requests and incidents, create and maintain runbooksParticipation in a structured 24*7 on-call rotation
Automation & Platform Engineering
Reduce operational toil through tooling and automation (Python or similar)Improve CI/CD reliability and deployment safety mechanismsBuild and maintain infrastructure-as-code (Terraform or equivalent)Enhance Kubernetes platform reliability (EKS, AKS, or similar)
Cross-Functional Leadership
Partner with business, engineering, security, and cloud teams to embed reliability early in the software development life cycleMentor mid-level engineers and help shape SRE best practicesChampioning a culture of ownership, accountability, and continuous improvement
Who You Are:
Minimum bachelor’s degree in computer science, Engineering, or a related field, or equivalent professional experienceExperience in either site reliability engineering, software engineering or related fields with production on-call experienceSolid experience with AWS and/or Azure, including setting up, monitoring, and maintaining cloud resources (incl.
Kubernetes, EKS, AKS, GKE, etc knowledge)Proficiency with observability toolsHands-on experience with incident management toolsProficiency in scripting languages for automation purposesDemonstrated proficiency in troubleshooting, especially in cloud and distributed system environmentsExcellent communication, teamwork and documentation skills, with a proactive and self-motivated approach to improving system reliability and operational efficienciesWe value and encourage candidates from diverse backgrounds and experiences, believing that diverse perspectives drive innovation and successExcelling in both spoken and written English communication
Where pay transparency applies, details are provided based on the primary posting location.
For this role, the primary location is Sant Cugat del Vallès.
If you are interested in additional locations where the role may be available, we will provide the relevant compensation details later in the hiring process.
Who we are
A healthier future drives us to innovate.
Together, more than 100’000 employees across the globe are dedicated to advance science, ensuring everyone has access to healthcare today and for generations to come.
Our efforts result in more than 26 million people treated with our medicines and over 30 billion tests conducted using our Diagnostics products.
We empower each other to explore new possibilities, foster creativity, and keep our ambitions high, so we can deliver life-changing healthcare solutions that make a global impact.
Let’s build a healthier future, together.
Roche is an Equal Opportunity Employer.
We have 68,618 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume — in under a minute we'll analyze all 68,618 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume