Site Reliability Engineer

Thrive It Systems Ltd — Poland · Posted ~5 hours ago

Senior Hybrid

Skills

Site Reliability Engineering DevOps Software Development Life Cycle (SDLC) Incident Management Root Cause Analysis Observability SLI/SLO Infrastructure Migration Disaster Recovery Automation Software Architecture SRE SDLC SLI SLO Infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join a global DevOps environment as a Site Reliability Engineer responsible for improving the availability, performance, security, and transparency of highly available 24x7 production services. The role involves applying SRE best practices, resolving incidents and conducting root cause analysis, contributing to software architecture, working across the full SDLC, defining SLIs and SLOs, building observability capabilities, executing infrastructure migrations and disaster recovery exercises, supporting upgrades, and expanding automation and self-service capabilities. This is a hybrid opportunity with regular in-office collaboration.

Highlights

Opportunity to work in a global DevOps environment supporting highly available 24x7 production services, with a strong focus on reliability, automation, observability, architecture, incident response, and continuous operational improvement.

Description

Role: Site Reliability Engineer Location: Poland, Krakow Mode: Hybrid - 3 day a week Description Serve as a Site Reliability Engineer within a global DevOps team supporting highly available 24x7 production servicesImplement solutions using SRE best practices to improve service availability performance security and transparencyResolve incidents conduct root cause analysis and facilitate postincident reviewsParticipating in software architecture designStrong experience with the Software Development Life Cycle SDLC including requirements gathering design development testing deployment and maintenanceDefine application SLIs and SLOs build maintain observability to continuously monitor the operational performance create execute action plans to address failures to meet the desired metricsPlan and execute application infrastructure migration disaster recovery exercise and product upgradeEnhance automation and develop self-service capability to improve user experience and reduce manual effortProvide oncall support as part of a rotation to ensure rapid response to critical incidentsParticipate in scheduled maintenance activities including those that may occur during weekends to ensure system reliability and minimal disruption to users Key Skills and Qualifications for this role Hands on years of professional experience in Production Application Support or Site Reliability Engineering demonstrating strong troubleshooting resolution and issue prevention skills in high-pressure environmentsProficiency with automation build and monitoring tools such as Ansible Jenkins Prometheus and GrafanaStrong analytical and troubleshooting skillsStrong engineering skills with some of the following Java Python NodeJS plus SQLKnowledge to the Software Development Life Cycle SDLC and practice its principlesGood communication skills with the ability to collaborate effectively with globally dispersed cross functional teams and vendorsPrior experience supporting large Atlassian Jira and Confluence Data Centre instances is preferred but not required Candidates who demonstrate a strong ability to learn quickly and adapt to new technologies will also be consideredSolid knowledge on observability and monitoring tools such as Grafana and Prometheus is an advantage