Senior Site Reliability Engineer

Thrive It Systems Ltd — Poland · Posted ~2 hours ago

Senior Hybrid

Skills

Site Reliability Engineering DevOps High availability Incident management Root cause analysis Software architecture SDLC SLI/SLO Observability Infrastructure migration Disaster recovery Automation SRE SLI SLO

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior Site Reliability Engineer is needed to support highly available 24x7 production services in a global DevOps environment. You will apply SRE practices, define SLIs and SLOs, build observability, handle incidents and root-cause analysis, participate in architecture, automate operational processes, and plan migrations and disaster-recovery activities.

Highlights

Work within a global DevOps organization supporting highly available 24x7 services. The role provides broad ownership across reliability, observability, security, architecture, automation, disaster recovery, and continuous operational improvement.

Description

Role: Site Reliability Engineer Location: Poland, Krakow Mode: Hybrid - 3 day a week Description Serve as a Site Reliability Engineer within a global DevOps team supporting highly available 24x7 production servicesImplement solutions using SRE best practices to improve service availability performance security and transparencyResolve incidents conduct root cause analysis and facilitate post-incident reviewsParticipating in software architecture designStrong experience with the Software Development Life Cycle SDLC including requirements gathering design development testing deployment and maintenanceDefine application SLIs and SLOs build maintain observability to continuously monitor the operational performance create execute action plans to address failures to meet the desired metricsPlan and execute application infrastructure migration disaster recovery exercise and product upgradeEnhance automation and develop self-service capability to improve user experience and reduce manual effortProvide on-call support as part of a rotation to ensure rapid response to critical incidentsParticipate in scheduled maintenance activities including those that may occur during weekends to ensure system reliability and minimal disruption to users Key Skills and Qualifications for this role Hands on years of professional experience in Production Application Support or Site Reliability Engineering demonstrating strong troubleshooting resolution and issue prevention skills in high-pressure environmentsProficiency with automation build and monitoring tools such as Ansible Jenkins Prometheus and GrafanaStrong analytical and troubleshooting skillsStrong engineering skills with some of the following Java Python NodeJS plus SQLKnowledge to the Software Development Life Cycle SDLC and practice its principlesGood communication skills with the ability to collaborate effectively with globally dispersed cross functional teams and vendorsPrior experience supporting large Atlassian Jira and Confluence Data Centre instances is preferred but not required Candidates who demonstrate a strong ability to learn quickly and adapt to new technologies will also be consideredSolid knowledge on observability and monitoring tools such as Grafana and Prometheus is an advantage Recruiter's Email : shikharsharma@thriveitsystems.com