Service Manager & Site Reliability Engineer

Gft Technologies Poland — Poland · Posted ~4 hours ago

Senior Full-time 18800-28900 PLN gross

Skills

Incident response Site Reliability Engineering Production monitoring Root cause analysis SLA management Python Kotlin Observability tools

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A reliability-focused role responsible for managing incidents, improving operational processes, enhancing system observability, and contributing to automation initiatives.

Highlights

Work at the intersection of reliability engineering and operations, improving system stability, incident processes, and automation capabilities.

Description

Type of contract: Employment contract Salary range: 18800 - 28900 PLN gross What will you do? You will manage real-time incident response, impact assessment, external communications, and coordination across production systems. Working at the intersection of Site Reliability Engineering, Incident Response, and Partner Operations, you will ensure timely, accurate, and SLA-compliant communication while supporting the scalability and reliability of global operations. Your tasks Monitor and respond to production incidentsCoordinate incident response activities across teamsAssess impact and determine incident severityManage external communications and status page updatesSupport incident reporting, RCA activities, and SLA trackingCollaborate with Engineering teams to improve reliability and observabilityDrive process improvements and automation initiativesContribute to internal reliability tooling using Python or Kotlin Your skills 5+ years of experience in Incident Operations, Site Reliability Engineering, Technical Operations, or a similar roleExperience working in on-call environments with SLA-driven responsibilitiesStrong understanding of distributed systems and production environmentsExperience with monitoring, alerting, and incident management toolsFamiliarity with APIs, system integrations, and observability platformsHands-on experience with Python or KotlinUnderstanding of SDLC and production reliability principlesStrong communication, stakeholder management, and decision-making skillsAbility to work effectively in high-pressure environments and manage multiple prioritiesStrong ownership mindset and cross-functional collaboration skills Nice to have Experience with Datadog or ChronosphereExperience with PagerDuty, Rootly, or Slack workflowsExperience managing external status pagesExperience with incident management automation and process improvementsExperience contributing to reliability tooling and platform engineering We offer Hybrid work in one of our locations: Lodz, Poznan, Krakow, Warsaw, Wroclaw (2 office days per week)Working in a highly experienced and dedicated teamBenefit package tailored to your needs (medical, sport, lunch subsidy, life insurance, etc.)Online training and certificationsAccess to e-learning platformWork From Anywhere (up to 140 days/year abroad)Social events