Lead Site Reliability Engineer

Bluerosetechnologies β€” Poland Β· Posted ~3 hours ago

Lead Contract Remote Visa History βœ“

Skills

SRE DevOps Incident management Distributed systems Python Kotlin Datadog PagerDuty APIs

πŸ”“ Log in to save this job, tailor your resume & track your apply process β€” 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A lead SRE/DevOps position focused on maintaining reliable distributed systems, managing incidents, improving automation, and supporting production environments in a global engineering team.

Highlights

Remote leadership role supporting mission-critical systems with focus on reliability, automation, and production excellence.

Description

🚨 Hiring: Lead Site Reliability Engineer (SRE / Devops) | Remote | Poland πŸ“ Location: Poland 🏠 Work model: Remote πŸ“† Job Type: Contract Are you a seasoned SRE / DevOps Engineer with strong incident management, distributed systems, and production operations experience? We’re looking for 2 Lead Site Reliability Engineers to join a global team supporting mission-critical systems in the Banking & Finance sector. πŸ§‘ πŸ’» What we're looking for βœ… 5+ years of experience in SRE / production operations βœ… Proven expertise in incident management & incident command βœ… Platform or software engineering background βœ… Hands-on Python and/or Kotlin experience βœ… Strong understanding of distributed systems & system design βœ… Experience with Datadog and/or Chronosphere βœ… Knowledge of PagerDuty and/or Rootly βœ… Strong troubleshooting, debugging & root cause analysis skills βœ… Experience with APIs, integrations & automation βœ… Excellent stakeholder management and communication skills βœ… Ability to perform effectively under pressure and ambiguity βœ… Fluent business & technical English πŸ› οΈ Tech Stack Python | Kotlin | Datadog | Chronosphere | PagerDuty | Rootly | APIs | Distributed Systems | Observability | Monitoring | Reliability Engineering πŸ”₯ What you'll do β€’ Lead and coordinate production incident response β€’ Drive incident resolution, impact assessment, and severity management β€’ Improve platform reliability, availability, and performance β€’ Manage external incident communications and status page updates β€’ Automate operational processes and reduce manual effort β€’ Collaborate with engineering and business stakeholders to improve system resilience β€’ Support SLA-driven operations and continuous service improvement β€’ Participate in global operational handovers and follow-the-sun support 🌍 This is an excellent opportunity to work on large-scale distributed systems in a globally distributed environment, focusing on reliability, resilience, observability, automation, and mission-critical production operations. πŸ“© Interested? Send me your updated CV via DM or reach out to discuss the opportunity. πŸ” Know an SRE/DevOps expert who could be a great fit? Tag them or share this post! #Hiring #SRE #SiteReliabilityEngineer #DevOps #LeadEngineer #DevOpsJobs #SREJobs #Python #Kotlin #DistributedSystems #IncidentManagement #Observability #Datadog #Chronosphere #PagerDuty #Rootly #CloudEngineering #BankingJobs #FinanceJobs #RemoteJobs #PolandJobs #TechJobs #NowHiring