Senior DevOps/SRE Engineer

Luxoft Poland — Poland · Posted ~2 hours ago

Senior Remote

Skills

DevOps Site Reliability Engineering SLO monitoring Alerting Application instrumentation Telemetry Observability Distributed systems Agile development SRE SLO

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Work from Poland as a senior DevOps/SRE engineer on a healthcare data platform. You will design and maintain SLO-based monitoring and alerting, improve application instrumentation and telemetry quality, and drive observability across distributed systems. The environment supports flexible work, Agile delivery, AI-assisted development, training, and career growth.

Highlights

Work from Poland on a challenging healthcare data platform, focusing on observability and reliability across distributed services, with flexible working, high-impact projects, training, and strong career support.

Description

📍work from Poland📍 Project Description: Join a healthcare data platform team building client-scoped SLO monitoring and alerting solutions. The project includes automated dashboard and alert generation across 110-140 services, as well as an internal application providing customer-specific SLA visibility. This role focuses on application instrumentation, telemetry quality, and observability adoption across distributed systems. AI-assisted development is expected and encouraged. In Luxoft, our culture is one that strives on solving difficult problems focusing on product engineering based on hypothesis testing to empower people to come up with ideas. Great Place to Work Institute certifies us as one of the Top 10 companies to work for in Mexico, and we do it with a truly flexible environment, high impact projects in Agile environments, a culture focused on results, training and strong support to grow your career. Responsibilities: - Design and maintain SLO-based monitoring and alerting solutions. - Create and optimize PromQL queries and multi-window burn-rate alerts. - Build and manage Grafana dashboards using configuration-as-code practices. - Develop automation and monitoring configuration generators in Python or TypeScript. - Contribute to Terraform-based observability infrastructure. - Validate monitoring signals and improve alert quality. Collaborate with engineering and platform teams to enhance reliability and operational visibility. Mandatory Skills Description: - Strong experience in Observability, SRE, or Platform Engineering. - Advanced PromQL (or equivalent) expertise, including the ability to identify misleading or incorrect query results. - Hands-on experience designing and implementing SLOs and multi-window burn-rate alerting. - Experience with Grafana provisioning and configuration as code. - Ability to build automation tools and code generators using Python or TypeScript. - Hands-on experience using AI-assisted development tools such as Claude Code, GitHub Copilot, or equivalent. AI-assisted engineering is an expected part of the development workflow. - Strong analytical mindset with a healthy skepticism toward telemetry data. - Experience validating signals through multiple independent sources before operationalizing alerts. Experience with cloud-native or distributed systems environments. Nice-to-Have Skills Description: - Groundcover - VictoriaMetrics - OpenTelemetry Collector - AWS EKS - Healthcare or regulated-industry experience - Experience building observability tooling and service health reporting solutions Domain: US Healthcare (PHI-adjacent) Okta and observability platform access AI-assisted engineering is a standard part of the development workflow. Languages: English: B2 Upper Intermediate