Senior DevOps/SRE Engineer

Luxoft Poland — Poland · Posted ~3 hours ago

Senior

Skills

DevOps SRE SLO monitoring Alerting Observability Application instrumentation Telemetry Distributed systems Monitoring automation SLA management SLO Distributed Systems Monitoring

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior DevOps/SRE opportunity focused on observability for a distributed healthcare data platform. You will design SLO-based monitoring and alerting, improve application instrumentation and telemetry quality, and automate dashboards and customer-specific visibility across more than one hundred services. AI-assisted development is encouraged.

Highlights

Design and maintain observability solutions across a large distributed service environment. The role provides high-impact work in SLO monitoring, telemetry quality, automated dashboards and alerting, with AI-assisted development encouraged and strong professional development support.

Description

Project description Join a healthcare data platform team building client-scoped SLO monitoring and alerting solutions. The project includes automated dashboard and alert generation across 110-140 services, as well as an internal application providing customer-specific SLA visibility. This role focuses on application instrumentation, telemetry quality, and observability adoption across distributed systems. AI-assisted development is expected and encouraged. In Luxoft, our culture is one that strives on solving difficult problems focusing on product engineering based on hypothesis testing to empower people to come up with ideas. Great Place to Work Institute certifies us as one of the Top 10 companies to work for in Mexico, and we do it with a truly flexible environment, high impact projects in Agile environments, a culture focused on results, training and strong support to grow your career. Responsibilities Design and maintain SLO-based monitoring and alerting solutions.Create and optimize PromQL queries and multi-window burn-rate alerts.Build and manage Grafana dashboards using configuration-as-code practices.Develop automation and monitoring configuration generators in Python or TypeScript.Contribute to Terraform-based observability infrastructure.Validate monitoring signals and improve alert quality.Collaborate with engineering and platform teams to enhance reliability and operational visibility. Skills Must have Strong experience in Observability, SRE, or Platform Engineering.Advanced PromQL (or equivalent) expertise, including the ability to identify misleading or incorrect query results.Hands-on experience designing and implementing SLOs and multi-window burn-rate alerting.Experience with Grafana provisioning and configuration as code.Ability to build automation tools and code generators using Python or TypeScript.Hands-on experience using AI-assisted development tools such as Claude Code, GitHub Copilot, or equivalent. AI-assisted engineering is an expected part of the development workflow.Strong analytical mindset with a healthy skepticism toward telemetry data.Experience validating signals through multiple independent sources before operationalizing alerts.Experience with cloud-native or distributed systems environments.Nice to have GroundcoverVictoriaMetricsOpenTelemetry CollectorAWS EKSHealthcare or regulated-industry experienceExperience building observability tooling and service health reporting solutionsDomain: US Healthcare (PHI-adjacent)Okta and observability platform accessAI-assisted engineering is a standard part of the development workflow. Languages English: B2 Upper Intermediate