Senior Site Reliability Engineer (SRE)

Luxoft — Poland · Posted ~23 hours ago

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Description

🔥Become a Luxoft employee🔥 Our Benefits: 💰Paid Referrals 💻Equipment: laptop and monitor 🩺Private Medical & Dental care & Life Insurance covered 🏋🏽 ♀️ MyBenefit program (sports card, well-being program etc.) 🌎 Internal Mobility program - possibility of rotation between projects, locations, accounts 🎓 LuxTalent platform (webinars, training, courses) ...and more! Project Description: We're building a Cloud Observability platform for market leading large Germany based company with 440,000 customers worldwide. Originally known for leadership in enterprise resource planning (ERP) software, the company has evolved to become a market leader in end-to-end enterprise application software, database, analytics, intelligent technologies, and experience management. A top cloud company with 200 million users worldwide, the company helps businesses of all sizes and in all industries to operate profitably, adapt continuously, and achieve their purpose. Responsibilities: • Independently resolve L3 operational issues using engineering expertise, AI tooling, and purpose-built automations • Build and maintain AI agents for L1 & L2 automated operations • Minimize escalation rate to development teams over time • Develop automation solutions that reduce manual operational effort • Participate in development tasks for run-operation • Monitor platform availability and manage incident response/resolution by severity (P1/P2/P3) • Perform root cause analysis and determine if issues require code-level fixes • Participate in 24x7 shift rotation during weekdays and on-call on weekends Mandatory Skills Description: • Strong Python development skills for automation and AI agent development • Kubernetes administration and troubleshooting • Experience with observability and monitoring platforms (Grafana, Prometheus, Jaeger, Splunk) • Incident management experience with severity-based response processes (P1/P2/P3) • CI/CD pipeline experience (GitHub Actions, ArgoCD) • Experience with AI/ML concepts for building intelligent automation • Strong root cause analysis skills • Willingness to work in 24x7 shift rotation Nice-to-Have Skills Description: • Experience with Generative AI / LLM for operational automation • Knowledge of OpenTelemetry and Kafka • Experience with OpenSearch / Elasticsearch • ITIL or similar service management certification Languages: English: C1