Senior Site Reliability Engineer (Python/Kubernetes)

Luxoft Poland — Poland · Posted ~23 hours ago

Senior Remote

Skills

Site Reliability Engineering Python Kubernetes Cloud observability L3 incident resolution Automation AI tooling AI agents Operational engineering Monitoring

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Work remotely from Poland as a Senior Site Reliability Engineer focused on cloud observability and intelligent operations. You will resolve complex L3 issues, build AI agents for automated operations, develop automation that reduces manual work, monitor systems, and collaborate on engineering and run-operation improvements.

Highlights

Senior SRE role working on a cloud observability platform with a strong focus on automation and AI-assisted operations. Responsibilities include resolving advanced operational issues, building AI agents, reducing manual effort, and improving operational reliability.

Description

📍work from Poland Project Description: We're building a Cloud Observability platform for market leading large Germany based company with 440,000 customers worldwide. Originally known for leadership in enterprise resource planning (ERP) software, the company has evolved to become a market leader in end-to-end enterprise application software, database, analytics, intelligent technologies, and experience management. A top cloud company with 200 million users worldwide, the company helps businesses of all sizes and in all industries to operate profitably, adapt continuously, and achieve their purpose. Responsibilities: • Independently resolve L3 operational issues using engineering expertise, AI tooling, and purpose-built automations • Build and maintain AI agents for L1 & L2 automated operations • Minimize escalation rate to development teams over time • Develop automation solutions that reduce manual operational effort • Participate in development tasks for run-operation • Monitor platform availability and manage incident response/resolution by severity (P1/P2/P3) • Perform root cause analysis and determine if issues require code-level fixes • Participate in 24x7 shift rotation during weekdays and on-call on weekends Mandatory Skills: Incident ManagementKubernetesPythonSecurity Monitoring & ObservabilityTelemetry Mandatory Skills Description: • Strong Python development skills for automation and AI agent development • Kubernetes administration and troubleshooting • Experience with observability and monitoring platforms (Grafana, Prometheus, Jaeger, Splunk) • Incident management experience with severity-based response processes (P1/P2/P3) • CI/CD pipeline experience (GitHub Actions, ArgoCD) • Experience with AI/ML concepts for building intelligent automation • Strong root cause analysis skills • Willingness to work in 24x7 shift rotation Nice-to-Have Skills Description: • Experience with Generative AI / LLM for operational automation • Knowledge of OpenTelemetry and Kafka • Experience with OpenSearch / Elasticsearch • ITIL or similar service management certification Languages: English: C1 Advanced