Site Reliability Engineer

Ubique Systems — Germany · Posted ~5 hours ago

Senior Full-time Remote

Skills

Kubernetes Linux Prometheus Grafana ELK Stack Monitoring Incident management Cloud-native platforms ELK Helm

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A site reliability engineering role focused on maintaining highly available platforms through monitoring, automation, troubleshooting, and cloud-native operations.

Highlights

Work on reliability, observability, automation, and security improvements for cloud-native systems in a remote operational environment.

Description

Site Reliability Engineer (SRE) – 24x7 Operational Support Location: Remote – Germany Working Pattern: 24x7 operational support / shift-based Role Overview We are looking for an experienced Site Reliability Engineer (SRE) with strong expertise in observability, monitoring, secure logging, automation and cloud-native platforms. The ideal candidate will have hands-on experience with Grafana, Prometheus and the ELK Stack, along with strong Linux, Kubernetes and DevOps skills. The role involves platform operations, incident management, scheduled maintenance, troubleshooting and engineering activities aimed at improving system reliability, performance and security. Key Responsibilities Manage and support Kubernetes and containerised environments, including Helm configurations.Maintain and optimise Prometheus, Grafana and Thanos monitoring solutions.Configure Prometheus scrape targets, alerting rules and PromQL queries.Administer and optimise the ELK Stack – Elasticsearch, Logstash and Kibana.Develop and maintain dashboards, alerts and secure logging solutions.Support CI/CD pipelines using Jenkins, ArgoCD and Git-based workflows.Develop automation scripts using Python, Bash and/or Go.Implement Infrastructure-as-Code (IaC) solutions.Participate in 24x7 on-call support, incident response and Major Incident Management (MIM).Troubleshoot platform, application, data, networking and performance issues.Execute scheduled maintenance activities in accordance with SOPs.Contribute to continuous improvement of operational procedures and platform reliability.Ensure systems comply with security, access control and data protection requirements.Mandatory Skills 5+ years of SRE / DevOps / Platform Engineering experience.Strong hands-on experience with Grafana and ELK Stack.Good experience with Prometheus and monitoring/observability platforms.Strong Linux administration and troubleshooting skills.Experience with Kubernetes and containers.Experience with Elasticsearch and/or OpenSearch.Knowledge of Helm and CI/CD tools such as Jenkins or ArgoCD.Strong scripting/programming skills in Python, Bash and/or Go.Good understanding of networking fundamentals and REST APIs.Proficiency in Git-based configuration and deployment workflows.Fluent English communication skills.Willingness to work 24x7 shifts, including weekends and public holidays. Good to Have: AWS and/or Azure experience.Elastic Certified Engineer.LPIC-2.Certified Kubernetes Administrator (CKA).Candidate Availability Candidates should ideally be immediately available or have a maximum notice period of 2 weeks. Note: Previously rejected candidates may also be reconsidered if they have strong Grafana + Prometheus + ELK experience. Please clearly mention their previous hiring history with HCLTech when resubmitting such profiles.