Summary
✨ AI‑Generated
An experienced Sovereign Cloud / Site Reliability Engineer is sought to operate and improve highly secure, business-critical cloud-native and Kubernetes platforms. The role offers the opportunity to work with an experienced DevOps/SRE team and take ownership of reliable, continuous platform operations. Applicants must meet strict citizenship, German residence and employment, and security-clearance requirements.
Highlights
Work on highly secure, business-critical cloud platforms within an experienced DevOps/SRE team, with significant responsibility for reliability and continuous platform operation.
Description
Before You Apply – Mandatory Requirements
Due to the security requirements of the role, candidates must meet all of the following:
Citizenship: Citizenship of a country that is a full member of both the EU and NATO.
If you hold multiple citizenships, all must meet this requirement.Residence & Employment: Residence in Germany and direct employment through a German legal entity under a German employment contract and German labor law.
Employment through foreign entities or subcontractors is not permitted.Security Clearance: A valid and verifiable Ü2 security clearance in accordance with the German Security Clearance Act (SÜG) is mandatory.
Overview
For our growing team, we are looking for an experienced Sovereign Cloud / Site Reliability Engineer (SRE) to support the secure, reliable, and continuous operation of modern cloud-native and Kubernetes platforms.
You will work with an experienced DevOps/SRE team on highly secure, business-critical platforms, taking responsibility for platform operations, automation, monitoring, incident management, security, and continuous improvement.
The role focuses particularly on Kubernetes, CI/CD, Infrastructure as Code, observability, logging, and operational excellence within a highly available Unified Observability Platform in a 24x7 operational environment.
Key Technologies
Kubernetes | Gardener | Helm | Jenkins | ArgoCD | Git | IaC | Python | Go | Bash | Prometheus | PromQL | Thanos | OpenTelemetry | Grafana | Elasticsearch | OpenSearch | Logstash | Kibana | ServiceNow | PagerDuty | REST APIs
Key Responsibilities
Kubernetes & Platform Operations
· Operate and maintain Kubernetes clusters with Gardener, including workloads, deployments, Helm charts and platform components.
· Troubleshoot availability, performance and deployment issues.
· Ensure secure, scalable, resilient and highly available platforms.
CI/CD & Automation
· Operate and optimize Jenkins and ArgoCD pipelines for automated deployments.
· Implement IaC and Git-based deployment workflows.
· Develop automation and operational tools using Python, Go and/or Bash.
· Automate provisioning, health/compliance checks, alerting and reporting.
Monitoring & Observability
· Manage Prometheus, Thanos and OpenTelemetry environments, including scrape jobs, alert rules and PromQL.
· Develop and maintain Grafana dashboards and continuously improve monitoring and alerting.
Logging & Log Management
· Operate and optimize Elasticsearch/OpenSearch, Logstash and Kibana.
· Monitor log ingestion, storage, performance and reliability.
· Support centralized troubleshooting, anomaly detection, security and compliance.
Integration & Operations
· Integrate observability and logging platforms with ServiceNow, PagerDuty and other enterprise tools via secure APIs.
· Handle operational requests and participate in Scrum, DevOps and service-improvement activities.
· Collaborate with internal teams, SAP, suppliers and stakeholders.
Incident & Problem Management
· Participate in a 24/7 on-call and shift rotation, including weekends and public holidays.
· Respond to platform, monitoring, logging and deployment incidents.
· Perform RCA, support Major Incident Management (MIM) and implement sustainable corrective actions.
· Continuously improve platform stability, resilience and operational processes.