Sovereign Cloud Engineer

Vaspp Technologies — Germany · Posted ~3 hours ago

Senior Full-time

Skills

Cloud engineering Site Reliability Engineering Kubernetes Cloud-native platforms DevOps Security engineering Ü2 security clearance German employment law and employment requirements Cloud-native SRE

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An experienced Sovereign Cloud / Site Reliability Engineer is sought to operate and improve highly secure, business-critical cloud-native and Kubernetes platforms. The role offers the opportunity to work with an experienced DevOps/SRE team and take ownership of reliable, continuous platform operations. Applicants must meet strict citizenship, German residence and employment, and security-clearance requirements.

Highlights

Work on highly secure, business-critical cloud platforms within an experienced DevOps/SRE team, with significant responsibility for reliability and continuous platform operation.

Description

Before You Apply – Mandatory Requirements Due to the security requirements of the role, candidates must meet all of the following: Citizenship: Citizenship of a country that is a full member of both the EU and NATO. If you hold multiple citizenships, all must meet this requirement.Residence & Employment: Residence in Germany and direct employment through a German legal entity under a German employment contract and German labor law. Employment through foreign entities or subcontractors is not permitted.Security Clearance: A valid and verifiable Ü2 security clearance in accordance with the German Security Clearance Act (SÜG) is mandatory. Overview For our growing team, we are looking for an experienced Sovereign Cloud / Site Reliability Engineer (SRE) to support the secure, reliable, and continuous operation of modern cloud-native and Kubernetes platforms. You will work with an experienced DevOps/SRE team on highly secure, business-critical platforms, taking responsibility for platform operations, automation, monitoring, incident management, security, and continuous improvement. The role focuses particularly on Kubernetes, CI/CD, Infrastructure as Code, observability, logging, and operational excellence within a highly available Unified Observability Platform in a 24x7 operational environment. Key Technologies Kubernetes | Gardener | Helm | Jenkins | ArgoCD | Git | IaC | Python | Go | Bash | Prometheus | PromQL | Thanos | OpenTelemetry | Grafana | Elasticsearch | OpenSearch | Logstash | Kibana | ServiceNow | PagerDuty | REST APIs Key Responsibilities Kubernetes & Platform Operations · Operate and maintain Kubernetes clusters with Gardener, including workloads, deployments, Helm charts and platform components. · Troubleshoot availability, performance and deployment issues. · Ensure secure, scalable, resilient and highly available platforms. CI/CD & Automation · Operate and optimize Jenkins and ArgoCD pipelines for automated deployments. · Implement IaC and Git-based deployment workflows. · Develop automation and operational tools using Python, Go and/or Bash. · Automate provisioning, health/compliance checks, alerting and reporting. Monitoring & Observability · Manage Prometheus, Thanos and OpenTelemetry environments, including scrape jobs, alert rules and PromQL. · Develop and maintain Grafana dashboards and continuously improve monitoring and alerting. Logging & Log Management · Operate and optimize Elasticsearch/OpenSearch, Logstash and Kibana. · Monitor log ingestion, storage, performance and reliability. · Support centralized troubleshooting, anomaly detection, security and compliance. Integration & Operations · Integrate observability and logging platforms with ServiceNow, PagerDuty and other enterprise tools via secure APIs. · Handle operational requests and participate in Scrum, DevOps and service-improvement activities. · Collaborate with internal teams, SAP, suppliers and stakeholders. Incident & Problem Management · Participate in a 24/7 on-call and shift rotation, including weekends and public holidays. · Respond to platform, monitoring, logging and deployment incidents. · Perform RCA, support Major Incident Management (MIM) and implement sustainable corrective actions. · Continuously improve platform stability, resilience and operational processes.