Senior DevOps Engineer – Cloud & AI Infrastructure

Guldberggmbh — Germany · Posted ~3 hours ago

Senior

Skills

DevOps Kubernetes Go containers cloud infrastructure CI/CD GitOps Infrastructure as Code monitoring logging alerting GPU workloads ArgoCD Helm GPU

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior DevOps role focused on building and operating Kubernetes infrastructure, automating deployments, and optimizing stable CI/CD pipelines. You will establish GitOps and Infrastructure-as-Code practices, implement monitoring and proactive alerting, and operate compute-intensive AI and GPU workloads while partnering with multiple engineering teams.

Highlights

Own and improve Kubernetes-based cloud infrastructure, automated delivery pipelines, and AI/GPU workloads. The role combines platform reliability, security, observability, GitOps, and Infrastructure as Code while collaborating closely with backend, AI, and front-end engineering teams.

Description

Ihre Aufgaben: Aufbau und Betrieb der Kubernetes-InfrastrukturAutomatisierung der Infrastruktur mithilfe von ArgoCD und HelmEntwicklung und Optimierung stabiler CI/CD-PipelinesEinführung und Weiterentwicklung von GitOps- und Infrastructure-as-Code-ProzessenEinrichtung einer durchgängigen Überwachung und proaktiven AlarmierungSicherstellung der Verfügbarkeit, Skalierbarkeit und Sicherheit der PlattformBetrieb und Optimierung rechenintensiver KI- und GPU-WorkloadsEnge Zusammenarbeit mit Backend-, KI- und Frontend-EntwicklungsteamsHinterfragen und Verbessern der aktuellen DevOps-Struktur Ihre Qualifikationen: Unabdingbare Voraussetzungen (für diesen Kunden sollten nach Möglichkeit ALLE erfüllt sein):Fundierte DevOps-Kenntnisse auf Senior-NiveauErfahrung im deutschen ArbeitsumfeldFundierte Kenntnisse in der Programmiersprache GoErfahrung mit Kubernetes, Containern und Cloud-InfrastrukturKenntnisse in CI/CD, GitOps und Infrastructure as CodeErfahrung mit Überwachung, Protokollierung und AlarmierungFundierte Kenntnisse in Linux, Netzwerken und PlattformbetriebStarker Fokus auf Automatisierung und zuverlässige SystemeFähigkeit, fundiertes Fachwissen einzubringen und den Wissenstransfer innerhalb des Teams zu fördernErfahrung mit Terraform, Pulumi oder Rust