DevOps Engineer

Thryvetalent — Germany · Posted ~2 days ago

Remote

Skills

Kubernetes or OpenShift Distributed Systems Microservices CI/CD Monitoring Observability Troubleshooting Incident Handling German English Kubernetes OpenShift Jenkins GitHub Actions Ansible ArgoCD Grafana Loki Prometheus OpenTelemetry Redis PostgreSQL MariaDB Kafka Elasticsearch MinIO

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary

A technology organization is seeking a DevOps engineer to operate and improve a large-scale, containerized AI platform deployed across complex customer environments. You will troubleshoot production issues, strengthen deployment and delivery processes, improve observability and runtime stability, and work closely with customer infrastructure teams. The role focuses on operational ownership, incident response, Kubernetes or OpenShift, distributed systems, and CI/CD. Candidates must communicate fluently in German and English and be eligible to work in Germany.

Highlights

Remote work across Germany, extended work-from-anywhere flexibility, deep exposure to large-scale distributed systems, meaningful operational ownership, and strong opportunities to develop expertise in complex production environments.

Description

DevOps Engineer — Kubernetes / Distributed Systems / AI Platform Remote anywhere in Germany | HQ in NRW | Work from anywhere for up to 180 days per year This is not a “keep the lights on” DevOps role... You’ll be part of the team responsible for running and operating a large-scale AI platform used in complex customer environments — including highly customised on-premise infrastructure deployments. The challenge here isn’t just Kubernetes. It’s making a highly distributed, containerised system reliably run in environments you don’t fully control. That means troubleshooting under pressure, improving deployment processes, working directly with customer-side infrastructure teams, and owning the operational reality of a production AI platform end-to-end. You’ll be working on systems running more than 1,000 containers in production across a large microservice architecture, helping improve everything from CI/CD pipelines and observability to runtime stability and deployment reliability. This role is heavily focused on runtime operations, incident handling, and delivery infrastructure — not feature development. The Engineering Muscle You Bring Experience with Kubernetes or OpenShiftStrong understanding of distributed systems and microservice architecturesExperience with CI/CD tooling such as Jenkins, GitHub Actions, Ansible, or ArgoCDExperience with monitoring and observability tooling such as Grafana, Loki, Prometheus, OpenTelemetry, Dynatrace, or InstanaKnowledge of technologies like Redis, Postgres, MariaDB, Kafka, Elastic, or MinioStrong troubleshooting and analytical skillsHands-on engineering mindset with a strong sense of ownershipFluent German and English communication skills Why This Role Appeals to People Who Like Complexity Large-scale production systems with 1,000+ containers running liveComplex Kubernetes and OpenShift environmentsReal operational ownership instead of pure maintenance workChallenging on-premise customer deploymentsExposure to modern AI platforms and distributed architecturesHigh-impact work with lots of technical depth and learning potential Khalifa@thryvetalent.com