Platform Engineer

Thryvetalent — Germany · Posted ~2 days ago

Mid Full-time Remote

Skills

Kubernetes OpenShift Distributed Systems CI/CD Jenkins GitHub Actions Ansible ArgoCD Grafana Prometheus Kafka OpenTelemetry Redis PostgreSQL

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary

Support mission-critical distributed infrastructure, improve deployment reliability, and optimize containerized production environments using modern platform engineering tools.

Highlights

Operate large-scale distributed infrastructure, solve complex production challenges, and work remotely with modern cloud-native technologies.

Description

Platform Engineer — Kubernetes / Distributed Systems / AI Platform Remote anywhere in Germany | HQ in NRW | Work from anywhere for up to 180 days per year This is not a “keep the lights on” DevOps role... You’ll be part of the team responsible for running and operating a large-scale AI platform used in complex customer environments — including highly customised on-premise infrastructure deployments. The challenge here isn’t just Kubernetes. It’s making a highly distributed, containerised system reliably run in environments you don’t fully control. That means troubleshooting under pressure, improving deployment processes, working directly with customer-side infrastructure teams, and owning the operational reality of a production AI platform end-to-end. You’ll be working on systems running more than 1,000 containers in production across a large microservice architecture, helping improve everything from CI/CD pipelines and observability to runtime stability and deployment reliability. This role is heavily focused on runtime operations, incident handling, and delivery infrastructure — not feature development. The Engineering Muscle You Bring Experience with Kubernetes or OpenShiftStrong understanding of distributed systems and microservice architecturesExperience with CI/CD tooling such as Jenkins, GitHub Actions, Ansible, or ArgoCDExperience with monitoring and observability tooling such as Grafana, Loki, Prometheus, OpenTelemetry, Dynatrace, or InstanaKnowledge of technologies like Redis, Postgres, MariaDB, Kafka, Elastic, or MinioStrong troubleshooting and analytical skillsHands-on engineering mindset with a strong sense of ownershipFluent German and English communication skills Why This Role Appeals to People Who Like Complexity Large-scale production systems with 1,000+ containers running liveComplex Kubernetes and OpenShift environmentsReal operational ownership instead of pure maintenance workChallenging on-premise customer deploymentsExposure to modern AI platforms and distributed architecturesHigh-impact work with lots of technical depth and learning potential Khalifa@thryvetalent.com