Senior Platform and Reliability Engineer

Contabo Gmbh β€” Germany Β· Posted ~5 hours ago

Senior Full-time Remote

Skills

Platform engineering Site reliability engineering Infrastructure architecture API gateways Ingress Persistent storage Secrets management Identity infrastructure WAF Observability Build vs buy evaluation Infrastructure troubleshooting Kong Nginx Ingress Ceph Longhorn Vault Keycloak Cloudflare

πŸ”“ Log in to save this job, tailor your resume & track your apply process β€” 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior platform and reliability engineering role with a remote-first setup and optional hybrid or on-site work. You will take architectural ownership of shared infrastructure services including API gateways, ingress, persistent storage, secrets and identity systems, edge security, WAF, and observability. The role involves improving mature systems, identifying weaknesses, evaluating aging components, and making pragmatic infrastructure decisions.

Highlights

Full-time permanent senior platform role with remote-first flexibility and architectural ownership across shared infrastructure, security, storage, identity, and observability systems.

Description

YOUR CREATIVE FIELD We are looking for a full-time, permanent Senior Platform & Reliability Engineer (all genders) to start as soon as possible. We live remote-first, but you have the freedom to choose whether you want to work hybrid or completely on-site due to your proximity to one of our locations (Berlin, Cologne, Hamburg, Munich). As a Senior Platform & Reliability Engineer at Contabo, you take architectural ownership of our shared infrastructure services – the foundation that multiple development teams build on: API gateway and ingress (Kong, Nginx Ingress), persistent storage (Ceph, Longhorn), secrets and identity infrastructure (Vault, Keycloak), edge security and WAF (Cloudflare), and our observability stack. You'll be taking over grown, partially under-documented systems – and that's exactly where the appeal of this role lies: you work your way deep into these systems, identify and remediate known weak points, evaluate aging components with solution-agnostic build-vs-buy reasoning, and decide which legacy pieces get fixed, replaced, or retired. One of your central mandates: you design and establish an on-call process, including runbooks for platform and infrastructure incidents – where no formal process exists today. In parallel, you mature our observability practices, rolling out distributed tracing, SLOs/SLIs, and meaningful dashboards across the stack, building on our existing tooling with Prometheus, Grafana, Alloy, and OpenTelemetry. In system design reviews, you bring your strong grounding in the fundamentals – load balancing, caching, sharding, replication, consistency trade-offs – and apply them to new and existing services alike. You won't be managing a team, but you will be the technical authority multiple teams rely on: you advise across teams on shared infrastructure services, document your knowledge consistently, and actively distribute it – so that critical know-how never again depends on a single person. Success in this role means: known risks are resolved, a solid on-call process is up and running, and single points of failure are measurably reduced platform-wide. The position is remote (Germany), with hybrid or on-site work optional; occasional travel for datacenter visits and team offsites is part of the role. WHAT CONVINCES US Your personality, paired with: 7+ years of experience in platform, infrastructure, or SRE roles, ideally with end-to-end responsibility for a private cloud or IaaS platform Hands-on production experience with distributed storage systems (Ceph) and Kubernetes persistent storage (Longhorn or comparable) Experience operating API gateways and ingress (Kong, Nginx Ingress, or comparable), including debugging cross-cutting concerns like CORS and rate limiting Experience with CDN/edge security, DDoS mitigation, and WAF configuration (e.g. Cloudflare), as well as firewall-rule design A strong grounding in system design fundamentals (load balancing, caching, sharding/replication, consistency models, message queues) – and the judgment to apply them to real-world trade-offs Experience designing or maturing observability (tracing, metrics, SLOs/SLIs) with tools such as Prometheus, Grafana, and OpenTelemetry is desirable Experience with secrets/identity infrastructure (Vault, Keycloak), messaging systems (NATS), and building on-call processes and incident runbooks from scratch is a plus Composure working with grown, incompletely documented systems – paired with the right mix of pragmatism and perfectionism: fix what's broken, retire what's dead, rather than rewriting everything Genuine enjoyment of acting as the technical go-to person across teams, actively sharing and documenting your knowledge Professional fluency in English (our working language); German language skills, certifications (CKA/CKS, Ceph training), and experience with virtualization platforms (Proxmox, OpenStack) are a plus AWESOME PROSPECTS At Contabo, we are constantly evolving. Our growth creates space for new ideas, ownership, and real impact for people who want to make a difference. What defines us is trust, direct communication, and an open feedback culture. We believe the best solutions are built together and that everyone has the opportunity to contribute, grow, and leave their own footprint. Additionally, we offer: Work the way that fits your life – Remote or hybrid with flexible hours for a healthy work-life balance and the same technical equipment at home as in the officeExperience real freedom – Workation across the EU and in our summer office in MallorcaStay active and healthy – Access to EGYM Wellpass and thousands of fitness and wellness facilitiesEnjoy exclusive perks – Attractive discounts on many products and services through our corporate benefits programRecharge your energy – 30 days of vacation plus additional days off on Christmas Eve and New Year’s EveGive back – An extra day off to get involved in social activities through our Volunteer DayMake an impact – Take ownership and bring your ideas to life with real creative freedomKeep growing – Individual professional and personal development opportunities in an innovative tech environmentFeel comfortable where you work – Modern, conveniently located offices across our European locationsBe part of a diverse team – Work in an international environment shaped by diversityCreate lasting memories – Company events and team activities you won’t forget HOW TO REACH US If you have any questions, you can reach out to Max by email at maximilian.kasemir@contabo.de. ABOUT US Contabo - High-quality Cloud Hosting Made in Germany! As a leading infrastructure service provider in the field of cloud instances and dedicated servers, we have stood for first-class quality, reliability and favorable prices since our foundation in Munich in 2003. Meanwhile, more than 300,000 customers in over 190 countries trust our services. We specialize in providing German quality to users worldwide, focusing on the stability and security of our data centers, the optimization of internal processes, and excellent customer support. Become a part of our future-oriented company, which plays a significant role in the development of digital cloud hosting, and support us with your ideas and commitment.