Staff Reliability Engineer

Ionos — Germany · Posted ~2 hours ago

Lead Full-time Visa History ✓

Skills

Ceph Linux networking distributed storage site reliability engineering infrastructure operations performance troubleshooting

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A staff-level reliability engineering role focused on keeping a large-scale distributed object-storage platform healthy and resilient. You will work deeply with Ceph, Linux, and networking, troubleshoot complex infrastructure issues, and improve reliability across multi-data-center systems handling very large data volumes.

Highlights

Staff-level infrastructure engineering opportunity working on a large-scale object storage platform spanning multiple data centers and petabytes of customer data. The role offers deep technical ownership across storage, Linux, networking, reliability, and high-scale infrastructure.

Description

About IONOS At IONOS, we don't just manage servers – we shape the digital future for more than 6.2 million customers worldwide with our solutions. As Europe's leading hosting provider and a pioneer in independent cloud solutions, we are building next-generation infrastructure – from sovereign cloud architectures and high-performance GPU clusters to integrated AI automation tools. What drives us is digital sovereignty for Europe and real impact for small and medium-sized businesses (SMBs) as well as large enterprises. We work in agile teams, rely on transparent structures, and believe that excellence and innovation only emerge through true team spirit. Ready for your next step? Become part of IONOS and grow with us.Our Object Storage platform runs on Ceph, spans multiple data centers, and holds double-digit petabytes of customer data. It is growing fast. We are looking for a staff-level engineer who knows Ceph, Linux and the network underneath it well enough to keep it healthy as it scales. We follow a "you build it, you run it" model. You will deploy, operate and improve the platform with the Network, SRE, Data Center and infrastructure support teams. A separate development team works on the product, and you will work closely with them. Your tasks Own the reliability, performance and capacity of our production Ceph clusters across locations Solid understanding of server hardware (disks, controllers, NICs) and how it affects storage performance, so you can diagnose hardware-related issues remotely Diagnose the hardest issues across the full stack, from disk and server through the Linux network stack to the Ceph service, and lead incident response and postmortems Automate deployments, upgrades and routine operations (Ansible) so they need less manual work Build monitoring and observability that catch problems before customers do Plan capacity and growth with the Data Center and Network teams Take part in the on-call rotation Follow Ceph releases and the community, and bring useful practices back to the team Your profile 5+ years as an SRE, Linux or storage engineer, with proven experience running Ceph in production (the more petabytes the better; 20 PB+ clusters are ideal) Strong knowledge of Linux and its network stack, plus good networking fundamentals Solid understanding of server hardware (disks, controllers, NICs) and how it affects storage performance Experience with object, block and file storage concepts Automation experience, ideally Ansible, and scripting in Python or Bash (you don't need to be a software engineer) Monitoring and observability experience (e.g., Prometheus, Grafana) Calm, structured troubleshooting when things break, and clear communication in English (German is a plus) Willingness to undergo extended security vetting (SÜ2) Benefits Hybrid working model. Flexible working hours through trust-based working hours. At some locations a subsidized canteen and various free drinks. Modern office space with very good transport connections. Various employee discounts for activities and products. Employee events such as summer and winter parties, as well as workshops. Numerous training and development opportunities. Various health offers, such as sports and health courses. Application Note We value diversity and welcome all applications – regardless of, for example, gender, nationality, ethnic or social origin, religion, disability, age as well as sexual orientation and identity, physical characteristics, marital status or any other irrelevant factor subject to applicable law.