Senior Storage Engineer - Ceph / Distributed Storage

1984 Hosting — Iceland · Posted ~7 hours ago

Senior Full-time

Skills

distributed storage Ceph Linux storage architecture production operations reliability engineering performance optimization data recovery data integrity cloud infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

An experienced Senior Storage Engineer is sought to design and operate reliable distributed storage infrastructure for a growing public cloud platform. You will work with Ceph and Linux, focusing on reliability, performance, recovery, and data integrity while designing new systems and operating production environments. The role offers meaningful influence over architecture and close collaboration with networking and cloud infrastructure engineers.

Highlights

Deep technical ownership over distributed storage infrastructure, combining architecture design with production operations. The role provides strong influence over storage architecture and close collaboration with Linux, networking, and cloud infrastructure specialists.

Description

1984 Hosting is looking for an experienced Storage Engineer to help build and operate the storage infrastructure behind our existing services and our new Icelandic public cloud. 1984 has operated internet infrastructure for more than 20 years and provides shared hosting, email hosting, VPS services and other infrastructure services to thousands of customers in Iceland and internationally. We are currently expanding our infrastructure significantly and building a new public cloud environment in Iceland. Storage is a critical part of that platform, and we are looking for someone with deep practical experience in distributed storage systems. The role involves both designing new systems and operating production environments where reliability, performance, recovery and data integrity matter. You will work closely with our Linux, networking and cloud infrastructure teams and have a real influence on how our storage architecture develops over time. Areas of responsibility Design, deploy and operate distributed storage systems. Operate and develop Ceph environments. Build and maintain block, object and file storage solutions. Design failure domains, replication and erasure-coding strategies. Carry out performance testing and optimization. Perform capacity planning and support infrastructure growth. Diagnose and resolve complex storage incidents. Design and test recovery and disaster-recovery procedures. Build and improve monitoring and alerting. Automate recurring operational tasks. Contribute to the storage architecture of 1984's new public cloud. Work closely with Linux, virtualization and networking engineers. Technological work environment Our environment includes technologies such as: Ceph.LINSTOR / DRBD.S3-compatible object storage.NVMe and SSD storage.KVM.Apache CloudStack.Kubernetes.Debian GNU/Linux.Ansible.Terraform.Prometheus / Grafana.High-speed Ethernet.Distributed infrastructure across multiple datacenters. Skills We are primarily looking for substantial hands-on experience. Strong knowledge and practical experience with: Ceph in production environments.Ceph RBD, CephFS and/or RGW.Distributed storage architecture.Linux.Filesystems and block storage.SSD and NVMe storage.Replication, recovery and rebalancing.Performance analysis and troubleshooting.Backup, restore and disaster recovery.Monitoring and operating production systems.Experience with the following is highly desirable: LINSTOR / DRBD.StorPool or similar distributed block-storage platforms.S3-compatible object storage.Kubernetes CSI.KVM / QEMU.Apache CloudStack or similar IaaS platforms.Ansible.Terraform.Prometheus / Grafana.Python.Shell scripting.High-speed Ethernet. We do not expect candidates to know every technology listed above. What matters most is a deep understanding of storage systems and experience operating environments where disk failures, node failures, network issues and performance problems have real consequences. Formal education and certifications are useful but are not requirements. Practical experience and technical ability matter more.