Site Resilience Engineer (SRE) - Hybrid Cloud Storage
Qumulo โ Ireland ยท Posted ~1 day ago
๐ Log in to save this job, tailor your resume & track your apply process โ 7 days free, no card needed.
Log in to add to target listDescription
Qumulo's cloud data platform manages exabytes of the world's most demanding data, unifying files, objects, and every workload across edge, core, and cloud.
Qumulo opened its European software R&D hub in Cork in March 2026, and this role is part of building that team out.
About the Role:
You'll be one of the first hires on a team with a single mandate: find out how Qumulo breaks before our customers do.
The platform manages exabytes of data for more than 1,100 customers across on-prem and every major cloud, and these are mission-critical workloads where a missed edge case becomes a customer's bad day.
This is a role for an engineer who thinks like a breaker.
You'll put on the customer's hat, work out how a feature will really get used, and design the tests that push it past its limits across hardware and cloud.
You'll automate the testing our principal engineers run by hand today, and decide what gets tested, how often, and why.
You'll help build this function from the ground up, including where we set the quality bar and which builds are good enough to ship.
As a Site Resilience Engineer at Qumulo, you will:
Design and operationalize testing for new features: work out how customers will actually use them, how to scale-test them, and how to break themAutomate the manual, repetitive testing our principal engineers run by hand today, using Python and our in-house frameworks on Jenkins and ArgoBuild a data-driven plan for which tests run, how often, and why, plus the framework to schedule and rerun themTroubleshoot build and test failures across VM instances and Qumulo-qualified hardware, from compile-time errors to integration failuresRead cluster output and C error logs to tell a test problem from an infrastructure problem from a real bugSet up monitoring and alerting so problems surface early (we use OpenMetrics, Grafana, InfluxDB, and Prometheus alongside home-grown tooling)Help set the quality bar for releases, including a real say in what shipsTake part in an on-call rotation for the systems your team owns
Our ideal SRE will have:
3+ years building and operating automated testing, validation, and/or certification for complex software systemsStrong programming ability in C.
Experience with distributed file systems, or parallel filesystems, would be a major plusA real breaker's instinct.
You go looking for edge cases and ask "what happens if I do this?" before anyone asks you toA track record of building tests yourself, not just running test plans handed to youHands-on experience across both on-premises infrastructure and cloud (AWS, GCP, or Azure), with a real grasp of where each one's limits areWorking fluency in Linux (we run Ubuntu) and PythonA data-driven approach to deciding what to test and how oftenExperience with orchestration tools (Ansible, Terraform), containers, and KubernetesSolid understanding of networks (routing, firewalls, security inspection devices, switch configuration) a plusStorage (IOPS, Latency, read/write patterns) or protocol experience (NFS, SMB, S3, etc) a strong plus
The annual pay range for the role is โฌ80,000โโฌ120,000 plus a comprehensive benefits package including equity, pension, and private healthcare.
Individual pay depends on various factors, such as role level, relevant experience, and skills.
We have 73,574 jobs that might be an even better fit for you
DontApply's real value goes far beyond a single job link or company name. Just upload your resume โ in under a minute we'll analyze all 73,574 jobs and tell you exactly which ones you should apply to right now.
Upload My Resume