Site Reliability Engineer

Saragossa โ€” United States ยท Posted ~1 day ago

Senior Onsite Visa History โœ“ $175000-$300000 base salary plus performance bonus

Skills

Site Reliability Engineering GCP Terraform Ansible CI/CD Datadog Python Distributed systems Observability Infrastructure automation Scalable systems

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Summary

Take ownership of reliability at scale as a Site Reliability Engineer within a highly engineering-driven organization. You will build internal tooling used by hundreds of engineers, work deeply across observability, automation, infrastructure, and distributed systems, and partner directly with engineering teams to establish strong reliability practices. The environment includes GCP, Terraform, Ansible, CI/CD, Datadog, and Python, with significant autonomy and a strong focus on engineering excellence.

Highlights

High-ownership SRE role focused on defining reliability practices across a large engineering organization. Offers substantial autonomy, exposure to modern cloud and observability tooling, close collaboration with engineering teams, and strong compensation with performance incentives.

Description

You're not here to maintain someone else's platform. You're here to define how reliability works across an entire engineering organization. This is a SRE role at one of the world's largest hedge funds, with a genuine engineering-first culture. Technology isn't treated as a support function here-it's a core part of the business, backed by significant investment and a track record that speaks for itself. You'll be building the tooling that hundreds of engineers rely on every day. That means getting deep into observability, automation, and infrastructure, and doing it at a scale where the decisions you make actually matter. You'll partner closely with engineering teams across the organization, driving automation and shaping reliability practices from the ground up. You'll have real ownership and the autonomy to move fast and do things properly. You'll be working across GCP, Terraform, Ansible, and CI/CD pipelines, with strong observability tooling like Datadog at the center of how you understand system health. You'll bring solid Python skills and a deep understanding of distributed systems and what it actually takes to keep them reliable at scale. Ideally you've already built or operated scalable, reliable systems in complex environments. You know how to work with engineering teams, not just support them, and you understand that the best reliability work happens when you're embedded in the product, not watching from the outside. The compensation reflects the seriousness of the role, starting at $175,000 - $300,000 base plus a highly competitive performance bonus, with the kind of exposure to best-in-class engineers and modern tooling that genuinely accelerates a career. Fully onsite Ready to own reliability at scale? Get in touch. No up-to-date CV required.