Summary
✨ AI‑Generated
A Site Reliability Engineer is sought to improve the reliability, scalability, and performance of critical production systems. You will own projects from inception through deployment and ongoing operations, provide expert operational support, implement Infrastructure-as-Code practices, anticipate potential failures, and build preventative solutions. The role suits an engineer who values autonomy, technical ownership, and measurable operational impact.
Highlights
High-autonomy SRE role focused on reliability, scalability, and performance of critical systems. The position combines expert operational support with ownership of infrastructure projects from design through production, while offering substantial technical impact, documentation responsibilities, and opportunities to solve complex reliability challenges.
Description
Our client, a world-renowned hedge fund, is looking for an SRE to join their agile, high-performing engineering team.This role offers a unique opportunity to drive the reliability, scalability, and performance of our core platform with a high degree of autonomy and ownership.
You will take full ownership of projects from inception through to production operations, including documentation and knowledge transfer.
The successful candidate will divide their time between providing expert operational support for our critical systems and leading exciting new infrastructure projects.
Our mindset is to find the right person, not simply someone whose experience perfectly matches our current technology stack.
The ability to anticipate problems before they arise and develop solutions to prevent them is key.
If you enjoy a challenging environment, implementing Infrastructure-as-Code principles, and seeing the direct impact of your work, this is the place for you.
What You Will Do:
* Design & Build: Architect, deploy, and maintain highly scalable and reliable infrastructure on Google Cloud Platform (GCP) using Kubernetes and Infrastructure-as-Code tools.
* Automation: Champion automation across the entire software development lifecycle (SDLC), utilizing IaC, Python, and Bash to reduce manual effort and improve operational efficiency.
* Infrastructure-as-Code (IaC): Own and evolve our declarative infrastructure using Terraform for cloud resources and Helm for Kubernetes application deployments.
* Monitoring & Observability: Implement and manage robust monitoring, alerting, and logging solutions to ensure clear system visibility and proactive issue detection.
* Reliability & Performance: Define, measure, and enforce Service Level Objectives (SLOs) and Service Level Indicators (SLIs).
Participate in the on-call rotation, where applicable, and lead post-incident reviews to drive continuous improvement.
Required Experience & Skills
Core Technical Stack
* Cloud Platform: 4+ years of hands-on experience with Google Cloud Platform (GCP) or a comparable cloud infrastructure provider.
* Container Orchestration: Expert-level proficiency in managing, scaling, and troubleshooting production Kubernetes environments.
* Infrastructure-as-Code: Deep expertise in Terraform for managing cloud and Kubernetes resources.
* Deployment: Strong experience with Helm for packaging and deploying applications on Kubernetes.
* Scripting/Programming: Proficiency in at least one major programming language, preferably Python, for automation and tool development.
Tooling & Concepts
* CI/CD: Experience building, maintaining, and improving modern CI/CD pipelines.
* Observability: Practical experience implementing and managing monitoring, alerting, and logging solutions.
* Networking: Solid understanding of TCP/IP, load balancing, DNS, and cloud-native networking within Kubernetes environments.
* Operating Systems: Strong command-line skills and hands-on experience with Linux systems.