Site Reliability Engineer

Hazeltree — United Kingdom · Posted ~23 hours ago

Mid Full-time

Skills

AWS infrastructure management monitoring containerization automation production operations reliability engineering EC2 VPC S3 RDS IAM Load Balancers Containers

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Join a mid-level reliability engineering team responsible for scalable, highly available cloud infrastructure. You’ll design and manage AWS environments, support containerized workloads, build monitoring and automation capabilities, and partner with developers to maintain stable and efficient production systems.

Highlights

Help design and operate highly available AWS systems with a strong focus on scalability, reliability, monitoring, automation, and efficient production operations while collaborating closely with software development teams.

Description

About Hazeltree Fund Services Inc. Hazeltree is the leading provider of financial treasury and data solutions to the alternative asset management industry, including hedge funds, private equity, pension and endowment funds, and traditional asset management firms. Hazeltree’s unique data networking value proposition delivers significant performance improvements and operational efficiency for its treasury clients. The firm’s global operations are based in New York, London, and Hong Kong. About the Role We are seeking a mid-level Site Reliability Engineer to support the design, scalability, and ongoing reliability of highly available systems on AWS. This role will be responsible for infrastructure reliability, monitoring, containerized workloads, and automation, working closely with development teams to ensure stable and efficient production operations. Key Responsibilities Design, deploy, and manage AWS infrastructure, including EC2, VPC, S3, RDS, IAM, Security Groups, Load Balancers, Lambda, and cost optimisation.Administer and troubleshoot Linux systems, including Ubuntu, RHEL, CentOS, and Amazon Linux, across production and staging environments.Deploy, manage, and troubleshoot containerized applications using Docker and Kubernetes, including deployments, services, ingress, Helm charts, and cluster upgrades.Build and maintain infrastructure monitoring solutions for CPU, memory, disk, network, and availability metrics using tools such as CloudWatch, Prometheus, Grafana, LogicMonitor, or Datadog.Implement application performance monitoring (APM) to track latency, errors, and throughput (e.g., Logic Monitor, ELK/EFK stack)Establish alerting and on-call escalation workflows to identify and address issues before they affect customers.Lead automation initiatives by replacing manual processes with scripts and Infrastructure as Code tools such as Terraform, Ansible, CloudFormation, Bash, or Python.Participate in incident response, root cause analysis, and post-incident reviews.Define and track SLIs/SLOs/error budgetsCollaborate with development teams to improve system reliability, scalability, and performance.Maintain clear documentation for runbooks, architecture, and operational procedures. Required Skills & Experience 3–5 years of hands-on experience in an SRE, DevOps, or infrastructure engineering roleStrong experience with AWS services (EC2, VPC, S3, IAM, RDS, ELB, Auto Scaling, Lambda)Solid Linux administration skills (shell scripting, systemd, networking, permissions, troubleshooting)Practical experience with containers and Kubernetes, including building images, writing Dockerfiles, deploying and scaling workloads, managing Kubernetes clusters, preferably EKS, and using Helm.Experience with infrastructure and application monitoring tools (CloudWatch, Prometheus, Grafana, LogicMonitor, New Relic, ELK/EFK, etc.)Proficiency in automation and scripting.Experience with Infrastructure as Code (Terraform, CloudFormation, or Ansible)Strong understanding of networking fundamentals, including DNS, load balancing, TCP/IP, and firewalls.Reasonable working knowledge of Windows Server administration, sufficient to support existing on premises/hybrid Windows infrastructure and assist internal IT operations as needed.Strong understanding of relational and non-relational databases (e.g., PostgreSQL, MySQL, MongoDB, DynamoDB), including query performance tuning, indexing, backups, and replication.Experience with incident management and on-call practicesStrong troubleshooting and problem-solving skills under pressure, with the ability to work independently on moderately complex issues. Nice to Have AWS certifications (Solutions Architect, SysOps Administrator) or CKA/CKADLinux certification (e.g., RHCSA/RHCE, LFCS/LFCE, or CompTIA Linux+)Experience with service mesh technologies such as AWS Lattice, Istio or Linkerd.Exposure to security best practices and compliance frameworksExperience working in a 24/7 production environmentBasic awareness of data warehousing concepts (e.g., ETL/ELT pipelines, data modeling) is a plus. What We Offer Competitive salary and performance-based bonuses.Comprehensive health, dental, and vision insurance plans.Retirement savings plan with company match.Opportunities for professional development and career advancement.A collaborative and supportive work environment. How to Apply Interested candidates are encouraged to apply directly through LinkedIn by clicking the Apply button on this posting and submitting their resume for the role. We are unable to consider candidates who require sponsorship or visa-related support, now or in the future. Hazeltree Fund Services Inc. is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.