Summary
✨ AI‑Generated
A senior SRE and DevOps engineering role focused on building secure, highly available cloud platforms. The position involves automation, infrastructure as code, reliability practices, and collaboration with engineering teams.
Highlights
Senior cloud reliability role with strong compensation, focus on scalable infrastructure, automation, observability, and advanced cloud practices.
Description
SRE / AWS DevOps Engineer
Location: Malvern, PA
Salary: $155,000 – $170,000 per year
Experience: Senior / Lead Level
Industry: Financial Services / Investment Management
The Opportunity
We are looking for an experienced SRE / AWS DevOps Engineer to join a large-scale financial services environment in Malvern, PA.
This role will focus on building and maintaining highly available, secure and scalable AWS platforms while driving Site Reliability Engineering, DevOps automation and modern observability practices.
You will work closely with engineering, cloud, data and machine learning teams, making this a particularly interesting opportunity for someone looking to gain further exposure to MLOps and AI/GenAI observability.
Key Responsibilities
Build, manage and optimise scalable and highly available AWS cloud infrastructureDrive Site Reliability Engineering (SRE) best practices across production environmentsDevelop and maintain Infrastructure as Code using Terraform, CloudFormation and AnsibleBuild and improve automated CI/CD pipelinesManage containerised environments using Kubernetes/EKS, Docker and HelmDefine and manage SLIs, SLOs and SLAsLead incident response, troubleshooting and Root Cause AnalysisImprove monitoring and observability across cloud platformsPartner with Data and ML teams around MLOps, model monitoring and AI observabilitySupport security and DevSecOps practices across AWS environments
Technical Experience
AWS: EC2, EKS, ECS, Lambda, S3, RDS, CloudWatch, IAM, VPC
IaC: Terraform, CloudFormation, Ansible
Containers: Kubernetes / EKS, Docker, Helm
CI/CD: Jenkins, GitHub Actions, GitLab CI/CD, AWS CodePipeline
Observability: Prometheus, Grafana, Datadog, ELK, OpenTelemetry, CloudWatch
Scripting: Python, Bash, PowerShell, SQL
Experience or knowledge of Arize would be beneficial but is not essential.
Exposure to MLOps, model monitoring, model drift, LLM observability or similar AI observability tooling would also be highly relevant.
Industry Experience
Experience within Financial Services, Investment Management, Banking, FinTech or Insurance would be highly desirable due to the scale, security and regulatory requirements of the environment.
Candidates coming from other large-scale, highly regulated or enterprise technology environments will also be considered where they can demonstrate strong production AWS and SRE experience.
What We're Looking For
Strong commercial experience across SRE, DevOps or Platform EngineeringDeep hands-on experience with AWSStrong Kubernetes/EKS and Terraform experienceExperience supporting large-scale, production-critical environmentsStrong understanding of reliability, availability and incident managementExcellent stakeholder communication and cross-functional collaborationPrevious technical leadership or Lead SRE experience would be advantageous
Package
$155,000 – $170,000 per year, alongside a comprehensive benefits package including medical, dental and vision coverage, 401(k), annual incentive, paid time off, parental leave and professional training/certification support.
Interested? Apply today or get in touch to discuss the opportunity in more detail.