Summary
An Application Site Reliability Engineer is sought to improve the reliability, scalability, and performance of critical enterprise production systems. The role combines cloud infrastructure, observability, automation, incident response, and DevOps practices while partnering with multiple engineering teams.
Highlights
Opportunity to support mission-critical production systems, lead reliability improvements through automation and observability, and collaborate across application, infrastructure, cloud, and engineering teams.
Description
Optomi, in partnership with a large national banking institution, is seeking a Application Site Reliability Engineer (SRE) to join their team.
This is an exciting opportunity to support and enhance the reliability, scalability, and performance of critical production systems within a highly available enterprise environment.
The ideal candidate will have experience with cloud infrastructure, observability, automation, incident response, and DevOps best practices, while partnering closely with application teams to improve platform stability and accelerate software delivery.
What The Right Professional Will Enjoy
Opportunity to support mission-critical production systems for a leading national financial institution.Driving reliability initiatives through automation, observability, and Site Reliability Engineering best practices.Collaborating with cross-functional application, infrastructure, cloud, and engineering teams to improve system health and operational excellence.Leading incident response efforts while promoting blameless postmortems and continuous service improvements.Leveraging modern AI-assisted tooling to accelerate incident triage and reduce mean-time-to-detect (MTTD).
Apply Today If Your Background Includes
5+ years of experience in Site Reliability Engineering, DevOps, Systems Engineering, or a related discipline supporting enterprise production environments.Experience supporting production applications, participating in on-call rotations, and maintaining high system availability and reliability.Hands-on experience with DevOps technologies such as GitLab, Git, and Infrastructure as Code tools including Terraform.Strong troubleshooting skills with experience building and enhancing observability using monitoring and logging platforms such as Dynatrace and Splunk.Experience developing and maintaining Service Level Objectives (SLOs), error budget tracking, production readiness standards, and operational health scoring.Proven ability to automate manual operational processes and identify opportunities to improve platform efficiency.Experience leading incident response efforts, conducting blameless postmortems, and partnering with development teams to improve logging, runbooks, and technical debt.Experience working with public cloud platforms, preferably AWS, including API Gateway technologies.Experience developing APIs, microservices, or frontend applications is a plus.Strong understanding of operational excellence, observability best practices, and continuous improvement within enterprise environments.