Summary
✨ AI‑Generated
An SRE position supporting a large-scale enterprise platform, focused on reliability, resilience, availability, and performance. You will build monitoring and alerting systems, lead incident response and root-cause analysis, and automate operational processes to improve platform stability.
Highlights
Own reliability, resiliency, availability, and performance for large-scale enterprise platforms. The role provides ownership of observability, incident response, root-cause analysis, automation, and service reliability improvements.
Description
Job Overview
We are seeking a Systems Reliability Engineer (SRE) to join a Technical Operations team supporting a large-scale enterprise platform.
In this role, you will be responsible for the reliability, resiliency, availability, and performance of critical systems.
You will design and maintain monitoring and alerting solutions, lead incident response activities, drive root cause analysis, and develop automation to improve operational efficiency and platform stability.
The ideal candidate has strong experience operating enterprise-scale environments and a solid understanding of reliability engineering and observability practices.
Responsibilities
Own the reliability, resiliency, and availability of large-scale enterprise platforms, proactively identifying and mitigating risks to service continuity.Design, implement, and maintain comprehensive monitoring, observability, and alerting frameworks using tools such as Splunk, Dynatrace, Grafana, and Datadog.Define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to measure and improve platform health.Lead and participate in incident response activities, serving as a technical driver during remediation efforts.Coordinate with engineering, product, operations, and risk teams during incidents and service-impacting events.Lead and improve the Root Cause Analysis (RCA) process by investigating incidents, documenting events and remediation actions, and identifying underlying causes to prevent recurrence.Ensure timely creation, management, and tracking of incident tickets using platforms such as ServiceNow.Monitor incident aging, trends, and reporting to identify opportunities for operational improvements.Build automation and operational tooling to reduce manual effort and improve Mean Time to Detection (MTTD) and Mean Time to Resolution (MTTR).Collaborate with engineering and product teams to incorporate reliability and resiliency best practices throughout the platform lifecycle.Proactively identify opportunities to improve system performance, availability, monitoring, and operational efficiency.
Required Qualifications
Hands-on experience with monitoring, observability, and alerting tools, including Splunk, Dynatrace, Grafana, and Datadog.Proven experience operating and supporting large-scale enterprise platform environments.Demonstrated experience with incident response and leading or contributing to Root Cause Analysis (RCA)processes.Strong understanding of Site Reliability Engineering (SRE) principles, including availability, resiliency, monitoring, alerting, and system performance.Experience with incident management and ticketing workflows, such as ServiceNow.Strong troubleshooting and problem-solving skills.Excellent communication skills with the ability to coordinate technical remediation efforts across multiple teams.Ability to work effectively in a fast-paced environment and respond to high-priority incidents.
Preferred Qualifications
Experience working within financial services, payments, or embedded finance environments.Proficiency with scripting or programming languages such as Python, Go, or Bash.Familiarity with cloud platforms, containerization, and CI/CD pipelines.Experience defining and managing SLOs, SLIs, and error budgets.Experience developing automation and tooling to improve operational reliability and reduce repetitive manual tasks.Knowledge of modern cloud-native architectures and distributed systems.