Systems Reliability Engineer (SRE) - Embedded Finance

Itmcsystem — United States · Posted ~7 hours ago

Mid Full-time

Skills

SRE Monitoring Splunk Dynatrace Grafana Datadog Incident Management Automation ServiceNow Root Cause Analysis Python Shell Scripting

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A financial technology company building an enterprise embedded finance platform needs an SRE to own system reliability, implement monitoring frameworks, lead incident response, and drive automation across distributed systems.

Highlights

Enterprise-scale embedded finance platform, industry-leading monitoring tools, impact critical financial infrastructure

Description

Position: Systems Reliability Engineer (SRE) – Embedded Finance Locations: Jacksonville, FL; Berkeley Heights, NJ; Alpharetta, GA; Toronto, ON Role Overview & Key Responsibilities: • Own the overall reliability, resiliency, and availability of an enterprise-scale Embedded Finance (EmFi) platform. • Design, implement, and maintain end-to-end monitoring and alerting frameworks using Splunk, Dynatrace, Grafana, and Datadog. • Define, track, and optimize platform health metrics including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. • Serve as the primary technical driver during incident response and high-severity outage remediation calls. • Lead the Root Cause Analysis (RCA) process—investigating incidents, documenting event sequences, and driving corrective actions to prevent recurrence. • Manage incident ticket lifecycles, tracking, aging, and reporting using ServiceNow. • Develop automation scripts and tools to reduce operational toil, MTTD, and MTTR. Qualifications & Required Skills: • Proven experience as a Systems Reliability Engineer (SRE) supporting large-scale enterprise platforms. • Hands-on proficiency with monitoring and observability tools: Splunk, Dynatrace, Grafana, and Datadog. • Demonstrated experience leading incident response calls and executing Root Cause Analysis (RCA). • Strong understanding of reliability engineering concepts (resiliency, fault tolerance, monitoring best practices). • Experience with ServiceNow or equivalent enterprise incident management ticketing systems. • Strong written and verbal technical communication skills. Nice to Have: • Background in Financial Services, Payments, or Embedded Finance. • Scripting proficiency in Python, Go, or Bash for operational automation. • Knowledge of cloud platforms, containers, and CI/CD deployment pipelines.