Sr. Site Reliability Engineer

Meridianlink โ€” United States ยท Posted ~5 hours ago

๐Ÿ”“ Log in to save this job, tailor your resume & track your apply process โ€” 7 days free, no card needed.

Log in to add to target list

Description

About The Role We are seeking a Senior Site Reliability Engineer to join our cloud engineering team. You will own the reliability, scalability, and observability of our critical financial SaaS applications and infrastructure, working across cloud platforms to ensure our customers experience is seamless, secure, and performant services. This is a high-impact role for someone who is passionate about building resilient systems and preventing outages before they happen. Key Responsibilities Design, implement, and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs) across all critical systems; ensure we meet or exceed targets consistentlyLead observability strategy by designing comprehensive monitoring, logging, and tracing architectures; select and deploy observability tools that provide deep visibility into system behaviorBuild and own runbooks, incident response procedures, and post-incident review processes; mentor the team on incident management and blameless postmortemsArchitect and deploy cloud infrastructure on AWS or Azure; implement infrastructure-as-code practices and ensure high availability, disaster recovery, and business continuityDevelop automation and AIOps capabilities to reduce toil, accelerate incident detection, and enable self-healing systems; implement intelligent alerting to minimize false positivesDrive reliability improvements through load testing, chaos engineering, and failure scenario analysis; identify and eliminate single points of failurePartner with application and backend teams to design reliable systems from inception; conduct architecture reviews and reliability assessmentsWrite production-grade Python tooling for automation, metrics collection, alert management, and operational workflowsChampion security and compliance in infrastructure; implement defense-in-depth principles for a regulated fintech environment Required Qualifications 7+ years in Site Reliability Engineering, DevOps, platform engineering, or closely related roles with significant responsibility for production systemsExpert-level experience with Azure or AWS (or both); deep knowledge of compute, networking, storage, and managed services; experience managing infrastructure at scaleDemonstrated expertise in observability: designing and implementing monitoring, alerting, logging, and distributed tracing solutions; hands-on with observability platforms (e.g., Prometheus, Grafana, ELK, Datadog, New Relic, or similar)Strong background in SLOs, SLIs, and SLAs; experience defining meaningful objectives and building systems to meet them; understanding of error budgets and their role in prioritizationProven experience designing and troubleshooting highly available, resilient, and scalable systems; deep understanding of distributed systems concepts and failure modesProficiency in Python, PowerShell, bash, etc. scripting languages for production automation, tooling, and systems programming; ability to write clean, maintainable code for operational workflowsHands-on experience with AIOps practices: event correlation, intelligent alerting, predictive analytics, and automated remediation; familiarity with AIOps platforms is a plusExperience with infrastructure-as-code tools (e.g., Terraform, CloudFormation, Ansible); version control and CI/CD pipeline designTrack record of incident management and on-call ownership; comfort with incident response and the ability to remain calm under pressureExcellent communication skills; ability to work cross-functionally and influence without authority; comfort mentoring junior engineers Preferred Qualifications Experience in the fintech, payments, banking, or other regulated industries; understanding of compliance requirements (SOC 2, PCI-DSS, etc.)Experience with Kubernetes and container orchestration; deep knowledge of containerized application deployment and managementProficiency with observability as code; experience building custom metrics, dashboards, and alerts programmaticallyBackground in chaos engineering or reliability testing; experience using tools like Gremlin or similar platformsContribution to open-source observability or infrastructure projectsExpertise in network security, application security, or infrastructure hardeningExperience with database optimization, query performance tuning, and backup/recovery strategies