Senior DevOps / Site Reliability Engineer

Robert Half International — United States · Posted ~1 day ago

Senior

Skills

DevOps Site Reliability Engineering Python Cloud Infrastructure Production Systems Incident Response Root Cause Analysis Observability Monitoring Alerting SLOs Distributed Systems

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A fast-growing technology organization is seeking a senior infrastructure engineer to own the reliability, availability, and performance of mission-critical production systems. You will build automation with Python, improve observability and SLOs, lead incident response, and optimize distributed cloud infrastructure while partnering closely with backend, data, and machine learning engineers.

Highlights

Highly visible, hands-on senior role with significant ownership of production reliability and scalability. Offers close collaboration with engineering teams, leadership in incident response, and opportunities to establish SRE best practices in a fast-growing environment.

Description

I'm partnering with a fast-growing Healthcare AI & Automation company that's looking to add a Sr DevOps as well as a Sr Site Reliability Engineer (SRE) to their team. This is a highly visible, hands-on role where you'll help scale and support production systems that power mission-critical solutions for enterprise customers. What You'll Be Doing Own the reliability, availability, and performance of production systems.Build automation, tooling, and operational workflows using Python.Partner closely with Backend, Data, and ML Engineering teams to improve platform scalability and resilience.Lead incident response efforts, root cause analysis, and post-mortem reviews.Establish and improve observability, monitoring, alerting, and SLOs.Help define and implement SRE best practices as the company continues to scale.Support and optimize distributed cloud infrastructure in a high-growth environment.What We're Looking For 7+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or related roles.Strong hands-on coding experience with Python.Deep understanding of Incident ManagementObservability & MonitoringPerformance TuningReliability EngineeringSLOs/SLIsAutomationExperience working in fast-paced environments with evolving requirements.Strong collaboration and communication skills.Tech Stack AWS (ECS, Lambda)TerraformDatadogGrafanaPostgreSQLModern CI/CD pipelinesPythonCompensation & Benefits Base Salary: Up to $180KStaff-Level Compensation: Up to $220KEquityUnlimited PTOComprehensive Benefits PackageFlexible Hybrid Schedule (Onsite Monday & Wednesday)This is an excellent opportunity for an engineer who enjoys ownership, solving complex reliability challenges, and helping build scalable infrastructure at a company where their impact will be felt immediately. Robert Half will consider qualified applicants with criminal histories in a manner consistent with the requirements of the San Francisco Fair Chance Ordinance.