Lead Site Reliability Engineer

Federal Reserve Bank Of San Francisco — United States · Posted ~4 hours ago

Lead Full-time

Skills

Cloud infrastructure DevOps Software engineering System reliability Scalability Monitoring SLI/SLO/SLA SLI SLO SLA

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A leadership-level reliability engineering position responsible for designing, improving, and operating resilient systems. The role combines software engineering, cloud infrastructure, and operational best practices to maintain high service availability.

Highlights

Lead a critical reliability engineering function focused on highly available systems, operational excellence, and modern cloud technologies.

Description

Company Federal Reserve Bank of San Francisco When you join the Federal Reserve—the nation's central bank—you’ll play a key role, collaborating with leading tech professionals to strengthen and protect our economic, financial and payments systems. We invest in contemporary and emerging technology each year to support the Federal Reserve and our economy, and we’re building a dynamic and diverse team for our future. We are seeking an experienced Lead Site Reliability Engineer to join our engineering team and drive the reliability, scalability, and performance of our critical systems. This role combines deep technical expertise in software engineering, cloud infrastructure, and DevOps practices to ensure our services meet the highest standards of availability and operational excellence. Responsibilities System Reliability & Performance Design, implement, and maintain highly available, scalable, and resilient systems across cloud infrastructure Establish and monitor SLIs, SLOs, and SLAs to ensure optimal system performance Lead incident response, conduct root cause analysis, and implement preventive measures Develop and maintain disaster recovery and business continuity plans Infrastructure & Automation Architect and manage cloud infrastructure on AWS using Infrastructure as Code (Terraform) Automate deployment pipelines, monitoring, and operational workflows Optimize cloud resource utilization and cost management Engineering & Development Build and maintain internal tools and services to improve operational efficiency Collaborate with development teams to implement reliability best practices Conduct code reviews and provide technical guidance on system design Develop monitoring solutions, alerting systems, and observability frameworks Security & Compliance Integrate security practices into CI/CD pipelines (SAST/DAST) Implement and maintain security controls across infrastructure and applications Ensure compliance with industry standards and regulatory requirements Conduct security assessments and vulnerability management Leadership & Collaboration Mentor junior SRE team members and promote SRE culture across the organization Partner with software engineering teams to improve system reliability Drive technical initiatives and contribute to architectural decisions Document processes, runbooks, and technical specifications Software Engineering: Strong proficiency in Java, Python, and Node.jsExperience with microservices architecture and distributed systemsSolid understanding of data structures, algorithms, and design patternsProficiency in writing clean, maintainable, and testable code Cloud Infrastructure (AWS): Extensive experience with AWS services including:Compute: Lambda, ECS, EC2, FargateStorage: S3, EBS, EFSDatabase: RDS, DynamoDB, AuroraNetworking: VPC, Route53, CloudFront, API GatewayMonitoring: CloudWatch, X-RayAWS certifications (Solutions Architect, DevOps Engineer) preferred DevOps & CI/CD: Expert-level knowledge of GitLab (CI/CD pipelines, runners, GitOps)Advanced Terraform skills for infrastructure provisioning and managementExperience with containerization (Docker) and orchestration (Kubernetes/ECS)Proficiency with configuration management tools Security: Hands-on experience with SAST (Static Application Security Testing) toolsKnowledge of DAST (Dynamic Application Security Testing) methodologiesUnderstanding of security best practices, OWASP Top 10, and compliance frameworksExperience with secrets management and identity access management (IAM) Monitoring & Observability: Experience with monitoring tools (Grafana, Datadog, New Relic, or similar)Log aggregation and analysis (CloudWatch Logs, Splunk)Distributed tracing with aws X-Ray Qualifications Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience7+ years of experience in Site Reliability Engineering, DevOps, or related roles3+ years in a lead or senior technical positionProven track record of managing large-scale production systemsExperience with on-call rotations and incident managementGenAI based Applications: Working knowledge of LLMs and agentic applications a plusExperience with serverless architectures and event-driven systemsFamiliarity with chaos engineering principles and practicesBackground in Agile/Scrum methodologiesExperience with multi-cloud or hybrid cloud environments The selected candidate will reside within a reasonable commuting distance, as defined by the employing Reserve Bank, and will work full-time onsite. Eligible Locations for Hire: Richmond, VA, San Francisco, CAThe following Reserve Bank locations are preferred due to the concentration of System IT team members in these locations: San Francisco, and Richmond, VA Base Salary Range: Min: $146,700 Mid: $190,500 Max: $234,300 (Location: San Francisco) The listed salary is applicable to 12th District/San Francisco. Final offers are determined by factors including the candidate’s qualifications, internal alignment considerations, district assignment, and geographic location. The Bank is committed to providing reasonable accommodations to individuals with disabilities to participate in the job application or interview process, perform essential job functions and receive other benefits and privileges of employment. The SF Fed is an Equal Opportunity Employer. If you need any assistance or accommodations due to a disability, please let us know at sf.hr.recruitment@sf.frb.org. Full Time / Part Time Full time Regular / Temporary Regular Job Exempt (Yes / No) Yes Job Category Information Technology Family Group Work Shift First (United States of America) The Federal Reserve Banks are committed to equal employment opportunity for employees and job applicants in compliance with applicable law and to an environment where employees are valued for their differences. Always verify and apply to jobs on Federal Reserve System Careers (https://rb.wd5.myworkdayjobs.com/FRS) or through verified Federal Reserve Bank social media channels. Privacy Notice