Lead Site Reliability Engineer, Infrastructure

Wellsfargo — United States · Posted ~2 hours ago

Lead

Skills

Site Reliability Engineering Infrastructure engineering Observability SLIs SLOs Error budgets Incident management Problem management Automation Operational resilience Technical leadership Mentoring SRE SLI SLO Agentic AI Low-code/No-code

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

Provide technical leadership for a team responsible for the reliability, availability, performance, and continuous improvement of enterprise workplace technology platforms. Establish SRE practices around observability, service objectives, incident management, automation, and resilience while using telemetry, intelligent automation, and AI-assisted operations to reduce recurring issues.

Highlights

Technical leadership, mentorship, enterprise-scale reliability work, focus on automation and resilience, and opportunities to apply AI to incident management and operational improvement.

Description

About This Role Wells Fargo is seeking a Lead Site Reliability Engineer to provide technical leadership, mentorship, and guidance to a team responsible for the stability, reliability, availability, performance, and continuous improvement of enterprise workplace technology platforms. In this role, you will establish and promote Site Reliability Engineering (SRE) standards and best practices across observability, service level indicators (SLIs), service level objectives (SLOs), error budgets, incident and problem management, automation, and operational resilience. You will leverage telemetry and data-driven insights to identify risks, reduce incident frequency and recurrence, and drive continuous reliability improvements. You will also help improve the stability of workplace technology platforms through intelligent automation, Agentic AI capabilities, and low-code/no-code solutions. This includes providing technical leadership for AI-assisted incident triage, root cause analysis, workflow automation, proactive remediation, and self-healing capabilities to improve operational efficiency and the end-user experience. In This Role, You Will Lead complex initiatives to develop infrastructure solutions that support business applications.Participate in projects intended to improve, modernize, and enhance technology infrastructure.Evaluate internal and external software solutions to support target-state architecture objectives.Review and analyze high-impact outages and implement processes to reduce future operational risk.Design, build, deploy, and maintain infrastructure solutions in collaboration with technology teams and third-party vendors.Design, code, test, debug, and document solutions using Agile development practices.Influence technical designs and implementation plans while identifying project risks and resource requirements.Provide technical leadership and guidance to engineers and partners across the organization.Direct risk and control activities by ensuring adherence to policies, procedures, and operational standards.Recommend solutions that improve efficiency, manage costs, and achieve business objectives.Collaborate with peers, leaders, customers, and vendors to resolve issues and deliver technology solutions. Required Qualifications: 5+ years of Technology Infrastructure Engineering and Solutions experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education5+ years of experience developing software, automation, or data processing solutions using Python5+ years of experience with observability and monitoring technologies such as Splunk, Grafana, Prometheus, Elastic, or similar tools Desired Qualifications: Experience providing technical leadership and mentoring engineersStrong knowledge of Site Reliability Engineering principles and practices, including observability, service level indicators (SLIs), service level objectives (SLOs), error budgets, incident management, problem management, and operational resilienceExperience developing Generative AI or Agentic AI solutionsKnowledge of large language models (LLMs), prompt engineering, retrieval-augmented generation (RAG), AI-assisted workflows, and API-based AI integrationsExperience with workflow automation and low-code/no-code platformsExperience using operational telemetry and data analytics to identify trends, anomalies, and opportunities for proactive remediationExperience with observability and monitoring platforms such as Splunk, Grafana, Prometheus, Elastic, or similar technologiesExperience with containers, Kubernetes, and Infrastructure as Code practicesStrong problem-solving, communication, and collaboration skillsAbility to provide technical direction and influence engineering practices across teams Job Expectations: Must be able to work onsite one of the posted locationsSponsorship is not available for this role Pay Range Reflected is the base pay range offered for this position. Pay may vary depending on factors including but not limited to demonstrated examples of prior performance, skills, experience, or work location. Employees may also be eligible for incentive opportunities. $119,000.00 - $224,000.00 Benefits Wells Fargo provides eligible employees with a comprehensive set of benefits, many of which are listed below. Visit Benefits - Wells Fargo Jobs for an overview of the following benefit plans and programs offered to employees. Health benefits401(k) PlanPaid time offDisability benefitsLife insurance, critical illness insurance, and accident insuranceParental leaveCritical caregiving leaveDiscounts and savingsCommuter benefitsTuition reimbursementScholarships for dependent childrenAdoption reimbursement Posting End Date: 1 Sep 2026 Job posting may come down early due to volume of applicants. We Value Equal Opportunity Wells Fargo is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other legally protected characteristic. Employees support our focus on building strong customer relationships balanced with a strong risk mitigating and compliance-driven culture which firmly establishes those disciplines as critical to the success of our customers and company. They are accountable for execution of all applicable risk programs (Credit, Market, Financial Crimes, Operational, Regulatory Compliance), which includes effectively following and adhering to applicable Wells Fargo policies and procedures, appropriately fulfilling risk and compliance obligations, timely and effective escalation and remediation of issues, and making sound risk decisions. There is emphasis on proactive monitoring, governance, risk identification and escalation, as well as making sound risk decisions commensurate with the business unit's risk appetite and all risk and compliance program requirements. Applicants With Disabilities To request a medical accommodation during the application or interview process, visit Disability Inclusion at Wells Fargo . Drug and Alcohol Policy Wells Fargo maintains a drug free workplace. Please see our Drug and Alcohol Policy to learn more. Wells Fargo Recruitment And Hiring Requirements Third-Party recordings are prohibited unless authorized by Wells Fargo. Wells Fargo requires you to directly represent your own experiences during the recruiting and hiring process. Reference Number R-569696-1