Senior Site Reliability Engineer Team Lead

Viveka Services — Nepal · Posted ~4 hours ago

Lead Full-time

Skills

Site Reliability Engineering Team leadership Incident response SLI/SLO management Automation Cloud infrastructure SRE SLI SLO

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A growing technology organization is seeking a senior reliability engineering leader to establish and scale an SRE function. The role includes team building, automation, service reliability practices, and improving engineering operations.

Highlights

Foundational leadership opportunity to build an SRE practice from the ground up, establish reliability standards, and lead engineering improvements.

Description

About the Role Viveka is building its first Site Reliability Engineering function, and we're looking for someone to found and lead it. You'll define what SRE means here, set the reliability standards our services must meet, and build a team that treats operations as an engineering problem to be solved with software. This is a greenfield role. Viveka has not had a formal SRE practice before, so you'll set up the foundations: service level objectives, incident response, on-call, and a blameless learning culture. What You'll Do Build and lead the SRE team. Hire, mentor, and grow engineers, set team norms, and establish how SRE works with product and platform engineering.Define and adopt SLIs, SLOs, and error budgets. Work with product and engineering leaders to agree on the level of reliability each service needs, and use error budgets to balance feature velocity against stability.Engineer reliability through software. Write code to automate manual operational work, build tooling, and improve systems. Keep the team's operational load in check so engineering work stays the priority.Eliminate toil. Identify repetitive, manual, automatable work and systematically remove it.Own monitoring and alerting. Design observability that gives actionable signals rather than noise, and keep on-call alerts meaningful.Lead incident response. Establish incident management practices, including roles, escalation paths, and communication, and serve as an escalation point during major incidents.Drive postmortem culture. Write and review blameless postmortems, track action items to completion, and share lessons across the organization.Design a sustainable on-call program that works across time zones and protects engineers from burnout.Plan capacity and manage change. Partner with engineering on release safety, capacity planning, and launch readiness for new services.Advise on reliability in system design. Review architectures and influence engineering decisions early, before problems reach production. Required Qualifications Prior experience working as an SRE (required). You have held an SRE role, ideally at a company with production systems at meaningful scale.Strong programming experience (required). You write production-quality code (e.g., Python, Go, Java) for automation, tooling, and systems work, not just scripts and glue. SRE is software engineering applied to operations, and this role requires you to work that way.Experience writing and reviewing postmortems (required). You've authored postmortems, reviewed others', run blameless reviews, and made sure the resulting action items got done.Experience with incident management and on-call, including leading incidents.Working knowledge of SLIs, SLOs, and error budgets, and experience putting them into practice.Hands-on experience with cloud infrastructure (AWS, GCP, or Azure), containers and orchestration (e.g., Kubernetes), and infrastructure as code.Experience with observability tooling (metrics, logging, tracing, alerting).Strong written and verbal communication. You can explain reliability trade-offs to technical and non-technical stakeholders. Preferred Qualifications Experience building or scaling a new SRE team, or introducing SRE practices into an organization that lacked them.Experience managing or leading distributed and cross-cultural teams.Experience in healthcare, health tech, or other regulated environments (e.g., HIPAA), where reliability, security, and compliance intersect.Experience with capacity planning, load testing, and chaos or resilience testing. What Success Looks Like First 90 days: Assess the current state of reliability, define initial SLOs for critical services, and establish an incident response process and postmortem practice. Build the team to at least 3 including yourself.6 months: On-call rotation running across time zones, first automation wins reducing toil, and the team fully staffed or well into hiring. Build the team to at least 6.12 months: SRE is an established, trusted partner to engineering, with error budgets informing roadmap decisions and measurable improvements in reliability and incident recovery times. Team size can handle 24/7 on-call without burnout.