Lead Site Reliability Engineer

Hconnectpro — Thailand · Posted ~3 hours ago

Lead Full-time Onsite

Skills

Site Reliability Engineering Cloud infrastructure Incident response On-call operations Observability Production systems management Cloud platforms Monitoring tools Observability systems Automation tools Kubernetes

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A growing technology organization is seeking a Lead Site Reliability Engineer to build and own reliability practices from the ground up. The role focuses on production stability, incident response, monitoring, automation, and improving engineering operations while collaborating with senior technical leaders.

Highlights

Leadership role establishing a new reliability function with ownership of critical production systems, incident management, and infrastructure improvements. Opportunity to work closely with engineering leadership in a growing international technology environment.

Description

LEAD SITE RELIABILITY ENGINEER(SRE)📍 Location: Bangkok, Thailand | On-site 🏢 Industry: B2B SaaS / HRTech 👤 Reporting to: CTO / Engineering Leadership Company OverviewOur client is a fast-growing international B2B SaaS / HRTech company headquartered in Bangkok, with a global customer base spanning more than 135 countries. The company provides AI-powered technology that helps businesses modernize and streamline their recruitment and staffing processes. As the platform and customer base continue to grow, the company is establishing a dedicated Site Reliability Engineering (SRE) function for the first time. Role OverviewThe Lead Site Reliability Engineer will be the first dedicated SRE hire in the organization. Working closely with the CTO, Director of Engineering, and Director of Product, this role will own incident response, on-call operations, and the observability stack. When production incidents occur, this person takes the lead. Between incidents, the focus is on improving systems, monitoring, alerting, reliability, and processes to detect and prevent future issues. This is a senior individual contributor role with significant cross-functional ownership. There are no direct reports initially, although the function may expand over time. Key ResponsibilitiesIncident Response• Lead production incidents from triage through resolution. • Coordinate response across engineering teams and communicate status to stakeholders. • Own incident management end-to-end: detection, escalation, communication, resolution, and post-mortem. • Create and maintain incident runbooks. • Conduct blameless post-mortems and track follow-up actions through completion. On-Call Operations• Own and manage the on-call rotation, schedules, escalation policies, and response expectations. • Reduce alert noise and false positives while ensuring alerts remain actionable. • Analyze recurring on-call issues and work with engineering teams to implement permanent fixes. Observability & Monitoring• Own the observability stack, primarily Datadog for infrastructure monitoring, APM, and logging, together with Sentry for error monitoring. • Build and maintain dashboards, alerts, SLOs, and monitoring coverage. • Ensure critical user flows and API endpoints are properly monitored. • Identify and close gaps in observability coverage. Security Incident Response• Lead responses to security incidents such as unauthorized access, credential compromise, and data breaches, in coordination with security leadership. • Ensure security alerts, escalation paths, and monitoring are properly configured. • Support penetration-test remediation validation and post-incident security reviews. Reliability & Process Improvement• Identify systemic reliability weaknesses such as single points of failure, insufficient redundancy, and inadequate failover. • Partner with engineering teams to improve deployment, rollback, and change-management practices. • Establish and monitor SLOs/SLIs for critical services. • Report on reliability trends and drive continuous improvement. Requirements• 5+ years of experience in SRE, DevOps, Production Engineering, or a similar role within a technology company. • Proven experience leading production incidents end-to-end, from initial triage through cross-functional coordination, resolution, and post-mortem. • Strong hands-on experience with Datadog or similar observability platforms, such as Dynatrace, New Relic, Grafana, or Prometheus. • Experience with infrastructure monitoring, APM, log management, dashboard creation, and alert configuration. • Strong production experience with Kubernetes, particularly AWS EKS. • Ability to independently troubleshoot pod failures, resource constraints, networking issues, and deployment problems. • Solid experience operating AWS managed services in production, including RDS, Amazon MQ, and OpenSearch. • Experience diagnosing performance issues, understanding failover behavior, and tuning configurations. • Experience with Sentry or equivalent application error-monitoring platforms. • Proven ability to design and manage on-call rotations, escalation policies, and alerting strategies. • Experience defining and tracking SLOs/SLIs to guide reliability decisions. • Excellent written and spoken English. • Willingness to participate in on-call outside standard working hours and take a leading role in maintaining platform availability. Nice to Have• Experience building an SRE function from scratch in a growing technology company. • Python proficiency, ideally with Django or FastAPI. • Experience with security incident response. • B2B SaaS or platform-company experience. • Experience with CI/CD and deployment automation, particularly ArgoCD or similar tools. Working Environment• Based at the company's Bangkok office in the city center. • International team with diverse cultures and nationalities. • English is the primary working language. • The team works on-site and values in-person collaboration. • International, fast-growing B2B SaaS environment. • Medical healthcare coverage. • Personal development allowance. • Work-from-anywhere benefit. • Regular team-building activities. Hiring ProcessThe recruitment process may include some or all of the following: Introduction interview with HRInterview with CTO and/or Director of EngineeringTechnical interview and assessmentCultural-fit interview with senior managementTech StackInfrastructure: AWS | Kubernetes / AWS EKS | ArgoCD / GitOps | GitHub Actions | Terraform Backend: Python | Django | FastAPI | PostgreSQL | MongoDB | OpenSearch / Elasticsearch | Redis | Celery | RabbitMQ Frontend: TypeScript | Vue | React