Lead Site Reliability Engineer

Hconnect Company Limited — Thailand · Posted ~6 hours ago

Lead Full-time Onsite

Skills

Site reliability engineering Incident response On-call operations Observability Production operations Infrastructure reliability Redis RabbitMQ SRE Incident management Cloud infrastructure

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A lead SRE role responsible for establishing and owning site reliability practices within a rapidly scaling technology organization. You will lead production incident response, on-call operations, and observability while partnering with engineering and product leadership to improve system reliability and operational maturity.

Highlights

Foundational SRE leadership opportunity with ownership of incident response, on-call operations, observability, and reliability improvements in a rapidly scaling international technology environment.

Description

LEAD SITE RELIABILITY ENGINEER (SRE) 📍 Location: Bangkok, Thailand | On-site 🏢 Industry: B2B SaaS / HRTech 👤 Reporting to: CTO / Engineering Leadership Company Overview Our client is a fast-growing international B2B SaaS / HRTech company headquartered in Bangkok, with a global customer base spanning more than 135 countries. The company provides AI-powered technology that helps businesses modernize and streamline their recruitment and staffing processes. As the platform and customer base continue to grow, the company is establishing a dedicated Site Reliability Engineering (SRE) function for the first time. Role Overview The Lead Site Reliability Engineer will be the first dedicated SRE hire in the organization. Working closely with the CTO, Director of Engineering, and Director of Product, this role will own incident response, on-call operations, and the observability stack. When production incidents occur, this person takes the lead. Between incidents, the focus is on improving systems, monitoring, alerting, reliability, and processes to detect and prevent future issues. This is a senior individual contributor role with significant cross-functional ownership. There are no direct reports initially, although the function may expand over time. Key Responsibilities Incident Response • Lead production incidents from triage through resolution. • Coordinate response across engineering teams and communicate status to stakeholders. • Own incident management end-to-end: detection, escalation, communication, resolution, and post-mortem. • Create and maintain incident runbooks. • Conduct blameless post-mortems and track follow-up actions through completion. On-Call Operations • Own and manage the on-call rotation, schedules, escalation policies, and response expectations. • Reduce alert noise and false positives while ensuring alerts remain actionable. • Analyze recurring on-call issues and work with engineering teams to implement permanent fixes. Observability & Monitoring • Own the observability stack, primarily Datadog for infrastructure monitoring, APM, and logging, together with Sentry for error monitoring. • Build and maintain dashboards, alerts, SLOs, and monitoring coverage. • Ensure critical user flows and API endpoints are properly monitored. • Identify and close gaps in observability coverage. Security Incident Response • Lead responses to security incidents such as unauthorized access, credential compromise, and data breaches, in coordination with security leadership. • Ensure security alerts, escalation paths, and monitoring are properly configured. • Support penetration-test remediation validation and post-incident security reviews. Reliability & Process Improvement • Identify systemic reliability weaknesses such as single points of failure, insufficient redundancy, and inadequate failover. • Partner with engineering teams to improve deployment, rollback, and change-management practices. • Establish and monitor SLOs/SLIs for critical services. • Report on reliability trends and drive continuous improvement. Requirements • 5+ years of experience in SRE, DevOps, Production Engineering, or a similar role within a technology company. • Proven experience leading production incidents end-to-end, from initial triage through cross-functional coordination, resolution, and post-mortem. • Strong hands-on experience with Datadog or similar observability platforms, such as Dynatrace, New Relic, Grafana, or Prometheus. • Experience with infrastructure monitoring, APM, log management, dashboard creation, and alert configuration. • Strong production experience with Kubernetes, particularly AWS EKS. • Ability to independently troubleshoot pod failures, resource constraints, networking issues, and deployment problems. • Solid experience operating AWS managed services in production, including RDS, Amazon MQ, and OpenSearch. • Experience diagnosing performance issues, understanding failover behavior, and tuning configurations. • Experience with Sentry or equivalent application error-monitoring platforms. • Proven ability to design and manage on-call rotations, escalation policies, and alerting strategies. • Experience defining and tracking SLOs/SLIs to guide reliability decisions. • Excellent written and spoken English. • Willingness to participate in on-call outside standard working hours and take a leading role in maintaining platform availability. Nice to Have • Experience building an SRE function from scratch in a growing technology company. • Python proficiency, ideally with Django or FastAPI. • Experience with security incident response. • B2B SaaS or platform-company experience. • Experience with CI/CD and deployment automation, particularly ArgoCD or similar tools. Working Environment • Based at the company's Bangkok office in the city center. • International team with diverse cultures and nationalities. • English is the primary working language. • The team works on-site and values in-person collaboration. • International, fast-growing B2B SaaS environment. • Medical healthcare coverage. • Personal development allowance. • Work-from-anywhere benefit. • Regular team-building activities. Hiring Process The recruitment process may include some or all of the following: 1. Introduction interview with HR 2. Interview with CTO and/or Director of Engineering 3. Technical interview and assessment 4. Cultural-fit interview with senior management Tech Stack Infrastructure: AWS | Kubernetes / AWS EKS | ArgoCD / GitOps | GitHub Actions | Terraform Backend: Python | Django | FastAPI | PostgreSQL | MongoDB | OpenSearch / Elasticsearch | Redis | Celery | RabbitMQ Frontend: TypeScript | Vue | React