Lead Site Reliability Engineer

Talentia Consulting8686 — Thailand · Posted ~6 hours ago

Lead Full-time

Skills

SRE DevOps Kubernetes AWS Incident management Observability EKS RDS Datadog Sentry Grafana Prometheus

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A lead SRE position responsible for production reliability, incident management, cloud infrastructure, monitoring, and improving operational excellence across large-scale systems.

Highlights

Lead reliability engineering initiatives, manage critical systems, and shape operational practices for a growing technology environment.

Description

Location: Bangkok, Thailand Salary: Up to THB 200,000 gross/month We’re supporting a fast-growing, global B2B SaaS / HRTech company backed by leading VCs, hiring its first dedicated Lead SRE. Key Responsibilities Lead production incident response, post-mortems & incident managementOwn on-call operations, escalation policies and alerting strategyOwn Datadog / Sentry observability, dashboards, monitoring & SLOs/SLIsTroubleshoot Kubernetes / AWS EKS production environmentsManage AWS services including RDS, Amazon MQ & OpenSearchDrive platform reliability, security incident response and process improvements Requirements 5+ years in SRE / DevOps / Production EngineeringStrong hands-on experience with Kubernetes & AWSProven experience leading production incidents end-to-endStrong observability experience: Datadog, Grafana, Prometheus, Dynatrace, New Relic, etc.Experience with SLO/SLI, on-call operations & incident managementStrong English communication skillsWilling to participate in an on-call rotation Nice to have: Python, Django/FastAPI, ArgoCD, CI/CD, security incident response, B2B SaaS experience. Benefits International & fast-growing SaaS environmentMedical healthcare planPersonal development allowance2 weeks work-from-anywhere/yearTeam-building activities