Summary
✨ AI‑Generated
A lead SRE role responsible for establishing and owning site reliability practices within a rapidly scaling technology organization. You will lead production incident response, on-call operations, and observability while partnering with engineering and product leadership to improve system reliability and operational maturity.
Highlights
Foundational SRE leadership opportunity with ownership of incident response, on-call operations, observability, and reliability improvements in a rapidly scaling international technology environment.
Description
LEAD SITE RELIABILITY ENGINEER
(SRE)
📍 Location: Bangkok, Thailand | On-site
🏢 Industry: B2B SaaS / HRTech
👤 Reporting to: CTO / Engineering Leadership
Company Overview
Our client is a fast-growing international B2B SaaS / HRTech company headquartered in
Bangkok, with a global customer base spanning more than 135 countries.
The company provides AI-powered technology that helps businesses modernize and streamline their
recruitment and staffing processes.
As the platform and customer base continue to grow, the company is establishing a dedicated Site
Reliability Engineering (SRE) function for the first time.
Role Overview
The Lead Site Reliability Engineer will be the first dedicated SRE hire in the organization.
Working closely with the CTO, Director of Engineering, and Director of Product, this role will own
incident response, on-call operations, and the observability stack.
When production incidents occur, this person takes the lead.
Between incidents, the focus is on
improving systems, monitoring, alerting, reliability, and processes to detect and prevent future
issues.
This is a senior individual contributor role with significant cross-functional ownership.
There
are no direct reports initially, although the function may expand over time.
Key Responsibilities
Incident Response
• Lead production incidents from triage through resolution.
• Coordinate response across engineering teams and communicate status to stakeholders.
• Own incident management end-to-end: detection, escalation, communication, resolution,
and post-mortem.
• Create and maintain incident runbooks.
• Conduct blameless post-mortems and track follow-up actions through completion.
On-Call Operations
• Own and manage the on-call rotation, schedules, escalation policies, and response
expectations.
• Reduce alert noise and false positives while ensuring alerts remain actionable.
• Analyze recurring on-call issues and work with engineering teams to implement permanent
fixes.
Observability & Monitoring
• Own the observability stack, primarily Datadog for infrastructure monitoring, APM, and
logging, together with Sentry for error monitoring.
• Build and maintain dashboards, alerts, SLOs, and monitoring coverage.
• Ensure critical user flows and API endpoints are properly monitored.
• Identify and close gaps in observability coverage.
Security Incident Response
• Lead responses to security incidents such as unauthorized access, credential compromise,
and data breaches, in coordination with security leadership.
• Ensure security alerts, escalation paths, and monitoring are properly configured.
• Support penetration-test remediation validation and post-incident security reviews.
Reliability & Process Improvement
• Identify systemic reliability weaknesses such as single points of failure, insufficient
redundancy, and inadequate failover.
• Partner with engineering teams to improve deployment, rollback, and change-management
practices.
• Establish and monitor SLOs/SLIs for critical services.
• Report on reliability trends and drive continuous improvement.
Requirements
• 5+ years of experience in SRE, DevOps, Production Engineering, or a similar role within a
technology company.
• Proven experience leading production incidents end-to-end, from initial triage through
cross-functional coordination, resolution, and post-mortem.
• Strong hands-on experience with Datadog or similar observability platforms, such as
Dynatrace, New Relic, Grafana, or Prometheus.
• Experience with infrastructure monitoring, APM, log management, dashboard creation, and
alert configuration.
• Strong production experience with Kubernetes, particularly AWS EKS.
• Ability to independently troubleshoot pod failures, resource constraints, networking issues,
and deployment problems.
• Solid experience operating AWS managed services in production, including RDS, Amazon
MQ, and OpenSearch.
• Experience diagnosing performance issues, understanding failover behavior, and tuning
configurations.
• Experience with Sentry or equivalent application error-monitoring platforms.
• Proven ability to design and manage on-call rotations, escalation policies, and alerting
strategies.
• Experience defining and tracking SLOs/SLIs to guide reliability decisions.
• Excellent written and spoken English.
• Willingness to participate in on-call outside standard working hours and take a leading
role in maintaining platform availability.
Nice to Have
• Experience building an SRE function from scratch in a growing technology company.
• Python proficiency, ideally with Django or FastAPI.
• Experience with security incident response.
• B2B SaaS or platform-company experience.
• Experience with CI/CD and deployment automation, particularly ArgoCD or similar tools.
Working Environment
• Based at the company's Bangkok office in the city center.
• International team with diverse cultures and nationalities.
• English is the primary working language.
• The team works on-site and values in-person collaboration.
• International, fast-growing B2B SaaS environment.
• Medical healthcare coverage.
• Personal development allowance.
• Work-from-anywhere benefit.
• Regular team-building activities.
Hiring Process
The recruitment process may include some or all of the following:
1.
Introduction interview with HR
2.
Interview with CTO and/or Director of Engineering
3.
Technical interview and assessment
4.
Cultural-fit interview with senior management
Tech Stack
Infrastructure: AWS | Kubernetes / AWS EKS | ArgoCD / GitOps | GitHub Actions | Terraform
Backend: Python | Django | FastAPI | PostgreSQL | MongoDB | OpenSearch / Elasticsearch | Redis
| Celery | RabbitMQ
Frontend: TypeScript | Vue | React